{"id":5497,"date":"2026-09-27T12:45:07","date_gmt":"2026-09-27T03:45:07","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/format-gguf-en\/"},"modified":"2026-09-27T12:45:07","modified_gmt":"2026-09-27T03:45:07","slug":"format-gguf-en","status":"publish","type":"page","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/","title":{"rendered":"GGUF Model Format Explained: Supported Engines and Models"},"content":{"rendered":"<h2>What Is GGUF?<\/h2>\n<p><strong>GGUF<\/strong> is the model file format built for the C\/C++ inference engine <a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">llama.cpp<\/a> and the ggml library underneath it. It stores the weights (tensors) together with the metadata a runtime needs \u2014 architecture, tokenizer, chat template \u2014 <strong>in a single file<\/strong>. Download one file and it runs, which is why GGUF is the most common format in local LLM use.<\/p>\n<h2>Why It Matters<\/h2>\n<ul>\n<li><strong>A fine-grained choice of quantizations.<\/strong> The same model usually comes as Q8_0, Q6_K, Q5_K_M, Q4_K_M, Q3_K_M, IQ2_XXS and more, so you can pick the one that fits your GPU memory. See our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">quantization and model-format glossary<\/a> for how to read the names.<\/li>\n<li><strong>Runs without a GPU.<\/strong> It works on CPU alone and on NVIDIA (CUDA), AMD (ROCm \/ Vulkan) and Apple Silicon (Metal) GPUs, and can offload only part of the layers to the GPU when the model does not fit.<\/li>\n<li><strong>Broad app support.<\/strong> Ollama, LM Studio, KoboldCpp, Jan and many others use llama.cpp internally and load GGUF files directly.<\/li>\n<\/ul>\n<h2>Tips for Running It Locally<\/h2>\n<ul>\n<li><strong>Q4_K_M is the usual starting point.<\/strong> It balances quality loss and file size well and is the default recommendation of many quantizers. Move up to Q5_K_M or Q6_K if you have memory to spare.<\/li>\n<li><strong>You need the file size plus room for context.<\/strong> The KV cache grows with conversation length, so a GPU that barely fits the file is not enough. The requirements tables in our articles are computed from the actual file sizes.<\/li>\n<li><strong>Vision models need a separate <code>mmproj<\/code> file.<\/strong> The vision encoder is usually shipped as its own GGUF; download it alongside the main file.<\/li>\n<li><strong>Many models have no official GGUF.<\/strong> In that case, use a build from a quantizer such as bartowski or unsloth. The list below separates articles whose repository is itself GGUF from those where we found a GGUF build.<\/li>\n<\/ul>\n<p><em>Sources: <a href=\"https:\/\/github.com\/ggml-org\/ggml\/blob\/master\/docs\/gguf.md\">the GGUF specification (docs\/gguf.md in ggml-org\/ggml)<\/a>, <a href=\"https:\/\/huggingface.co\/docs\/hub\/gguf\">Hugging Face Hub documentation on GGUF<\/a> and <a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\">the llama.cpp README<\/a> (all as of 2026-09-27).<\/em><\/p>\n<h2>Our Coverage and Data<\/h2>\n<p>Local Model Watch has published 22 article(s) on models available in GGUF: 12 where the repository itself is in GGUF, and 10 where we found a GGUF build of the model. The lists below only include builds we have checked (the publisher&#8217;s organization and well-known quantizers); a model missing here may still have a GGUF build elsewhere. Part of our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/formats-en\/\">model format index<\/a>.<\/p>\n<h2>Main Engines That Load This Format<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Engine<\/th>\n<th>Overview<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a><\/td>\n<td>LLM inference engine written in C\/C++. Runs GGUF models on CPU and GPU (CUDA \/ Metal \/ Vulkan \/ ROCm) and underpins much of the local-LLM ecosystem, including Ollama, LM Studio and KoboldCpp.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a><\/td>\n<td>Local LLM runtime that pulls and runs models with a single command. Ships an OpenAI-compatible API server for macOS, Linux and Windows.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-koboldcpp-en\/\">KoboldCpp<\/a><\/td>\n<td>Single-file runtime that bundles llama.cpp with a web UI and API. Download a GGUF model and launch.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llamafile-en\/\">llamafile<\/a><\/td>\n<td>Packages model weights and llama.cpp into one executable that runs on any OS without installation.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-localai-en\/\">LocalAI<\/a><\/td>\n<td>Self-hosted, OpenAI-API-compatible inference server that fronts multiple backends for text, image and audio.<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-jan-en\/\">Jan<\/a><\/td>\n<td>Offline desktop chat app with llama.cpp built in; doubles as a local API server.<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Models Available in GGUF<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Published<\/th>\n<th>Model<\/th>\n<th>Where to get it<\/th>\n<th>Quantizations<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-26<\/td>\n<td>MiniMaxAI\/MiniMax-H3<\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/MiniMax-H3-GGUF\">unsloth\/MiniMax-H3-GGUF<\/a><\/td>\n<td>Q2_K, Q2_K_XL, Q3_K, Q3_K_XL, Q4_K, Q2_K_M \u2026<\/td>\n<td>32GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/minimax-h3-open-omnimodal-video-generation\/\">MiniMax-H3 Audio-Visual Video Generation Model: 32GB+ VRAM, File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>deepseek-ai\/DeepSeek-V4-Pro-0813<\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/DeepSeek-V4-Pro-0813-GGUF\">unsloth\/DeepSeek-V4-Pro-0813-GGUF<\/a>, <a href=\"https:\/\/huggingface.co\/DevQuasar\/deepseek-ai.DeepSeek-V4-Pro-0813-GGUF\">DevQuasar\/deepseek-ai.DeepSeek-V4-Pro-0813-GGUF<\/a><\/td>\n<td>Q4_K_XL, Q8_K_XL, Q2_K, Q3_K_M, Q4_K_M<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/deepseek-v4-pro-0813-released\/\">DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>Qwen\/Qwen3.8-27B<\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/Qwen3.8-27B-GGUF\">unsloth\/Qwen3.8-27B-GGUF<\/a>, <a href=\"https:\/\/huggingface.co\/lmstudio-community\/Qwen3.8-27B-GGUF\">lmstudio-community\/Qwen3.8-27B-GGUF<\/a><\/td>\n<td>IQ1_S, IQ1_M, IQ2_XXS, IQ2_S, Q2_K_XL, IQ3_XXS \u2026<\/td>\n<td>8GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B Multimodal Vision-Language Model: 8GB+ VRAM, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-25<\/td>\n<td>LiquidAI\/LFM2.5-VL-3B-DSpark<\/td>\n<td><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B-DSpark-GGUF\">LiquidAI\/LFM2.5-VL-3B-DSpark-GGUF<\/a><\/td>\n<td>F16<\/td>\n<td>4GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/liquidai-lfm2-5-vl-dspark-released\/\">LFM2.5-VL-3B-DSpark Draft Model for Vision-Language Models: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>pfnet\/plamo-3-610m-fin-instruct<\/td>\n<td><a href=\"https:\/\/huggingface.co\/mradermacher\/plamo-3-610m-fin-instruct-GGUF\">mradermacher\/plamo-3-610m-fin-instruct-GGUF<\/a><\/td>\n<td>Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S \u2026<\/td>\n<td>4GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/plamo-3-610m-fin-instruct-2\/\">plamo-3-610m-fin-instruct Text Generation Model: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>togethercomputer\/Tev1-4B-experimental<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/togethercomputer_Tev1-4B-experimental-GGUF\">bartowski\/togethercomputer_Tev1-4B-experimental-GGUF<\/a><\/td>\n<td>IQ2_M, Q2_K, IQ3_XXS, Q3_K_S, IQ3_XS, Q3_K_M \u2026<\/td>\n<td>4GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/tev1-4b-experimental-released\/\">Tev1-4B-experimental Text Generation Model: 4GB+ VRAM, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td>ggml-org\/MiMo-V2.6-Flash-RL-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf-released\/\">MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td>ggml-org\/MiMo-V2.6-Distill-Qwen-9B-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>12GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/\">MiMo-V2.6-Distill-Qwen-9B-GGUF Vision-Language Model: 12GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-21<\/td>\n<td>abenzerps\/Qwen-Image-2.1-Uncensored-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>16GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/uncensored-qwen-image-2-1-gguf-released\/\">Qwen-Image-2.1-Uncensored-GGUF Image Generation Model: 16GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>prism-ml\/Ternary-Bonsai-2-27B-gguf<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>8GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf\/\">Ternary-Bonsai-2-27B-gguf Text Generation Model: 8GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>tencent\/WeVisDoc-4B<\/td>\n<td><a href=\"https:\/\/huggingface.co\/mradermacher\/WeVisDoc-4B-GGUF\">mradermacher\/WeVisDoc-4B-GGUF<\/a>, <a href=\"https:\/\/huggingface.co\/mradermacher\/WeVisDoc-4B-i1-GGUF\">mradermacher\/WeVisDoc-4B-i1-GGUF<\/a><\/td>\n<td>Q2_K, Q3_K_S, Q3_K_M, Q3_K_L, IQ4_XS, Q4_K_S \u2026<\/td>\n<td>4GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/tencent-releases-wevisdoc-document-parsing-models\/\">Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td>bartowski\/vectionlabs_Salience-27B-R6-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>12GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">vectionlabs_Salience-27B-R6-GGUF Vision-Language Model: 12GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td>bartowski\/Intern-S2-397B-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B-GGUF Vision-Language Model: ~102GB Memory<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-14<\/td>\n<td>bartowski\/TheDrummer_Orion-26B-A4B-v1.1-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>12GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/\">TheDrummer_Orion-26B-A4B-v1.1-GGUF Multimodal Model: 12GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-13<\/td>\n<td>DavidAU\/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>12GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/13\/qwen3-8-27b-twin-turbo-fable-cold-fusion-709-l-gguf\/\">Qwen3.8-27B TWIN-TURBO Uncensored GGUF Released<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-12<\/td>\n<td>agentionai\/Signal-3.8-27B-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>16GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF Token-Efficient Optimized GGUF Model: 16GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-11<\/td>\n<td>bartowski\/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF<\/td>\n<td>This repository<\/td>\n<td>\u2014<\/td>\n<td>12GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/pantheon-reasoning-26b-gguf\/\">Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF: 12GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-09<\/td>\n<td>nex-agi\/Nex-N2.5-Pro<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/Nex-N2.5-Pro-GGUF\">bartowski\/Nex-N2.5-Pro-GGUF<\/a>, <a href=\"https:\/\/huggingface.co\/DevQuasar\/nex-agi.Nex-N2.5-Pro-GGUF\">DevQuasar\/nex-agi.Nex-N2.5-Pro-GGUF<\/a><\/td>\n<td>IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ2_M \u2026<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-agi-announces-nex-n25-pro-agent-model\/\">Nex-N2.5-Pro Long-Horizon Agent Model: ~98GB Memory, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-09<\/td>\n<td>nex-agi\/Nex-N2.5-mini<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/nex-agi_Nex-N2.5-mini-GGUF\">bartowski\/nex-agi_Nex-N2.5-mini-GGUF<\/a>, <a href=\"https:\/\/huggingface.co\/mradermacher\/Nex-N2.5-mini-GGUF\">mradermacher\/Nex-N2.5-mini-GGUF<\/a><\/td>\n<td>IQ2_XXS, IQ2_XS, IQ2_S, IQ2_M, Q2_K, IQ3_XXS \u2026<\/td>\n<td>16GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-n25-mini-released\/\">Nex-N2.5-mini Agent Model for Long-Horizon Tasks: 16GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-08<\/td>\n<td>openbmb\/MiniCPM5-2B<\/td>\n<td><a href=\"https:\/\/huggingface.co\/openbmb\/MiniCPM5-2B-GGUF\">openbmb\/MiniCPM5-2B-GGUF<\/a>, <a href=\"https:\/\/huggingface.co\/bartowski\/MiniCPM5-2B-GGUF\">bartowski\/MiniCPM5-2B-GGUF<\/a><\/td>\n<td>Q4_K_M, Q8_0, F16, IQ2_M, Q2_K, IQ3_XXS \u2026<\/td>\n<td>4GB<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">MiniCPM5-2B On-Device Model Strong in Code and Math: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>&#8220;Smallest VRAM tier&#8221; is the smallest tier in the requirements table of each article (for other builds, of those builds; for image, video and audio models, always the article&#8217;s own table, which counts every component). Leave headroom for context length.<\/em><\/p>\n<p><em>Last updated 2026-09-26 (JST). The explanation at the top of this page was written with the help of AI from the primary sources it cites. The tables under &#8220;Our Coverage and Data&#8221; are assembled by code from our article log.<\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>What Is GGUF? GGUF is the model file format built for the C\/C++ inference engine llama.cpp and the ggml library underneath it. It stores the weights (tensors) together with the metadata a runtime needs \u2014 architecture, tokenizer, chat template \u2014 in a single file. Download one file and it runs, which is why GGUF is [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-5497","page","type-page","status-publish","hentry"],"lang":"en","translations":{"en":5497,"ja":5496},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/5497","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=5497"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/5497\/revisions"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=5497"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}