{"id":2416,"date":"2026-09-22T07:10:43","date_gmt":"2026-09-21T22:10:43","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/"},"modified":"2026-09-22T07:10:43","modified_gmt":"2026-09-21T22:10:43","slug":"mimo-v2-6-distill-qwen-9b-gguf","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/","title":{"rendered":"MiMo-V2.6-Distill-Qwen-9B GGUF Released"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/ggml-org\/MiMo-V2.6-Distill-Qwen-9B-GGUF\">ggml-org\/MiMo-V2.6-Distill-Qwen-9B-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-22<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>mit<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>GGUF<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>On September 21, 2026, ggml-org released the GGUF version of &#8220;MiMo-V2.6-Distill-Qwen-9B&#8221;, an agent-oriented model developed by Xiaomi MiMo. Based on Qwen3.5-9B, this 9.4B parameter model has undergone distillation and supervised fine-tuning (SFT) using datasets generated by MiMo. It is designed to target four main domains: coding, general agent tasks, visual coding, and cybersecurity.<\/p>\n<p>According to the model card, it was released as an open research starting point for agentic reinforcement learning. The GGUF distribution also includes a Q8_0 mmproj file for the vision encoder, providing support for multimodal tasks. Readers can try out this agent-specialized model in GGUF-compatible environments such as llama.cpp and Ollama.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameters: 9.4B<\/li>\n<li>Architecture: Qwen3_5ForConditionalGeneration (Qwen 3.5)<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>The performance evaluations shown below are figures based on the technical report of the original model (XiaomiMiMo\/MiMo-V2.6-Distill-Qwen-9B) released by Xiaomi MiMo, featuring comparisons with the base model Qwen3.5-9B. Note that in the GGUF quantization format, these figures may vary slightly due to the quantization process.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Domain<\/th>\n<th>Benchmark<\/th>\n<th>Metric<\/th>\n<th style=\"text-align: right;\">Qwen3.5-9B<\/th>\n<th style=\"text-align: right;\">MiMo-V2.6-Distill-Qwen-9B (SFT)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Code<\/td>\n<td>SWE Verified<\/td>\n<td>avg@3<\/td>\n<td style=\"text-align: right;\">60.0<\/td>\n<td style=\"text-align: right;\">61.1<\/td>\n<\/tr>\n<tr>\n<td>Code<\/td>\n<td>SWE Pro<\/td>\n<td>avg@3<\/td>\n<td style=\"text-align: right;\">32.0<\/td>\n<td style=\"text-align: right;\">44.6<\/td>\n<\/tr>\n<tr>\n<td>Code<\/td>\n<td>MiMo Code (mini)\u2020<\/td>\n<td>avg@3<\/td>\n<td style=\"text-align: right;\">19.5<\/td>\n<td style=\"text-align: right;\">51.6<\/td>\n<\/tr>\n<tr>\n<td>Cyber<\/td>\n<td>MiMo Cyber (mini)\u2020<\/td>\n<td>avg@3<\/td>\n<td style=\"text-align: right;\">5.7<\/td>\n<td style=\"text-align: right;\">31.3<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td>AutomationBench v1.0.6<\/td>\n<td>avg@1<\/td>\n<td style=\"text-align: right;\">5.0<\/td>\n<td style=\"text-align: right;\">30.3<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td>Terminal Bench 2.1<\/td>\n<td>avg@1<\/td>\n<td style=\"text-align: right;\">27.0<\/td>\n<td style=\"text-align: right;\">37.1<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td>Toolathlon-Verified<\/td>\n<td>avg@1<\/td>\n<td style=\"text-align: right;\">25.9<\/td>\n<td style=\"text-align: right;\">35.2<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td>OfficeQA<\/td>\n<td>avg@1<\/td>\n<td style=\"text-align: right;\">9.0<\/td>\n<td style=\"text-align: right;\">19.5<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td>JobBench<\/td>\n<td>avg@1<\/td>\n<td style=\"text-align: right;\">2.6<\/td>\n<td style=\"text-align: right;\">18.3<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td>MiMo General (mini)\u2020<\/td>\n<td>avg@1<\/td>\n<td style=\"text-align: right;\">28.5<\/td>\n<td style=\"text-align: right;\">62.2<\/td>\n<\/tr>\n<tr>\n<td>Visual<\/td>\n<td>MiMo Visual Coding (mini)\u2020<\/td>\n<td>avg@1<\/td>\n<td style=\"text-align: right;\">61.7<\/td>\n<td style=\"text-align: right;\">64.0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>\u2020 Figures from internal evaluation sets<\/p>\n<p>From the publisher&#8217;s measurement results, it can be seen that this model improves scores across almost all evaluation items compared to the base model Qwen3.5-9B. Notably, the score increase on &#8220;Terminal Bench 2.1&#8221; from 27.0 to 37.1 is worth highlighting. Terminal-Bench is an indicator that measures practical ability as an agent\u2014whether it can actually execute commands on a terminal and see a given task through to completion\u2014suggesting that this model is strong in tasks involving terminal operations.<\/p>\n<p>In addition, dramatic improvements are seen in specific specialized areas, such as moving from 32.0 to 44.6 on SWE Pro, a challenging benchmark in the coding field, and from 5.7 to 31.3 on the internal cyber security evaluation (MiMo Cyber mini). On the other hand, items like SWE Verified showed only a marginal improvement from 60.0 to 61.1, indicating that it does not show a uniformly overwhelming difference across all coding tasks. Overall, it can be said to be a model that demonstrates particular strength in areas requiring agentic behavior, such as automation (AutomationBench) and tool use (Toolathlon-Verified).<\/p>\n<p>For reference, the composition of the data used to train this model is reported as follows. Out of a total 77.4B tokens, loss-bearing tokens contributing to learning account for 27.2B tokens.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Domain<\/th>\n<th style=\"text-align: right;\">Total tokens (B)<\/th>\n<th style=\"text-align: right;\">Token share (%)<\/th>\n<th style=\"text-align: right;\">Loss-bearing tokens (B)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Code<\/td>\n<td style=\"text-align: right;\">23.2<\/td>\n<td style=\"text-align: right;\">29.9<\/td>\n<td style=\"text-align: right;\">7.3<\/td>\n<\/tr>\n<tr>\n<td>Cyber<\/td>\n<td style=\"text-align: right;\">11.0<\/td>\n<td style=\"text-align: right;\">14.2<\/td>\n<td style=\"text-align: right;\">4.8<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td style=\"text-align: right;\">22.0<\/td>\n<td style=\"text-align: right;\">28.5<\/td>\n<td style=\"text-align: right;\">5.7<\/td>\n<\/tr>\n<tr>\n<td>Visual<\/td>\n<td style=\"text-align: right;\">21.2<\/td>\n<td style=\"text-align: right;\">27.4<\/td>\n<td style=\"text-align: right;\">9.4<\/td>\n<\/tr>\n<tr>\n<td><strong>Total<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>77.4<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>100.0<\/strong><\/td>\n<td style=\"text-align: right;\"><strong>27.2<\/strong><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Strengths and Use Cases<\/h2>\n<p>The greatest identity of this model lies in its advanced autonomy as an &#8220;agent&#8221; that goes beyond mere text generation. As developer Xiaomi MiMo positions this model as a research foundation for agentic reinforcement learning, it excels in the ability to interpret complex instructions, break them down into concrete steps, and execute them. By applying supervised fine-tuning (SFT) using high-quality datasets generated by MiMo to the base model Qwen3.5-9B, it has achieved dramatic evolution in specific specialized domains.<\/p>\n<p>Particularly noteworthy is its adaptability to tasks involving terminal operations. The high score on &#8220;Terminal-Bench 2.1&#8221; mentioned in the performance section indicates that this model can accurately simulate or direct command-line operations in local or server environments. It holds the potential to directly support engineers&#8217; daily workflows, such as creating system administration automation scripts, building complex deployment procedures, and even interactive system troubleshooting. Regarding tool-use, as indicated by the score improvement in Toolathlon-Verified, its capability to properly invoke external tools to complete tasks has been enhanced.<\/p>\n<p>Furthermore, the model is designed to explicitly output a &#8220;thinking&#8221; process. This means the model takes internal reasoning steps before producing a final answer, which has the effect of preventing logical breakdowns, especially in mathematical calculations and complex algorithm construction. Users can enable the thinking function via API to check <code>reasoning_content<\/code> and see what logical progression the model took to reach a conclusion. This transparency is extremely useful in debugging tasks or critical scenes where humans need to verify the grounds for the model&#8217;s decisions.<\/p>\n<p>The breadth of supported domains and specialization is also noteworthy. The training data strategically blends the following four areas:<\/p>\n<ul>\n<li><strong>Fusion of Coding and Vision<\/strong>: Through the enhancement of the &#8220;Visual Coding&#8221; domain, it supports development tasks starting from visual information, such as reading website screenshots or UI blueprints and generating HTML\/CSS or React code to realize them. 27.4% of the training data is allocated to visual-related content (Visual), expecting advanced coordination between image recognition and code generation.<\/li>\n<li><strong>Cybersecurity<\/strong>: Cybersecurity-related data amounting to 11 billion tokens (11.0B tokens) has been fed into the model. This anticipates applications in highly specialized security tasks that are difficult for general LLMs, such as assisting with vulnerability diagnostics, analyzing security logs, and creating penetration testing scenarios.<\/li>\n<li><strong>General Agent Tasks<\/strong>: As shown by score improvements in benchmarks like OfficeQA and JobBench, practical agent functions such as office automation and job process management have been enhanced.<\/li>\n<\/ul>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 9.4B parameters (taken from the base model XiaomiMiMo\/MiMo-V2.6-Distill-Qwen-9B)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>12GB (RTX 4070 \/ 3060 12GB, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>8.9GB<\/td>\n<td>10.6GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model is released in GGUF format and can be downloaded from the <code>ggml-org<\/code> repository. The adoption of the GGUF format enables operations that flexibly combine CPU inference and GPU offloading, allowing the 9.4B parameter model to run efficiently even in local environments with limited memory resources.<\/p>\n<p>Tools from the <code>llama.cpp<\/code> ecosystem are recommended for acquisition and execution. For example, when using <code>llama.app<\/code>, you can load the model directly from Hugging Face and launch an API server with the following command:<\/p>\n<pre><code class=\"language-bash\">llama serve -hf ggml-org\/MiMo-V2.6-Distill-Qwen-9B-GGUF\n<\/code><\/pre>\n<p>Since this model is multimodal, an <code>mmproj<\/code> file for the vision encoder is required in addition to the main model file for text. The repository includes a <code>Q8_0<\/code> quantized <code>mmproj<\/code> file, which should be specified and loaded when using image inputs. Even in the quantized version, efforts are made to maintain visual information recognition accuracy by allocating a high-precision bit count of Q8_0 to the vision encoder.<\/p>\n<p>Additionally, based on the specifications of the original model, it is recommended to use the MiMo v2.6 exclusive chat template during inference. When using engines such as SGLang, passing the <code>--reasoning-parser mimo<\/code> flag allows you to properly parse and utilize the aforementioned thinking process. It is a very easy-to-handle package for engineers looking to build a &#8220;thinking agent&#8221; in a local environment.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/ggml-v0-24-0-released-2\/\">ggml v0.24.0 Released with Backend Improvements and API Updates<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/ggml-org\/MiMo-V2.6-Distill-Qwen-9B-GGUF\">https:\/\/huggingface.co\/ggml-org\/MiMo-V2.6-Distill-Qwen-9B-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/XiaomiMiMo\/MiMo-V2.6-Distill-Qwen-9B\">https:\/\/huggingface.co\/XiaomiMiMo\/MiMo-V2.6-Distill-Qwen-9B<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Download the GGUF version of MiMo-V2.6-Distill-Qwen-9B, a 9.4B agentic model for coding and cybersecurity.<\/p>\n","protected":false},"author":1,"featured_media":2415,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[161,163,1950,165,520,996,1547,1952,1954],"class_list":["post-2416","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-ggml-en","tag-gguf-en","tag-mimo-v2-6-distill-qwen-9b-en","tag-moe-en","tag-qwen-en","tag-qwen3-5-en","tag-verified","tag-vlm-en","tag-xiaomi-en"],"lang":"en","translations":{"en":2416,"ja":2414},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2416","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=2416"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2416\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/2415"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=2416"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=2416"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=2416"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}