Open-Weight Models Weekly: The MiMo-V2.6 Family & Transformers GGUF Support

Open-Weight Models Weekly: The MiMo-V2.6 Family & Transformers GGUF Support

This Week in Numbers

Counted by Local Model Watch from the articles published this week.

Item Count
Articles published 47
 New Models 12
 Engines and Tools 10
 Technical Reports 9
 Image, Video and Audio 7
 Companies and Funding 6
 Community 3
New models covered 19
 fit in 8GB of VRAM (est.) 5
 fit in 12GB of VRAM (est.) 7
 fit in 16GB of VRAM (est.) 8
 fit in 24GB of VRAM (est.) 8
 parameters 4–15B 5
 parameters up to 4B 3
 parameters 15–40B 1
 parameters over 40B 1
Converted builds appended to earlier articles 17 (MLX 3, GGUF 9, FP8 3, NVFP4 1, MXFP4 1)
Most active publishers nvidia developer (5), hugging face blog (3), xiaomimimo (3)
Articles still marked unverified 2

Trending models we did not cover separately

Hugging Face repositories that trended this week but did not get their own article. Listed for reference; we have not reviewed their model cards.

Model Likes Downloads Why no article
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B 522 8,839 publisher not on our notable list
XiaomiMiMo/MiMo-V2.6-Flash-RL 491 25,661 publisher not on our notable list
Contrastive-LM/CLM-v0.1-8B 408 766 publisher not on our notable list
pottokao/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF 299 145,246 publisher not on our notable list
apple/LensVLM-9B 243 1,740 publisher not on our notable list
orcarouter/OrcaSAQ-2-27B 169 1,330 publisher not on our notable list
ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF 117 28,802 publisher not on our notable list
cua-ai/cua-s1-forms 111 0 publisher not on our notable list
bottlecapai/ThinkingCap-Qwen3.8-27B 110 609 publisher not on our notable list
orcarouter/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF 99 313 publisher not on our notable list

Patch releases we did not cover separately

Releases of watched projects that were patch-level or had short notes. Each project’s page lists every version.

Project Version Release notes
LostRuins/koboldcpp v1.122.1 GitHub
ggml-org/ggml v0.25.1 GitHub
ggml-org/ggml v0.25.3 GitHub
ollama/ollama v0.34.4 GitHub
turboderp-org/exllamav3 v1.5.2 GitHub
turboderp-org/exllamav3 v1.5.3 GitHub
unslothai/unsloth prebuilt-wheels-cu13 GitHub
unslothai/unsloth v0.1.815-beta GitHub

Weekly Highlights

The most visible development this week was Xiaomi’s MiMo-V2.6 family arriving together. Alongside “MiMo-V2.6-Pro-RL", a 1.02T-parameter multimodal MoE model, ggml-org published GGUF builds in “MiMo-V2.6-Flash-RL-GGUF" (about 141GB of memory required) and “MiMo-V2.6-Distill-Qwen-9B-GGUF" (9.4B, 12GB+ VRAM), followed by MiMo-V2.6-MOPD and “MiMo-V2.6-RL-oss", a reinforcement-learning environment for agents. With everything from a server-class flagship to a distilled model that fits in the 12GB tier landing in the same week, and GGUF builds from llama.cpp’s own organization available early, this matters to readers who want to try the models locally.

On the inference infrastructure side, Hugging Face transformers added support for directly running GGUF models (Hugging Face transformers adds direct execution support for GGUF models). Since the GGUF format, which has traditionally tended to be siloed within the llama.cpp ecosystem, has begun to seamlessly integrate with major Python frameworks, the scope for development and application is expanding significantly.

Furthermore, the rise of lightweight decision-making models focused on specific tasks cannot be overlooked. Together AI announced “Tev1-4B-experimental" and “Tev1-0.8B-experimental", which are designed to perform structured data decision-making with extremely small resource requirements starting from a minimum of 4GB VRAM. Rather than a pure focus on scaling up, a miniaturization and specialized approach for practical local execution has become a clear trend.

Trends by Domain

Text Generation

In the text generation domain, the growing range of small models that run on a local PC or a single GPU stood out. Four of the models we covered this week fit within 4GB of VRAM, showing a strong focus on edge and local deployment.

Particularly notable was the trend of small models specialized for decision-making and classification. In addition to “Tev1-4B-experimental" and the 873M “Tev1-0.8B-experimental", the community also saw the release of “Kev", a derivative of Qwen3.5. Additionally, Together AI shared a method for building low-cost models based on Qwen3.5 4B (Together AI shares method to train and build Jev-style classification model for about $17 based on Qwen3.5 4B). Echoing these developments, initiatives like the local runtime for decision models “Ollaya" have emerged, though Ollaya contains unverified community-sourced information, and its future trajectory needs to be monitored carefully.

As a language- and domain-specific model, Preferred Networks released “plamo-3-610m-fin-instruct" (890M parameters, requiring 4GB+ VRAM), a Japanese finance-focused model that adds a practical option for lightweight Japanese LLMs.

Meanwhile, in the server-class tier, large models combining MoE architectures with newer quantization formats kept coming, including “K2-Horizon-375B-A23B-NVFP4" and “K2-Horizon-32B-NVFP4" (both NVFP4) and the MiMo-V2.6-Pro-RL mentioned above.

This week we also published our articles on the base models “Qwen3.8-27B" and “DeepSeek-V4-Pro-0813", both released in August. We had already covered their derivatives, and neither is a new arrival this week.

Image, Video and Audio

In the media generation domain, the rapid expansion of the Qwen-Image ecosystem and the diversification of task-specific architectures are progressing simultaneously.

Regarding Qwen-Image-2.1, deployments included “Qwen-Image-2.1-Uncensored-GGUF" (requiring 16GB+ VRAM), a GGUF quantized version easy to handle in ComfyUI, and “Qwen-Image-2.1-viggle-turbo" (requiring 48GB+ VRAM) which performs high-speed generation in just 4 steps. Furthermore, specialization beyond mere general-purpose image generation is advancing, with the introduction of “Ming-Image" as a model specialized in UI design generation.

In the audio realm as well, it was a week where practical local models were established across visual and audio modalities, marked by Google DeepMind’s publication of the text-to-speech model “Gemini 3.8 Flash TTS" and the release of “Audio8-ASR-Infinite" (4.1B parameters, supporting streaming speech recognition, requiring 12GB+ VRAM).

Engines and Tools

Updates to inference engines and development tools supporting model diversification also followed one after another. What they have in common is lightweight operation, high speed, and prompt model compatibility.

In inference infrastructure, updates to C/C++ libraries were prominent. Along with the release of “llama.cpp v0.5.0", its foundational ggml library also saw successive releases of “ggml v0.25.0" and “ggml v0.25.2", strengthening FlashAttention, MoE, and optimizations for various backends. Additionally, the frontend tool “KoboldCpp v1.122" added agent features, and the container-based model runner “ramalama v0.25.0" has also incorporated the latest llama.cpp.

In UI and framework environments, “ComfyUI v0.37.0" implemented automatic fast-disk optimization and Qwen support, while “Unsloth" added Qwen-Image-2.1 support and Agent Skills features. Furthermore, the WebUI environment “Open WebUI v0.11.4" significantly reduced its Docker image size to around 175MB, representing a steady accumulation of overall improvements that reduce friction when building and operating local AI environments.

Related Articles

This Week’s Articles

Compiled by Local Model Watch from the article log. Articles about the same story are merged into one row.

New Models

→ Scroll horizontally to see all columns

Date Model Params Smallest VRAM tier License Converted builds Article
2026-09-28 XiaomiMiMo/MiMo-V2.6-Flash-MOPD — — mit — Xiaomi Releases MiMo-V2.6-Flash-MOPD and Pro-MOPD Models
2026-09-27 XiaomiMiMo/MiMo-V2.6-Pro-RL — — mit 1 Xiaomi Releases MiMo-V2.6-Pro-RL: 1.02T MoE Flagship
2026-09-26 deepseek-ai/DeepSeek-V4-Pro-0813 1650.5B — mit 3 DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds
2026-09-26 Qwen/Qwen3.8-27B 27.8B 8GB apache-2.0 6 Qwen3.8-27B Multimodal Vision-Language Model: 8GB+ VRAM, GGUF Builds
2026-09-26 togethercomputer/Tev1-0.8B-experimental 873M 4GB — — Tev1-0.8B-experimental Text Generation Model: 4GB+ VRAM
2026-09-25 LiquidAI/LFM2.5-VL-3B-DSpark 279M 4GB — 1 LFM2.5-VL-3B-DSpark Draft Model for Vision-Language Models: 4GB+ VRAM
2026-09-24 pfnet/plamo-3-610m-fin-instruct 890M 4GB other 1 plamo-3-610m-fin-instruct Text Generation Model: 4GB+ VRAM
2026-09-24 togethercomputer/Tev1-4B-experimental 4.7B 4GB — 1 Tev1-4B-experimental Text Generation Model: 4GB+ VRAM, GGUF Builds
2026-09-23 IFM/K2-Horizon-32B-NVFP4 — 32GB apache-2.0 — K2-Horizon-32B-NVFP4 Long-Context Reasoning Model: 32GB+ VRAM
2026-09-23 IFM/K2-Horizon-375B-A23B-NVFP4 — — apache-2.0 — K2-Horizon-375B-A23B-NVFP4 Text Generation Model: ~257GB Memory
2026-09-22 ggml-org/MiMo-V2.6-Flash-RL-GGUF — — mit — MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory
2026-09-22 ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF 9.4B 12GB mit — MiMo-V2.6-Distill-Qwen-9B-GGUF Vision-Language Model: 12GB+ VRAM

Image, Video and Audio

→ Scroll horizontally to see all columns

Date Model Params Smallest VRAM tier License Converted builds Article
2026-09-26 MiniMaxAI/MiniMax-H3 — 32GB other 1 MiniMax-H3 Audio-Visual Video Generation Model: 32GB+ VRAM, File List
2026-09-24 Comfy-Org/Ming-Image — — mit — Ming-Image UI Design-Specialized Image Generation Model: ComfyUI Paths
2026-09-24 Edge0/Audio8-ASR-Infinite 4.1B 12GB apache-2.0 — Audio8-ASR-Infinite Speech Recognition Model: 12GB+ VRAM, File List
2026-09-24 — — — — — Google DeepMind Announces Gemini 3.8 Flash TTS Models
2026-09-23 Viggle/Qwen-Image-2.1-viggle-turbo 7.1B 48GB other — Qwen-Image-2.1-viggle-turbo Image Generation Model: 48GB+ VRAM
2026-09-23 SupraLabs/Supra2-IMG — — apache-2.0 — Supra2-IMG Lightweight Text-to-Image Model: File List
2026-09-21 abenzerps/Qwen-Image-2.1-Uncensored-GGUF 7.1B 16GB other 2 Qwen-Image-2.1-Uncensored-GGUF Image Generation Model: 16GB+ VRAM

Engines and Tools

Date Project Version Article
2026-09-28 invoke-ai/InvokeAI v6.14.2 InvokeAI v6.14.2 Released with Custom Fonts and Model Fixes
2026-09-26 LostRuins/koboldcpp v1.122 KoboldCpp v1.122 Released with Built-in Agentic Framework
2026-09-25 containers/ramalama v0.25.0 RamaLama v0.25.0 Released: Security and Engine Updates
2026-09-24 ggml-org/ggml v0.25.2 ggml v0.25.2 Released with Hardware Backend Optimizations
2026-09-24 ggml-org/llama.cpp v0.5.0 llama.cpp v0.5.0 Released with Backend and Server Upgrades
2026-09-23 ggml-org/ggml v0.25.0 ggml v0.25.0 Released: FlashAttention and MoE Optimizations
2026-09-23 unslothai/unsloth v0.1.812-beta Unsloth Update: Qwen-Image-2.1 Support and Agent Skills Added
2026-09-22 vllm-project/vllm v0.30.0 vLLM v0.30.0 Released: Fast Start Weight Caching and New Models
2026-09-22 open-webui/open-webui v0.11.4 Open WebUI v0.11.4 Released: Dramatic Docker Image Slimming
2026-09-21 Comfy-Org/ComfyUI v0.37.0

Community

Companies and Funding

Technical Reports