{"id":6078,"date":"2026-09-28T09:18:52","date_gmt":"2026-09-28T00:18:52","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/open-weight-models-weekly-2026-09-26\/"},"modified":"2026-09-28T11:10:49","modified_gmt":"2026-09-28T02:10:49","slug":"open-weight-models-weekly-2026-09-26","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/open-weight-models-weekly-2026-09-26\/","title":{"rendered":"Open-Weight Models Weekly: The MiMo-V2.6 Family &#038; Transformers GGUF Support"},"content":{"rendered":"<p><!-- lmw:weekly-stats --><\/p>\n<h2>This Week in Numbers<\/h2>\n<p><em>Counted by Local Model Watch from the articles published this week.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Count<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Articles published<\/td>\n<td>47<\/td>\n<\/tr>\n<tr>\n<td>\u3000New Models<\/td>\n<td>12<\/td>\n<\/tr>\n<tr>\n<td>\u3000Engines and Tools<\/td>\n<td>10<\/td>\n<\/tr>\n<tr>\n<td>\u3000Technical Reports<\/td>\n<td>9<\/td>\n<\/tr>\n<tr>\n<td>\u3000Image, Video and Audio<\/td>\n<td>7<\/td>\n<\/tr>\n<tr>\n<td>\u3000Companies and Funding<\/td>\n<td>6<\/td>\n<\/tr>\n<tr>\n<td>\u3000Community<\/td>\n<td>3<\/td>\n<\/tr>\n<tr>\n<td>New models covered<\/td>\n<td>19<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 8GB of VRAM (est.)<\/td>\n<td>5<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 12GB of VRAM (est.)<\/td>\n<td>7<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 16GB of VRAM (est.)<\/td>\n<td>8<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 24GB of VRAM (est.)<\/td>\n<td>8<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters 4\u201315B<\/td>\n<td>5<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters up to 4B<\/td>\n<td>3<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters 15\u201340B<\/td>\n<td>1<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters over 40B<\/td>\n<td>1<\/td>\n<\/tr>\n<tr>\n<td>Converted builds appended to earlier articles<\/td>\n<td>17 (MLX 3, GGUF 9, FP8 3, NVFP4 1, MXFP4 1)<\/td>\n<\/tr>\n<tr>\n<td>Most active publishers<\/td>\n<td>nvidia developer (5), hugging face blog (3), xiaomimimo (3)<\/td>\n<\/tr>\n<tr>\n<td>Articles still marked unverified<\/td>\n<td>2<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Trending models we did not cover separately<\/h3>\n<p><em>Hugging Face repositories that trended this week but did not get their own article. Listed for reference; we have not reviewed their model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Likes<\/th>\n<th>Downloads<\/th>\n<th>Why no article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/XiaomiMiMo\/MiMo-V2.6-Distill-Qwen-9B\">XiaomiMiMo\/MiMo-V2.6-Distill-Qwen-9B<\/a><\/td>\n<td>522<\/td>\n<td>8,839<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/XiaomiMiMo\/MiMo-V2.6-Flash-RL\">XiaomiMiMo\/MiMo-V2.6-Flash-RL<\/a><\/td>\n<td>491<\/td>\n<td>25,661<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/Contrastive-LM\/CLM-v0.1-8B\">Contrastive-LM\/CLM-v0.1-8B<\/a><\/td>\n<td>408<\/td>\n<td>766<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/pottokao\/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF\">pottokao\/Qwen-Image-2.1-Text-Encoder-Heretic-GGUF<\/a><\/td>\n<td>299<\/td>\n<td>145,246<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/apple\/LensVLM-9B\">apple\/LensVLM-9B<\/a><\/td>\n<td>243<\/td>\n<td>1,740<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/orcarouter\/OrcaSAQ-2-27B\">orcarouter\/OrcaSAQ-2-27B<\/a><\/td>\n<td>169<\/td>\n<td>1,330<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/ukisai\/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF\">ukisai\/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF<\/a><\/td>\n<td>117<\/td>\n<td>28,802<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/cua-ai\/cua-s1-forms\">cua-ai\/cua-s1-forms<\/a><\/td>\n<td>111<\/td>\n<td>0<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/bottlecapai\/ThinkingCap-Qwen3.8-27B\">bottlecapai\/ThinkingCap-Qwen3.8-27B<\/a><\/td>\n<td>110<\/td>\n<td>609<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/orcarouter\/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF\">orcarouter\/OrcaSAQ-2-Cyber-27B-Uncensored-GGUF<\/a><\/td>\n<td>99<\/td>\n<td>313<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Patch releases we did not cover separately<\/h3>\n<p><em>Releases of watched projects that were patch-level or had short notes. Each project&#8217;s page lists every version.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Project<\/th>\n<th>Version<\/th>\n<th>Release notes<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>LostRuins\/koboldcpp<\/td>\n<td>v1.122.1<\/td>\n<td><a href=\"https:\/\/github.com\/LostRuins\/koboldcpp\/releases\/tag\/v1.122.1\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>ggml-org\/ggml<\/td>\n<td>v0.25.1<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/ggml\/releases\/tag\/v0.25.1\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>ggml-org\/ggml<\/td>\n<td>v0.25.3<\/td>\n<td><a href=\"https:\/\/github.com\/ggml-org\/ggml\/releases\/tag\/v0.25.3\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>ollama\/ollama<\/td>\n<td>v0.34.4<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.34.4\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>turboderp-org\/exllamav3<\/td>\n<td>v1.5.2<\/td>\n<td><a href=\"https:\/\/github.com\/turboderp-org\/exllamav3\/releases\/tag\/v1.5.2\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>turboderp-org\/exllamav3<\/td>\n<td>v1.5.3<\/td>\n<td><a href=\"https:\/\/github.com\/turboderp-org\/exllamav3\/releases\/tag\/v1.5.3\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>unslothai\/unsloth<\/td>\n<td>prebuilt-wheels-cu13<\/td>\n<td><a href=\"https:\/\/github.com\/unslothai\/unsloth\/releases\/tag\/prebuilt-wheels-cu13\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>unslothai\/unsloth<\/td>\n<td>v0.1.815-beta<\/td>\n<td><a href=\"https:\/\/github.com\/unslothai\/unsloth\/releases\/tag\/v0.1.815-beta\">GitHub<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:weekly-stats --><\/p>\n<h2>Weekly Highlights<\/h2>\n<p>The most visible development this week was Xiaomi&#8217;s MiMo-V2.6 family arriving together. Alongside <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/27\/xiaomimimo-v2-6-pro-rl\/\">&#8220;MiMo-V2.6-Pro-RL&#8221;<\/a>, a 1.02T-parameter multimodal MoE model, ggml-org published GGUF builds in <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf\/\">&#8220;MiMo-V2.6-Flash-RL-GGUF&#8221;<\/a> (about 141GB of memory required) and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/22\/mimo-v26-distill-qwen-9b-gguf\/\">&#8220;MiMo-V2.6-Distill-Qwen-9B-GGUF&#8221;<\/a> (9.4B, 12GB+ VRAM), followed by <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/mimo-v26-mopd\/\">MiMo-V2.6-MOPD<\/a> and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/27\/mimo-v26-rl-oss\/\">&#8220;MiMo-V2.6-RL-oss&#8221;<\/a>, a reinforcement-learning environment for agents. With everything from a server-class flagship to a distilled model that fits in the 12GB tier landing in the same week, and GGUF builds from llama.cpp&#8217;s own organization available early, this matters to readers who want to try the models locally.<\/p>\n<p>On the inference infrastructure side, Hugging Face transformers added support for directly running GGUF models (<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/22\/transformers-llama-cpp-quants\/\">Hugging Face transformers adds direct execution support for GGUF models<\/a>). Since the GGUF format, which has traditionally tended to be siloed within the llama.cpp ecosystem, has begun to seamlessly integrate with major Python frameworks, the scope for development and application is expanding significantly.<\/p>\n<p>Furthermore, the rise of lightweight decision-making models focused on specific tasks cannot be overlooked. Together AI announced <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/tev1-4b-experimental\/\">&#8220;Tev1-4B-experimental&#8221;<\/a> and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/tev1-08b-experimental\/\">&#8220;Tev1-0.8B-experimental&#8221;<\/a>, which are designed to perform structured data decision-making with extremely small resource requirements starting from a minimum of 4GB VRAM. Rather than a pure focus on scaling up, a miniaturization and specialized approach for practical local execution has become a clear trend.<\/p>\n<h2>Trends by Domain<\/h2>\n<h3>Text Generation<\/h3>\n<p>In the text generation domain, the growing range of small models that run on a local PC or a single GPU stood out. Four of the models we covered this week fit within 4GB of VRAM, showing a strong focus on edge and local deployment.<\/p>\n<p>Particularly notable was the trend of small models specialized for decision-making and classification. In addition to <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/tev1-4b-experimental\/\">&#8220;Tev1-4B-experimental&#8221;<\/a> and the 873M <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/tev1-08b-experimental\/\">&#8220;Tev1-0.8B-experimental&#8221;<\/a>, the community also saw the release of <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/21\/kev-qwen35-decision-models\/\">&#8220;Kev&#8221;<\/a>, a derivative of Qwen3.5. Additionally, Together AI shared a method for building low-cost models based on Qwen3.5 4B (<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/train-your-own-jev-together-ai\/\">Together AI shares method to train and build Jev-style classification model for about $17 based on Qwen3.5 4B<\/a>). Echoing these developments, initiatives like the local runtime for decision models <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/ollaya\/\">&#8220;Ollaya&#8221;<\/a> have emerged, though Ollaya contains unverified community-sourced information, and its future trajectory needs to be monitored carefully.<\/p>\n<p>As a language- and domain-specific model, Preferred Networks released <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/plamo-3-610m-fin-instruct\/\">&#8220;plamo-3-610m-fin-instruct&#8221;<\/a> (890M parameters, requiring 4GB+ VRAM), a Japanese finance-focused model that adds a practical option for lightweight Japanese LLMs.<\/p>\n<p>Meanwhile, in the server-class tier, large models combining MoE architectures with newer quantization formats kept coming, including <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/23\/k2-horizon-375b-a23b-nvfp4\/\">&#8220;K2-Horizon-375B-A23B-NVFP4&#8221;<\/a> and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/23\/ifm-k2-horizon-32b-nvfp4\/\">&#8220;K2-Horizon-32B-NVFP4&#8221;<\/a> (both NVFP4) and the MiMo-V2.6-Pro-RL mentioned above.<\/p>\n<p>This week we also published our articles on the base models <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/qwen3-8-27b\/\">&#8220;Qwen3.8-27B&#8221;<\/a> and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/deepseek-v4-pro-0813\/\">&#8220;DeepSeek-V4-Pro-0813&#8221;<\/a>, both released in August. We had already covered their derivatives, and neither is a new arrival this week.<\/p>\n<h3>Image, Video and Audio<\/h3>\n<p>In the media generation domain, the rapid expansion of the Qwen-Image ecosystem and the diversification of task-specific architectures are progressing simultaneously.<\/p>\n<p>Regarding Qwen-Image-2.1, deployments included <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/21\/abenzerps-qwen-image-21-uncensored-gguf\/\">&#8220;Qwen-Image-2.1-Uncensored-GGUF&#8221;<\/a> (requiring 16GB+ VRAM), a GGUF quantized version easy to handle in ComfyUI, and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/23\/viggle-qwen-image-2-1-viggle-turbo\/\">&#8220;Qwen-Image-2.1-viggle-turbo&#8221;<\/a> (requiring 48GB+ VRAM) which performs high-speed generation in just 4 steps. Furthermore, specialization beyond mere general-purpose image generation is advancing, with the introduction of <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/comfy-org-ming-image\/\">&#8220;Ming-Image&#8221;<\/a> as a model specialized in UI design generation.<\/p>\n<p>In the audio realm as well, it was a week where practical local models were established across visual and audio modalities, marked by Google DeepMind&#8217;s publication of the text-to-speech model <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/gemini-3-8-flash-tts\/\">&#8220;Gemini 3.8 Flash TTS&#8221;<\/a> and the release of <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/edge0-audio8-asr-infinite\/\">&#8220;Audio8-ASR-Infinite&#8221;<\/a> (4.1B parameters, supporting streaming speech recognition, requiring 12GB+ VRAM).<\/p>\n<h3>Engines and Tools<\/h3>\n<p>Updates to inference engines and development tools supporting model diversification also followed one after another. What they have in common is lightweight operation, high speed, and prompt model compatibility.<\/p>\n<p>In inference infrastructure, updates to C\/C++ libraries were prominent. Along with the release of <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/llama-cpp-v0-5-0\/\">&#8220;llama.cpp v0.5.0&#8221;<\/a>, its foundational ggml library also saw successive releases of <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/23\/ggml-v0-25-0\/\">&#8220;ggml v0.25.0&#8221;<\/a> and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/24\/ggml-v0252-released\/\">&#8220;ggml v0.25.2&#8221;<\/a>, strengthening FlashAttention, MoE, and optimizations for various backends. Additionally, the frontend tool <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/koboldcpp-v1122\/\">&#8220;KoboldCpp v1.122&#8221;<\/a> added agent features, and the container-based model runner <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/25\/ramalama-v0250-released\/\">&#8220;ramalama v0.25.0&#8221;<\/a> has also incorporated the latest llama.cpp.<\/p>\n<p>In UI and framework environments, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/21\/comfyui-v0370-update\/\">&#8220;ComfyUI v0.37.0&#8221;<\/a> implemented automatic fast-disk optimization and Qwen support, while <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/23\/unsloth-qwen-image-21-skills\/\">&#8220;Unsloth&#8221;<\/a> added Qwen-Image-2.1 support and Agent Skills features. Furthermore, the WebUI environment <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/22\/open-webui-v0114\/\">&#8220;Open WebUI v0.11.4&#8221;<\/a> significantly reduced its Docker image size to around 175MB, representing a steady accumulation of overall improvements that reduce friction when building and operating local AI environments.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF Token-Efficient Optimized GGUF Model: 16GB+ VRAM<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:weekly-articles --><\/p>\n<h2>This Week&#8217;s Articles<\/h2>\n<p><em>Compiled by Local Model Watch from the article log. Articles about the same story are merged into one row.<\/em><\/p>\n<h3>New Models<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Model<\/th>\n<th>Params<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Converted builds<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-28<\/td>\n<td>XiaomiMiMo\/MiMo-V2.6-Flash-MOPD<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>mit<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/xiaomi-mimo-v2-6-mopd-models\/\">Xiaomi Releases MiMo-V2.6-Flash-MOPD and Pro-MOPD Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-27<\/td>\n<td>XiaomiMiMo\/MiMo-V2.6-Pro-RL<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>mit<\/td>\n<td>1<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/27\/mimo-v2-6-pro-rl-released\/\">Xiaomi Releases MiMo-V2.6-Pro-RL: 1.02T MoE Flagship<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>deepseek-ai\/DeepSeek-V4-Pro-0813<\/td>\n<td>1650.5B<\/td>\n<td>\u2014<\/td>\n<td>mit<\/td>\n<td>3<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/deepseek-v4-pro-0813-released\/\">DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>Qwen\/Qwen3.8-27B<\/td>\n<td>27.8B<\/td>\n<td>8GB<\/td>\n<td>apache-2.0<\/td>\n<td>6<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B Multimodal Vision-Language Model: 8GB+ VRAM, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>togethercomputer\/Tev1-0.8B-experimental<\/td>\n<td>873M<\/td>\n<td>4GB<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/tev1-08b-experimental-2\/\">Tev1-0.8B-experimental Text Generation Model: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-25<\/td>\n<td>LiquidAI\/LFM2.5-VL-3B-DSpark<\/td>\n<td>279M<\/td>\n<td>4GB<\/td>\n<td>\u2014<\/td>\n<td>1<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/liquidai-lfm2-5-vl-dspark-released\/\">LFM2.5-VL-3B-DSpark Draft Model for Vision-Language Models: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>pfnet\/plamo-3-610m-fin-instruct<\/td>\n<td>890M<\/td>\n<td>4GB<\/td>\n<td>other<\/td>\n<td>1<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/plamo-3-610m-fin-instruct-2\/\">plamo-3-610m-fin-instruct Text Generation Model: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>togethercomputer\/Tev1-4B-experimental<\/td>\n<td>4.7B<\/td>\n<td>4GB<\/td>\n<td>\u2014<\/td>\n<td>1<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/tev1-4b-experimental-released\/\">Tev1-4B-experimental Text Generation Model: 4GB+ VRAM, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>IFM\/K2-Horizon-32B-NVFP4<\/td>\n<td>\u2014<\/td>\n<td>32GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-32b-nvfp4-released\/\">K2-Horizon-32B-NVFP4 Long-Context Reasoning Model: 32GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>IFM\/K2-Horizon-375B-A23B-NVFP4<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-375b-a23b-nvfp4-released-2\/\">K2-Horizon-375B-A23B-NVFP4 Text Generation Model: ~257GB Memory<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td>ggml-org\/MiMo-V2.6-Flash-RL-GGUF<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>mit<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf-released\/\">MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td>ggml-org\/MiMo-V2.6-Distill-Qwen-9B-GGUF<\/td>\n<td>9.4B<\/td>\n<td>12GB<\/td>\n<td>mit<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/\">MiMo-V2.6-Distill-Qwen-9B-GGUF Vision-Language Model: 12GB+ VRAM<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Image, Video and Audio<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Model<\/th>\n<th>Params<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Converted builds<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-26<\/td>\n<td>MiniMaxAI\/MiniMax-H3<\/td>\n<td>\u2014<\/td>\n<td>32GB<\/td>\n<td>other<\/td>\n<td>1<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/minimax-h3-open-omnimodal-video-generation\/\">MiniMax-H3 Audio-Visual Video Generation Model: 32GB+ VRAM, File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>Comfy-Org\/Ming-Image<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>mit<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/comfy-org-ming-image-released\/\">Ming-Image UI Design-Specialized Image Generation Model: ComfyUI Paths<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>Edge0\/Audio8-ASR-Infinite<\/td>\n<td>4.1B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/edge0-audio8-asr-infinite-2\/\">Audio8-ASR-Infinite Speech Recognition Model: 12GB+ VRAM, File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/google-deepmind-announces-gemini-38-flash-tts-models\/\">Google DeepMind Announces Gemini 3.8 Flash TTS Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>Viggle\/Qwen-Image-2.1-viggle-turbo<\/td>\n<td>7.1B<\/td>\n<td>48GB<\/td>\n<td>other<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/qwen-image-2-1-viggle-turbo-v0-1-preview\/\">Qwen-Image-2.1-viggle-turbo Image Generation Model: 48GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>SupraLabs\/Supra2-IMG<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/supra2-img-lightweight-104m-text-to-image-model\/\">Supra2-IMG Lightweight Text-to-Image Model: File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-21<\/td>\n<td>abenzerps\/Qwen-Image-2.1-Uncensored-GGUF<\/td>\n<td>7.1B<\/td>\n<td>16GB<\/td>\n<td>other<\/td>\n<td>2<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/uncensored-qwen-image-2-1-gguf-released\/\">Qwen-Image-2.1-Uncensored-GGUF Image Generation Model: 16GB+ VRAM<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Engines and Tools<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Project<\/th>\n<th>Version<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-28<\/td>\n<td>invoke-ai\/InvokeAI<\/td>\n<td>v6.14.2<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/invokeai-v6-14-2-released\/\">InvokeAI v6.14.2 Released with Custom Fonts and Model Fixes<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>LostRuins\/koboldcpp<\/td>\n<td>v1.122<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/koboldcpp-v1-122-released\/\">KoboldCpp v1.122 Released with Built-in Agentic Framework<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-25<\/td>\n<td>containers\/ramalama<\/td>\n<td>v0.25.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/ramalama-v0-25-0-released\/\">RamaLama v0.25.0 Released: Security and Engine Updates<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>ggml-org\/ggml<\/td>\n<td>v0.25.2<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/ggml-v0-25-2-released\/\">ggml v0.25.2 Released with Hardware Backend Optimizations<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td>ggml-org\/llama.cpp<\/td>\n<td>v0.5.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/llama-cpp-v0-5-0-released\/\">llama.cpp v0.5.0 Released with Backend and Server Upgrades<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>ggml-org\/ggml<\/td>\n<td>v0.25.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ggml-v025-0-released\/\">ggml v0.25.0 Released: FlashAttention and MoE Optimizations<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td>unslothai\/unsloth<\/td>\n<td>v0.1.812-beta<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/unsloth-qwen-image-2-1-agent-skills-update\/\">Unsloth Update: Qwen-Image-2.1 Support and Agent Skills Added<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td>vllm-project\/vllm<\/td>\n<td>v0.30.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/vllm-v0-30-0-released\/\">vLLM v0.30.0 Released: Fast Start Weight Caching and New Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td>open-webui\/open-webui<\/td>\n<td>v0.11.4<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/open-webui-v0-11-4\/\">Open WebUI v0.11.4 Released: Dramatic Docker Image Slimming<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-21<\/td>\n<td>Comfy-Org\/ComfyUI<\/td>\n<td>v0.37.0<\/td>\n<td><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Community<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-26<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/ollaya-local-runtime-decision-models\/\">Ollaya: Local Runtime for Open-Source Decision Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-21<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/mini-agi-byte-level-continual-learning-model\/\">Mini-AGI: A Byte-Level Continual Learning Model for 8GB VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-21<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/kev-lightweight-local-decision-making-model-qwen35\/\">Kev: Lightweight Local Decision-Making Model Based on Qwen3.5<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Companies and Funding<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-28<\/td>\n<td>(unverified)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/inclusionai-publishes-training-content-summaries\/\">inclusionAI Publishes Training Content Summaries for EU Compliance<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/train-your-own-jev-classifier-together-ai\/\">Train Your Own Jev-style Classifier for $17 with Together AI<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/nvidia-topograph-cluster-topology-toolkit\/\">NVIDIA Topograph: Open Source Cluster Topology Toolkit<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/transformers-llama-cpp-gguf-support\/\">Run GGUF Models Directly in Transformers with llama.cpp Support<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-22<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/nvidia-dynamo-triton-26-07-multi-device-inference\/\">NVIDIA Dynamo-Triton 26.07 Adds Multi-Device Inference<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Technical Reports<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-27<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/27\/xiaomimimo-agentic-rl-training-environment\/\">XiaomiMiMo Releases Agentic RL Training Environment and Dataset<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-27<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/27\/deepseek-elastic-compute-dsec-sandbox-infrastructure\/\">DeepSeek Elastic Compute (DSec): Agentic Training Sandbox<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/smoldataenvs-rl-tasks-small-model-optimization\/\">SmolDataEnvs: RL Tasks for Small Model Optimization<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-25<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/nvidia-efficient-moe-training-biological-foundation-models\/\">NVIDIA Optimizes MoE Training for Biological Foundation Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/offloaded-inference-for-real-world-physical-ai-robotics\/\">Offloaded Inference for Real-World Physical AI Robotics<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-24<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/nvidia-releases-swe-serve-benchmark\/\">NVIDIA Releases SWE-Serve Benchmark for AI Coding Agents<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/together-ai-canary-rollouts-2\/\">Together AI Announces Canary Rollouts for Zero-Downtime Updates<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-23<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/nvidia-blackwell-confidential-computing-ai-inference\/\">AI Inference Performance with NVIDIA Blackwell Confidential Computing<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-21<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/pruning-llms-like-a-physicist-cbo\/\">Pruning LLMs Like a Physicist: Constrained Binary Optimization<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:weekly-articles --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Weekly roundup of open-weight generative models. Highlights include Qwen3.8-27B, native GGUF support in Transformers, and new efficient decision models.<\/p>\n","protected":false},"author":1,"featured_media":6099,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[159],"tags":[677,163,1073,2523,522,2525,1547],"class_list":["post-6078","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-weekly-roundup","tag-comfyui-en","tag-gguf-en","tag-llama-cpp-en","tag-localmodelwatch-en","tag-qwen3-8-en","tag-tev1-en","tag-verified"],"lang":"en","translations":{"en":6078,"ja":6076},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6078","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=6078"}],"version-history":[{"count":4,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6078\/revisions"}],"predecessor-version":[{"id":6164,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/6078\/revisions\/6164"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/6099"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=6078"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=6078"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=6078"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}