{"id":2337,"date":"2026-09-21T08:54:24","date_gmt":"2026-09-20T23:54:24","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/21\/open-weight-model-weekly-2026-09-22\/"},"modified":"2026-09-21T08:54:24","modified_gmt":"2026-09-20T23:54:24","slug":"open-weight-model-weekly-2026-09-22","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/open-weight-model-weekly-2026-09-22\/","title":{"rendered":"Open-Weight Model Weekly: Ternary 27B Models and Apple Silicon Updates"},"content":{"rendered":"<p><!-- lmw:weekly-stats --><\/p>\n<h2>This Week in Numbers<\/h2>\n<p><em>Counted by Local Model Watch from the articles published this week.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Count<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Articles published<\/td>\n<td>39<\/td>\n<\/tr>\n<tr>\n<td>\u3000New Models<\/td>\n<td>12<\/td>\n<\/tr>\n<tr>\n<td>\u3000Engines and Tools<\/td>\n<td>10<\/td>\n<\/tr>\n<tr>\n<td>\u3000Technical Reports<\/td>\n<td>7<\/td>\n<\/tr>\n<tr>\n<td>\u3000Community<\/td>\n<td>4<\/td>\n<\/tr>\n<tr>\n<td>\u3000Image, Video and Audio<\/td>\n<td>3<\/td>\n<\/tr>\n<tr>\n<td>\u3000Companies and Funding<\/td>\n<td>3<\/td>\n<\/tr>\n<tr>\n<td>New models covered<\/td>\n<td>15<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 8GB of VRAM (est.)<\/td>\n<td>4<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 12GB of VRAM (est.)<\/td>\n<td>10<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 16GB of VRAM (est.)<\/td>\n<td>10<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 24GB of VRAM (est.)<\/td>\n<td>12<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters up to 4B<\/td>\n<td>3<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters 15\u201340B<\/td>\n<td>5<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters over 40B<\/td>\n<td>1<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters 4\u201315B<\/td>\n<td>2<\/td>\n<\/tr>\n<tr>\n<td>Converted builds appended to earlier articles<\/td>\n<td>15 (GGUF 9, MLX 4, GPTQ 1, NVFP4 1)<\/td>\n<\/tr>\n<tr>\n<td>Most active publishers<\/td>\n<td>bartowski (3), unslothai\/unsloth (3), prism-ml (3)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Trending models we did not cover separately<\/h3>\n<p><em>Hugging Face repositories that trended this week but did not get their own article. Listed for reference; we have not reviewed their model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Likes<\/th>\n<th>Downloads<\/th>\n<th>Why no article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/XingChen-AGI\/Xing4.0-29B-A4B\">XingChen-AGI\/Xing4.0-29B-A4B<\/a><\/td>\n<td>872<\/td>\n<td>12,617<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/dealignai\/Bonsai-2-27B-Ternary-CRACK-GGUF\">dealignai\/Bonsai-2-27B-Ternary-CRACK-GGUF<\/a><\/td>\n<td>108<\/td>\n<td>25,385<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/empero-ai\/Qwen3.8-35B-A3B-Distill-GGUF\">empero-ai\/Qwen3.8-35B-A3B-Distill-GGUF<\/a><\/td>\n<td>107<\/td>\n<td>52,476<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/Altworld\/Hemmingway-1\">Altworld\/Hemmingway-1<\/a><\/td>\n<td>97<\/td>\n<td>0<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Patch releases we did not cover separately<\/h3>\n<p><em>Releases of watched projects that were patch-level or had short notes. Each project&#8217;s page lists every version.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Project<\/th>\n<th>Version<\/th>\n<th>Release notes<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>ollama\/ollama<\/td>\n<td>v0.34.2<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.34.2\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>turboderp-org\/exllamav3<\/td>\n<td>v1.5.1<\/td>\n<td><a href=\"https:\/\/github.com\/turboderp-org\/exllamav3\/releases\/tag\/v1.5.1\">GitHub<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:weekly-stats --><\/p>\n<h2>Highlights of the Week<\/h2>\n<ul>\n<li>\n<p><strong>Succession of Ternary Quantization Models in the 27B Class<\/strong><br \/>\n  Prism ML released a batch of ternary quantization models based on Qwen3.8-27B across multiple formats, significantly expanding the options for running medium-sized models locally in environments with limited VRAM capacity.<\/p>\n<\/li>\n<li>\n<p><strong>Expansion of Dedicated Inference and Training Environments for Apple Silicon and ARM64<\/strong><br \/>\n  Local execution environments for edge and non-x86 environments have been further improved, such as the release of the Qwen3.8 acceleration engine &#8220;Splash Engine&#8221; by LM Studio, and Windows ARM64 binaries by Unsloth.<\/p>\n<\/li>\n<li>\n<p><strong>Synchronized Updates of Major Local Inference Backends and Infrastructure Tools<\/strong><br \/>\n  Core engines such as ggml, llama.cpp, llamafile, SGLang, and LocalAI were updated one after another, adding precision control API features and optimizing loading methods.<\/p>\n<\/li>\n<\/ul>\n<h2>Trends by Category<\/h2>\n<h3>Text Generation<\/h3>\n<p>This week stood out for high-efficiency model deployments incorporating extreme quantization and sparse attention techniques, centered around the 15B\u201340B medium parameter band.<\/p>\n<p>Of particular note is the concentrated release of models using ternary weights. Starting with <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/18\/bonsai-2-27b-gguf\/\">Ternary-quantized 27B model &#8220;Bonsai 2 27B GGUF&#8221; released<\/a>, reports include <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/18\/bonsai-2-27b-mlx-2bit\/\">Qwen3.8-27B-based ternary weight model &#8220;Bonsai 2 27B&#8221; MLX 2bit version released<\/a> and the development version <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/18\/ternary-bonsai-2-27b-gguf-dev\/\">Ternary weight model test version &#8220;prism-ml\/Ternary-Bonsai-2-27B-gguf-dev&#8221; released<\/a>, advancing the practical application of 27B models capable of running in VRAM environments of 12GB or less. As a theoretical background, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/17\/breaking-the-158-bit-barrier-for-ternary-llms\/\">&#8220;Breaking the 1.58-Bit Barrier for Ternary LLMs&#8221; paper released<\/a>, which outlines methods challenging the 1.58-bit limit, also became a hot topic (unconfirmed information).<\/p>\n<p>Additionally, quantized versions of multimodal models and those optimized for reasoning have expanded, with <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf\/\">Orion-26B-A4B-v1.1 GGUF quantized version released, multimodal support<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/15\/salience-27b-r6-gguf\/\">Salience-27B-R6 GGUF quantized version released, reasoning-optimized 27B multimodal model<\/a>, and in the ultra-large band, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/15\/intern-s2-397b-gguf\/\">Multimodal foundational model &#8220;Intern-S2-397B&#8221; GGUF quantized version released<\/a> being provided.<\/p>\n<p>For on-device or specific-purpose lightweight models, releases included <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/17\/harshatheg-qwen-25-1b-rlcd\/\">harshatheg\/Qwen-2.5-1B-RLCD released: parallel constrained decoding for Apple Silicon<\/a>, ultra-small <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/19\/cactus-needle-3\/\">Cactus Compute releases 8\u201329MB automated model &#8220;Needle 3&#8221;<\/a>, and OCR-focused <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/18\/tencent-wevisdoc-2b-4b\/\">Tencent releases document analysis model &#8220;WeVisDoc-2B\/4B&#8221;<\/a>.<\/p>\n<h3>Image, Video, and Audio<\/h3>\n<p>In the media generation domain, enhancements to real-time performance and conversational capabilities are progressing.<\/p>\n<p>In the audio sector, sound quality was improved via <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/15\/yue2-mothersuperior-realaudio-tokenizer-v4\/\">Real-audio tokenizer v4 for music generation model YuE2-3B released<\/a>, alongside the provision of real-time conversational models integrating voice and video, such as <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/19\/realtime-venus\/\">inclusionAI releases full-duplex voice and video dialogue model &#8220;Realtime-Venus&#8221;<\/a>.<\/p>\n<p>In the image sector, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/20\/qwen-image-21\/\">Image generation and editing model &#8220;Qwen-Image-2.1&#8221; released, supporting transparency processing and advanced multi-image editing<\/a> and its smaller model <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/21\/qwen-image-21-2\/\">Image generation model &#8220;Qwen-Image-2.1&#8221; released, scaled down to 7B with transparency processing support<\/a> were released, improving open-weight image editing environments.<\/p>\n<h3>Engines and Tools<\/h3>\n<p>It was a week where environments were improved across a wide range of layers, from inference infrastructure to UI tools.<\/p>\n<p>In low-layer inference engines, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/14\/ggml-v0-24-0-released\/\">ggml v0.24.0 released: addition of precision control API and backend enhancements<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/15\/llama-cpp-v0-4-1-released\/\">llama.cpp v0.4.1 released: loading method changes and new model support<\/a>, and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/16\/llamafile-0106\/\">llamafile 0.10.6 released, llama.cpp updates and enhanced CUDA\/HIP graph features<\/a> were successively published. On the serving side, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/19\/sglang-v0520-released\/\">Large-scale serving framework &#8220;SGLang v0.5.20&#8221; released<\/a> was also updated.<\/p>\n<p>In terms of hardware optimization, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/20\/splash-engine-qwen3-apple-silicon\/\">Local inference engine for Apple Silicon &#8220;Splash Engine&#8221; released, accelerates Qwen3.8<\/a> drew attention. Furthermore, Unsloth, the standard fine-tuning tool, announced <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/16\/unsloth-windows-arm64\/\">Unsloth releases Windows ARM64 binary version<\/a> in addition to <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/18\/unsloth-v0-1-810-beta-update\/\">&#8220;unsloth&#8221; v0.1.810-beta released: multi-user support and Docker overhaul<\/a> and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/19\/unsloth-v0-1-811-beta\/\">Unsloth v0.1.811-beta released, Docker, AMD, and multi-user support<\/a>, improving ARM support and ease of use in container environments.<\/p>\n<p>In frontend and UI-related news, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/16\/koboldcpp-v1-121\/\">koboldcpp v1.121 released: supports Minimax H3 media reference and SDUI LoRA selector<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/16\/comfyui-v0360\/\">ComfyUI v0.36.0 released: supports general-purpose loops, Yue2, and Marigold v2<\/a>, and <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/18\/localai-v4100\/\">Local AI execution engine &#8220;LocalAI v4.10.0&#8221; released, dashboard overhauled<\/a> were released.<\/p>\n<p><!-- lmw:weekly-articles --><\/p>\n<h2>This Week&#8217;s Articles<\/h2>\n<p><em>Compiled by Local Model Watch from the article log. Articles about the same story are merged into one row.<\/em><\/p>\n<h3>New Models<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Model<\/th>\n<th>Params<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Converted builds<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-19<\/td>\n<td>Cactus-Compute\/needle3<\/td>\n<td>\u2014<\/td>\n<td>4GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/cactus-compute-needle-3\/\">Cactus Compute Releases Needle 3: 8\u201329 MB On-Device Automation Model<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>prism-ml\/Ternary-Bonsai-2-27B-gguf-dev<\/td>\n<td>27.8B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf-dev-released\/\">Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>prism-ml\/Ternary-Bonsai-2-27B-mlx-2bit<\/td>\n<td>27.4B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>prism-ml\/Ternary-Bonsai-2-27B-gguf<\/td>\n<td>27.8B<\/td>\n<td>8GB<\/td>\n<td>apache-2.0<\/td>\n<td>1<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf\/\">Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>tencent\/WeVisDoc-4B<\/td>\n<td>4.4B<\/td>\n<td>4GB<\/td>\n<td>apache-2.0<\/td>\n<td>2<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/tencent-releases-wevisdoc-document-parsing-models\/\">Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-17<\/td>\n<td>harshatheg\/Qwen-2.5-1B-RLCD<\/td>\n<td>1.5B<\/td>\n<td>4GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/17\/qwen-25-1b-rlcd-mlx-constrained-decoding\/\">Fast Structured Generation on Apple Silicon with MLX and Qwen<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td>EleutherAI\/olmo3-7b-sdf-sft-clean150<\/td>\n<td>7.3B<\/td>\n<td>24GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/eleutherai-olmo-3-7b-reward-hacking-models\/\">EleutherAI Releases OLMo-3-7B Models for Reward-Hacking Research<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td>bartowski\/vectionlabs_Salience-27B-R6-GGUF<\/td>\n<td>27.8B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">Salience-27B-R6 GGUF Released by bartowski<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td>bartowski\/Intern-S2-397B-GGUF<\/td>\n<td>403.4B<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B GGUF Quantized Models Released by bartowski<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td>EleutherAI\/bergson-wikitext-gpt2-leaderboard<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/eleutherai-releases-bergson-leaderboard-baseline-gpt2-model\/\">EleutherAI Releases Bergson Leaderboard Baseline GPT-2 Model<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-14<\/td>\n<td>bartowski\/TheDrummer_Orion-26B-A4B-v1.1-GGUF<\/td>\n<td>25.8B<\/td>\n<td>12GB<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/\">Orion-26B-A4B-v1.1 GGUF Quants Released by bartowski<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-14<\/td>\n<td>tencent\/Simple-Attention-Sparsification<\/td>\n<td>4.0B<\/td>\n<td>12GB<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/tencent-simple-attention-sparsification-qwen3\/\">Tencent Releases Simple Attention Sparsification for Qwen3<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Image, Video and Audio<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Model<\/th>\n<th>Params<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Converted builds<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-20<\/td>\n<td>Qwen\/Qwen-Image-2.1-PE-I2I<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>other<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/20\/qwen-image-2-1-released\/\">Qwen-Image-2.1 Released: Open-Weight Image Gen &amp; Editing<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-19<\/td>\n<td>inclusionAI\/Realtime-Venus<\/td>\n<td>\u2014<\/td>\n<td>24GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/realtime-venus-multimodal-conversational-ai\/\">Realtime-Venus: Multimodal Conversational AI with Full-Duplex<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td>Mothersuperior\/yue2-mothersuperior-realaudio-tokenizer-v4<\/td>\n<td>3.6B<\/td>\n<td>12GB<\/td>\n<td>cc-by-nc-4.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/mothersuperior-yue2-realaudio-toolkit\/\">Mothersuperior Releases Real Audio Toolkit for YuE2-3B<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Engines and Tools<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Project<\/th>\n<th>Version<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-19<\/td>\n<td>sgl-project\/sglang<\/td>\n<td>v0.5.20<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/sglang-v0-5-20-released\/\">SGLang v0.5.20 Released: New Models and Optimizations<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-19<\/td>\n<td>unslothai\/unsloth<\/td>\n<td>v0.1.811-beta<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/unsloth-v0-1-811-beta-released\/\">Unsloth v0.1.811-beta Released with AMD &amp; NVIDIA Docker Images<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>mudler\/LocalAI<\/td>\n<td>v4.10.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/localai-v4-10-0-released\/\">LocalAI v4.10.0 Released: Fleet Dashboard &amp; M5 Support<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td>unslothai\/unsloth<\/td>\n<td>v0.1.810-beta<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/unsloth-v0-1-810-beta-released\/\">Unsloth v0.1.810-beta Released with Multi-User and AMD Support<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td>unslothai\/unsloth<\/td>\n<td>Windows-ARM64<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/unsloth-windows-arm64-release\/\">Unsloth Releases Windows ARM64 Binary Version<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td>comfyanonymous\/ComfyUI<\/td>\n<td>v0.36.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/comfyui-v0360-released\/\">ComfyUI v0.36.0 Released: Generic Loops, Yue2 &amp; Marigold v2<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td>Mozilla-Ocho\/llamafile<\/td>\n<td>0.10.6<\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td>LostRuins\/koboldcpp<\/td>\n<td>v1.121<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/koboldcpp-v1-121-released\/\">koboldcpp v1.121 Released: New Features &amp; Bug Fixes<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td>ggml-org\/llama.cpp<\/td>\n<td>v0.4.1<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/llama-cpp-v041-released\/\">llama.cpp v0.4.1 Released with New Models and Load Modes<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-14<\/td>\n<td>ggml-org\/ggml<\/td>\n<td>v0.24.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/ggml-v0-24-0-released-2\/\">ggml v0.24.0 Released with Backend Improvements and API Updates<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Community<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-21<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/hacker-news-discussion-on-qwen-image-21\/\">Hacker News Discussion on Open-Weight Qwen-Image-2.1<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-17<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/17\/breaking-the-1-58-bit-barrier-for-ternary-llms\/\">Breaking the 1.58-bit Barrier for Ternary LLMs with BITCOS<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/why-tech-community-is-debating-bearish-views-on-llms\/\">Why Tech Community is Debating Bearish Views on LLMs<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/migrating-large-prompts-to-local-ollama\/\">Challenges of Migrating Large Prompts to Local Ollama<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Companies and Funding<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-20<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/20\/inco-ai-releases-splash-engine-for-mac\/\">Inco AI Releases Splash Engine for Fast Local LLMs on Mac<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/2026-ai-inference-hardware-revolution\/\">The 2026 AI Inference Hardware Revolution and Local LLM Impact<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/mistral-mozilla-firefox-smart-window-2\/\">Mistral AI and Mozilla Partner for Firefox Smart Window<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Technical Reports<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-19<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/scaling-coding-agent-traffic-glm-52-together-ai\/\">Scaling Coding Agent Traffic with GLM-5.2 and Dedicated Inference<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-19<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/nvidia-aiperf-benchmarking-llm-inference\/\">NVIDIA AIPerf: Benchmarking LLM Inference at Scale<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-18<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/lm-studio-bionic-introspection-session-references\/\">LM Studio Announces Session References and Introspection in Bionic<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-17<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/17\/glm-custom-inference-infrastructure-and-local-moe-tests\/\">GLM Announces Custom Inference Infrastructure and Local MoE Tests<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-17<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/17\/tensorrt-edge-llm-mlperf-edge-agentic-jetson-thor\/\">TensorRT Edge-LLM Accelerates MLPerf Agentic Benchmark<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-16<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/nvidia-groq-3-lpx-deterministic-execution\/\">NVIDIA Groq 3 LPX Details and Deterministic Execution<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-15<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/accelerating-dropless-moe-training-in-jax\/\">NVIDIA Accelerates Dropless MoE Training in JAX<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:weekly-articles --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Weekly roundup of open-weight generative models: ternary quantization for 27B models, Apple Silicon optimizations, and core engine updates.<\/p>\n","protected":false},"author":1,"featured_media":2336,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[159],"tags":[505,1846,163,1073,522,509,1547],"class_list":["post-2337","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-weekly-roundup","tag-apple-silicon-en","tag-bonsai-en","tag-gguf-en","tag-llama-cpp-en","tag-qwen3-8-en","tag-unsloth-en","tag-verified"],"lang":"en","translations":{"en":2337,"ja":2335},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2337","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=2337"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2337\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/2336"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=2337"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=2337"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=2337"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}