{"id":9915,"date":"2026-10-05T09:29:52","date_gmt":"2026-10-05T00:29:52","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/05\/local-ai-weekly-2026-10-week-1\/"},"modified":"2026-10-05T09:29:52","modified_gmt":"2026-10-05T00:29:52","slug":"local-ai-weekly-2026-10-week-1","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/05\/local-ai-weekly-2026-10-week-1\/","title":{"rendered":"Local AI Model Weekly Highlights: Structured Output &#038; Video LoRAs"},"content":{"rendered":"<p><!-- lmw:weekly-stats --><\/p>\n<h2>This Week in Numbers<\/h2>\n<p><em>Counted by Local Model Watch from the articles published this week.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Count<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Articles published<\/td>\n<td>36<\/td>\n<\/tr>\n<tr>\n<td>\u3000Image, Video and Audio<\/td>\n<td>13<\/td>\n<\/tr>\n<tr>\n<td>\u3000New Models<\/td>\n<td>12<\/td>\n<\/tr>\n<tr>\n<td>\u3000Engines and Tools<\/td>\n<td>10<\/td>\n<\/tr>\n<tr>\n<td>\u3000Community<\/td>\n<td>1<\/td>\n<\/tr>\n<tr>\n<td>New models covered<\/td>\n<td>25<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 8GB of VRAM (est.)<\/td>\n<td>5<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 12GB of VRAM (est.)<\/td>\n<td>7<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 16GB of VRAM (est.)<\/td>\n<td>8<\/td>\n<\/tr>\n<tr>\n<td>\u3000fit in 24GB of VRAM (est.)<\/td>\n<td>11<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters 15\u201340B<\/td>\n<td>6<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters 4\u201315B<\/td>\n<td>5<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters up to 4B<\/td>\n<td>2<\/td>\n<\/tr>\n<tr>\n<td>\u3000parameters over 40B<\/td>\n<td>2<\/td>\n<\/tr>\n<tr>\n<td>Converted builds appended to earlier articles<\/td>\n<td>18 (GGUF 10, FP8 1, MLX 7)<\/td>\n<\/tr>\n<tr>\n<td>Most active publishers<\/td>\n<td>efficient-large-model (5), ggml-org (3), unslothai\/unsloth (2)<\/td>\n<\/tr>\n<tr>\n<td>Articles still marked unverified<\/td>\n<td>1<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Trending models we did not cover separately<\/h3>\n<p><em>Hugging Face repositories that trended this week but did not get their own article. Listed for reference; we have not reviewed their model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Likes<\/th>\n<th>Downloads<\/th>\n<th>Why no article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/PSRben\/VisionHOPE\">PSRben\/VisionHOPE<\/a><\/td>\n<td>402<\/td>\n<td>1,516<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/Aleph-Alpha\/Kolibri-1\">Aleph-Alpha\/Kolibri-1<\/a><\/td>\n<td>388<\/td>\n<td>1,135<\/td>\n<td>converted build without a parent article<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/ISTA-DASLab\/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF\">ISTA-DASLab\/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF<\/a><\/td>\n<td>260<\/td>\n<td>351,230<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/NaiveAI\/Naive-N0.5-Flash\">NaiveAI\/Naive-N0.5-Flash<\/a><\/td>\n<td>165<\/td>\n<td>1,920<\/td>\n<td>publisher not on our notable list<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/FrogNano-4B-2609-GGUF\">bartowski\/FrogNano-4B-2609-GGUF<\/a><\/td>\n<td>11<\/td>\n<td>7,519<\/td>\n<td>converted build without a parent article<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/bytkim_Qwen3.8-27B-pi-GGUF\">bartowski\/bytkim_Qwen3.8-27B-pi-GGUF<\/a><\/td>\n<td>7<\/td>\n<td>3,909<\/td>\n<td>converted build without a parent article<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/GLM-5.3-Flash-BF16-GGUF\">bartowski\/GLM-5.3-Flash-BF16-GGUF<\/a><\/td>\n<td>4<\/td>\n<td>2,288<\/td>\n<td>converted build without a parent article<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/cosmicoptima_computer-10-GGUF\">bartowski\/cosmicoptima_computer-10-GGUF<\/a><\/td>\n<td>3<\/td>\n<td>6,853<\/td>\n<td>converted build without a parent article<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/OmniJev_OneJev-9B-GGUF\">bartowski\/OmniJev_OneJev-9B-GGUF<\/a><\/td>\n<td>3<\/td>\n<td>3,333<\/td>\n<td>converted build without a parent article<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/Lythri_Lythri-4B-A2B-GGUF\">bartowski\/Lythri_Lythri-4B-A2B-GGUF<\/a><\/td>\n<td>2<\/td>\n<td>2,819<\/td>\n<td>converted build without a parent article<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Patch releases we did not cover separately<\/h3>\n<p><em>Releases of watched projects that were patch-level or had short notes. Each project&#8217;s page lists every version.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Project<\/th>\n<th>Version<\/th>\n<th>Release notes<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>magnitudedev\/magnitude<\/td>\n<td>@magnitudedev\/cli@0.2.2<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.2\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>magnitudedev\/magnitude<\/td>\n<td>@magnitudedev\/cli@0.2.3<\/td>\n<td><a href=\"https:\/\/github.com\/magnitudedev\/magnitude\/releases\/tag\/%40magnitudedev\/cli%400.2.3\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>ollama\/ollama<\/td>\n<td>v0.35.1<\/td>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\/tag\/v0.35.1\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>turboderp-org\/exllamav3<\/td>\n<td>v1.5.4<\/td>\n<td><a href=\"https:\/\/github.com\/turboderp-org\/exllamav3\/releases\/tag\/v1.5.4\">GitHub<\/a><\/td>\n<\/tr>\n<tr>\n<td>unslothai\/unsloth<\/td>\n<td>v0.1.902-beta<\/td>\n<td><a href=\"https:\/\/github.com\/unslothai\/unsloth\/releases\/tag\/v0.1.902-beta\">GitHub<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Models Gaining the Most Likes<\/h3>\n<p><em>Compiled by Local Model Watch from weekly snapshots of models we&#8217;ve covered. Likes are cumulative on Hugging Face; downloads are its trailing-30-day count.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Likes gained<\/th>\n<th>Total likes<\/th>\n<th>Downloads (30d)<\/th>\n<th>Period<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/Edge0\/Audio8-ASR-Infinite\">Edge0\/Audio8-ASR-Infinite<\/a><\/td>\n<td>+1,401<\/td>\n<td>2,426<\/td>\n<td>40,004<\/td>\n<td>2026-W40 \u2192 2026-W41<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/Qwen\/Qwen3.8-27B\">Qwen\/Qwen3.8-27B<\/a><\/td>\n<td>+515<\/td>\n<td>16,939<\/td>\n<td>6,821,761<\/td>\n<td>2026-W40 \u2192 2026-W41<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4.1-Flash\">deepseek-ai\/DeepSeek-V4.1-Flash<\/a><\/td>\n<td>+283<\/td>\n<td>4,093<\/td>\n<td>798,422<\/td>\n<td>2026-W40 \u2192 2026-W41<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/Viggle\/Qwen-Image-2.1-viggle-turbo\">Viggle\/Qwen-Image-2.1-viggle-turbo<\/a><\/td>\n<td>+246<\/td>\n<td>587<\/td>\n<td>272,896<\/td>\n<td>2026-W40 \u2192 2026-W41<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/huggingface.co\/prism-ml\/Ternary-Bonsai-2-27B-gguf\">prism-ml\/Ternary-Bonsai-2-27B-gguf<\/a><\/td>\n<td>+228<\/td>\n<td>2,416<\/td>\n<td>4,045,810<\/td>\n<td>2026-W40 \u2192 2026-W41<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:weekly-stats --><\/p>\n<h2>Highlights of the Week<\/h2>\n<p>The most notable highlights surrounding local AI models this week are the following three points.<\/p>\n<p>First, there is the rise of models specialized in structured output and decision-making, along with advances in their GGUF support. Models optimized for specific tasks and structured data processing, such as &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/cloudflare-clef\/\">clef<\/a>&#8221; and &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/cloudflare-clef-flash\/\">clef-flash<\/a>&#8221; released by Cloudflare, and &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/ggml-org-openjev-gguf-released\/\">OpenJev-GGUF<\/a>&#8221; released by ggml-org, have appeared one after another. This significantly improves the practicality of agents and automation tools that can be run on local PCs.<\/p>\n<p>Second, LoRA adapters for video generation models have been released very actively. Methods for practically customizing and controlling existing powerful models have been enriched, such as the &#8220;LongLive-Plug&#8221; series by Efficient-Large-Model and LoRAs for video inpainting and quality conversion by Lightricks.<\/p>\n<p>Third, competition to accelerate local inference engines and expand their features is intensifying. The emergence of the new Rust-based engine &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/01\/magnitude\/\">Magnitude<\/a>&#8221; and the addition of new APIs for decision-making models in &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/30\/ollama-v0350\/\">Ollama v0.35.0<\/a>&#8221; are rapidly establishing the infrastructure for developers to run models more comfortably in their local environments.<\/p>\n<h2>Trends by Category<\/h2>\n<h3>Text Generation<\/h3>\n<p>This week saw the release of numerous practical models specialized for specific use cases and reasoning capabilities. Optimization is progressing not only for general conversation, but also for tasks requiring structured data output and complex thought processes.<\/p>\n<p>Cloudflare released &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/cloudflare-clef\/\">clef<\/a>&#8221; (27.4B) and its lightweight version &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/cloudflare-clef-flash\/\">clef-flash<\/a>&#8221; (9.4B), which are specialized for structured decision-making tasks. These exhibit high accuracy in use cases such as API integration and data extraction. Additionally, domestic Japanese developer ELYZA released the 32B and 33B models of &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/elyza-thinking-1-0-32b-33b\/\">ELYZA-Thinking-1.0<\/a>&#8220;, an inference model capable of outputting thought processes, enabling local execution of advanced reasoning tasks in Japanese environments. Furthermore, previously released models like &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/xing40-29b-a4b\/\">Xing4.0-29B-A4B<\/a>` and &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/28\/altworld-hemmingway-1\/\">Hemmingway-1<\/a>&#8221; have also been covered on this site, further enriching the options in the mid-size tier.<\/p>\n<h3>Image, Video, and Audio<\/h3>\n<p>It was a week notable for advancements in video generation control technology and the release of unique models specialized in audio processing.<\/p>\n<p>In the field of video generation, numerous LoRAs were released to efficiently control underlying large-scale models. In particular, LoRA adapters targeting Wan2.1 and Wan2.2, such as &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/29\/longlive-plug-wan22-ti2v-5b-cfg\/\">LongLive-Plug-Wan2.2-TI2V-5B-cfg<\/a>&#8220;, enable high-quality generation with fewer steps and camera angle control. In the audio field, Google released &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/05\/google-diarizationlm-gemma-4-e4b-v1\/\">DiarizationLM-Gemma-4-E4B-v1<\/a>&#8220;, which is suitable for speech transcription and speaker diarization, supporting the local execution of practical audio processing. Note that the release of the Turkish synthetic speech dataset &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/02\/alania-synthetic-speech-tr\/\">alania-synthetic-speech-tr<\/a>&#8221; has also been reported, but please be aware that this is unverified information that has not been officially confirmed.<\/p>\n<h3>Engines and Tools<\/h3>\n<p>Development competition continues with the goal of maximizing execution speed in local environments and rapidly supporting the latest models.<\/p>\n<p>The image generation GUI &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/30\/comfyui-v0380-released\/\">ComfyUI v0.38.0<\/a>&#8221; was released, adding support for the latest Hunyuan Image 3.5 and Qwen-Image 2.1. In addition, &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/30\/ollama-v0350\/\">Ollama v0.35.0<\/a>&#8220;, a popular local LLM execution environment, implemented a new API to make decision-making models easier to handle. Furthermore, a new inference engine written in Rust called &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/01\/magnitude\/\">Magnitude<\/a>&#8221; appeared, followed by rapid-fire updates such as &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/03\/magnitude-cli-0-2-4\/\">Magnitude CLI 0.2.4<\/a>&#8221; and &#8220;<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/03\/magnitudedev-cli-v0-2-5-released\/\">@magnitudedev\/cli v0.2.5<\/a>&#8221; which advance memory reduction and Apple Silicon optimization. The evolution of these tools is building an environment where the latest models can be run efficiently even with limited hardware resources.<\/p>\n<h2>Industry News<\/h2>\n<p>These are announcements from companies and research institutions that were not made into standalone articles because they are not topics about running locally. Only the key points are listed.<\/p>\n<ul>\n<li><strong>ankitjh4<\/strong>: The dataset &#8220;Bharat Guide&#8221; containing 55,050 text documents extracted and normalized from Indian government public documents has been released. (<a href=\"https:\/\/huggingface.co\/datasets\/ankitjh4\/bharat-government-documents\">Announcement<\/a>)<\/li>\n<li><strong>Kuyawa<\/strong>: The desktop app &#8220;DeepSeek Harness Desktop&#8221; is being developed for macOS and Windows, allowing users to create and extend plugins through chat. (<a href=\"https:\/\/www.deepseek.com\/en\/harness\/\">Announcement<\/a>)<\/li>\n<li><strong>NVIDIA Developer<\/strong>: &#8220;DIN Deploy&#8221; has been released, a C++ sample code combining ONNX Runtime and NVIDIA TensorRT RTX on Windows and Linux to accelerate local AI inference. (<a href=\"https:\/\/developer.nvidia.com\/blog\/build-local-ai-apps-with-c-and-nvidia-tensorrt-rtx-samples\/\">Announcement<\/a>)<\/li>\n<li><strong>NVIDIA Developer<\/strong>: A fine-tuning method was announced for the speech recognition model &#8220;NVIDIA Nemotron 3.5 ASR&#8221; to support regional dialects in Saudi Arabia such as Najdi and Hijazi. (<a href=\"https:\/\/developer.nvidia.com\/blog\/fine-tuning-nvidia-nemotron-for-saudi-arabic-dialects-with-a-path-to-other-languages\/\">Announcement<\/a>)<\/li>\n<li><strong>Hugging Face Blog<\/strong>: The &#8220;Open TTS Leaderboard&#8221; has been published on Hugging Face to evaluate multilingual text-to-speech (TTS) and voice cloning models using a standardized methodology. (<a href=\"https:\/\/huggingface.co\/blog\/open-tts-leaderboard\">Announcement<\/a>)<\/li>\n<\/ul>\n<p><!-- lmw:weekly-articles --><\/p>\n<h2>This Week&#8217;s Articles<\/h2>\n<p><em>Compiled by Local Model Watch from the article log. Articles about the same story are merged into one row.<\/em><\/p>\n<h3>New Models<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Model<\/th>\n<th>Params<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Converted builds<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-10-05<\/td>\n<td>Qwen\/Qwen3.8-Flash-Next<\/td>\n<td>180.0B<\/td>\n<td>\u2014<\/td>\n<td>qwen-community-1.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/05\/qwen-3-8-flash-next-released\/\">Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-04<\/td>\n<td>LiquidAI\/LFM2.5-350M-Diffusion-Exp<\/td>\n<td>425M<\/td>\n<td>4GB<\/td>\n<td>lfm1.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/04\/liquidai-lfm2-5-350m-diffusion-exp\/\">LFM2.5-350M-Diffusion-Exp Text Generation Model: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-04<\/td>\n<td>ggml-org\/GLM-5.3-Flash-GGUF<\/td>\n<td>321.3B<\/td>\n<td>\u2014<\/td>\n<td>other<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/04\/glm-5-3-flash-gguf-released\/\">GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-03<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/astabrief-8b-open-source-scientific-report-generator\/\">Ai2 Releases AstaBrief 8B: Fast Open-Source Scientific Report\u2026<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-02<\/td>\n<td>elyza\/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b<\/td>\n<td>32.1B<\/td>\n<td>80GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/elyza-thinking-1-0-32b-33b-2\/\">ELYZA Releases ELYZA-Thinking-1.0 32B\/33B Reasoning Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-02<\/td>\n<td>Cloudflare\/clef-flash<\/td>\n<td>9.4B<\/td>\n<td>4GB<\/td>\n<td>apache-2.0<\/td>\n<td>3<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/clef-flash-specs-performance-and-use-cases\/\">clef-flash Vision-Language Model: Our Test Answers, 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-28<\/td>\n<td>orcarouter\/OrcaSAQ-2-27B<\/td>\n<td>27.8B<\/td>\n<td>16GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/orcasaq2-27b-qwen3-8-27b-quantized\/\">OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM<\/a> (follow-up: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/cloudflare-clef-27b-multimodal-decision-model\/\">clef Structured Decision-Making Model: 12GB+ VRAM, GGUF Builds<\/a>)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-28<\/td>\n<td>Altworld\/Hemmingway-1<\/td>\n<td>26.9B<\/td>\n<td>12GB<\/td>\n<td>cc-by-nc-4.0<\/td>\n<td>3<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/hemmingway-1-open-27b-model-specialized-for-human-like-writing\/\">Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-28<\/td>\n<td>apple\/LensVLM-9B<\/td>\n<td>9.4B<\/td>\n<td>4GB<\/td>\n<td>apple-amlr<\/td>\n<td>3<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/apple-lensvlm-9b-2\/\">LensVLM-9B Vision-Language Model: 4GB+ VRAM, GGUF Builds<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-28<\/td>\n<td>XingChen-AGI\/Xing4.0-29B-A4B<\/td>\n<td>31.2B<\/td>\n<td>24GB<\/td>\n<td>apache-2.0<\/td>\n<td>3<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/xing4-0-29b-a4b-china-telecom\/\">Xing4.0-29B-A4B Text Generation Model: Our Test Answers, 24GB+ VRAM<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Image, Video and Audio<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Model<\/th>\n<th>Params<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Converted builds<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-10-05<\/td>\n<td>google\/DiarizationLM-Gemma-4-E4B-v1<\/td>\n<td>8.0B<\/td>\n<td>8GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/05\/diarizationlm-gemma-4-e4b-v1\/\">DiarizationLM-Gemma-4-E4B-v1 Vision-Language Model: 8GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-02<\/td>\n<td>nvidia\/PixelUMM<\/td>\n<td>8.2B<\/td>\n<td>24GB<\/td>\n<td>nvidia-one-way-noncommercial-license<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/nvidia-releases-pixelumm\/\">PixelUMM Multimodal Model: 24GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-02<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/alania-synthetic-speech-tr-dataset\/\">Alania Synthetic Speech TR: Turkish Speech Dataset Released<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-01<\/td>\n<td>FermionResearch\/Phonon-2<\/td>\n<td>627M<\/td>\n<td>4GB<\/td>\n<td>cc-by-4.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/01\/phonon-2-open-weight-english-asr-model\/\">Phonon-2 Speech Recognition Model: 4GB+ VRAM<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-01<\/td>\n<td>Lightricks\/LTX-2.5-22b-IC-LoRA-SDR-To-HDR<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>ltx-2.x-community-license<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/01\/ltx-25-sdr-to-hdr-ic-lora\/\">LTX-2.5-22b-IC-LoRA-SDR-To-HDR Video Generation Model: File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-30<\/td>\n<td>Lightricks\/LTX-2.5-22b-IC-LoRA-Restore<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>ltx-2.x-community-license<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/lightricks-ltx-2-5-22b-ic-lora-restore\/\">LTX-2.5-22b-IC-LoRA-Restore Video Generation Model: File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-30<\/td>\n<td>lilylilith\/QI_2.1_AnyAngle<\/td>\n<td>7.1B<\/td>\n<td>48GB<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/qi-2-1-anyangle-camera-control-lora\/\">QI_2.1_AnyAngle Camera Angle Control LoRA: 48GB+ VRAM, File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-30<\/td>\n<td>Efficient-Large-Model\/LongLive-Plug-Wan2.1-T2V-14B-few-step<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/longlive-plug-wan21-t2v-14b-few-step-2\/\">LongLive-Plug-Wan2.1-T2V-14B-few-step: Our Generated Video, File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-30<\/td>\n<td>Efficient-Large-Model\/LongLive-Plug-Wan2.1-T2V-14B-cfg<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/longlive-plug-wan2-1-t2v-14b-cfg\/\">LongLive-Plug-Wan2.1-T2V-14B-cfg: Our Generated Video, File List<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-30<\/td>\n<td>Efficient-Large-Model\/LongLive-Plug-MiniMax-H3-cfg<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>minimax-h3-community-license-agreement<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/30\/minimax-h3-lora-adapters\/\">Japanese article only<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-29<\/td>\n<td>Efficient-Large-Model\/LongLive-Plug-Wan2.2-TI2V-5B-few-step<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/29\/longlive-plug-wan2-2-ti2v-5b-few-step-2\/\">LongLive-Plug-Wan2.2-TI2V-5B-few-step: Our Generated Video, File List<\/a> (follow-up: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/29\/longlive-plug-wan21-ti2v-5b-cfg-lora\/\">LongLive-Plug-Wan2.2-TI2V-5B-cfg: Our Generated Video, File List<\/a>)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-28<\/td>\n<td>akatz-ai\/MiniMax-H3-Character-Swap-LoRA<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>minimax-h3-community-license-agreement<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/minimax-h3-character-swap-lora\/\">MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Engines and Tools<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Project<\/th>\n<th>Version<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-10-03<\/td>\n<td>magnitudedev\/magnitude<\/td>\n<td>@magnitudedev\/cli@0.2.5<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/magnitude-cli-v0-2-5-released\/\">Magnitude CLI v0.2.5 Released: M5 Mac &amp; MoE Speedups<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-03<\/td>\n<td>mudler\/LocalAI<\/td>\n<td>v4.11.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/localai-v4-11-0-released\/\">LocalAI v4.11.0 Released with Failover and NeMo Audio Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-03<\/td>\n<td>magnitudedev\/magnitude<\/td>\n<td>@magnitudedev\/cli@0.2.4<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/03\/magnitude-0-2-4-released\/\">Magnitude 0.2.4 Released: Faster Inference and Lower Memory<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-02<\/td>\n<td>sgl-project\/sglang<\/td>\n<td>v0.5.21<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/sglang-v0-5-21-released-2\/\">SGLang v0.5.21 Released with Dynamic PD Role Switching<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-02<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/testing-openjev-gguf-tiny-model-for-runtime-loaders\/\">Testing OpenJev GGUF: A Tiny Model for Runtime Loaders<\/a> (follow-up: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/ggml-org-openjev-gguf-released-2\/\">OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM<\/a>)<\/td>\n<\/tr>\n<tr>\n<td>2026-10-01<\/td>\n<td>unslothai\/unsloth<\/td>\n<td>v0.1.901-beta<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/01\/unsloth-v0-1-901-beta-released-2\/\">Unsloth v0.1.901-beta Released: Major LoRA Memory Reductions<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-10-01<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/01\/magnitude-self-optimizing-inference-engine-for-agents\/\">Magnitude: Up to 2x Faster Than llama.cpp Officially, 21x Slower on Our CPU<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-30<\/td>\n<td>ollama\/ollama<\/td>\n<td>v0.35.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/ollama-v0-35-0-released\/\">Ollama v0.35.0 Released: New API for Decision Models<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-30<\/td>\n<td>Comfy-Org\/ComfyUI<\/td>\n<td>v0.38.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/30\/comfyui-v0380-released-2\/\">ComfyUI v0.38.0 Released: Hunyuan Image 3.5 and More<\/a><\/td>\n<\/tr>\n<tr>\n<td>2026-09-29<\/td>\n<td>unslothai\/unsloth<\/td>\n<td>v0.1.900-beta<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/29\/unsloth-desktop-v01900-beta-release\/\">unsloth Desktop v0.1.900-beta Released with Laya and Speedups<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Community<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Date<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-10-04<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/04\/aleph-alpha-kolibri-open-weight-moe\/\">Aleph Alpha Releases Kolibri: A New Open-Weight MoE Model<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:weekly-articles --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Weekly roundup of local generative AI models: structured output models, video generation LoRAs, inference engines, and industry news.<\/p>\n","protected":false},"author":1,"featured_media":9914,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[159],"tags":[2808,2886,163,2791,1093,1547],"class_list":["post-9915","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-weekly-roundup","tag-clef-en","tag-elyza-thinking-1-0-en","tag-gguf-en","tag-magnitude-en","tag-ollama-en","tag-verified"],"lang":"en","translations":{"en":9915,"ja":9913},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9915","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9915"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9915\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9914"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9915"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9915"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9915"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}