Local AI Model Weekly Highlights: Structured Output & Video LoRAs

This Week in Numbers
Counted by Local Model Watch from the articles published this week.
| Item | Count |
|---|---|
| Articles published | 36 |
| Image, Video and Audio | 13 |
| New Models | 12 |
| Engines and Tools | 10 |
| Community | 1 |
| New models covered | 25 |
| fit in 8GB of VRAM (est.) | 5 |
| fit in 12GB of VRAM (est.) | 7 |
| fit in 16GB of VRAM (est.) | 8 |
| fit in 24GB of VRAM (est.) | 11 |
| parameters 15–40B | 6 |
| parameters 4–15B | 5 |
| parameters up to 4B | 2 |
| parameters over 40B | 2 |
| Converted builds appended to earlier articles | 18 (GGUF 10, FP8 1, MLX 7) |
| Most active publishers | efficient-large-model (5), ggml-org (3), unslothai/unsloth (2) |
| Articles still marked unverified | 1 |
Trending models we did not cover separately
Hugging Face repositories that trended this week but did not get their own article. Listed for reference; we have not reviewed their model cards.
| Model | Likes | Downloads | Why no article |
|---|---|---|---|
| PSRben/VisionHOPE | 402 | 1,516 | publisher not on our notable list |
| Aleph-Alpha/Kolibri-1 | 388 | 1,135 | converted build without a parent article |
| ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF | 260 | 351,230 | publisher not on our notable list |
| NaiveAI/Naive-N0.5-Flash | 165 | 1,920 | publisher not on our notable list |
| bartowski/FrogNano-4B-2609-GGUF | 11 | 7,519 | converted build without a parent article |
| bartowski/bytkim_Qwen3.8-27B-pi-GGUF | 7 | 3,909 | converted build without a parent article |
| bartowski/GLM-5.3-Flash-BF16-GGUF | 4 | 2,288 | converted build without a parent article |
| bartowski/cosmicoptima_computer-10-GGUF | 3 | 6,853 | converted build without a parent article |
| bartowski/OmniJev_OneJev-9B-GGUF | 3 | 3,333 | converted build without a parent article |
| bartowski/Lythri_Lythri-4B-A2B-GGUF | 2 | 2,819 | converted build without a parent article |
Patch releases we did not cover separately
Releases of watched projects that were patch-level or had short notes. Each project’s page lists every version.
| Project | Version | Release notes |
|---|---|---|
| magnitudedev/magnitude | @magnitudedev/cli@0.2.2 | GitHub |
| magnitudedev/magnitude | @magnitudedev/cli@0.2.3 | GitHub |
| ollama/ollama | v0.35.1 | GitHub |
| turboderp-org/exllamav3 | v1.5.4 | GitHub |
| unslothai/unsloth | v0.1.902-beta | GitHub |
Models Gaining the Most Likes
Compiled by Local Model Watch from weekly snapshots of models we’ve covered. Likes are cumulative on Hugging Face; downloads are its trailing-30-day count.
→ Scroll horizontally to see all columns
| Model | Likes gained | Total likes | Downloads (30d) | Period |
|---|---|---|---|---|
| Edge0/Audio8-ASR-Infinite | +1,401 | 2,426 | 40,004 | 2026-W40 → 2026-W41 |
| Qwen/Qwen3.8-27B | +515 | 16,939 | 6,821,761 | 2026-W40 → 2026-W41 |
| deepseek-ai/DeepSeek-V4.1-Flash | +283 | 4,093 | 798,422 | 2026-W40 → 2026-W41 |
| Viggle/Qwen-Image-2.1-viggle-turbo | +246 | 587 | 272,896 | 2026-W40 → 2026-W41 |
| prism-ml/Ternary-Bonsai-2-27B-gguf | +228 | 2,416 | 4,045,810 | 2026-W40 → 2026-W41 |
Highlights of the Week
The most notable highlights surrounding local AI models this week are the following three points.
First, there is the rise of models specialized in structured output and decision-making, along with advances in their GGUF support. Models optimized for specific tasks and structured data processing, such as “clef" and “clef-flash" released by Cloudflare, and “OpenJev-GGUF" released by ggml-org, have appeared one after another. This significantly improves the practicality of agents and automation tools that can be run on local PCs.
Second, LoRA adapters for video generation models have been released very actively. Methods for practically customizing and controlling existing powerful models have been enriched, such as the “LongLive-Plug" series by Efficient-Large-Model and LoRAs for video inpainting and quality conversion by Lightricks.
Third, competition to accelerate local inference engines and expand their features is intensifying. The emergence of the new Rust-based engine “Magnitude" and the addition of new APIs for decision-making models in “Ollama v0.35.0" are rapidly establishing the infrastructure for developers to run models more comfortably in their local environments.
Trends by Category
Text Generation
This week saw the release of numerous practical models specialized for specific use cases and reasoning capabilities. Optimization is progressing not only for general conversation, but also for tasks requiring structured data output and complex thought processes.
Cloudflare released “clef" (27.4B) and its lightweight version “clef-flash" (9.4B), which are specialized for structured decision-making tasks. These exhibit high accuracy in use cases such as API integration and data extraction. Additionally, domestic Japanese developer ELYZA released the 32B and 33B models of “ELYZA-Thinking-1.0“, an inference model capable of outputting thought processes, enabling local execution of advanced reasoning tasks in Japanese environments. Furthermore, previously released models like “Xing4.0-29B-A4B` and “Hemmingway-1" have also been covered on this site, further enriching the options in the mid-size tier.
Image, Video, and Audio
It was a week notable for advancements in video generation control technology and the release of unique models specialized in audio processing.
In the field of video generation, numerous LoRAs were released to efficiently control underlying large-scale models. In particular, LoRA adapters targeting Wan2.1 and Wan2.2, such as “LongLive-Plug-Wan2.2-TI2V-5B-cfg“, enable high-quality generation with fewer steps and camera angle control. In the audio field, Google released “DiarizationLM-Gemma-4-E4B-v1“, which is suitable for speech transcription and speaker diarization, supporting the local execution of practical audio processing. Note that the release of the Turkish synthetic speech dataset “alania-synthetic-speech-tr" has also been reported, but please be aware that this is unverified information that has not been officially confirmed.
Engines and Tools
Development competition continues with the goal of maximizing execution speed in local environments and rapidly supporting the latest models.
The image generation GUI “ComfyUI v0.38.0" was released, adding support for the latest Hunyuan Image 3.5 and Qwen-Image 2.1. In addition, “Ollama v0.35.0“, a popular local LLM execution environment, implemented a new API to make decision-making models easier to handle. Furthermore, a new inference engine written in Rust called “Magnitude" appeared, followed by rapid-fire updates such as “Magnitude CLI 0.2.4" and “@magnitudedev/cli v0.2.5" which advance memory reduction and Apple Silicon optimization. The evolution of these tools is building an environment where the latest models can be run efficiently even with limited hardware resources.
Industry News
These are announcements from companies and research institutions that were not made into standalone articles because they are not topics about running locally. Only the key points are listed.
- ankitjh4: The dataset “Bharat Guide" containing 55,050 text documents extracted and normalized from Indian government public documents has been released. (Announcement)
- Kuyawa: The desktop app “DeepSeek Harness Desktop" is being developed for macOS and Windows, allowing users to create and extend plugins through chat. (Announcement)
- NVIDIA Developer: “DIN Deploy" has been released, a C++ sample code combining ONNX Runtime and NVIDIA TensorRT RTX on Windows and Linux to accelerate local AI inference. (Announcement)
- NVIDIA Developer: A fine-tuning method was announced for the speech recognition model “NVIDIA Nemotron 3.5 ASR" to support regional dialects in Saudi Arabia such as Najdi and Hijazi. (Announcement)
- Hugging Face Blog: The “Open TTS Leaderboard" has been published on Hugging Face to evaluate multilingual text-to-speech (TTS) and voice cloning models using a standardized methodology. (Announcement)
This Week’s Articles
Compiled by Local Model Watch from the article log. Articles about the same story are merged into one row.
New Models
→ Scroll horizontally to see all columns
| Date | Model | Params | Smallest VRAM tier | License | Converted builds | Article |
|---|---|---|---|---|---|---|
| 2026-10-05 | Qwen/Qwen3.8-Flash-Next | 180.0B | — | qwen-community-1.0 | — | Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory |
| 2026-10-04 | LiquidAI/LFM2.5-350M-Diffusion-Exp | 425M | 4GB | lfm1.0 | — | LFM2.5-350M-Diffusion-Exp Text Generation Model: 4GB+ VRAM |
| 2026-10-04 | ggml-org/GLM-5.3-Flash-GGUF | 321.3B | — | other | — | GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory |
| 2026-10-03 | — | — | — | — | — | Ai2 Releases AstaBrief 8B: Fast Open-Source Scientific Report… |
| 2026-10-02 | elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b | 32.1B | 80GB | apache-2.0 | — | ELYZA Releases ELYZA-Thinking-1.0 32B/33B Reasoning Models |
| 2026-10-02 | Cloudflare/clef-flash | 9.4B | 4GB | apache-2.0 | 3 | clef-flash Vision-Language Model: Our Test Answers, 4GB+ VRAM |
| 2026-09-28 | orcarouter/OrcaSAQ-2-27B | 27.8B | 16GB | apache-2.0 | — | OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM (follow-up: clef Structured Decision-Making Model: 12GB+ VRAM, GGUF Builds) |
| 2026-09-28 | Altworld/Hemmingway-1 | 26.9B | 12GB | cc-by-nc-4.0 | 3 | Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds |
| 2026-09-28 | apple/LensVLM-9B | 9.4B | 4GB | apple-amlr | 3 | LensVLM-9B Vision-Language Model: 4GB+ VRAM, GGUF Builds |
| 2026-09-28 | XingChen-AGI/Xing4.0-29B-A4B | 31.2B | 24GB | apache-2.0 | 3 | Xing4.0-29B-A4B Text Generation Model: Our Test Answers, 24GB+ VRAM |
Image, Video and Audio
→ Scroll horizontally to see all columns
| Date | Model | Params | Smallest VRAM tier | License | Converted builds | Article |
|---|---|---|---|---|---|---|
| 2026-10-05 | google/DiarizationLM-Gemma-4-E4B-v1 | 8.0B | 8GB | apache-2.0 | — | DiarizationLM-Gemma-4-E4B-v1 Vision-Language Model: 8GB+ VRAM |
| 2026-10-02 | nvidia/PixelUMM | 8.2B | 24GB | nvidia-one-way-noncommercial-license | — | PixelUMM Multimodal Model: 24GB+ VRAM |
| 2026-10-02 | — | — | — | — | — | Alania Synthetic Speech TR: Turkish Speech Dataset Released |
| 2026-10-01 | FermionResearch/Phonon-2 | 627M | 4GB | cc-by-4.0 | — | Phonon-2 Speech Recognition Model: 4GB+ VRAM |
| 2026-10-01 | Lightricks/LTX-2.5-22b-IC-LoRA-SDR-To-HDR | — | — | ltx-2.x-community-license | — | LTX-2.5-22b-IC-LoRA-SDR-To-HDR Video Generation Model: File List |
| 2026-09-30 | Lightricks/LTX-2.5-22b-IC-LoRA-Restore | — | — | ltx-2.x-community-license | — | LTX-2.5-22b-IC-LoRA-Restore Video Generation Model: File List |
| 2026-09-30 | lilylilith/QI_2.1_AnyAngle | 7.1B | 48GB | apache-2.0 | — | QI_2.1_AnyAngle Camera Angle Control LoRA: 48GB+ VRAM, File List |
| 2026-09-30 | Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-few-step | — | — | apache-2.0 | — | LongLive-Plug-Wan2.1-T2V-14B-few-step: Our Generated Video, File List |
| 2026-09-30 | Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-cfg | — | — | apache-2.0 | — | LongLive-Plug-Wan2.1-T2V-14B-cfg: Our Generated Video, File List |
| 2026-09-30 | Efficient-Large-Model/LongLive-Plug-MiniMax-H3-cfg | — | — | minimax-h3-community-license-agreement | — | Japanese article only |
| 2026-09-29 | Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-few-step | — | — | apache-2.0 | — | LongLive-Plug-Wan2.2-TI2V-5B-few-step: Our Generated Video, File List (follow-up: LongLive-Plug-Wan2.2-TI2V-5B-cfg: Our Generated Video, File List) |
| 2026-09-28 | akatz-ai/MiniMax-H3-Character-Swap-LoRA | — | — | minimax-h3-community-license-agreement | — | MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List |
Engines and Tools
Community
| Date | Article |
|---|---|
| 2026-10-04 | Aleph Alpha Releases Kolibri: A New Open-Weight MoE Model |
