Local AI Model Weekly Highlights: Structured Output & Video LoRAs

Local AI Model Weekly Highlights: Structured Output & Video LoRAs

This Week in Numbers

Counted by Local Model Watch from the articles published this week.

Item Count
Articles published 36
 Image, Video and Audio 13
 New Models 12
 Engines and Tools 10
 Community 1
New models covered 25
 fit in 8GB of VRAM (est.) 5
 fit in 12GB of VRAM (est.) 7
 fit in 16GB of VRAM (est.) 8
 fit in 24GB of VRAM (est.) 11
 parameters 15–40B 6
 parameters 4–15B 5
 parameters up to 4B 2
 parameters over 40B 2
Converted builds appended to earlier articles 18 (GGUF 10, FP8 1, MLX 7)
Most active publishers efficient-large-model (5), ggml-org (3), unslothai/unsloth (2)
Articles still marked unverified 1

Trending models we did not cover separately

Hugging Face repositories that trended this week but did not get their own article. Listed for reference; we have not reviewed their model cards.

Model Likes Downloads Why no article
PSRben/VisionHOPE 402 1,516 publisher not on our notable list
Aleph-Alpha/Kolibri-1 388 1,135 converted build without a parent article
ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF 260 351,230 publisher not on our notable list
NaiveAI/Naive-N0.5-Flash 165 1,920 publisher not on our notable list
bartowski/FrogNano-4B-2609-GGUF 11 7,519 converted build without a parent article
bartowski/bytkim_Qwen3.8-27B-pi-GGUF 7 3,909 converted build without a parent article
bartowski/GLM-5.3-Flash-BF16-GGUF 4 2,288 converted build without a parent article
bartowski/cosmicoptima_computer-10-GGUF 3 6,853 converted build without a parent article
bartowski/OmniJev_OneJev-9B-GGUF 3 3,333 converted build without a parent article
bartowski/Lythri_Lythri-4B-A2B-GGUF 2 2,819 converted build without a parent article

Patch releases we did not cover separately

Releases of watched projects that were patch-level or had short notes. Each project’s page lists every version.

Project Version Release notes
magnitudedev/magnitude @magnitudedev/cli@0.2.2 GitHub
magnitudedev/magnitude @magnitudedev/cli@0.2.3 GitHub
ollama/ollama v0.35.1 GitHub
turboderp-org/exllamav3 v1.5.4 GitHub
unslothai/unsloth v0.1.902-beta GitHub

Models Gaining the Most Likes

Compiled by Local Model Watch from weekly snapshots of models we’ve covered. Likes are cumulative on Hugging Face; downloads are its trailing-30-day count.

→ Scroll horizontally to see all columns

Model Likes gained Total likes Downloads (30d) Period
Edge0/Audio8-ASR-Infinite +1,401 2,426 40,004 2026-W40 → 2026-W41
Qwen/Qwen3.8-27B +515 16,939 6,821,761 2026-W40 → 2026-W41
deepseek-ai/DeepSeek-V4.1-Flash +283 4,093 798,422 2026-W40 → 2026-W41
Viggle/Qwen-Image-2.1-viggle-turbo +246 587 272,896 2026-W40 → 2026-W41
prism-ml/Ternary-Bonsai-2-27B-gguf +228 2,416 4,045,810 2026-W40 → 2026-W41

Highlights of the Week

The most notable highlights surrounding local AI models this week are the following three points.

First, there is the rise of models specialized in structured output and decision-making, along with advances in their GGUF support. Models optimized for specific tasks and structured data processing, such as “clef" and “clef-flash" released by Cloudflare, and “OpenJev-GGUF" released by ggml-org, have appeared one after another. This significantly improves the practicality of agents and automation tools that can be run on local PCs.

Second, LoRA adapters for video generation models have been released very actively. Methods for practically customizing and controlling existing powerful models have been enriched, such as the “LongLive-Plug" series by Efficient-Large-Model and LoRAs for video inpainting and quality conversion by Lightricks.

Third, competition to accelerate local inference engines and expand their features is intensifying. The emergence of the new Rust-based engine “Magnitude" and the addition of new APIs for decision-making models in “Ollama v0.35.0" are rapidly establishing the infrastructure for developers to run models more comfortably in their local environments.

Trends by Category

Text Generation

This week saw the release of numerous practical models specialized for specific use cases and reasoning capabilities. Optimization is progressing not only for general conversation, but also for tasks requiring structured data output and complex thought processes.

Cloudflare released “clef" (27.4B) and its lightweight version “clef-flash" (9.4B), which are specialized for structured decision-making tasks. These exhibit high accuracy in use cases such as API integration and data extraction. Additionally, domestic Japanese developer ELYZA released the 32B and 33B models of “ELYZA-Thinking-1.0“, an inference model capable of outputting thought processes, enabling local execution of advanced reasoning tasks in Japanese environments. Furthermore, previously released models like “Xing4.0-29B-A4B` and “Hemmingway-1" have also been covered on this site, further enriching the options in the mid-size tier.

Image, Video, and Audio

It was a week notable for advancements in video generation control technology and the release of unique models specialized in audio processing.

In the field of video generation, numerous LoRAs were released to efficiently control underlying large-scale models. In particular, LoRA adapters targeting Wan2.1 and Wan2.2, such as “LongLive-Plug-Wan2.2-TI2V-5B-cfg“, enable high-quality generation with fewer steps and camera angle control. In the audio field, Google released “DiarizationLM-Gemma-4-E4B-v1“, which is suitable for speech transcription and speaker diarization, supporting the local execution of practical audio processing. Note that the release of the Turkish synthetic speech dataset “alania-synthetic-speech-tr" has also been reported, but please be aware that this is unverified information that has not been officially confirmed.

Engines and Tools

Development competition continues with the goal of maximizing execution speed in local environments and rapidly supporting the latest models.

The image generation GUI “ComfyUI v0.38.0" was released, adding support for the latest Hunyuan Image 3.5 and Qwen-Image 2.1. In addition, “Ollama v0.35.0“, a popular local LLM execution environment, implemented a new API to make decision-making models easier to handle. Furthermore, a new inference engine written in Rust called “Magnitude" appeared, followed by rapid-fire updates such as “Magnitude CLI 0.2.4" and “@magnitudedev/cli v0.2.5" which advance memory reduction and Apple Silicon optimization. The evolution of these tools is building an environment where the latest models can be run efficiently even with limited hardware resources.

Industry News

These are announcements from companies and research institutions that were not made into standalone articles because they are not topics about running locally. Only the key points are listed.

  • ankitjh4: The dataset “Bharat Guide" containing 55,050 text documents extracted and normalized from Indian government public documents has been released. (Announcement)
  • Kuyawa: The desktop app “DeepSeek Harness Desktop" is being developed for macOS and Windows, allowing users to create and extend plugins through chat. (Announcement)
  • NVIDIA Developer: “DIN Deploy" has been released, a C++ sample code combining ONNX Runtime and NVIDIA TensorRT RTX on Windows and Linux to accelerate local AI inference. (Announcement)
  • NVIDIA Developer: A fine-tuning method was announced for the speech recognition model “NVIDIA Nemotron 3.5 ASR" to support regional dialects in Saudi Arabia such as Najdi and Hijazi. (Announcement)
  • Hugging Face Blog: The “Open TTS Leaderboard" has been published on Hugging Face to evaluate multilingual text-to-speech (TTS) and voice cloning models using a standardized methodology. (Announcement)

This Week’s Articles

Compiled by Local Model Watch from the article log. Articles about the same story are merged into one row.

New Models

→ Scroll horizontally to see all columns

Date Model Params Smallest VRAM tier License Converted builds Article
2026-10-05 Qwen/Qwen3.8-Flash-Next 180.0B — qwen-community-1.0 — Qwen3.8-Flash-Next Multimodal MoE Model: ~402GB Memory
2026-10-04 LiquidAI/LFM2.5-350M-Diffusion-Exp 425M 4GB lfm1.0 — LFM2.5-350M-Diffusion-Exp Text Generation Model: 4GB+ VRAM
2026-10-04 ggml-org/GLM-5.3-Flash-GGUF 321.3B — other — GLM-5.3-Flash-GGUF Vision-Language Model: ~150GB Memory
2026-10-03 — — — — — Ai2 Releases AstaBrief 8B: Fast Open-Source Scientific Report…
2026-10-02 elyza/ELYZA-Thinking-1.0-llm-jp-4-32b-a3b 32.1B 80GB apache-2.0 — ELYZA Releases ELYZA-Thinking-1.0 32B/33B Reasoning Models
2026-10-02 Cloudflare/clef-flash 9.4B 4GB apache-2.0 3 clef-flash Vision-Language Model: Our Test Answers, 4GB+ VRAM
2026-09-28 orcarouter/OrcaSAQ-2-27B 27.8B 16GB apache-2.0 — OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM (follow-up: clef Structured Decision-Making Model: 12GB+ VRAM, GGUF Builds)
2026-09-28 Altworld/Hemmingway-1 26.9B 12GB cc-by-nc-4.0 3 Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds
2026-09-28 apple/LensVLM-9B 9.4B 4GB apple-amlr 3 LensVLM-9B Vision-Language Model: 4GB+ VRAM, GGUF Builds
2026-09-28 XingChen-AGI/Xing4.0-29B-A4B 31.2B 24GB apache-2.0 3 Xing4.0-29B-A4B Text Generation Model: Our Test Answers, 24GB+ VRAM

Image, Video and Audio

→ Scroll horizontally to see all columns

Date Model Params Smallest VRAM tier License Converted builds Article
2026-10-05 google/DiarizationLM-Gemma-4-E4B-v1 8.0B 8GB apache-2.0 — DiarizationLM-Gemma-4-E4B-v1 Vision-Language Model: 8GB+ VRAM
2026-10-02 nvidia/PixelUMM 8.2B 24GB nvidia-one-way-noncommercial-license — PixelUMM Multimodal Model: 24GB+ VRAM
2026-10-02 — — — — — Alania Synthetic Speech TR: Turkish Speech Dataset Released
2026-10-01 FermionResearch/Phonon-2 627M 4GB cc-by-4.0 — Phonon-2 Speech Recognition Model: 4GB+ VRAM
2026-10-01 Lightricks/LTX-2.5-22b-IC-LoRA-SDR-To-HDR — — ltx-2.x-community-license — LTX-2.5-22b-IC-LoRA-SDR-To-HDR Video Generation Model: File List
2026-09-30 Lightricks/LTX-2.5-22b-IC-LoRA-Restore — — ltx-2.x-community-license — LTX-2.5-22b-IC-LoRA-Restore Video Generation Model: File List
2026-09-30 lilylilith/QI_2.1_AnyAngle 7.1B 48GB apache-2.0 — QI_2.1_AnyAngle Camera Angle Control LoRA: 48GB+ VRAM, File List
2026-09-30 Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-few-step — — apache-2.0 — LongLive-Plug-Wan2.1-T2V-14B-few-step: Our Generated Video, File List
2026-09-30 Efficient-Large-Model/LongLive-Plug-Wan2.1-T2V-14B-cfg — — apache-2.0 — LongLive-Plug-Wan2.1-T2V-14B-cfg: Our Generated Video, File List
2026-09-30 Efficient-Large-Model/LongLive-Plug-MiniMax-H3-cfg — — minimax-h3-community-license-agreement — Japanese article only
2026-09-29 Efficient-Large-Model/LongLive-Plug-Wan2.2-TI2V-5B-few-step — — apache-2.0 — LongLive-Plug-Wan2.2-TI2V-5B-few-step: Our Generated Video, File List (follow-up: LongLive-Plug-Wan2.2-TI2V-5B-cfg: Our Generated Video, File List)
2026-09-28 akatz-ai/MiniMax-H3-Character-Swap-LoRA — — minimax-h3-community-license-agreement — MiniMax-H3-Character-Swap-LoRA Character Swap LoRA Adapter: File List

Engines and Tools

Date Project Version Article
2026-10-03 magnitudedev/magnitude @magnitudedev/cli@0.2.5 Magnitude CLI v0.2.5 Released: M5 Mac & MoE Speedups
2026-10-03 mudler/LocalAI v4.11.0 LocalAI v4.11.0 Released with Failover and NeMo Audio Models
2026-10-03 magnitudedev/magnitude @magnitudedev/cli@0.2.4 Magnitude 0.2.4 Released: Faster Inference and Lower Memory
2026-10-02 sgl-project/sglang v0.5.21 SGLang v0.5.21 Released with Dynamic PD Role Switching
2026-10-02 — — Testing OpenJev GGUF: A Tiny Model for Runtime Loaders (follow-up: OpenJev-GGUF 27.4B Decision-Making Specialized Model: 24GB+ VRAM)
2026-10-01 unslothai/unsloth v0.1.901-beta Unsloth v0.1.901-beta Released: Major LoRA Memory Reductions
2026-10-01 — — Magnitude: Up to 2x Faster Than llama.cpp Officially, 21x Slower on Our CPU
2026-09-30 ollama/ollama v0.35.0 Ollama v0.35.0 Released: New API for Decision Models
2026-09-30 Comfy-Org/ComfyUI v0.38.0 ComfyUI v0.38.0 Released: Hunyuan Image 3.5 and More
2026-09-29 unslothai/unsloth v0.1.900-beta unsloth Desktop v0.1.900-beta Released with Laya and Speedups

Community