Local Models by Task: Image, Audio, Video and Text
Open-weight models covered by Local Model Watch, grouped by what they do. The task is the one each publisher declares on Hugging Face (pipeline_tag), sorted by code — models without a declared task are not listed. The smallest VRAM tier is computed by this site from the actual file sizes; to find everything that fits your GPU, see the VRAM quick reference.
38 models listed.
- Text Generation (Chat and Code) (17)
- Vision-Language and Multimodal Models (12)
- Image Generation and Editing (4)
- Video Generation (2)
- Audio: Speech Synthesis, Music and Speech Recognition (3)
Text Generation (Chat and Code)
→ Scroll horizontally to see all columns
| Model | Task | Parameters | Smallest VRAM tier | License | Published | Article |
|---|---|---|---|---|---|---|
| togethercomputer/Tev1-0.8B-experimental | text-generation | 873M | 4GB | — | 2026-09-26 | Read |
| pfnet/plamo-3-610m-fin-instruct | text-generation | 890M | 4GB | other | 2026-09-24 | Read |
| togethercomputer/Tev1-4B-experimental | text-generation | 4.7B | 12GB | — | 2026-09-24 | Read |
| IFM/K2-Horizon-32B-NVFP4 | text-generation | — | 32GB | apache-2.0 | 2026-09-23 | Read |
| IFM/K2-Horizon-375B-A23B-NVFP4 | text-generation | — | Over 80GB | apache-2.0 | 2026-09-23 | Read |
| Cactus-Compute/needle3 | text-generation | — | 4GB | apache-2.0 | 2026-09-19 | Read |
| prism-ml/Ternary-Bonsai-2-27B-gguf | text-generation | 27.8B | 8GB | apache-2.0 | 2026-09-18 | Read |
| harshatheg/Qwen-2.5-1B-RLCD | text-generation | 1.5B | 4GB | apache-2.0 | 2026-09-17 | Read |
| EleutherAI/olmo3-7b-sdf-sft-clean150 | text-generation | 7.3B | 24GB | apache-2.0 | 2026-09-16 | Read |
| tencent/Simple-Attention-Sparsification | text-generation | 4.0B | 12GB | — | 2026-09-14 | Read |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | text-generation | 26.5B | 12GB | apache-2.0 | 2026-09-11 | Read |
| Edge0/Edge0-35B-A3B-preview | text-generation | 36.0B | 24GB | apache-2.0 | 2026-09-11 | Read |
| nvidia/DeepSeek-V4-Pro-0813-nvfp4-DSpark | text-generation | 1650.5B | Over 80GB | mit | 2026-09-10 | Read |
| nex-agi/Nex-N2.5-Pro | text-generation | 396.8B | Over 80GB | apache-2.0 | 2026-09-09 | Read |
| nex-agi/Nex-N2.5-mini | text-generation | 35.1B | 16GB | apache-2.0 | 2026-09-09 | Read |
| openbmb/MiniCPM5-2B | text-generation | 2.5B | 4GB | apache-2.0 | 2026-09-08 | Read |
| IFM/K2-Horizon-MoVA-36B-A4B-GGUF | text-generation | 37.4B | 32GB | apache-2.0 | 2026-09-06 | Read (Japanese) |
Vision-Language and Multimodal Models
→ Scroll horizontally to see all columns
| Model | Task | Parameters | Smallest VRAM tier | License | Published | Article |
|---|---|---|---|---|---|---|
| LiquidAI/LFM2.5-VL-3B-DSpark | image-text-to-text | 279M | 4GB | — | 2026-09-25 | Read |
| ggml-org/MiMo-V2.6-Flash-RL-GGUF | image-text-to-text | — | Over 80GB | mit | 2026-09-22 | Read |
| ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF | image-text-to-text | 9.4B | 12GB | mit | 2026-09-22 | Read |
| inclusionAI/Realtime-Venus | any-to-any | — | 24GB | apache-2.0 | 2026-09-19 | Read |
| bartowski/vectionlabs_Salience-27B-R6-GGUF | image-text-to-text | 27.8B | 12GB | apache-2.0 | 2026-09-15 | Read |
| bartowski/Intern-S2-397B-GGUF | image-text-to-text | 403.4B | Over 80GB | apache-2.0 | 2026-09-15 | Read |
| bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF | any-to-any | 25.8B | 12GB | — | 2026-09-14 | Read |
| dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 | image-text-to-text | 763.2B | Over 80GB | mit | 2026-09-13 | Read |
| DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF | image-text-to-text | 27.8B | 12GB | apache-2.0 | 2026-09-13 | Read |
| agentionai/Signal-3.8-27B-GGUF | image-text-to-text | 27.8B | 16GB | apache-2.0 | 2026-09-12 | Read |
| deepseek-ai/DeepSeek-V4.1-Flash | image-text-to-text | 763.2B | Over 80GB | mit | 2026-09-10 | Read |
| Jackrong/Qwopus3.8-27B-Flash-GGUF | image-text-to-text | 27.8B | 16GB | apache-2.0 | 2026-09-06 | Read (Japanese) |
Image Generation and Editing
→ Scroll horizontally to see all columns
| Model | Task | Parameters | Smallest VRAM tier | License | Published | Article |
|---|---|---|---|---|---|---|
| Viggle/Qwen-Image-2.1-viggle-turbo | text-to-image | 7.1B | — | other | 2026-09-23 | Read |
| SupraLabs/Supra2-IMG | text-to-image | — | — | apache-2.0 | 2026-09-23 | Read |
| abenzerps/Qwen-Image-2.1-Uncensored-GGUF | text-to-image | 7.1B | 8GB | other | 2026-09-21 | Read |
| Qwen/Qwen-Image-2.1-PE-I2I | text-to-image | — | — | other | 2026-09-20 | Read |
Video Generation
→ Scroll horizontally to see all columns
| Model | Task | Parameters | Smallest VRAM tier | License | Published | Article |
|---|---|---|---|---|---|---|
| Lightricks/LTX-2.5-22b-IC-LoRA-Ingredients | video-to-video | — | — | other | 2026-09-11 | Read |
| WarmBloodAban/Minimax-h3_Singularity | image-to-video | 33.1B | 16GB | apache-2.0 | 2026-09-07 | Read (Japanese) |
Audio: Speech Synthesis, Music and Speech Recognition
→ Scroll horizontally to see all columns
| Model | Task | Parameters | Smallest VRAM tier | License | Published | Article |
|---|---|---|---|---|---|---|
| Edge0/Audio8-ASR-Infinite | automatic-speech-recognition | 4.1B | 12GB | apache-2.0 | 2026-09-24 | Read |
| m-a-p/YuE2-3B | text-to-audio | 3.6B | 12GB | cc-by-nc-4.0 | 2026-09-10 | Read |
| phasefield-audio/Irodori-TTS-v4.1-Anime | text-to-speech | 766M | 4GB | mit | 2026-09-07 | Read |