Local Models That Run in 12GB of VRAM (RTX 4070 / 3060 12GB, etc.)
Every model covered by Local Model Watch that fits in 12GB of GPU memory, with the best build that fits. The “Best build for 12GB" column is the largest (highest-quality) quantization or precision whose estimated memory stays within 12GB — often better than the smallest build listed in the VRAM quick reference. Estimates are computed by this site from actual file sizes plus runtime overhead (method); long contexts need more.
21 models listed.
Everything that fits your GPU: 8GB / 12GB / 16GB / 24GB
Text Generation (Chat and Code)
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 12GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| prism-ml/Ternary-Bonsai-2-27B-gguf | 27.8B | Q2_0 (6.7GB) | 8.1GB | 8GB | Ollama / LM Studio, MLX | Read |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | 26.5B | IQ2_S (10.0GB) | 12.0GB | 12GB | Ollama / LM Studio | Read |
| togethercomputer/Tev1-4B-experimental | 4.7B | BF16 (8.7GB) | 10.4GB | 12GB | vLLM | Read |
| tencent/Simple-Attention-Sparsification | 4.0B | BF16 (7.5GB) | 9.0GB | 12GB | – | Read |
| openbmb/MiniCPM5-2B | 2.5B | F16 (4.7GB) | 5.6GB | 4GB | Ollama / LM Studio, vLLM, MLX | Read |
| harshatheg/Qwen-2.5-1B-RLCD | 1.5B | BF16 (2.9GB) | 3.5GB | 4GB | MLX | Read |
| pfnet/plamo-3-610m-fin-instruct | 890M | BF16 (1.7GB) | 2.0GB | 4GB | vLLM | Read |
| togethercomputer/Tev1-0.8B-experimental | 873M | F32 (1.6GB) | 2.0GB | 4GB | vLLM | Read |
| Cactus-Compute/needle3 | – | Original precision (0.2GB) | 0.3GB | 4GB | – | Read |
Vision-Language and Multimodal Models
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 12GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| bartowski/vectionlabs_Salience-27B-R6-GGUF | 27.8B | IQ2_M (9.8GB) | 11.8GB | 12GB | Ollama / LM Studio | Read |
| DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF | 27.8B | IQ2_M (9.3GB) | 11.2GB | 12GB | Ollama / LM Studio | Read |
| bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF | 25.8B | IQ2_S (9.6GB) | 11.5GB | 12GB | Ollama / LM Studio | Read |
| ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF | 9.4B | Q8_0 (8.9GB) | 10.6GB | 12GB | Ollama / LM Studio | Read |
| LiquidAI/LFM2.5-VL-3B-DSpark | 279M | F16 (0.5GB) | 0.6GB | 4GB | – | Read |
Image Generation and Editing
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 12GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| abenzerps/Qwen-Image-2.1-Uncensored-GGUF | 7.1B | Q8_0 (7.1GB) | 8.5GB | 8GB | – | Read |
Audio: Speech Synthesis, Music and Speech Recognition
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 12GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| Edge0/Audio8-ASR-Infinite | 4.1B | BF16 (7.6GB) | 9.1GB | 12GB | – | Read |
| m-a-p/YuE2-3B | 3.6B | BF16 (6.8GB) | 8.1GB | 12GB | – | Read |
| phasefield-audio/Irodori-TTS-v4.1-Anime | 766M | F32 (2.9GB) | 3.4GB | 4GB | – | Read |
Task not declared
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 12GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| tencent/WeVisDoc-4B | 4.4B | F16 (8.2GB) | 9.9GB | 4GB | – | Read |
| Comfy-Org/YuE2 | 3.6B | BF16 (6.8GB) | 8.1GB | 12GB | – | Read |
| Efficient-Large-Model/H3-to-LTX-Latent-Adapter | 195M | BF16 (0.4GB) | 0.4GB | 4GB | – | Read |