Local Models That Run in 12GB of VRAM (RTX 4070 / 3060 12GB, etc.)

September 26, 2026

Every model covered by Local Model Watch that fits in 12GB of GPU memory, with the best build that fits. The “Best build for 12GB" column is the largest (highest-quality) quantization or precision whose estimated memory stays within 12GB — often better than the smallest build listed in the VRAM quick reference. Estimates are computed by this site from actual file sizes plus runtime overhead (method); long contexts need more.

21 models listed.

Everything that fits your GPU: 8GB / 12GB / 16GB / 24GB

Text Generation (Chat and Code)

→ Scroll horizontally to see all columns

Model Parameters Best build for 12GB Est. memory Smallest tier Runs on Article
prism-ml/Ternary-Bonsai-2-27B-gguf 27.8B Q2_0 (6.7GB) 8.1GB 8GB Ollama / LM Studio, MLX Read
bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF 26.5B IQ2_S (10.0GB) 12.0GB 12GB Ollama / LM Studio Read
togethercomputer/Tev1-4B-experimental 4.7B BF16 (8.7GB) 10.4GB 12GB vLLM Read
tencent/Simple-Attention-Sparsification 4.0B BF16 (7.5GB) 9.0GB 12GB – Read
openbmb/MiniCPM5-2B 2.5B F16 (4.7GB) 5.6GB 4GB Ollama / LM Studio, vLLM, MLX Read
harshatheg/Qwen-2.5-1B-RLCD 1.5B BF16 (2.9GB) 3.5GB 4GB MLX Read
pfnet/plamo-3-610m-fin-instruct 890M BF16 (1.7GB) 2.0GB 4GB vLLM Read
togethercomputer/Tev1-0.8B-experimental 873M F32 (1.6GB) 2.0GB 4GB vLLM Read
Cactus-Compute/needle3 – Original precision (0.2GB) 0.3GB 4GB – Read

Vision-Language and Multimodal Models

→ Scroll horizontally to see all columns

Model Parameters Best build for 12GB Est. memory Smallest tier Runs on Article
bartowski/vectionlabs_Salience-27B-R6-GGUF 27.8B IQ2_M (9.8GB) 11.8GB 12GB Ollama / LM Studio Read
DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF 27.8B IQ2_M (9.3GB) 11.2GB 12GB Ollama / LM Studio Read
bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF 25.8B IQ2_S (9.6GB) 11.5GB 12GB Ollama / LM Studio Read
ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF 9.4B Q8_0 (8.9GB) 10.6GB 12GB Ollama / LM Studio Read
LiquidAI/LFM2.5-VL-3B-DSpark 279M F16 (0.5GB) 0.6GB 4GB – Read

Image Generation and Editing

→ Scroll horizontally to see all columns

Model Parameters Best build for 12GB Est. memory Smallest tier Runs on Article
abenzerps/Qwen-Image-2.1-Uncensored-GGUF 7.1B Q8_0 (7.1GB) 8.5GB 8GB – Read

Audio: Speech Synthesis, Music and Speech Recognition

→ Scroll horizontally to see all columns

Model Parameters Best build for 12GB Est. memory Smallest tier Runs on Article
Edge0/Audio8-ASR-Infinite 4.1B BF16 (7.6GB) 9.1GB 12GB – Read
m-a-p/YuE2-3B 3.6B BF16 (6.8GB) 8.1GB 12GB – Read
phasefield-audio/Irodori-TTS-v4.1-Anime 766M F32 (2.9GB) 3.4GB 4GB – Read

Task not declared

→ Scroll horizontally to see all columns

Model Parameters Best build for 12GB Est. memory Smallest tier Runs on Article
tencent/WeVisDoc-4B 4.4B F16 (8.2GB) 9.9GB 4GB – Read
Comfy-Org/YuE2 3.6B BF16 (6.8GB) 8.1GB 12GB – Read
Efficient-Large-Model/H3-to-LTX-Latent-Adapter 195M BF16 (0.4GB) 0.4GB 4GB – Read