Local Models That Run in 8GB of VRAM (RTX 4060 / 3060 Ti, etc.)

September 26, 2026

Every model covered by Local Model Watch that fits in 8GB of GPU memory, with the best build that fits. The “Best build for 8GB" column is the largest (highest-quality) quantization or precision whose estimated memory stays within 8GB — often better than the smallest build listed in the VRAM quick reference. Estimates are computed by this site from actual file sizes plus runtime overhead (method); long contexts need more.

11 models listed.

Everything that fits your GPU: 8GB / 12GB / 16GB / 24GB

Text Generation (Chat and Code)

→ Scroll horizontally to see all columns

Model Parameters Best build for 8GB Est. memory Smallest tier Runs on Article
prism-ml/Ternary-Bonsai-2-27B-gguf 27.8B Q1_0 (5.5GB) 6.6GB 8GB Ollama / LM Studio, MLX Read
openbmb/MiniCPM5-2B 2.5B F16 (4.7GB) 5.6GB 4GB Ollama / LM Studio, vLLM, MLX Read
harshatheg/Qwen-2.5-1B-RLCD 1.5B BF16 (2.9GB) 3.5GB 4GB MLX Read
pfnet/plamo-3-610m-fin-instruct 890M BF16 (1.7GB) 2.0GB 4GB vLLM Read
togethercomputer/Tev1-0.8B-experimental 873M F32 (1.6GB) 2.0GB 4GB vLLM Read
Cactus-Compute/needle3 – Original precision (0.2GB) 0.3GB 4GB – Read

Vision-Language and Multimodal Models

→ Scroll horizontally to see all columns

Model Parameters Best build for 8GB Est. memory Smallest tier Runs on Article
LiquidAI/LFM2.5-VL-3B-DSpark 279M F16 (0.5GB) 0.6GB 4GB – Read

Image Generation and Editing

→ Scroll horizontally to see all columns

Model Parameters Best build for 8GB Est. memory Smallest tier Runs on Article
abenzerps/Qwen-Image-2.1-Uncensored-GGUF 7.1B Q6_K (5.5GB) 6.6GB 8GB – Read

Audio: Speech Synthesis, Music and Speech Recognition

→ Scroll horizontally to see all columns

Model Parameters Best build for 8GB Est. memory Smallest tier Runs on Article
phasefield-audio/Irodori-TTS-v4.1-Anime 766M F32 (2.9GB) 3.4GB 4GB – Read

Task not declared

→ Scroll horizontally to see all columns

Model Parameters Best build for 8GB Est. memory Smallest tier Runs on Article
tencent/WeVisDoc-4B 4.4B Q8_0 (4.4GB) 5.2GB 4GB – Read
Efficient-Large-Model/H3-to-LTX-Latent-Adapter 195M BF16 (0.4GB) 0.4GB 4GB – Read