Local Models That Run in 16GB of VRAM (RTX 5060 Ti 16GB / 4060 Ti 16GB, etc.)

September 26, 2026

Every model covered by Local Model Watch that fits in 16GB of GPU memory, with the best build that fits. The “Best build for 16GB" column is the largest (highest-quality) quantization or precision whose estimated memory stays within 16GB — often better than the smallest build listed in the VRAM quick reference. Estimates are computed by this site from actual file sizes plus runtime overhead (method); long contexts need more.

25 models listed.

Everything that fits your GPU: 8GB / 12GB / 16GB / 24GB

Text Generation (Chat and Code)

→ Scroll horizontally to see all columns

Model Parameters Best build for 16GB Est. memory Smallest tier Runs on Article
nex-agi/Nex-N2.5-mini 35.1B Q2_K (12.9GB) 15.5GB 16GB Ollama / LM Studio, vLLM, MLX Read
prism-ml/Ternary-Bonsai-2-27B-gguf 27.8B Q2_0 (6.7GB) 8.1GB 8GB Ollama / LM Studio, MLX Read
bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF 26.5B IQ3_XS (12.7GB) 15.2GB 12GB Ollama / LM Studio Read
togethercomputer/Tev1-4B-experimental 4.7B BF16 (8.7GB) 10.4GB 12GB vLLM Read
tencent/Simple-Attention-Sparsification 4.0B BF16 (7.5GB) 9.0GB 12GB – Read
openbmb/MiniCPM5-2B 2.5B F16 (4.7GB) 5.6GB 4GB Ollama / LM Studio, vLLM, MLX Read
harshatheg/Qwen-2.5-1B-RLCD 1.5B BF16 (2.9GB) 3.5GB 4GB MLX Read
pfnet/plamo-3-610m-fin-instruct 890M BF16 (1.7GB) 2.0GB 4GB vLLM Read
togethercomputer/Tev1-0.8B-experimental 873M F32 (1.6GB) 2.0GB 4GB vLLM Read
Cactus-Compute/needle3 – Original precision (0.2GB) 0.3GB 4GB – Read

Vision-Language and Multimodal Models

→ Scroll horizontally to see all columns

Model Parameters Best build for 16GB Est. memory Smallest tier Runs on Article
bartowski/vectionlabs_Salience-27B-R6-GGUF 27.8B Q3_K_L (13.2GB) 15.8GB 12GB Ollama / LM Studio Read
DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF 27.8B IQ3_M (11.7GB) 14.1GB 12GB Ollama / LM Studio Read
agentionai/Signal-3.8-27B-GGUF 27.8B IQ4_XS (13.3GB) 15.9GB 16GB Ollama / LM Studio Read
Jackrong/Qwopus3.8-27B-Flash-GGUF 27.8B Q3_K_M (12.6GB) 15.1GB 16GB Ollama / LM Studio Read (Japanese)
bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF 25.8B IQ3_M (13.4GB) 16.0GB 12GB Ollama / LM Studio Read
ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF 9.4B Q8_0 (8.9GB) 10.6GB 12GB Ollama / LM Studio Read
LiquidAI/LFM2.5-VL-3B-DSpark 279M F16 (0.5GB) 0.6GB 4GB – Read

Image Generation and Editing

→ Scroll horizontally to see all columns

Model Parameters Best build for 16GB Est. memory Smallest tier Runs on Article
abenzerps/Qwen-Image-2.1-Uncensored-GGUF 7.1B BF16 (13.3GB) 15.9GB 8GB – Read

Video Generation

→ Scroll horizontally to see all columns

Model Parameters Best build for 16GB Est. memory Smallest tier Runs on Article
WarmBloodAban/Minimax-h3_Singularity 33.1B W4A8 (11.0GB) 13.2GB 16GB – Read (Japanese)

Audio: Speech Synthesis, Music and Speech Recognition

→ Scroll horizontally to see all columns

Model Parameters Best build for 16GB Est. memory Smallest tier Runs on Article
Edge0/Audio8-ASR-Infinite 4.1B BF16 (7.6GB) 9.1GB 12GB – Read
m-a-p/YuE2-3B 3.6B BF16 (6.8GB) 8.1GB 12GB – Read
phasefield-audio/Irodori-TTS-v4.1-Anime 766M F32 (2.9GB) 3.4GB 4GB – Read

Task not declared

→ Scroll horizontally to see all columns

Model Parameters Best build for 16GB Est. memory Smallest tier Runs on Article
tencent/WeVisDoc-4B 4.4B F16 (8.2GB) 9.9GB 4GB – Read
Comfy-Org/YuE2 3.6B BF16 (6.8GB) 8.1GB 12GB – Read
Efficient-Large-Model/H3-to-LTX-Latent-Adapter 195M BF16 (0.4GB) 0.4GB 4GB – Read