Local Models That Run in 16GB of VRAM (RTX 5060 Ti 16GB / 4060 Ti 16GB, etc.)
Every model covered by Local Model Watch that fits in 16GB of GPU memory, with the best build that fits. The “Best build for 16GB" column is the largest (highest-quality) quantization or precision whose estimated memory stays within 16GB — often better than the smallest build listed in the VRAM quick reference. Estimates are computed by this site from actual file sizes plus runtime overhead (method); long contexts need more.
25 models listed.
Everything that fits your GPU: 8GB / 12GB / 16GB / 24GB
Text Generation (Chat and Code)
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 16GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| nex-agi/Nex-N2.5-mini | 35.1B | Q2_K (12.9GB) | 15.5GB | 16GB | Ollama / LM Studio, vLLM, MLX | Read |
| prism-ml/Ternary-Bonsai-2-27B-gguf | 27.8B | Q2_0 (6.7GB) | 8.1GB | 8GB | Ollama / LM Studio, MLX | Read |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | 26.5B | IQ3_XS (12.7GB) | 15.2GB | 12GB | Ollama / LM Studio | Read |
| togethercomputer/Tev1-4B-experimental | 4.7B | BF16 (8.7GB) | 10.4GB | 12GB | vLLM | Read |
| tencent/Simple-Attention-Sparsification | 4.0B | BF16 (7.5GB) | 9.0GB | 12GB | – | Read |
| openbmb/MiniCPM5-2B | 2.5B | F16 (4.7GB) | 5.6GB | 4GB | Ollama / LM Studio, vLLM, MLX | Read |
| harshatheg/Qwen-2.5-1B-RLCD | 1.5B | BF16 (2.9GB) | 3.5GB | 4GB | MLX | Read |
| pfnet/plamo-3-610m-fin-instruct | 890M | BF16 (1.7GB) | 2.0GB | 4GB | vLLM | Read |
| togethercomputer/Tev1-0.8B-experimental | 873M | F32 (1.6GB) | 2.0GB | 4GB | vLLM | Read |
| Cactus-Compute/needle3 | – | Original precision (0.2GB) | 0.3GB | 4GB | – | Read |
Vision-Language and Multimodal Models
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 16GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| bartowski/vectionlabs_Salience-27B-R6-GGUF | 27.8B | Q3_K_L (13.2GB) | 15.8GB | 12GB | Ollama / LM Studio | Read |
| DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF | 27.8B | IQ3_M (11.7GB) | 14.1GB | 12GB | Ollama / LM Studio | Read |
| agentionai/Signal-3.8-27B-GGUF | 27.8B | IQ4_XS (13.3GB) | 15.9GB | 16GB | Ollama / LM Studio | Read |
| Jackrong/Qwopus3.8-27B-Flash-GGUF | 27.8B | Q3_K_M (12.6GB) | 15.1GB | 16GB | Ollama / LM Studio | Read (Japanese) |
| bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF | 25.8B | IQ3_M (13.4GB) | 16.0GB | 12GB | Ollama / LM Studio | Read |
| ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF | 9.4B | Q8_0 (8.9GB) | 10.6GB | 12GB | Ollama / LM Studio | Read |
| LiquidAI/LFM2.5-VL-3B-DSpark | 279M | F16 (0.5GB) | 0.6GB | 4GB | – | Read |
Image Generation and Editing
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 16GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| abenzerps/Qwen-Image-2.1-Uncensored-GGUF | 7.1B | BF16 (13.3GB) | 15.9GB | 8GB | – | Read |
Video Generation
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 16GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| WarmBloodAban/Minimax-h3_Singularity | 33.1B | W4A8 (11.0GB) | 13.2GB | 16GB | – | Read (Japanese) |
Audio: Speech Synthesis, Music and Speech Recognition
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 16GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| Edge0/Audio8-ASR-Infinite | 4.1B | BF16 (7.6GB) | 9.1GB | 12GB | – | Read |
| m-a-p/YuE2-3B | 3.6B | BF16 (6.8GB) | 8.1GB | 12GB | – | Read |
| phasefield-audio/Irodori-TTS-v4.1-Anime | 766M | F32 (2.9GB) | 3.4GB | 4GB | – | Read |
Task not declared
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 16GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| tencent/WeVisDoc-4B | 4.4B | F16 (8.2GB) | 9.9GB | 4GB | – | Read |
| Comfy-Org/YuE2 | 3.6B | BF16 (6.8GB) | 8.1GB | 12GB | – | Read |
| Efficient-Large-Model/H3-to-LTX-Latent-Adapter | 195M | BF16 (0.4GB) | 0.4GB | 4GB | – | Read |