Local Models by VRAM: Quick Reference
Open-weight models covered by Local Model Watch, grouped by the smallest amount of GPU memory (VRAM) they are estimated to run in. Each figure is computed by this site from the actual size of the distributed files plus a runtime overhead — not quoted from model cards. See our editorial policy for the method.
Models listed under a tier do not fit in the smaller tiers. If you have a 24GB GPU, the 8GB, 12GB, 16GB and 24GB tables all apply to you. Models that don’t fit even the largest tier (80GB) are listed separately at the bottom, since that too is something readers need to know.
Last updated 2026-09-19 (JST). 28 models listed.
4GB (laptop iGPU / phone class)
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest build | Est. memory | Article |
|---|---|---|---|---|
| Cactus Compute、8〜29MBの自動化モデル「Needle 3」公開 | – | そのままの精度 (0.2GB) | 0.3GB | Read |
| Lightricks/LTX-2.5-22b-IC-LoRA-Ingredients | – | そのままの精度 (1.2GB) | 1.5GB | Read |
| openbmb/MiniCPM5-2B | 2.5B | Q8_0 (2.5GB) | 3.0GB | Read |
8GB (RTX 4060 / 3060 Ti, etc.)
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest build | Est. memory | Article |
|---|---|---|---|---|
| prism-ml/Ternary-Bonsai-2-27B-gguf | 27.8B | Q1_0 (5.5GB) | 6.6GB | Read |
| harshatheg/Qwen-2.5-1B-RLCD | 1.5B | BF16 (2.9GB) | 3.5GB | Read |
| EleutherAI/bergson-wikitext-gpt2-leaderboard | – | そのままの精度 (6.5GB) | 7.8GB | Read |
| bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF | 25.8B | F16 (1.1GB) | 1.3GB | Read |
| openbmb/MiniCPM5-2B-GGUF | – | F16 (4.7GB) | 5.6GB | Read |
12GB (RTX 4070 / 3060 12GB, etc.)
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest build | Est. memory | Article |
|---|---|---|---|---|
| tencent/WeVisDoc-4B | 4.4B | BF16 (9.0GB) | 10.8GB | Read |
| prism-ml/Ternary-Bonsai-2-27B-mlx-2bit | 27.4B | MLX 2bit (8.0GB) | 9.6GB | Read |
| prism-ml/Ternary-Bonsai-2-27B-gguf-dev | 27.8B | Q2_0 (7.1GB) | 8.5GB | Read |
| Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4 | 3.6B | BF16 (6.8GB) | 8.1GB | Read |
| bartowski/vectionlabs_Salience-27B-R6-GGUF | 27.8B | IQ2_M (9.8GB) | 11.8GB | Read |
| tencent/Simple-Attention-Sparsification | 4.0B | BF16 (7.5GB) | 9.0GB | Read |
| Comfy-Org/YuE2 | 3.6B | BF16 (6.8GB) | 8.1GB | Read |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | 26.5B | IQ2_S (10.0GB) | 12.0GB | Read |
| m-a-p/YuE2-3B | 3.6B | BF16 (6.8GB) | 8.1GB | Read |
16GB (RTX 5060 Ti 16GB / 4060 Ti 16GB, etc.)
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest build | Est. memory | Article |
|---|---|---|---|---|
| bartowski/nex-agi_Nex-N2.5-mini-GGUF | 35.1B | Q2_K (12.9GB) | 15.5GB | Read |
| agentionai/Signal-3.8-27B-GGUF | 27.8B | IQ4_XS (13.3GB) | 15.9GB | Read |
| nex-agi/Nex-N2.5-mini | 35.1B | Q2_K (12.9GB) | 15.5GB | Read |
| Jackrong/Qwopus3.8-27B-Flash-GGUF | 27.8B | Q3_K_M (12.6GB) | 15.1GB | Read |
24GB (RTX 4090 / 3090, etc.)
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest build | Est. memory | Article |
|---|---|---|---|---|
| EleutherAI/olmo3-7b-sdf-sft-clean150 | 7.3B | BF16 (13.6GB) | 16.3GB | Read |
80GB class (A100 / H100)
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest build | Est. memory | Article |
|---|---|---|---|---|
| Edge0/Edge0-35B-A3B-preview | 34.7B | U32 (64.6GB) | 77.5GB | Read |
Over 80GB (datacenter GPU / multi-GPU required)
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest build | Est. memory | Article |
|---|---|---|---|---|
| bartowski/Intern-S2-397B-GGUF | 403.4B | IQ1_S (85.3GB) | 102.4GB | Read |
| dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 | 763.2B | BF16 (902.8GB) | 1083.3GB | Read |
| deepseek-ai/DeepSeek-V4.1-Flash | 763.2B | BF16 (902.8GB) | 1083.3GB | Read |
| nex-agi/Nex-N2.5-Pro | 396.8B | IQ1_S (81.8GB) | 98.2GB | Read |
| IFM/K2-Horizon-MoVA-36B-A4B-GGUF | 37.4B | BF16 (69.8GB) | 83.7GB | Read |