Local Models by VRAM: Quick Reference

September 19, 2026

Open-weight models covered by Local Model Watch, grouped by the smallest amount of GPU memory (VRAM) they are estimated to run in. Each figure is computed by this site from the actual size of the distributed files plus a runtime overhead — not quoted from model cards. See our editorial policy for the method.

Models listed under a tier do not fit in the smaller tiers. If you have a 24GB GPU, the 8GB, 12GB, 16GB and 24GB tables all apply to you. Models that don’t fit even the largest tier (80GB) are listed separately at the bottom, since that too is something readers need to know.

Last updated 2026-09-19 (JST). 28 models listed.

4GB (laptop iGPU / phone class)

→ Scroll horizontally to see all columns

Model Parameters Smallest build Est. memory Article
Cactus Compute、8〜29MBの自動化モデル「Needle 3」公開 そのままの精度 (0.2GB) 0.3GB Read
Lightricks/LTX-2.5-22b-IC-LoRA-Ingredients そのままの精度 (1.2GB) 1.5GB Read
openbmb/MiniCPM5-2B 2.5B Q8_0 (2.5GB) 3.0GB Read

8GB (RTX 4060 / 3060 Ti, etc.)

→ Scroll horizontally to see all columns

Model Parameters Smallest build Est. memory Article
prism-ml/Ternary-Bonsai-2-27B-gguf 27.8B Q1_0 (5.5GB) 6.6GB Read
harshatheg/Qwen-2.5-1B-RLCD 1.5B BF16 (2.9GB) 3.5GB Read
EleutherAI/bergson-wikitext-gpt2-leaderboard そのままの精度 (6.5GB) 7.8GB Read
bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF 25.8B F16 (1.1GB) 1.3GB Read
openbmb/MiniCPM5-2B-GGUF F16 (4.7GB) 5.6GB Read

12GB (RTX 4070 / 3060 12GB, etc.)

→ Scroll horizontally to see all columns

Model Parameters Smallest build Est. memory Article
tencent/WeVisDoc-4B 4.4B BF16 (9.0GB) 10.8GB Read
prism-ml/Ternary-Bonsai-2-27B-mlx-2bit 27.4B MLX 2bit (8.0GB) 9.6GB Read
prism-ml/Ternary-Bonsai-2-27B-gguf-dev 27.8B Q2_0 (7.1GB) 8.5GB Read
Mothersuperior/yue2-mothersuperior-realaudio-tokenizer-v4 3.6B BF16 (6.8GB) 8.1GB Read
bartowski/vectionlabs_Salience-27B-R6-GGUF 27.8B IQ2_M (9.8GB) 11.8GB Read
tencent/Simple-Attention-Sparsification 4.0B BF16 (7.5GB) 9.0GB Read
Comfy-Org/YuE2 3.6B BF16 (6.8GB) 8.1GB Read
bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF 26.5B IQ2_S (10.0GB) 12.0GB Read
m-a-p/YuE2-3B 3.6B BF16 (6.8GB) 8.1GB Read

16GB (RTX 5060 Ti 16GB / 4060 Ti 16GB, etc.)

→ Scroll horizontally to see all columns

Model Parameters Smallest build Est. memory Article
bartowski/nex-agi_Nex-N2.5-mini-GGUF 35.1B Q2_K (12.9GB) 15.5GB Read
agentionai/Signal-3.8-27B-GGUF 27.8B IQ4_XS (13.3GB) 15.9GB Read
nex-agi/Nex-N2.5-mini 35.1B Q2_K (12.9GB) 15.5GB Read
Jackrong/Qwopus3.8-27B-Flash-GGUF 27.8B Q3_K_M (12.6GB) 15.1GB Read

24GB (RTX 4090 / 3090, etc.)

→ Scroll horizontally to see all columns

Model Parameters Smallest build Est. memory Article
EleutherAI/olmo3-7b-sdf-sft-clean150 7.3B BF16 (13.6GB) 16.3GB Read

80GB class (A100 / H100)

→ Scroll horizontally to see all columns

Model Parameters Smallest build Est. memory Article
Edge0/Edge0-35B-A3B-preview 34.7B U32 (64.6GB) 77.5GB Read

Over 80GB (datacenter GPU / multi-GPU required)

→ Scroll horizontally to see all columns

Model Parameters Smallest build Est. memory Article
bartowski/Intern-S2-397B-GGUF 403.4B IQ1_S (85.3GB) 102.4GB Read
dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8 763.2B BF16 (902.8GB) 1083.3GB Read
deepseek-ai/DeepSeek-V4.1-Flash 763.2B BF16 (902.8GB) 1083.3GB Read
nex-agi/Nex-N2.5-Pro 396.8B IQ1_S (81.8GB) 98.2GB Read
IFM/K2-Horizon-MoVA-36B-A4B-GGUF 37.4B BF16 (69.8GB) 83.7GB Read