Local Models That Run in 24GB of VRAM (RTX 4090 / 3090, etc.)
Every model covered by Local Model Watch that fits in 24GB of GPU memory, with the best build that fits. The “Best build for 24GB" column is the largest (highest-quality) quantization or precision whose estimated memory stays within 24GB — often better than the smallest build listed in the VRAM quick reference. Estimates are computed by this site from actual file sizes plus runtime overhead (method); long contexts need more.
28 models listed.
Everything that fits your GPU: 8GB / 12GB / 16GB / 24GB
Text Generation (Chat and Code)
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 24GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| Edge0/Edge0-35B-A3B-preview | 36.0B | MLX 4bit (18.2GB) | 21.8GB | 24GB | vLLM, MLX | Read |
| nex-agi/Nex-N2.5-mini | 35.1B | Q4_K_S (19.5GB) | 23.4GB | 16GB | Ollama / LM Studio, vLLM, MLX | Read |
| prism-ml/Ternary-Bonsai-2-27B-gguf | 27.8B | Q2_0 (6.7GB) | 8.1GB | 8GB | Ollama / LM Studio, MLX | Read |
| bartowski/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | 26.5B | Q5_K_M (19.4GB) | 23.2GB | 12GB | Ollama / LM Studio | Read |
| EleutherAI/olmo3-7b-sdf-sft-clean150 | 7.3B | BF16 (13.6GB) | 16.3GB | 24GB | vLLM | Read |
| togethercomputer/Tev1-4B-experimental | 4.7B | BF16 (8.7GB) | 10.4GB | 12GB | vLLM | Read |
| tencent/Simple-Attention-Sparsification | 4.0B | BF16 (7.5GB) | 9.0GB | 12GB | – | Read |
| openbmb/MiniCPM5-2B | 2.5B | F16 (4.7GB) | 5.6GB | 4GB | Ollama / LM Studio, vLLM, MLX | Read |
| harshatheg/Qwen-2.5-1B-RLCD | 1.5B | BF16 (2.9GB) | 3.5GB | 4GB | MLX | Read |
| pfnet/plamo-3-610m-fin-instruct | 890M | BF16 (1.7GB) | 2.0GB | 4GB | vLLM | Read |
| togethercomputer/Tev1-0.8B-experimental | 873M | F32 (1.6GB) | 2.0GB | 4GB | vLLM | Read |
| Cactus-Compute/needle3 | – | Original precision (0.2GB) | 0.3GB | 4GB | – | Read |
Vision-Language and Multimodal Models
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 24GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| bartowski/vectionlabs_Salience-27B-R6-GGUF | 27.8B | Q5_K_M (19.5GB) | 23.4GB | 12GB | Ollama / LM Studio | Read |
| DavidAU/Qwen3.8-27B-TWIN-TURBO-Fable-Cold-Fusion-709-L-Uncensored-NM-DAU-NEO-MTP-GGUF | 27.8B | Q5_K_M (18.2GB) | 21.8GB | 12GB | Ollama / LM Studio | Read |
| agentionai/Signal-3.8-27B-GGUF | 27.8B | Q5_K_M (18.2GB) | 21.8GB | 16GB | Ollama / LM Studio | Read |
| Jackrong/Qwopus3.8-27B-Flash-GGUF | 27.8B | Q5_K_M (18.2GB) | 21.8GB | 16GB | Ollama / LM Studio | Read (Japanese) |
| bartowski/TheDrummer_Orion-26B-A4B-v1.1-GGUF | 25.8B | Q5_K_M (18.6GB) | 22.4GB | 12GB | Ollama / LM Studio | Read |
| ggml-org/MiMo-V2.6-Distill-Qwen-9B-GGUF | 9.4B | Q8_0 (8.9GB) | 10.6GB | 12GB | Ollama / LM Studio | Read |
| LiquidAI/LFM2.5-VL-3B-DSpark | 279M | F16 (0.5GB) | 0.6GB | 4GB | – | Read |
| inclusionAI/Realtime-Venus | – | Original precision(Realtime-Venus-Audio) (17.5GB) | 20.9GB | 24GB | – | Read |
Image Generation and Editing
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 24GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| abenzerps/Qwen-Image-2.1-Uncensored-GGUF | 7.1B | BF16 (13.3GB) | 15.9GB | 8GB | – | Read |
Video Generation
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 24GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| WarmBloodAban/Minimax-h3_Singularity | 33.1B | INT8 (19.5GB) | 23.4GB | 16GB | – | Read (Japanese) |
Audio: Speech Synthesis, Music and Speech Recognition
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 24GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| Edge0/Audio8-ASR-Infinite | 4.1B | BF16 (7.6GB) | 9.1GB | 12GB | – | Read |
| m-a-p/YuE2-3B | 3.6B | BF16 (6.8GB) | 8.1GB | 12GB | – | Read |
| phasefield-audio/Irodori-TTS-v4.1-Anime | 766M | F32 (2.9GB) | 3.4GB | 4GB | – | Read |
Task not declared
→ Scroll horizontally to see all columns
| Model | Parameters | Best build for 24GB | Est. memory | Smallest tier | Runs on | Article |
|---|---|---|---|---|---|---|
| tencent/WeVisDoc-4B | 4.4B | F16 (8.2GB) | 9.9GB | 4GB | – | Read |
| Comfy-Org/YuE2 | 3.6B | BF16 (6.8GB) | 8.1GB | 12GB | – | Read |
| Efficient-Large-Model/H3-to-LTX-Latent-Adapter | 195M | BF16 (0.4GB) | 0.4GB | 4GB | – | Read |