{"id":3543,"date":"2026-09-25T11:21:29","date_gmt":"2026-09-25T02:21:29","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/vram-guide-8gb-en\/"},"modified":"2026-09-26T04:11:18","modified_gmt":"2026-09-25T19:11:18","slug":"vram-guide-8gb-en","status":"publish","type":"page","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-8gb-en\/","title":{"rendered":"Local Models That Run in 8GB of VRAM (RTX 4060 \/ 3060 Ti, etc.)"},"content":{"rendered":"<p>Every model covered by Local Model Watch that fits in <strong>8GB<\/strong> of GPU memory, with the best build that fits. The &#8220;Best build for 8GB&#8221; column is the largest (highest-quality) quantization or precision whose estimated memory stays within 8GB \u2014 often better than the smallest build listed in the <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. Estimates are computed by this site from actual file sizes plus runtime overhead (<a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/editorial-policy-en\/\">method<\/a>); long contexts need more.<\/p>\n<p><em>11 models listed.<\/em><\/p>\n<p>Everything that fits your GPU: <strong>8GB<\/strong> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-12gb-en\/\">12GB<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-16gb-en\/\">16GB<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-24gb-en\/\">24GB<\/a><\/p>\n<h2>Text Generation (Chat and Code)<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Best build for 8GB<\/th>\n<th>Est. memory<\/th>\n<th>Smallest tier<\/th>\n<th>Runs on<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>prism-ml\/Ternary-Bonsai-2-27B-gguf<\/td>\n<td>27.8B<\/td>\n<td>Q1_0 (5.5GB)<\/td>\n<td>6.6GB<\/td>\n<td>8GB<\/td>\n<td>Ollama \/ LM Studio, MLX<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf\/\">Read<\/a><\/td>\n<\/tr>\n<tr>\n<td>openbmb\/MiniCPM5-2B<\/td>\n<td>2.5B<\/td>\n<td>F16 (4.7GB)<\/td>\n<td>5.6GB<\/td>\n<td>4GB<\/td>\n<td>Ollama \/ LM Studio, vLLM, MLX<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">Read<\/a><\/td>\n<\/tr>\n<tr>\n<td>harshatheg\/Qwen-2.5-1B-RLCD<\/td>\n<td>1.5B<\/td>\n<td>BF16 (2.9GB)<\/td>\n<td>3.5GB<\/td>\n<td>4GB<\/td>\n<td>MLX<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/17\/qwen-25-1b-rlcd-mlx-constrained-decoding\/\">Read<\/a><\/td>\n<\/tr>\n<tr>\n<td>pfnet\/plamo-3-610m-fin-instruct<\/td>\n<td>890M<\/td>\n<td>BF16 (1.7GB)<\/td>\n<td>2.0GB<\/td>\n<td>4GB<\/td>\n<td>vLLM<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/plamo-3-610m-fin-instruct-2\/\">Read<\/a><\/td>\n<\/tr>\n<tr>\n<td>togethercomputer\/Tev1-0.8B-experimental<\/td>\n<td>873M<\/td>\n<td>F32 (1.6GB)<\/td>\n<td>2.0GB<\/td>\n<td>4GB<\/td>\n<td>vLLM<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/tev1-08b-experimental-2\/\">Read<\/a><\/td>\n<\/tr>\n<tr>\n<td>Cactus-Compute\/needle3<\/td>\n<td>&#8211;<\/td>\n<td>Original precision (0.2GB)<\/td>\n<td>0.3GB<\/td>\n<td>4GB<\/td>\n<td>&#8211;<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/19\/cactus-compute-needle-3\/\">Read<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Vision-Language and Multimodal Models<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Best build for 8GB<\/th>\n<th>Est. memory<\/th>\n<th>Smallest tier<\/th>\n<th>Runs on<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>LiquidAI\/LFM2.5-VL-3B-DSpark<\/td>\n<td>279M<\/td>\n<td>F16 (0.5GB)<\/td>\n<td>0.6GB<\/td>\n<td>4GB<\/td>\n<td>&#8211;<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/liquidai-lfm2-5-vl-dspark-released\/\">Read<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Image Generation and Editing<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Best build for 8GB<\/th>\n<th>Est. memory<\/th>\n<th>Smallest tier<\/th>\n<th>Runs on<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>abenzerps\/Qwen-Image-2.1-Uncensored-GGUF<\/td>\n<td>7.1B<\/td>\n<td>Q6_K (5.5GB)<\/td>\n<td>6.6GB<\/td>\n<td>8GB<\/td>\n<td>&#8211;<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/21\/uncensored-qwen-image-2-1-gguf-released\/\">Read<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Audio: Speech Synthesis, Music and Speech Recognition<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Best build for 8GB<\/th>\n<th>Est. memory<\/th>\n<th>Smallest tier<\/th>\n<th>Runs on<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>phasefield-audio\/Irodori-TTS-v4.1-Anime<\/td>\n<td>766M<\/td>\n<td>F32 (2.9GB)<\/td>\n<td>3.4GB<\/td>\n<td>4GB<\/td>\n<td>&#8211;<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/07\/irodori-tts-v41-anime-released\/\">Read<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Task not declared<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Best build for 8GB<\/th>\n<th>Est. memory<\/th>\n<th>Smallest tier<\/th>\n<th>Runs on<\/th>\n<th>Article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>tencent\/WeVisDoc-4B<\/td>\n<td>4.4B<\/td>\n<td>Q8_0 (4.4GB)<\/td>\n<td>5.2GB<\/td>\n<td>4GB<\/td>\n<td>&#8211;<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/tencent-releases-wevisdoc-document-parsing-models\/\">Read<\/a><\/td>\n<\/tr>\n<tr>\n<td>Efficient-Large-Model\/H3-to-LTX-Latent-Adapter<\/td>\n<td>195M<\/td>\n<td>BF16 (0.4GB)<\/td>\n<td>0.4GB<\/td>\n<td>4GB<\/td>\n<td>&#8211;<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/h3-to-ltx-latent-adapter-released\/\">Read<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>Every model covered by Local Model Watch that fits in 8GB of GPU memory, with the best build that fits. The &#8220;Best build for 8GB&#8221; column is the largest (highest-quality) quantization or precision whose estimated memory stays within 8GB \u2014 often better than the smallest build listed in the VRAM quick reference. Estimates are computed [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-3543","page","type-page","status-publish","hentry"],"lang":"en","translations":{"en":3543,"ja":3542},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/3543","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=3543"}],"version-history":[{"count":2,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/3543\/revisions"}],"predecessor-version":[{"id":4452,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/3543\/revisions\/4452"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=3543"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}