{"id":4626,"date":"2026-09-26T18:16:29","date_gmt":"2026-09-26T09:16:29","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/26\/deepseek-v4-pro-0813-released\/"},"modified":"2026-09-27T20:08:46","modified_gmt":"2026-09-27T11:08:46","slug":"deepseek-v4-pro-0813-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/deepseek-v4-pro-0813-released\/","title":{"rendered":"DeepSeek-V4-Pro-0813 Text Generation Model: ~998GB Memory, GGUF Builds"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Pro-0813\">deepseek-ai\/DeepSeek-V4-Pro-0813<\/a><\/td>\n<\/tr>\n<tr>\n<td>Family guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-deepseek-ai-deepseek-v4-en\/\">DeepSeek-V4 guide (2 articles)<\/a><\/td>\n<\/tr>\n<tr>\n<td>Publisher guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-deepseek-en\/\">DeepSeek: models and licenses<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-08-13<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>mit<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Paper<\/td>\n<td><a href=\"https:\/\/arxiv.org\/abs\/2606.19348\">arXiv:2606.19348<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>DeepSeek has released the official version of its large-scale MoE model, <strong>DeepSeek-V4-Pro-0813<\/strong>, on Hugging Face. According to the model card, it replaces the previous preview version (DeepSeek-V4-Pro (Preview)) with significantly enhanced agent capabilities, bringing noticeable performance improvements especially in production environments. The structure remains the same as the preview version, with the addition of the <strong>DSpark<\/strong> module for speculative decoding.<\/p>\n<p>According to the summary of the technical report for the DeepSeek-V4 series, the series is designed to be &#8220;a highly efficient model capable of routinely handling 1-million-token contexts.&#8221; At a 1-million-token context, it is reported to require 27% of the inference compute (FLOPs) per token and 10% of the KV cache compared to DeepSeek-V3.2.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameter Count: According to the technical report summary, DeepSeek-V4-Pro (Preview) is a MoE model with 1.6T total parameters and 49B active parameters. The official 0813 version uses this exact structure.<\/li>\n<li>Architecture: MoE. To efficiently handle long contexts, it uses a hybrid attention combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA). Residual connections are enhanced with Manifold-Constrained Hyper-Connections (mHC), and training uses the Muon optimizer.<\/li>\n<li>Training Data: Pre-trained on 32T+ tokens, followed by extensive post-training (from the preview version technical report summary).<\/li>\n<li>Context Length: 1 million tokens.<\/li>\n<li>Speculative Decoding: The DSpark module is bundled into the same checkpoint, eliminating the need for a separate draft model.<\/li>\n<li>Reasoning Depth: Three levels can be selected via <code>reasoning_effort<\/code>: <code>low<\/code>, <code>high<\/code>, and <code>max<\/code>.<\/li>\n<li>Weight Precision: FP8 (from Hugging Face tags).<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>ursery The model card provides a comparison table featuring the official version, the concurrent DeepSeek-V4-Flash-0731, both preview versions, and four models from other companies. The figures are measurements by the publishers and have not been independently verified.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th style=\"text-align: left;\">Benchmark<\/th>\n<th style=\"text-align: center;\">DeepSeek-V4-Pro-0813<\/th>\n<th style=\"text-align: center;\">DeepSeek-V4-Flash-0731<\/th>\n<th style=\"text-align: center;\">DeepSeek-V4-Pro (Preview)<\/th>\n<th style=\"text-align: center;\">DeepSeek-V4-Flash (Preview)<\/th>\n<th style=\"text-align: center;\">GLM-5.2<\/th>\n<th style=\"text-align: center;\">Kimi K3<\/th>\n<th style=\"text-align: center;\">Opus-4.8<\/th>\n<th style=\"text-align: center;\">Fable-5 (w\/ fallback)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td style=\"text-align: left;\">HLE (wo \/ w tools)<\/td>\n<td style=\"text-align: center;\">42.7 \/ 60.0<\/td>\n<td style=\"text-align: center;\">37.8 \/ 51.5<\/td>\n<td style=\"text-align: center;\">37.7 \/ 48.2<\/td>\n<td style=\"text-align: center;\">34.8 \/ 45.1<\/td>\n<td style=\"text-align: center;\">40.5 \/ 54.7<\/td>\n<td style=\"text-align: center;\">43.5 \/ 56.0<\/td>\n<td style=\"text-align: center;\">49.8 \/ 57.9<\/td>\n<td style=\"text-align: center;\">53.3 \/ 63.0<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Terminal Bench 2.1<\/td>\n<td style=\"text-align: center;\">87.9<\/td>\n<td style=\"text-align: center;\">82.7<\/td>\n<td style=\"text-align: center;\">72.1<\/td>\n<td style=\"text-align: center;\">61.8<\/td>\n<td style=\"text-align: center;\">81.0<\/td>\n<td style=\"text-align: center;\">88.3<\/td>\n<td style=\"text-align: center;\">85.0<\/td>\n<td style=\"text-align: center;\">88.0<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">NL2Repo<\/td>\n<td style=\"text-align: center;\">61.5<\/td>\n<td style=\"text-align: center;\">54.2<\/td>\n<td style=\"text-align: center;\">38.5<\/td>\n<td style=\"text-align: center;\">39.4<\/td>\n<td style=\"text-align: center;\">48.9<\/td>\n<td style=\"text-align: center;\">&#8211;<\/td>\n<td style=\"text-align: center;\">69.7<\/td>\n<td style=\"text-align: center;\">&#8211;<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Cybergym<\/td>\n<td style=\"text-align: center;\">83.3<\/td>\n<td style=\"text-align: center;\">76.7<\/td>\n<td style=\"text-align: center;\">52.7<\/td>\n<td style=\"text-align: center;\">38.7<\/td>\n<td style=\"text-align: center;\">&#8211;<\/td>\n<td style=\"text-align: center;\">80.0<\/td>\n<td style=\"text-align: center;\">78.3<\/td>\n<td style=\"text-align: center;\">83.1<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">DeepSWE<\/td>\n<td style=\"text-align: center;\">62.7<\/td>\n<td style=\"text-align: center;\">54.4<\/td>\n<td style=\"text-align: center;\">12.8<\/td>\n<td style=\"text-align: center;\">7.3<\/td>\n<td style=\"text-align: center;\">46.2<\/td>\n<td style=\"text-align: center;\">67.5<\/td>\n<td style=\"text-align: center;\">58.0<\/td>\n<td style=\"text-align: center;\">70.0<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Toolathlon-Verified<\/td>\n<td style=\"text-align: center;\">74.1<\/td>\n<td style=\"text-align: center;\">70.3<\/td>\n<td style=\"text-align: center;\">55.9<\/td>\n<td style=\"text-align: center;\">49.7<\/td>\n<td style=\"text-align: center;\">59.9<\/td>\n<td style=\"text-align: center;\">76.5<\/td>\n<td style=\"text-align: center;\">76.2<\/td>\n<td style=\"text-align: center;\">77.9<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">Agents&#8217; Last Exam<\/td>\n<td style=\"text-align: center;\">25.7<\/td>\n<td style=\"text-align: center;\">25.2<\/td>\n<td style=\"text-align: center;\">16.5<\/td>\n<td style=\"text-align: center;\">15.8<\/td>\n<td style=\"text-align: center;\">23.8<\/td>\n<td style=\"text-align: center;\">27.6<\/td>\n<td style=\"text-align: center;\">25.7<\/td>\n<td style=\"text-align: center;\">&#8211;<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">AutomationBench (Public)<\/td>\n<td style=\"text-align: center;\">31.8<\/td>\n<td style=\"text-align: center;\">25.1<\/td>\n<td style=\"text-align: center;\">12.8<\/td>\n<td style=\"text-align: center;\">10.8<\/td>\n<td style=\"text-align: center;\">12.9<\/td>\n<td style=\"text-align: center;\">30.8<\/td>\n<td style=\"text-align: center;\">27.2<\/td>\n<td style=\"text-align: center;\">29.1<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">DSBench-FullStack \u2020<\/td>\n<td style=\"text-align: center;\">71.1<\/td>\n<td style=\"text-align: center;\">68.7<\/td>\n<td style=\"text-align: center;\">41.8<\/td>\n<td style=\"text-align: center;\">37.0<\/td>\n<td style=\"text-align: center;\">61.8<\/td>\n<td style=\"text-align: center;\">73.7<\/td>\n<td style=\"text-align: center;\">71.6<\/td>\n<td style=\"text-align: center;\">77.2<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left;\">DSBench-Hard \u2020<\/td>\n<td style=\"text-align: center;\">67.2<\/td>\n<td style=\"text-align: center;\">59.6<\/td>\n<td style=\"text-align: center;\">31.1<\/td>\n<td style=\"text-align: center;\">25.8<\/td>\n<td style=\"text-align: center;\">54.5<\/td>\n<td style=\"text-align: center;\">63.0<\/td>\n<td style=\"text-align: center;\">71.7<\/td>\n<td style=\"text-align: center;\">68.3<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Gains from the preview version are substantial across all rows.<\/strong> The gap widens particularly in agent-type tasks, with DeepSWE rising from 12.8 to 62.7, AutomationBench (Public) from 12.8 to 31.8, and internal DSBench-Hard from 31.1 to 67.2. The model card&#8217;s description that &#8220;agent capabilities have been significantly enhanced&#8221; is borne out by this table.<\/p>\n<p><strong>It trades blows with top models from other companies depending on the metric.<\/strong> Terminal Bench 2.1, which measures the ability to complete tasks in a terminal, scores 87.9\u2014a close margin within 0.4 points of Kimi K3 (88.3) and Fable-5 (88.0), and outperforming Opus-4.8 (85.0). Cybergym (83.3) and AutomationBench (Public) (31.8) are the highest in the table. On the other hand, HLE, a collection of extremely difficult questions created by domain experts, stops at 42.7 without tools, showing a 7.1 to 10.6 point gap behind Fable-5 (53.3) and Opus-4.8 (49.8) (though with tools it reaches 60.0, ranking second behind Fable-5&#8217;s 63.0). NL2Repo falls short of Opus-4.8 by 8.2 points, DeepSWE falls short of Fable-5 by 7.3 points and Kimi K3 by 4.8 points. Against Kimi K3, it scores lower on 6 out of the 10 numeric metrics where both have values. It is fair to read this as <strong>&#8220;it ranks among the very top in many agent-type tasks, but a gap remains in knowledge-intensive difficult questions and some coding tasks.&#8221;<\/strong><\/p>\n<p>Attention must also be paid to the comparison conditions. According to the notes in the model card, the code agent tasks in the table are measured using DeepSeek&#8217;s own agent infrastructure (DeepSeek Harness in minimal mode) with a reasoning depth of <code>max<\/code>, <code>temperature = 1.0<\/code>, and <code>top_p = 0.95<\/code>. DSBench-FullStack and DSBench-Hard (marked with \u2020) are DeepSeek&#8217;s internal test sets and are not evaluations that third parties can verify under the same conditions. The Fable-5 column includes a &#8220;w\/ fallback&#8221; condition.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>The model card emphasizes agent capabilities, and the tasks showing the largest gains in the table are those requiring &#8220;hands-on execution through to completion,&#8221; such as terminal operations, repository generation, software patching, workflow automation, and tool usage. The efficiency of &#8220;routinely handling 1-million-token contexts&#8221; highlighted in the technical report summary is said to make multi-step workflows and heavy token usage for reasoning (test-time scaling) practical.<\/p>\n<p>Reasoning depth can be selected from three levels of <code>reasoning_effort<\/code>. The model card recommends maximum output lengths of up to <strong>384K tokens<\/strong> for <code>high<\/code> and <code>max<\/code>, presupposing such long outputs for deep reasoning.<\/p>\n<p>On the other hand, it is not suited for running on a reader&#8217;s local hardware. The vLLM example in the model card is for execution on <strong>a single node equipped with four GB300s<\/strong>, and the SGLang example also assumes tensor parallelism of 4. This is a model meant for data-center-grade environments or provider-operated APIs, rather than something to be tested on consumer GPUs.<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<p>Our site previously covered the quantized text generation model <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/nvidia-releases-deepseek-v4-pro-nvfp4\/\">DeepSeek-V4-Pro-0813-nvfp4-DSpark: ~1005GB Memory<\/a> quantized by NVIDIA. That version quantized everything down to NVFP4, including the DSpark draft portion, so it could run on a single checkpoint targeting NVIDIA environments. Today&#8217;s 0813 release is DeepSeek&#8217;s own original weights upon which that was based.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 1650.5B parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>More than 998GB of VRAM (multi-GPU or CPU offload required)<\/td>\n<td>FP8<\/td>\n<td>831.4GB<\/td>\n<td>997.7GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-09-27): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): not registered. &#8220;Not registered&#8221; means the name is absent from that registry today, not that the model cannot run.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Runs in Ollama, LM Studio and llama.cpp via a converted build.<\/strong><\/p>\n<p>The publisher ships safetensors, but <a href=\"https:\/\/huggingface.co\/unsloth\/DeepSeek-V4-Pro-0813-GGUF\">unsloth\/DeepSeek-V4-Pro-0813-GGUF<\/a> provides a GGUF build you can use.<\/p>\n<p><strong>License \u2014 <code>mit<\/code> (Commercial use allowed):<\/strong> Permits commercial use, modification and redistribution, provided the copyright notice and license text are retained.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<h2>How to Get It<\/h2>\n<ul>\n<li>Distribution Format: Safetensors weights (FP8) are available at Hugging Face&#8217;s <code>deepseek-ai\/DeepSeek-V4-Pro-0813<\/code>. Downloading does not require agreeing to a terms of use agreement.<\/li>\n<li>Example Download Command: <code>huggingface-cli download deepseek-ai\/DeepSeek-V4-Pro-0813<\/code><\/li>\n<li>Supported Engines: The model card guides users toward vLLM and SGLang. Both can enable speculative decoding via DSpark. In vLLM, add <code>--speculative-config '{\"method\":\"dspark\",\"num_speculative_tokens\":7,\"draft_sample_method\":\"greedy\"}'<\/code> to the launch command, and in SGLang, specify <code>--speculative-algorithm DSPARK<\/code>. SGLang reads the draft weights from the same checkpoint, so <code>--speculative-draft-model-path<\/code> should not be specified.<\/li>\n<li>Chat Template: This release <strong>does not include<\/strong> the Jinja-format chat template used by many tools. Instead, the repository&#8217;s <code>encoding<\/code> folder contains Python scripts and test cases to convert OpenAI-compatible messages into input strings for the model and parse the outputs. Users must check individual engine compatibility before proceeding.<\/li>\n<li>Running Locally: The repository&#8217;s <code>inference<\/code> folder contains procedures for weight conversion and interactive demos. The model card recommends <code>temperature = 1.0<\/code>, with <code>top_p<\/code> set to 0.95 for agent use cases and 1.0 otherwise.<\/li>\n<li>GGUF versions or similar formats for running on consumer hardware are not mentioned in the model card.<\/li>\n<\/ul>\n<p><!-- lmw:variants --><\/p>\n<h2>Quantized and Converted Variants<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Added<\/th>\n<th>Publisher<\/th>\n<th>Format<\/th>\n<th>Repository<\/th>\n<th>Smallest VRAM tier (build, est. memory)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-09-26<\/td>\n<td>unsloth<\/td>\n<td>GGUF<\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/DeepSeek-V4-Pro-0813-GGUF\">unsloth\/DeepSeek-V4-Pro-0813-GGUF<\/a><\/td>\n<td>Q4_K_XL 949.6GB (does not fit a single consumer GPU)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>unsloth<\/td>\n<td>FP8<\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/DeepSeek-V4-Pro-0813\">unsloth\/DeepSeek-V4-Pro-0813<\/a><\/td>\n<td>FP8 997.7GB (does not fit a single consumer GPU)<\/td>\n<\/tr>\n<tr>\n<td>2026-09-26<\/td>\n<td>DevQuasar<\/td>\n<td>GGUF<\/td>\n<td><a href=\"https:\/\/huggingface.co\/DevQuasar\/deepseek-ai.DeepSeek-V4-Pro-0813-GGUF\">DevQuasar\/deepseek-ai.DeepSeek-V4-Pro-0813-GGUF<\/a><\/td>\n<td>Q2_K 636.4GB (does not fit a single consumer GPU)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>File sizes of each build:<\/p>\n<ul>\n<li>Available builds in unsloth\/DeepSeek-V4-Pro-0813-GGUF: Q4_K_XL 791.3GB \/ Q8_K_XL 813.5GB<\/li>\n<li>Available builds in unsloth\/DeepSeek-V4-Pro-0813: FP8 831.4GB<\/li>\n<li>Available builds in DevQuasar\/deepseek-ai.DeepSeek-V4-Pro-0813-GGUF: Q2_K 530.3GB \/ Q3_K_M 697.0GB \/ Q4_K_M 885.6GB<\/li>\n<\/ul>\n<p>In addition, 8 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model&#8217;s publisher or established quantization maintainers.<\/p>\n<p><em>This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:variants --><\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/nvidia-releases-deepseek-v4-pro-nvfp4\/\">DeepSeek-V4-Pro-0813-nvfp4-DSpark: ~1005GB Memory<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/deepseek-v4-1-flash-released\/\">DeepSeek-V4.1-Flash 552B Multimodal MoE Model: ~570GB Memory<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-32b-nvfp4-released\/\">K2-Horizon-32B-NVFP4 Long-Context Reasoning Model: 32GB+ VRAM<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-375b-a23b-nvfp4-released-2\/\">K2-Horizon-375B-A23B-NVFP4 Text Generation Model: ~257GB Memory<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model does not fit even in 80GB; its smallest build needs about 998GB) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference, including models over 80GB<\/a><\/li>\n<li><strong>Explore the same model family<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-deepseek-ai-deepseek-v4-en\/\">DeepSeek-V4 family overview (2 articles, 3 converted builds)<\/a><\/li>\n<li><strong>Engines that run this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-llama-cpp-en\/\">llama.cpp<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-ollama-en\/\">Ollama<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/engine-vllm-en\/\">vLLM<\/a><\/li>\n<li><strong>What FP8 mean and where to get this model<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-gguf-en\/\">GGUF format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-fp8-en\/\">FP8 format guide and models<\/a> \/ <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">Quantization and model-format glossary<\/a><\/li>\n<li><strong>Learn about the publisher<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/publisher-deepseek-en\/\">DeepSeek: models, licenses and articles<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Pro-0813\">https:\/\/huggingface.co\/deepseek-ai\/DeepSeek-V4-Pro-0813<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-09-26: Added converted builds to \u201cQuantized and Converted Variants\u201d: unsloth\/DeepSeek-V4-Pro-0813-GGUF, unsloth\/DeepSeek-V4-Pro-0813, DevQuasar\/deepseek-ai.DeepSeek-V4-Pro-0813-GGUF<\/li>\n<li>2026-09-26: Changed the title to show what the article covers (VRAM requirements, file list, etc.).<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Discover DeepSeek-V4-Pro-0813, a 1.6T MoE model on Hugging Face featuring enhanced agent capabilities, 1M context, and DSpark speculative decoding.<\/p>\n","protected":false},"author":1,"featured_media":4630,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[604,811,2367,165,169,1547,592,1555],"class_list":["post-4626","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-deepseek-en","tag-deepseek-v4-pro-en","tag-deepseek-v4-pro-0813-en","tag-moe-en","tag-sglang-en","tag-verified","tag-vllm-en","tag--en"],"lang":"en","translations":{"en":4626,"ja":4624},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/4626","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=4626"}],"version-history":[{"count":9,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/4626\/revisions"}],"predecessor-version":[{"id":5894,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/4626\/revisions\/5894"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/4630"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=4626"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=4626"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=4626"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}