{"id":2968,"date":"2026-09-23T08:13:59","date_gmt":"2026-09-22T23:13:59","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/23\/ifm-k2-horizon-375b-a23b-nvfp4-released-2\/"},"modified":"2026-09-23T08:13:59","modified_gmt":"2026-09-22T23:13:59","slug":"ifm-k2-horizon-375b-a23b-nvfp4-released-2","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-375b-a23b-nvfp4-released-2\/","title":{"rendered":"IFM Releases K2-Horizon-375B-A23B-NVFP4 Quantized Model"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/IFM\/K2-Horizon-375B-A23B-NVFP4\">IFM\/K2-Horizon-375B-A23B-NVFP4<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-22<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>safetensors<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>IFM has released &#8220;K2-Horizon-375B-A23B-NVFP4&#8221;, which is the NVFP4 quantized version of its flagship model &#8220;K2-Horizon-375B-A23B&#8221; built on the MoE (Mixture-of-Experts) architecture. This model reduces memory usage and speeds up inference by quantizing the linear layers (weights and activations) of the routed experts into the NVFP4 format. Meanwhile, the attention mechanism, shared experts, router, the first three dense layers, and lm_head maintain BF16 precision. This model is intended for use on NVIDIA Blackwell generation (B series) and newer GPUs that natively support NVFP4.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Total parameters: 375B<\/li>\n<li>Activated parameters: 23B<\/li>\n<li>Architecture: MoE<\/li>\n<li>Context length: 512K (524,288 tokens)<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<h3>NVFP4 vs. BF16<\/h3>\n<p>Regarding the impact of quantization on performance, evaluation results in the model card indicate that the NVFP4 version experiences a very slight performance decrease compared to the original BF16 version. Specifically, the average score drops from 91.7 to 91.2, and GSM8K in the mathematics domain slightly decreases from 96.06 to 95.53. However, this performance drop is extremely minor, and the design compensates for this slight accuracy difference with the benefits of reduced memory footprint and improved inference speed on NVFP4-supported hardware.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>K2-Horizon-375B-A23B<\/th>\n<th>IFEval (Prompt)<\/th>\n<th>GSM8K<\/th>\n<th>MBPP<\/th>\n<th>MMLU-Pro<\/th>\n<th>GPQA-Diamond<\/th>\n<th>BBH (3-shot)<\/th>\n<th>AIME 26 (avg @ 32)<\/th>\n<th>Average<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>BF16<\/td>\n<td>90.02<\/td>\n<td>96.06<\/td>\n<td>97.00<\/td>\n<td>84.22<\/td>\n<td>85.80<\/td>\n<td>94.73<\/td>\n<td>94.38<\/td>\n<td>91.7<\/td>\n<\/tr>\n<tr>\n<td>NVFP4<\/td>\n<td>88.72<\/td>\n<td>95.53<\/td>\n<td>96.60<\/td>\n<td>83.98<\/td>\n<td>85.45<\/td>\n<td>94.26<\/td>\n<td>93.65<\/td>\n<td>91.2<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h3>Benchmark Results<\/h3>\n<p>Based on measurement results published by the creators for the original model &#8220;K2-Horizon-375B-A23B&#8221;, a table focusing on major comparison targets is shown below.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th><\/th>\n<th>K2-Horizon-375B-A23B<\/th>\n<th>Open-weight models \/ Nemotron 3 Ultra<\/th>\n<th>Open-weight models \/ Inkling (xhigh)<\/th>\n<th>Open-weight models \/ MiniMax-M3<\/th>\n<th>Closed models \/ Claude Sonnet5 (max)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td># Params<\/td>\n<td>375B<\/td>\n<td>550B<\/td>\n<td>975B<\/td>\n<td>428B<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<tr>\n<td># Activated params<\/td>\n<td>23B<\/td>\n<td>55B<\/td>\n<td>41B<\/td>\n<td>23B<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<tr>\n<td>Architecture<\/td>\n<td>MoE<\/td>\n<td>MoE<\/td>\n<td>MoE<\/td>\n<td>MoE<\/td>\n<td>Closed<\/td>\n<\/tr>\n<tr>\n<td>Agents<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>GDPVal-AA Real-world professional tasks (Elo)<\/td>\n<td>1,441<\/td>\n<td>1,162<\/td>\n<td>1,234<\/td>\n<td>1,380<\/td>\n<td>1,584<\/td>\n<\/tr>\n<tr>\n<td>tau3-Banking Agentic tool use<\/td>\n<td>34.0<\/td>\n<td>14.2<\/td>\n<td>29.1<\/td>\n<td>15.3<\/td>\n<td>37.3<\/td>\n<\/tr>\n<tr>\n<td>Coding<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Terminal-Bench 2.1 Agentic terminal use<\/td>\n<td>70.2<\/td>\n<td>53.9<\/td>\n<td>55.1<\/td>\n<td>65.2<\/td>\n<td>80.5<\/td>\n<\/tr>\n<tr>\n<td>SciCode Scientific coding<\/td>\n<td>42.7<\/td>\n<td>39.9<\/td>\n<td>46.1<\/td>\n<td>45.4<\/td>\n<td>53.6<\/td>\n<\/tr>\n<tr>\n<td>Scientific Reasoning<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Humanity&#8217;s Last Exam (without tools) Expert-level reasoning<\/td>\n<td>32.0<\/td>\n<td>28.4<\/td>\n<td>31.9<\/td>\n<td>39.0<\/td>\n<td>41.3<\/td>\n<\/tr>\n<tr>\n<td>GPQA Diamond Graduate-level science QA<\/td>\n<td>87.3<\/td>\n<td>86.7<\/td>\n<td>87.2<\/td>\n<td>92.9<\/td>\n<td>91.1<\/td>\n<\/tr>\n<tr>\n<td>CritPt Frontier physics reasoning<\/td>\n<td>8.6<\/td>\n<td>3.1<\/td>\n<td>5.4<\/td>\n<td>3.7<\/td>\n<td>16.9<\/td>\n<\/tr>\n<tr>\n<td>General<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>AA-LCR Long-context reasoning<\/td>\n<td>76.0<\/td>\n<td>71.0<\/td>\n<td>73.3<\/td>\n<td>80.3<\/td>\n<td>77.0<\/td>\n<\/tr>\n<tr>\n<td>AA-Omniscience Accuracy Factual accuracy<\/td>\n<td>23.0<\/td>\n<td>23.0<\/td>\n<td>42.0<\/td>\n<td>17.0<\/td>\n<td>40.0<\/td>\n<\/tr>\n<tr>\n<td>AA-Omniscience Non-Hallucination Non-hallucination rate<\/td>\n<td>74.7<\/td>\n<td>70.0<\/td>\n<td>32.0<\/td>\n<td>82.0<\/td>\n<td>61.0<\/td>\n<\/tr>\n<tr>\n<td>Agentic Evaluations<\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Toolathlon Verified Agentic tool use<\/td>\n<td>65.3<\/td>\n<td>34.3<\/td>\n<td>45.5<\/td>\n<td>53.7<\/td>\n<td>71.6<\/td>\n<\/tr>\n<tr>\n<td>Automation Bench Public Workflow automation<\/td>\n<td>25.3<\/td>\n<td>8.0<\/td>\n<td>12.8<\/td>\n<td>20.5<\/td>\n<td>34.7<\/td>\n<\/tr>\n<tr>\n<td>Apex-Agents (pass@1) Long-horizon professional workflows<\/td>\n<td>24.8<\/td>\n<td>9.0<\/td>\n<td>19.0<\/td>\n<td>23.8<\/td>\n<td>31.7<\/td>\n<\/tr>\n<tr>\n<td>MCPMark MCP tool use<\/td>\n<td>67.7<\/td>\n<td>45.7<\/td>\n<td>51.2<\/td>\n<td>48.8<\/td>\n<td>65.3<\/td>\n<\/tr>\n<tr>\n<td>BrowseComp Deep web research<\/td>\n<td>72.8<\/td>\n<td>44.4<\/td>\n<td>77.1<\/td>\n<td>83.5<\/td>\n<td>84.7<\/td>\n<\/tr>\n<tr>\n<td>WildClawBench In-the-wild agentic tasks<\/td>\n<td>50.9<\/td>\n<td>34.2<\/td>\n<td>52.3<\/td>\n<td>56.4<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<tr>\n<td>SWE-Atlas-QnA Repo-level code Q&amp;A (strict)<\/td>\n<td>48.4<\/td>\n<td>&#8212;<\/td>\n<td>25.5<\/td>\n<td>42.3<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<tr>\n<td>SWE Bench Pro Software engineering (strict)<\/td>\n<td>42.6<\/td>\n<td>38.7<\/td>\n<td>43.1<\/td>\n<td>43.8<\/td>\n<td>&#8212;<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>This table shows that this model achieves a very high standard in agent performance. In particular, it records 1,441 in GDPVal-AA (Elo) under the &#8220;Agents&#8221; category, making it extremely powerful among open-weight MoE models. It scores 70.2 in Terminal-Bench 2.1 in the &#8220;Coding&#8221; field, demonstrating high capability in agentic tasks involving terminal operations. Additionally, it records a high score of 87.3 in GPQA Diamond for &#8220;Scientific Reasoning&#8221;.<\/p>\n<p>On the other hand, it falls behind closed models in certain metrics. For example, in Toolathlon, this model scores 65.3 compared to 71.6 for Claude Sonnet5 (max), suggesting that there is still room for improvement in agent tool utilization capabilities when compared to top-tier closed models. However, it displays performance that surpasses open-weight MoE models up to 2.6 times its size, making it an exceptionally efficient model.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Based on the characteristics of the original model, this model has specialized capabilities in advanced agent task execution, coding, and long-form reading comprehension. Specifically, it is expected to be utilized in the following applications:<\/p>\n<ul>\n<li><strong>Agent and Tool Use<\/strong>: Demonstrates high performance in terminal operations, agent tasks involving complex workflows, and tool utilization using MCP (Model Context Protocol).<\/li>\n<li><strong>Coding<\/strong>: Suitable for agentic coding tasks involving terminal operations, repository-level code Q&amp;A, and software engineering tasks.<\/li>\n<li><strong>Scientific and Advanced Reasoning<\/strong>: Capable of handling questions requiring doctoral-level scientific knowledge, expert-level reasoning tasks, and mathematical reasoning.<\/li>\n<li><strong>Long Context Processing<\/strong>: Natively supports an extremely long context window of 512K (524,288 tokens), enabling tasks that handle massive amounts of information at once.<\/li>\n<\/ul>\n<p>Additionally, the model is recommended for use with specific sampling parameters such as <code>reasoning_effort=\"high\"<\/code>, temperature of 1.0, and top_p of 0.95. The reasoning process (thinking) is structured to be output as <code>reasoning_content<\/code>, and the final answer as <code>content<\/code>.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 218.2B parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>More than 257GB of VRAM (multi-GPU or CPU offload required)<\/td>\n<td>NVFP4<\/td>\n<td>213.8GB<\/td>\n<td>256.6GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-09-22): llama.cpp: not registered, vLLM: registered, MLX (mlx-lm): not registered. &#8220;Not registered&#8221; means the name is absent from that registry today, not that the model cannot run.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:peers --><\/p>\n<h2>Recent Models in the Same Size Class<\/h2>\n<p><em>Models with <\/em><em>over 40B<\/em><em> parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site&#8217;s estimates; licenses are as stated on the model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>nvidia\/DeepSeek-V4-Pro-0813-nvfp4-DSpark<\/td>\n<td>1650.5B<\/td>\n<td>\u2014<\/td>\n<td>mit<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/nvidia-releases-deepseek-v4-pro-nvfp4\/\">NVIDIA Releases NVFP4 Quantized DeepSeek-V4-Pro<\/a> (2026-09-10)<\/td>\n<\/tr>\n<tr>\n<td>nex-agi\/Nex-N2.5-Pro<\/td>\n<td>396.8B<\/td>\n<td>\u2014<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-agi-announces-nex-n25-pro-agent-model\/\">Nex-AGI Releases Agent Model Nex-N2.5-Pro<\/a> (2026-09-09)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:peers --><\/p>\n<h2>How to Get It<\/h2>\n<ul>\n<li><strong>Distribution format<\/strong>: safetensors<\/li>\n<li><strong>Supported engines<\/strong>: vLLM, SGLang, Transformers<\/li>\n<\/ul>\n<p>Verified recipes for SGLang and serving methods for vLLM have been published. Note that this model is released under the Apache 2.0 license.<\/p>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/nvidia-releases-deepseek-v4-pro-nvfp4\/\">NVIDIA Releases NVFP4 Quantized DeepSeek-V4-Pro<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">OpenBMB Releases MiniCPM5-2B: A SOTA 2B On-Device Model<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/IFM\/K2-Horizon-375B-A23B-NVFP4\">https:\/\/huggingface.co\/IFM\/K2-Horizon-375B-A23B-NVFP4<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Explore the NVFP4 quantized version of IFM&#8217;s K2-Horizon-375B-A23B MoE model, designed for NVIDIA Blackwell GPUs with reduced memory usage.<\/p>\n","protected":false},"author":1,"featured_media":2967,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[2051,2053,165,813,169,1547,592,1555],"class_list":["post-2968","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-ifm-en","tag-k2-horizon-en","tag-moe-en","tag-nvfp4-en","tag-sglang-en","tag-verified","tag-vllm-en","tag--en"],"lang":"en","translations":{"en":2968,"ja":2966},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2968","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=2968"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2968\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/2967"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=2968"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=2968"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=2968"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}