{"id":2992,"date":"2026-09-23T16:14:50","date_gmt":"2026-09-23T07:14:50","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/23\/ifm-k2-horizon-32b-nvfp4-released\/"},"modified":"2026-09-23T16:14:50","modified_gmt":"2026-09-23T07:14:50","slug":"ifm-k2-horizon-32b-nvfp4-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-32b-nvfp4-released\/","title":{"rendered":"IFM Releases K2-Horizon-32B-NVFP4 with Native Blackwell Support"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/IFM\/K2-Horizon-32B-NVFP4\">IFM\/K2-Horizon-32B-NVFP4<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-22<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>safetensors<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>IFM has released &#8220;IFM\/K2-Horizon-32B-NVFP4&#8221;, a model obtained by quantizing the Stage1 checkpoint of the open-weight large language model &#8220;K2-Horizon-32B&#8221; into the NVFP4 format.<\/p>\n<p>This model is a 32B dense-configuration, decoder-only model featuring a massive context length of 512K (524,288 tokens). Both weights and activations are quantized to NVFP4 across all linear layers except for <code>lm_head<\/code>, making it optimized for NVIDIA Blackwell generation (B\u30b7\u30ea\u30fc\u30ba) and later GPUs with native NVFP4 support.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameter count: 32B (Total parameters 32B \/ Activated parameters 32B)<\/li>\n<li>Architecture: Dense (decoder-only)<\/li>\n<li>Context length: 512K (524,288 tokens)<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>The model card published by the release origin features a comparison table against other open-weight dense models, as well as a quantization comparison table between the BF16 version and the NVFP4 version.<\/p>\n<p>First, the comparison results with open-weight dense models in the same scale category are as follows:<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th><\/th>\n<th>K2-Horizon-32B-Stage1<\/th>\n<th>Open-weight dense models \/ Qwen3.8-27B<\/th>\n<th>Open-weight dense models \/ Muse Glimmer-30B<\/th>\n<th>Open-weight dense models \/ IBM Granite 4.2 30B<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td># Params<\/td>\n<td>32B<\/td>\n<td>27B<\/td>\n<td>30B<\/td>\n<td>30B<\/td>\n<\/tr>\n<tr>\n<td># Activated params<\/td>\n<td>32B<\/td>\n<td>27B<\/td>\n<td>30B<\/td>\n<td>30B<\/td>\n<\/tr>\n<tr>\n<td>Architecture<\/td>\n<td>Dense<\/td>\n<td>Dense<\/td>\n<td>Dense<\/td>\n<td>Dense<\/td>\n<\/tr>\n<tr>\n<td><strong>Agents<\/strong><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>tau3-Banking Agentic tool use<\/td>\n<td>22.5<\/td>\n<td>48.0<\/td>\n<td>23.5<\/td>\n<td>14.4<\/td>\n<\/tr>\n<tr>\n<td><strong>Coding<\/strong><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Terminal-Bench 2.1 Agentic terminal use<\/td>\n<td>36.6<\/td>\n<td>79.8<\/td>\n<td>51.7<\/td>\n<td>26.6<\/td>\n<\/tr>\n<tr>\n<td>SciCode Scientific coding<\/td>\n<td>30.2<\/td>\n<td>44.7<\/td>\n<td>43.6<\/td>\n<td>36.6<\/td>\n<\/tr>\n<tr>\n<td><strong>Scientific Reasoning<\/strong><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>Humanity&#8217;s Last Exam (without tools) Expert-level reasoning<\/td>\n<td>22.8<\/td>\n<td>33.9<\/td>\n<td>22.0<\/td>\n<td>11.2<\/td>\n<\/tr>\n<tr>\n<td>GPQA Diamond Graduate-level science QA<\/td>\n<td>82.3<\/td>\n<td>90.5<\/td>\n<td>83.5<\/td>\n<td>64.4<\/td>\n<\/tr>\n<tr>\n<td>CritPt Frontier physics reasoning<\/td>\n<td>1.4<\/td>\n<td>5.4<\/td>\n<td>2.6<\/td>\n<td>0.3<\/td>\n<\/tr>\n<tr>\n<td><strong>General<\/strong><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<td><\/td>\n<\/tr>\n<tr>\n<td>AA-LCR Long-context reasoning<\/td>\n<td>65.3<\/td>\n<td>77.3<\/td>\n<td>80.0<\/td>\n<td>46.7<\/td>\n<\/tr>\n<tr>\n<td>AA-Omniscience Accuracy Factual accuracy<\/td>\n<td>16.8<\/td>\n<td>15.6<\/td>\n<td>27.0<\/td>\n<td>10.1<\/td>\n<\/tr>\n<tr>\n<td>AA-Omniscience Non-Hallucination Non-hallucination rate<\/td>\n<td>58.3<\/td>\n<td>69.7<\/td>\n<td>18.1<\/td>\n<td>74.4<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>From this table, it can be seen that K2-Horizon-32B-Stage1 demonstrates high-level reasoning capabilities comparable to Muse Glimmer-30B and outperforms IBM Granite 4.2 30B in Humanity&#8217;s Last Exam (22.8%), which collects expert-level ultra-difficult questions, and GPQA Diamond (82.3%), which tests PhD-level science questions. On the other hand, in agent capabilities and coding metrics such as Terminal-Bench 2.1 (36.6%) measuring terminal operation completion, tau3-Banking (22.5%), and SciCode (30.2%), it falls short of Qwen3.8-27B (79.8%, 48.0%, 44.7%), showing a clear gap.<\/p>\n<p>Next is the direct comparison data between the original model&#8217;s BF16 version and this NVFP4 version (65,536 token context length, 0-shot evaluation).<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>K2-Horizon-32B-Stage1<\/th>\n<th>IFEval (Prompt)<\/th>\n<th>GSM8K<\/th>\n<th>MBPP<\/th>\n<th>MMLU-Pro<\/th>\n<th>GPQA-Diamond<\/th>\n<th>BBH (3-shot)<\/th>\n<th>AIME 26 (avg @ 32)<\/th>\n<th>Average<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>BF16<\/td>\n<td>86.69<\/td>\n<td>96.21<\/td>\n<td>94.40<\/td>\n<td>81.52<\/td>\n<td>81.76<\/td>\n<td>93.20<\/td>\n<td>92.60<\/td>\n<td>89.5<\/td>\n<\/tr>\n<tr>\n<td>NVFP4<\/td>\n<td>85.40<\/td>\n<td>96.59<\/td>\n<td>93.20<\/td>\n<td>80.47<\/td>\n<td>78.82<\/td>\n<td>93.30<\/td>\n<td>91.15<\/td>\n<td>88.4<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>According to measurements by the release origin, the average score dropped by a mere 1.1 points from 89.5% for the BF16 version to 88.4% for the NVFP4 version. While maintaining figures equal to or higher than BF16 in GSM8K (96.59%), which solves math word problems, and BBH (93.30%), a reasoning task, slight score drops are observed in challenging fields such as GPQA-Diamond (81.76% \u2192 78.82%), AIME 26 (92.60% \u2192 91.15%) testing Math Olympiad preliminary levels, MBPP (94.40% \u2192 93.20%) solving Python problems, and IFEval (86.69% \u2192 85.40%) measuring instruction-following. Note that evaluation of the NVFP4 version is currently limited to non-agent tasks, and agent task results are scheduled to be released at a later date.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>While possessing an easy-to-manage scale of 32B, &#8220;IFM\/K2-Horizon-32B-NVFP4&#8221; combines extremely long-context processing capabilities, advanced reasoning abilities, and flexible tool use functions, making it particularly strong in the following use cases:<\/p>\n<h3>Long-Context Processing (Supports 512K Tokens)<\/h3>\n<p>This model supports a native context window of 524,288 tokens (512K) from the midtraining stage onwards. This makes it possible to input bundles of academic papers, large-scale source code bases, or lengthy contracts all at once, and perform advanced analysis, summarization, and Q&amp;A based on their contents.<\/p>\n<h3>Advanced Reasoning with Thought Processes<\/h3>\n<p>This model features the capability to output thought processes (Reasoning). When used via the API, by specifying the recommended setting <code>reasoning_effort=\"high\"<\/code>, the model outputs the thought process leading to the answer in <code>reasoning_content<\/code> and the final answer in <code>content<\/code> separately. This enables obtaining high-quality answers that build accurate step-by-step thoughts in difficult tasks such as mathematics, science, and complex logic puzzles.<\/p>\n<h3>Flexible Tool Use and Agent Integration<\/h3>\n<p>This model supports multiple formats for integration with external tools. Specifically, it supports three types of tool-calling formats: <code>json<\/code>, <code>xml<\/code>, and <code>xml_typed<\/code>, with <code>xml<\/code> used by default. These can be toggled per request via <code>chat_template_kwargs<\/code>, allowing flexible agent construction tailored to system requirements.<\/p>\n<h3>Fast and Memory-Efficient Inference on Blackwell GPUs<\/h3>\n<p>The greatest feature of this model is that the weights and activations of all linear layers except for <code>lm_head<\/code> are quantized into the NVFP4 format. This allows for extremely fast inference while drastically reducing memory usage on GPUs with native NVFP4 support from the NVIDIA Blackwell generation (B\u30b7\u30ea\u30fc\u30ba) and later. Because the performance drop compared to the BF16 version is suppressed to a mere minimum (a 1.1-point drop in average score), it becomes an extremely practical choice for engineers who own Blackwell hardware.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 20.7B parameters<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>32GB (RTX 5090, etc.)<\/td>\n<td>NVFP4<\/td>\n<td>21.7GB<\/td>\n<td>26.0GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-09-22): llama.cpp: not registered, vLLM: registered, MLX (mlx-lm): not registered. &#8220;Not registered&#8221; means the name is absent from that registry today, not that the model cannot run.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:peers --><\/p>\n<h2>Recent Models in the Same Size Class<\/h2>\n<p><em>Models with <\/em><em>15\u201340B<\/em><em> parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site&#8217;s estimates; licenses are as stated on the model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>prism-ml\/Ternary-Bonsai-2-27B-gguf<\/td>\n<td>27.8B<\/td>\n<td>8GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf\/\">Prism ML Releases Ternary-Bonsai-2-27B-gguf: A 1.72 bpw Ternary Model<\/a> (2026-09-18)<\/td>\n<\/tr>\n<tr>\n<td>prism-ml\/Ternary-Bonsai-2-27B-gguf-dev<\/td>\n<td>27.8B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf-dev-released\/\">Prism ML Releases Ternary-Bonsai-2-27B-gguf-dev Build<\/a> (2026-09-18)<\/td>\n<\/tr>\n<tr>\n<td>Edge0\/Edge0-35B-A3B-preview<\/td>\n<td>34.7B<\/td>\n<td>24GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/edge0-35b-a3b-preview-sparse-moe\/\">Edge0-35B-A3B-Preview: Sparse MoE for Phone-Class Memory<\/a> (2026-09-11)<\/td>\n<\/tr>\n<tr>\n<td>bartowski\/Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF<\/td>\n<td>26.5B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/pantheon-reasoning-26b-gguf\/\">Pantheon-Reasoning-26B-A4B-1.1-V2 GGUF Quantizations<\/a> (2026-09-11)<\/td>\n<\/tr>\n<tr>\n<td>nex-agi\/Nex-N2.5-mini<\/td>\n<td>35.1B<\/td>\n<td>16GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-n25-mini-released\/\">Nex-AGI Releases Open-Weight Long-Task Model Nex-N2.5-mini<\/a> (2026-09-09)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:peers --><\/p>\n<h2>How to Get It<\/h2>\n<p>This model can be freely downloaded and used from the Hugging Face repository. It is published under the &#8220;Apache-2.0&#8221; license and is not a gated model requiring prior consent for use, meaning anyone can obtain it immediately.<\/p>\n<p>The distribution format is <code>safetensors<\/code>, and it is compatible with vLLM, SGLang, and Hugging Face&#8217;s Transformers library as serving engines.<\/p>\n<p>Below are usage methods and code examples for each engine provided by the release origin. Note that the following code examples are written based on the BF16 version (<code>IFM\/K2-Horizon-32B<\/code>), but similar settings and parser designations are also recommended when running the NVFP4 version.<\/p>\n<h3>Serving with vLLM<\/h3>\n<p>Here is an example command for launching a server using vLLM. It is recommended to enable the <code>k2_horizon<\/code> reasoning parser for chat and the tool call parser for agents.<\/p>\n<pre><code class=\"language-shell\">vllm serve IFM\/K2-Horizon-32B \\\n  --revision main \\\n  --model-impl vllm \\\n  --tensor-parallel-size 2 \\\n  --trust-remote-code \\\n  --dtype bfloat16 \\\n  --max-model-len 131072 \\\n  --reasoning-parser k2_horizon \\\n  --enable-auto-tool-choice \\\n  --tool-call-parser k2_horizon\n<\/code><\/pre>\n<h3>Serving with SGLang<\/h3>\n<p>Here is an example command when using SGLang. This recipe is verified on a 2\u00d7 H200 environment.<\/p>\n<pre><code class=\"language-shell\">python3 -m sglang.launch_server \\\n  --model-path IFM\/K2-Horizon-32B \\\n  --revision main \\\n  --tp 2 \\\n  --dtype bfloat16 \\\n  --attention-backend fa3 \\\n  --reasoning-parser k2_horizon \\\n  --tool-call-parser k2_horizon \\\n  --host 0.0.0.0 --port 30000\n<\/code><\/pre>\n<h3>API Usage Example (Python)<\/h3>\n<p>Here is a Python code example to obtain answers with thought processes via the OpenAI-compatible API after starting the server. Recommended settings such as <code>temperature=1.0<\/code>, <code>top_p=0.95<\/code>, and <code>reasoning_effort=\"high\"<\/code> are specified.<\/p>\n<pre><code class=\"language-python\">from openai import OpenAI\n\nclient = OpenAI(base_url=&quot;http:\/\/localhost:30000\/v1&quot;, api_key=&quot;EMPTY&quot;)\nresponse = client.chat.completions.create(\n    model=&quot;IFM\/K2-Horizon-32B&quot;,\n    messages=[{&quot;role&quot;: &quot;user&quot;, &quot;content&quot;: &quot;Explain the result step by step.&quot;}],\n    temperature=1.0,\n    top_p=0.95,\n    max_tokens=32768,\n    extra_body={&quot;chat_template_kwargs&quot;: {&quot;reasoning_effort&quot;: &quot;high&quot;, &quot;tool_call_format&quot;: &quot;xml&quot;}},\n)\nmessage = response.choices[0].message\nprint(&quot;Reasoning:&quot;, getattr(message, &quot;reasoning_content&quot;, None))\nprint(&quot;Answer:&quot;, message.content)\n<\/code><\/pre>\n<h3>Usage Example with Transformers<\/h3>\n<p>Here is a code example to load the model directly and generate text using Transformers (verified with version 5.15.0, PyTorch 2.13.0, and Safetensors 0.8.0).<\/p>\n<pre><code class=\"language-python\">from transformers import AutoModelForCausalLM, AutoTokenizer\n\nmodel_id = &quot;IFM\/K2-Horizon-32B&quot;\ntokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)\nmodel = AutoModelForCausalLM.from_pretrained(\n    model_id, device_map=&quot;auto&quot;, dtype=&quot;bfloat16&quot;, low_cpu_mem_usage=True, trust_remote_code=True\n)\n\ninputs = tokenizer(&quot;Explain why long-context evaluation is difficult.&quot;, return_tensors=&quot;pt&quot;).to(model.device)\ninputs.pop(&quot;token_type_ids&quot;, None)\noutputs = model.generate(**inputs, max_new_tokens=32768, temperature=1.0, top_p=0.95, do_sample=True)\nprint(tokenizer.decode(outputs[0], skip_special_tokens=True))\n<\/code><\/pre>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-375b-a23b-nvfp4-released-2\/\">IFM Releases K2-Horizon-375B-A23B-NVFP4 Quantized Model<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/nvidia-releases-deepseek-v4-pro-nvfp4\/\">NVIDIA Releases NVFP4 Quantized DeepSeek-V4-Pro<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">OpenBMB Releases MiniCPM5-2B: A SOTA 2B On-Device Model<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/IFM\/K2-Horizon-32B-NVFP4\">IFM\/K2-Horizon-32B-NVFP4 (Hugging Face)<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>IFM releases K2-Horizon-32B-NVFP4, an NVFP4 quantized version of the 32B long-context open-weight LLM optimized for NVIDIA Blackwell GPUs.<\/p>\n","protected":false},"author":1,"featured_media":2991,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[968,2051,2082,2053,813,169,1547,592,1555],"class_list":["post-2992","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-blackwell-en","tag-ifm-en","tag-ifm-k2-horizon-32b-nvfp4-en","tag-k2-horizon-en","tag-nvfp4-en","tag-sglang-en","tag-verified","tag-vllm-en","tag--en"],"lang":"en","translations":{"en":2992,"ja":2990},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2992","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=2992"}],"version-history":[{"count":0,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/2992\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/2991"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=2992"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=2992"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=2992"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}