{"id":571,"date":"2026-09-12T20:11:45","date_gmt":"2026-09-12T11:11:45","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/09\/12\/bartowski-nex-agi-nex-n2-5-mini-gguf-2\/"},"modified":"2026-09-18T21:42:00","modified_gmt":"2026-09-18T12:42:00","slug":"bartowski-nex-agi-nex-n2-5-mini-gguf-2","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/bartowski-nex-agi-nex-n2-5-mini-gguf-2\/","title":{"rendered":"bartowski\/nex-agi_Nex-N2.5-mini-GGUF: Specs and Hardware"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/nex-agi_Nex-N2.5-mini-GGUF\">bartowski\/nex-agi_Nex-N2.5-mini-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-11<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>apache-2.0<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td>GGUF<\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code at collection time. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<h2>Overview<\/h2>\n<p>It is reported that <code>bartowski\/nex-agi_Nex-N2.5-mini-GGUF<\/code>, published by bartowski, consists of GGUF quantization files for the multimodal agent model &#8220;Nex-N2.5-mini&#8221; developed by nex-agi, designed to run with llama.cpp. Since it contains information that has not been officially confirmed at this time, caution is required when actually deploying it.<\/p>\n<h2>Specifications<\/h2>\n<ul>\n<li>Parameters: 35B<\/li>\n<li>License: apache-2.0<\/li>\n<\/ul>\n<h2>Performance<\/h2>\n<p>According to the model card, the following text and multimodal benchmark scores have been published for the original model, Nex-N2.5-mini (excerpting Nex-N2.5-Pro, Nex-N2.5-Max, Claude Opus 5, GPT-5.6 Sol, Kimi-K3, GLM-5.3, DeepSeek-V4-Pro-0813, and Qwen3.8-Max for comparison).<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Benchmark \/ CODING 3<\/th>\n<th>Nex-N2.5-mini \/ CODING 3<\/th>\n<th>Nex-N2.5-Pro \/ CODING 3<\/th>\n<th>Nex-N2.5-Max \/ CODING 3<\/th>\n<th>Claude Opus 5 \/ CODING 3<\/th>\n<th>GPT-5.6 Sol \/ CODING 3<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Terminal-Bench 2.1<\/td>\n<td>73.4<\/td>\n<td>82.7<\/td>\n<td>86.1<\/td>\n<td>89.1<\/td>\n<td>88.8<\/td>\n<\/tr>\n<tr>\n<td>SWE-Bench Pro<\/td>\n<td>43.8<\/td>\n<td>61.2<\/td>\n<td>65.7<\/td>\n<td>79.2<\/td>\n<td>64.6<\/td>\n<\/tr>\n<tr>\n<td>DeepSWE v1.1<\/td>\n<td>36.1<\/td>\n<td>55.8<\/td>\n<td>65.6<\/td>\n<td>73.7<\/td>\n<td>72.7<\/td>\n<\/tr>\n<tr>\n<td>AutomationBench v1.0.6 5<\/td>\n<td>32.3<\/td>\n<td>44.2<\/td>\n<td>50.2<\/td>\n<td>50.3<\/td>\n<td>45.8<\/td>\n<\/tr>\n<tr>\n<td>Toolathlon Verified<\/td>\n<td>54.6<\/td>\n<td>68.5<\/td>\n<td>74.7<\/td>\n<td>76.5<\/td>\n<td>74.9<\/td>\n<\/tr>\n<tr>\n<td>GDPval-AA v2<\/td>\n<td>1446<\/td>\n<td>1628<\/td>\n<td>1713<\/td>\n<td>1831<\/td>\n<td>1711<\/td>\n<\/tr>\n<tr>\n<td>Job Bench<\/td>\n<td>28.5<\/td>\n<td>41.4<\/td>\n<td>53.6<\/td>\n<td>65.7<\/td>\n<td>45.4<\/td>\n<\/tr>\n<tr>\n<td>BrowseComp 6<\/td>\n<td>83.4<\/td>\n<td>89.7<\/td>\n<td>92.6<\/td>\n<td>90.8<\/td>\n<td>90.4<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>Nex-N2.5-mini<\/th>\n<th>Nex-N2.5-Pro<\/th>\n<th>MiniMax-M3<\/th>\n<th>Claude Opus 5<\/th>\n<th>GPT-5.6 Sol<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>OSWorld-Verified 8<\/td>\n<td>71.2<\/td>\n<td>82.2<\/td>\n<td>75.2<\/td>\n<td>83.4<\/td>\n<td>83.2<\/td>\n<\/tr>\n<tr>\n<td>OSWorld-2<\/td>\n<td>30.5<\/td>\n<td>56.4<\/td>\n<td>22.3<\/td>\n<td>68.3<\/td>\n<td>62.7<\/td>\n<\/tr>\n<tr>\n<td>WebTest 8, 9<\/td>\n<td>48.6<\/td>\n<td>52.8<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>54.0<\/td>\n<\/tr>\n<tr>\n<td>WebArena-Verified 8<\/td>\n<td>63.4<\/td>\n<td>67.6<\/td>\n<td>\u2014<\/td>\n<td>\u2014<\/td>\n<td>69.7<\/td>\n<\/tr>\n<tr>\n<td>OSWorld-G<\/td>\n<td>82.9<\/td>\n<td>87.4<\/td>\n<td>\u2014<\/td>\n<td>76.8<\/td>\n<td>77.7<\/td>\n<\/tr>\n<tr>\n<td>Vision2Web 7<\/td>\n<td>52.9<\/td>\n<td>68.2<\/td>\n<td>59.0<\/td>\n<td>\u2014<\/td>\n<td>79.8<\/td>\n<\/tr>\n<tr>\n<td>SWE-MM<\/td>\n<td>25.5<\/td>\n<td>38.2<\/td>\n<td>\u2014<\/td>\n<td>59.4<\/td>\n<td>40.2<\/td>\n<\/tr>\n<tr>\n<td>OmniDoc<\/td>\n<td>89.7<\/td>\n<td>92.2<\/td>\n<td>91.6<\/td>\n<td>\u2014<\/td>\n<td>92.9<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>Looking at the published scores of the original model, it records 73.4 on Terminal-Bench 2.1, showing solid performance in terminal operation agent tasks. On the other hand, compared to higher-end models like Nex-N2.5-Pro or large-scale models from other companies (such as Claude Opus 5 and GPT-5.6 Sol), the numbers are lower, meaning it is not at the state-of-the-art level. The gap widens further on SWE-Bench Pro (43.8) and real-repository\/real-task metrics, indicating it falls short of top-tier proprietary models or higher-end variants in the same series.<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>The original model, Nex-N2.5-mini, is said to be part of a next-generation agent model family specialized in computer operation, web browsing, and visually grounded agent capabilities. It supports image inputs in addition to text inputs, allowing it to handle tasks involving image recognition when used together with a multimodal projector file (<code>mmproj<\/code>).<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong> \u2014 35.1B parameters (taken from the base model nex-agi\/Nex-N2.5-mini)<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>16GB (RTX 5060 Ti 16GB \/ 4060 Ti 16GB, etc.)<\/td>\n<td>Q2_K<\/td>\n<td>12.9GB<\/td>\n<td>15.5GB<\/td>\n<\/tr>\n<tr>\n<td>24GB (RTX 4090 \/ 3090, etc.)<\/td>\n<td>Q4_K_S<\/td>\n<td>19.5GB<\/td>\n<td>23.4GB<\/td>\n<\/tr>\n<tr>\n<td>32GB (RTX 5090, etc.)<\/td>\n<td>Q5_K_M<\/td>\n<td>25.1GB<\/td>\n<td>30.2GB<\/td>\n<\/tr>\n<tr>\n<td>48GB (RTX 6000 Ada \/ A6000, etc.)<\/td>\n<td>Q8_0<\/td>\n<td>34.4GB<\/td>\n<td>41.3GB<\/td>\n<\/tr>\n<tr>\n<td>80GB class (A100 \/ H100)<\/td>\n<td>BF16<\/td>\n<td>64.6GB<\/td>\n<td>77.5GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:peers --><\/p>\n<h2>Recent Models in the Same Size Class<\/h2>\n<p><em>Models with <\/em><em>15\u201340B<\/em><em> parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site&#8217;s estimates; licenses are as stated on the model cards.<\/em><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Parameters<\/th>\n<th>Smallest VRAM tier<\/th>\n<th>License<\/th>\n<th>Our article<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>bartowski\/vectionlabs_Salience-27B-R6-GGUF<\/td>\n<td>27.8B<\/td>\n<td>12GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">Salience-27B-R6 GGUF Released by bartowski<\/a> (2026-09-15)<\/td>\n<\/tr>\n<tr>\n<td>agentionai\/Signal-3.8-27B-GGUF<\/td>\n<td>27.8B<\/td>\n<td>16GB<\/td>\n<td>apache-2.0<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF: Faster and More Token-Efficient<\/a> (2026-09-12)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><!-- \/lmw:peers --><\/p>\n<h2>How to Get It<\/h2>\n<p>Distributed in GGUF format, it is compatible with engines such as llama.cpp, LM Studio, koboldcpp, ramalama, Jan AI, Text Generation Web UI, LoLLMs, and Atomic Chat. An example download command using Hugging Face CLI is as follows:<\/p>\n<pre><code>hf download bartowski\/nex-agi_Nex-N2.5-mini-GGUF --include &quot;nex-agi_Nex-N2.5-mini-Q4_K_M.gguf&quot; --local-dir.\/\n<\/code><\/pre>\n<p><!-- lmw:related --><\/p>\n<h2>Related Articles<\/h2>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-n25-mini-released\/\">Nex-AGI Releases Open-Weight Long-Task Model Nex-N2.5-mini<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B GGUF Quantized Models Released by bartowski<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/\">Orion-26B-A4B-v1.1 GGUF Quants Released by bartowski<\/a><\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">Salience-27B-R6 GGUF Released by bartowski<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:related --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/bartowski\/nex-agi_Nex-N2.5-mini-GGUF\">https:\/\/huggingface.co\/bartowski\/nex-agi_Nex-N2.5-mini-GGUF<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/nex-agi\/Nex-N2.5-mini\">https:\/\/huggingface.co\/nex-agi\/Nex-N2.5-mini<\/a><\/li>\n<\/ul>\n","protected":false},"excerpt":{"rendered":"<p>Explore the GGUF quantization of nex-agi&#8217;s multimodal agent model Nex-N2.5-mini by bartowski, including hardware requirements and benchmarks.<\/p>\n","protected":false},"author":1,"featured_media":661,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[1030,163,1073,549,551,896,117],"class_list":["post-571","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-bartowski-en","tag-gguf-en","tag-llama-cpp-en","tag-nex-agi-en","tag-nex-n2-5-mini-en","tag--en"],"lang":"en","translations":{"en":571,"ja":569},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/571","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=571"}],"version-history":[{"count":8,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/571\/revisions"}],"predecessor-version":[{"id":1653,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/571\/revisions\/1653"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/661"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=571"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=571"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=571"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}