{"id":9968,"date":"2026-10-05T18:25:31","date_gmt":"2026-10-05T09:25:31","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/2026\/10\/05\/naive-n05-flash-released\/"},"modified":"2026-10-06T01:40:37","modified_gmt":"2026-10-05T16:40:37","slug":"naive-n05-flash-released","status":"publish","type":"post","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/05\/naive-n05-flash-released\/","title":{"rendered":"Naive-N0.5-Flash MoE Model for AI Research and Coding: ~690GB Memory"},"content":{"rendered":"<p><!-- lmw:facts --><\/p>\n<h2>At a Glance<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Item<\/th>\n<th>Value<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Repository<\/td>\n<td><a href=\"https:\/\/huggingface.co\/NaiveAI\/Naive-N0.5-Flash\">NaiveAI\/Naive-N0.5-Flash<\/a><\/td>\n<\/tr>\n<tr>\n<td>Family guide<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-naiveai-naive-n0-5-en\/\">Naive-N0.5 guide (1 articles)<\/a><\/td>\n<\/tr>\n<tr>\n<td>Published<\/td>\n<td>2026-09-27<\/td>\n<\/tr>\n<tr>\n<td>License<\/td>\n<td>mit<\/td>\n<\/tr>\n<tr>\n<td>Formats<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-safetensors-en\/\">safetensors<\/a><\/td>\n<\/tr>\n<tr>\n<td>Source type<\/td>\n<td>Primary source (the publisher itself)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><em>Values determined by this site&#8217;s code when the information was collected. Dates are JST.<\/em><\/p>\n<p><!-- \/lmw:facts --><\/p>\n<p><!-- lmw:lab-summary --><\/p>\n<p><strong>What we checked ourselves<\/strong><\/p>\n<ul>\n<li>The same Japanese text takes about as many tokens as with the Qwen3 tokenizer.<\/li>\n<\/ul>\n<p>Details and conditions are in \u201cOur Own Measurements\u201d below.<\/p>\n<p><!-- \/lmw:lab-summary --><\/p>\n<h2>Overview<\/h2>\n<p>Emerging research organization NaiveAI has released its first model, <strong>Naive-N0.5-Flash<\/strong>, on Hugging Face. Out of its 309B total parameters, 15.5B are active per token in this Mixture of Experts (MoE) model, which is built for coding and &#8220;AI R&amp;D&#8221; tasks. It natively handles up to 1 million tokens of context. The model is licensed under MIT, allowing free use of its weights and inference code, including for commercial purposes.<\/p>\n<p>The foundation is <strong>MiMo-V2.5 Base<\/strong>, released by Xiaomi. NaiveAI modified parts of this model&#8217;s attention weights and performed 3.25T (3.25 trillion) tokens of additional training. Its defining feature is that <strong>it does not contain a single standard full attention layer<\/strong>. Most layers use Sliding Window Attention (SWA), looking only at the immediately preceding 128 tokens, while the remaining layers are replaced with DeepSeek Sparse Attention (DSA), devised by DeepSeek. Even with a 1 million token context, the computation required to generate a single token is structured to be largely independent of context length.<\/p>\n<p>Another distinctive aspect is how it was built. According to the technical blog, NaiveAI pursued a &#8220;AI builds AI&#8221; approach, having AI models handle architecture candidate implementation and comparative experiments, as well as finding and fixing bugs in the training system, while researchers handled goals, constraints, and final decisions. Alongside the main model, the publisher has also released an FP8 version (<code>NaiveAI\/Naive-N0.5-Flash-FP8<\/code>) and an FP8 repository with &#8220;Draft&#8221; in the name.<\/p>\n<p>Regarding Xiaomi models of the same 309B \/ 15B active scale, our site has covered <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf-released\/\">MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory<\/a>.<\/p>\n<h2>Specifications<\/h2>\n<p>Extracted from the model card&#8217;s architecture table and config.json.<\/p>\n<ul>\n<li>Parameter count: 309B total \/ 15.5B active (MoE)<\/li>\n<li>Number of layers: 48 (39 SWA layers, 9 DSA layers)<\/li>\n<li>Layer arrangement: There are 8 modules consisting of 6 layers each, where standard modules place 1 DSA layer after 5 SWA layers. Only the first module has a DSA layer at the very beginning<\/li>\n<li>SWA window: 128 tokens<\/li>\n<li>DSA: A lightweight &#8220;indexer&#8221; scores all past tokens, and only the top 2,048 tokens are used for the main attention computation. The indexer has 16 query heads<\/li>\n<li>DSA KV: The original DSA&#8217;s MLA is replaced with GQA (GQA4), which groups KVs into 4 groups<\/li>\n<li>Experts: Uses 8 out of 256 experts (config.json). The MoE layers number 47 out of 48, with the first layer being a dense FFN<\/li>\n<li>Hidden layer dimension: 4,096 (config.json)<\/li>\n<li>Context length: 1M tokens natively<\/li>\n<li>Weight precision: The main model is distributed in BF16. Inference supports mixed FP8 precision<\/li>\n<\/ul>\n<p>It should be noted that what DSA reduces is <strong>computation and memory read operations<\/strong>, not the KV cache itself. The model card explicitly states: &#8220;The indexer scans the entire past, and all KV caches are retained.&#8221; Even if it becomes faster with long contexts, KV cache memory scales with context length.<\/p>\n<h2>Performance<\/h2>\n<p>Model card evaluations are presented in charts. Figures were read from the included PDF and converted into a table. &#8220;-&#8221; indicates no result for that item. Values for Naive-N0.5-Flash were measured by the publisher using Claude Code (2.1.207, 1M token context, file operations and Bash only configuration), while values for other models are cited from company announcements or leaderboards. Measurement conditions are not necessarily identical.<\/p>\n<h3>Coding<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>Naive-N0.5-Flash<\/th>\n<th>DeepSeek-V4.1-Flash<\/th>\n<th>Kimi-K3<\/th>\n<th>GLM-5.3<\/th>\n<th>Hy4-preview<\/th>\n<th>Opus-5<\/th>\n<th>Opus-5.5<\/th>\n<th>GPT-5.6-Sol<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>DeepSWE v1.1<\/td>\n<td>67.8<\/td>\n<td>74.2<\/td>\n<td>67.5<\/td>\n<td>66.9<\/td>\n<td>64.3<\/td>\n<td>73.6<\/td>\n<td>74.2<\/td>\n<td>72.7<\/td>\n<\/tr>\n<tr>\n<td>Agents&#8217; Last Exam (ALE-CLI)<\/td>\n<td>32.4<\/td>\n<td>&#8211;<\/td>\n<td>28.3<\/td>\n<td>28.5<\/td>\n<td>22.8<\/td>\n<td>29.5<\/td>\n<td>34.3<\/td>\n<td>28.6<\/td>\n<\/tr>\n<tr>\n<td>Terminal-Bench 2.1<\/td>\n<td>86.7<\/td>\n<td>90.6<\/td>\n<td>88.3<\/td>\n<td>88.2<\/td>\n<td>85.4<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<td>88.8<\/td>\n<\/tr>\n<tr>\n<td>SWE-bench Pro<\/td>\n<td>73.6<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<td>65.7<\/td>\n<td>79.2<\/td>\n<td>89.9<\/td>\n<td>64.6<\/td>\n<\/tr>\n<tr>\n<td>ProgramBench<\/td>\n<td>17.5<\/td>\n<td>20.3<\/td>\n<td>&#8211;<\/td>\n<td>19.0<\/td>\n<td>17.5<\/td>\n<td>37.0<\/td>\n<td>&#8211;<\/td>\n<td>23.0<\/td>\n<\/tr>\n<tr>\n<td>NL2Repo-Bench<\/td>\n<td>71.9<\/td>\n<td>64.0<\/td>\n<td>&#8211;<\/td>\n<td>58.0<\/td>\n<td>58.9<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<\/tr>\n<tr>\n<td>FrontierSWE v1<\/td>\n<td>78.2<\/td>\n<td>&#8211;<\/td>\n<td>81.2<\/td>\n<td>78.1<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<td>&#8211;<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The charts also list Muse-Spark-1.3 (DeepSWE 75.4, ALE-CLI 33.3, Terminal-Bench 2.1 88.8), Step-5-Preview (DeepSWE 67.7, ALE-CLI 29.5), GLM-5.3-Flash (DeepSWE 63.4, ALE-CLI 26.3, Terminal-Bench 2.1 84.3), Qwen-3.8-Max (DeepSWE 56.6, SWE-bench Pro 67.7, NL2Repo-Bench 55.9, FrontierSWE 73.5), GPT-6-Astra (ALE-CLI 33.3), and Fable-5 (with fallback: DeepSWE 69.7, ALE-CLI 23.8, Terminal-Bench 2.1 88.0, ProgramBench 33.0, FrontierSWE 88.2).<\/p>\n<h3>AI R&amp;D<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Benchmark<\/th>\n<th>Naive-N0.5-Flash<\/th>\n<th>Comparison Targets<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>PostTrainBench<\/td>\n<td>37.5<\/td>\n<td>GPT-5.6-Sol 41.8 \/ GLM-5.3 39.8 \/ Kimi-K3 36.6 \/ Fable-5 (with fallback) 36.2 \/ Hy4-preview 35.6<\/td>\n<\/tr>\n<tr>\n<td>MLE-bench-30<\/td>\n<td>73.7%<\/td>\n<td>Sonnet-5 66.9% \/ Gemini-3.6-Flash 63.9% \/ Gemini-3.5-Flash 49.7% \/ GPT-5.6-Luna 47.6% \/ Grok-4.5 43.2%<\/td>\n<\/tr>\n<tr>\n<td>PaperBench<\/td>\n<td>63.2<\/td>\n<td>Opus-4.7 58.5 \/ GPT-5.5 57.5 \/ MiniMax-M3 52.6 \/ Gemini-3.1-Pro 46.7<\/td>\n<\/tr>\n<tr>\n<td>SOL-ExecBench (Mean SOL Score, higher is better)<\/td>\n<td>72.81<\/td>\n<td>Recursive Superintelligence, Inc. (unreleased third-party model) 67.56<\/td>\n<\/tr>\n<tr>\n<td>NanoChat AutoResearch (Validation BPB, lower is better)<\/td>\n<td>0.9051<\/td>\n<td>Same as above 0.9109<\/td>\n<\/tr>\n<tr>\n<td>NanoGPT SpeedRun (training time in seconds, lower is better)<\/td>\n<td>73.8<\/td>\n<td>Same as above 77.5<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>For tasks that require carrying out long procedures to completion, it ranks near the top among open models.<\/strong> NL2Repo-Bench, which builds an entire repository from natural language specifications, scores 71.9, the highest in the chart (2nd place is DeepSeek-V4.1-Flash at 64.0). ALE-CLI, which collects long operational tasks across diverse fields, scores 32.4, outperforming all open models listed in the chart (Step-5-Preview 29.5, GLM-5.3 28.5, Kimi-K3 28.3, etc.), scoring higher than Opus-5 (29.5) and within 1.9 points of Opus-5.5 (34.3). SWE-bench Pro, which solves issues in actual GitHub repositories, scores 73.6, outperforming GPT-5.6-Sol (64.6) and Qwen-3.8-Max (67.7).<\/p>\n<p><strong>On the other hand, there are clear areas where it is weaker compared to models of similar size.<\/strong> In the chart legend, DeepSeek-V4.1-Flash falls into the same &#8220;under 600B&#8221; category, but scores higher than Naive-N0.5-Flash on DeepSWE (74.2 vs 67.8), Terminal-Bench 2.1 (90.6 vs 86.7), and ProgramBench (20.3 vs 17.5). Its Terminal-Bench 2.1 score of 86.7 ranks 7th out of 9 models in the chart, and its ProgramBench score of 17.5 (recreating programs from executables and documents) is the lowest in the chart (tied with Hy4-preview). The table reveals a trend of being strong on tasks that &#8220;assemble large new things from specifications&#8221; and weaker on tasks that &#8220;analyze and reproduce existing things.&#8221;<\/p>\n<p><strong>The AI R&amp;D metrics require caution regarding comparison partners and measurement methods.<\/strong> While MLE-bench-30 (73.7%) and PaperBench (63.2) lead the chart, their comparison targets are third-party models from model cards like Gemini 3.6 Flash and MiniMax M3, and do not include the latest top-tier models like Opus-5.5. MLE-bench-30 is the average position score following Gemini 3.6 Flash&#8217;s evaluation procedure. The three items SOL-ExecBench, NanoChat AutoResearch, and NanoGPT SpeedRun are results from running Naive-N0.5-Flash in NaiveAI&#8217;s internal harness, compared against only a single unreleased third-party model. While Naive-N0.5-Flash performs better on all three, the margins are small (0.0058 in NanoChat BPB, 3.7 seconds in NanoGPT training time) and have not been replicated by third parties under the same conditions. PostTrainBench at 37.5 falls short of GPT-5.6-Sol (41.8) and GLM-5.3 (39.8).<\/p>\n<h2>Strengths and Use Cases<\/h2>\n<p>Model card and technical blog targets indicate <strong>coding agents that read entire large repositories or long work histories<\/strong> and <strong>research agents that run machine learning experiments autonomously<\/strong>. The 3.25T additional training tokens focus primarily on data centered around AI R&amp;D and coding, with the entire training conducted within a 1-million-token context.<\/p>\n<p>Additional training occurs in three stages:<\/p>\n<ul>\n<li><strong>Indexer warmup (50B tokens)<\/strong>: Trains only the newly added DSA indexer while keeping other weights frozen. At this stage, layers replaced by DSA are still computed using full attention, matching the indexer selection method to that attention distribution via KL divergence<\/li>\n<li><strong>Sparse attention training (3T tokens)<\/strong>: Switches to sparse attention and performs continued pre-training with a fixed learning rate<\/li>\n<li><strong>Decay stage (200B tokens)<\/strong>: Lowers the learning rate while performing Supervised Fine-Tuning (SFT)<\/li>\n<\/ul>\n<p>The technical blog also cites specific examples where AI models contributed to improving the training system: finding and fixing numerical instability in how the indexer selected the top 2,048 items, proposing separate parallelization schemes for DSA and SWA layers (Ulysses for DSA, neighboring split and window-exchange for SWA) to reduce communication volume, and discovering positional encoding precision issues at 1 million tokens. The indexer reportedly reduces selection time by 44% compared to the original DSA implementation. The degree of this AI involvement is described by the publisher and cannot be verified externally.<\/p>\n<p>Regarding inference speed, the model card reports 50 tokens per user per second in standard mode and up to 2,000 tokens per second in Ultrafast mode on their proprietary inference system, NaiveRT. These figures were obtained using their system on datacenter GPUs and do not serve as a benchmark for local hardware.<\/p>\n<h2>How It Differs from Similar Models<\/h2>\n<ul>\n<li><strong>MiMo-V2.5 Base (Foundation)<\/strong>: MIT-licensed weights released by Xiaomi. According to the model card, most layers use SWA, with a few full-attention layers preserving distant information. At a 1-million-token context, these full-attention layers account for the majority of generation overhead, so Naive-N0.5-Flash replaced them with DSA<\/li>\n<li><strong>MiMo-V2.6-Flash-RL<\/strong>: The next generation in the same Xiaomi lineage, with the same scale of 309B total \/ 15B active parameters. This version focuses on scaling agent capabilities via reinforcement learning, while retaining the original MiMo attention architecture. Our site&#8217;s GGUF version article: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf-released\/\">MiMo-V2.6-Flash-RL-GGUF Multimodal MoE Model: ~141GB Memory<\/a>. Upper model article: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/27\/mimo-v2-6-pro-rl-released\/\">MiMo-V2.6-Pro-RL Text Generation Model: ~417GB Memory, GGUF Builds<\/a><\/li>\n<li><strong>DeepSeek-V4.1-Flash<\/strong>: A model falling into the same &#8220;under 600B&#8221; category in evaluation charts. As shown in the table above, DeepSeek-V4.1-Flash outperforms on DeepSWE, Terminal-Bench 2.1, and ProgramBench, while Naive-N0.5-Flash outperforms on NL2Repo-Bench<\/li>\n<\/ul>\n<p>Viewing this as an example of modifying existing weight attention mechanisms to alternative schemes and using continued training to maintain performance while making long contexts lightweight clarifies the value of this model. Rather than being trained from scratch, it is a derivative model where an organization extended publicly released weights in its own direction.<\/p>\n<p><!-- lmw:hardware --><\/p>\n<h2>Hardware Requirements<\/h2>\n<p><strong>Estimated requirements (calculated by Local Model Watch)<\/strong><\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Your VRAM<\/th>\n<th>Quantization<\/th>\n<th>File size<\/th>\n<th>Est. memory needed<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>More than 690GB of VRAM (multi-GPU or CPU offload required)<\/td>\n<td>Original precision<\/td>\n<td>575.4GB<\/td>\n<td>690.5GB<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p><strong>Inference engine support<\/strong> (architecture name matched against each project&#8217;s own model registry in its source code, checked 2026-10-06): llama.cpp: not registered, vLLM: not registered, MLX (mlx-lm): not registered. &#8220;Not registered&#8221; means the name is absent from that registry today, not that the model cannot run.<\/p>\n<p><em>Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model&#8217;s authors. Compare with other models in our <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference<\/a>. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:hardware --><\/p>\n<p><!-- lmw:runnability --><\/p>\n<h2>Can You Run It Locally?<\/h2>\n<p><strong>Not usable in Ollama, LM Studio and llama.cpp yet.<\/strong><\/p>\n<p>The publisher ships safetensors only, and llama.cpp&#8217;s registry does not list this architecture. llama.cpp would need to add support before these tools can run it. 4 converted build(s) from other uploaders exist. Today it can be run with transformers, using the memory figures in the table above.<\/p>\n<p><strong>License \u2014 <code>mit<\/code> (Commercial use allowed):<\/strong> Permits commercial use, modification and redistribution, provided the copyright notice and license text are retained.<\/p>\n<p><em>Compiled by this site&#8217;s code from the published formats, converted builds we have found, and each engine&#8217;s own model registry. &#8220;Not found&#8221; means we have not seen such a build, not that none exists. License summaries are not legal advice \u2014 check the publisher&#8217;s original terms before relying on them.<\/em><\/p>\n<p><!-- \/lmw:runnability --><\/p>\n<p><!-- lmw:lab --><\/p>\n<h2>Our Own Measurements<\/h2>\n<p>Values we measured ourselves on our server (no GPU) by actually reading and running this model&#8217;s files \u2014 not figures copied from the model card. How we measure, and the results for every model: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/\">Our Measurements<\/a>.<\/p>\n<h3>Japanese Token Efficiency<\/h3>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Tokenizer<\/th>\n<th>Tokens per 1,000 Japanese characters<\/th>\n<th>Ratio to the same text in English<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><strong>This model<\/strong><\/td>\n<td><strong>688<\/strong><\/td>\n<td><strong>1.26\u00d7<\/strong><\/td>\n<\/tr>\n<tr>\n<td>Qwen3<\/td>\n<td>688<\/td>\n<td>1.26\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Llama 3.2<\/td>\n<td>744<\/td>\n<td>1.36\u00d7<\/td>\n<\/tr>\n<tr>\n<td>Gemma 3<\/td>\n<td>564<\/td>\n<td>1.03\u00d7<\/td>\n<\/tr>\n<tr>\n<td>gpt-oss<\/td>\n<td>795<\/td>\n<td>1.45\u00d7<\/td>\n<\/tr>\n<tr>\n<td>LLM-jp-3<\/td>\n<td>497<\/td>\n<td>0.85\u00d7<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>The same Japanese text takes about as many tokens as with the Qwen3 tokenizer.<\/p>\n<p>Counted with the <code>tokenizer.json<\/code> of <a href=\"https:\/\/huggingface.co\/NaiveAI\/Naive-N0.5-Flash\">NaiveAI\/Naive-N0.5-Flash<\/a> on a fixed text we wrote ourselves (876 Japanese characters across news, conversation, technical docs, a formal email, travel writing and a recipe) and its English translation. Fewer tokens mean more Japanese fits in the context window.<\/p>\n<p><!-- \/lmw:lab --><\/p>\n<h2>How to Get It<\/h2>\n<ul>\n<li>Distribution format: BF16 safetensors (49 files) at Hugging Face&#8217;s <code>NaiveAI\/Naive-N0.5-Flash<\/code>, and an FP8 version at <code>NaiveAI\/Naive-N0.5-Flash-FP8<\/code>. Downloading does not require agreeing to terms of use<\/li>\n<li>Example download command: <code>huggingface-cli download NaiveAI\/Naive-N0.5-Flash-FP8<\/code><\/li>\n<li>Requirements: The model card states that &#8220;an NVIDIA GPU supporting FP8 is required,&#8221; with the FP8 weights alone taking up about 315GB, and inference requiring additional GPU memory<\/li>\n<li>Usage method: The model card example loads the FP8 version using Transformers (5.17.0 or higher) with <code>trust_remote_code=True<\/code>. Because the architecture is custom (<code>NaiveN05FlashForCausalLM<\/code>), it runs by loading the distributor&#8217;s model code. While SGLang is mentioned in the acknowledgments, startup examples for SGLang and vLLM are not provided in the model card<\/li>\n<li>Recommended sampling settings: <code>temperature=1.0<\/code>, <code>top_p=0.95<\/code><\/li>\n<li>Running on local hardware: The model card does not mention llama.cpp, Ollama, or LM Studio. Quantized versions by third parties (such as GGUF, EXL3, NVFP4) have emerged. However, for example, Baekpica&#8217;s GGUF version specifies a custom runtime (<code>ds4-dfm-rs<\/code>) in its target execution environment rather than llama.cpp, meaning it may not necessarily run in llama.cpp despite the GGUF format. At present, trying this on personal hardware is difficult in scale and form<\/li>\n<li>API: The model card states an API will be provided, with pricing indicated at $0.10 per 1M input tokens, $0.40 per 1M output tokens, and $0.01 per 1M cache read tokens<\/li>\n<\/ul>\n<p><!-- lmw:variants --><\/p>\n<h2>Quantized and Converted Variants<\/h2>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Added<\/th>\n<th>Publisher<\/th>\n<th>Format<\/th>\n<th>Repository<\/th>\n<th>Smallest VRAM tier (build, est. memory)<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>2026-10-05<\/td>\n<td>NaiveAI<\/td>\n<td>FP8<\/td>\n<td><a href=\"https:\/\/huggingface.co\/NaiveAI\/Naive-N0.5-Flash-FP8-Draft\">NaiveAI\/Naive-N0.5-Flash-FP8-Draft<\/a><\/td>\n<td>FP8 1.5GB (fits in 4GB VRAM)<\/td>\n<\/tr>\n<tr>\n<td>2026-10-05<\/td>\n<td>NaiveAI<\/td>\n<td>FP8<\/td>\n<td><a href=\"https:\/\/huggingface.co\/NaiveAI\/Naive-N0.5-Flash-FP8\">NaiveAI\/Naive-N0.5-Flash-FP8<\/a><\/td>\n<td>FP8 352.2GB (does not fit a single consumer GPU)<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<p>File sizes of each build:<\/p>\n<ul>\n<li>Available builds in NaiveAI\/Naive-N0.5-Flash-FP8-Draft: FP8 1.2GB<\/li>\n<li>Available builds in NaiveAI\/Naive-N0.5-Flash-FP8: FP8 293.5GB<\/li>\n<\/ul>\n<p>In addition, 4 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model&#8217;s publisher or established quantization maintainers.<\/p>\n<p><em>This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files. What the quantization names mean: <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-quantization-en\/\">glossary<\/a>.<\/em><\/p>\n<p><!-- \/lmw:variants --><\/p>\n<p><!-- lmw:same-task --><\/p>\n<h2>Other Models for the Same Task<\/h2>\n<p><em>Recent text generation models covered by Local Model Watch, newest first. Grouped by the task each publisher declares on Hugging Face (pipeline_tag); the smallest VRAM tier is this site&#8217;s estimate.<\/em><\/p>\n<ul>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/04\/liquidai-lfm2-5-350m-diffusion-exp\/\">LFM2.5-350M-Diffusion-Exp Text Generation Model: 4GB+ VRAM<\/a> (425M, 4GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/10\/02\/elyza-thinking-1-0-32b-33b-2\/\">ELYZA Releases ELYZA-Thinking-1.0 32B\/33B Reasoning Models<\/a> (32.1B, 80GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/orcasaq2-27b-qwen3-8-27b-quantized\/\">OrcaSAQ-2-27B Text Generation Model: Our Test Answers, 16GB+ VRAM<\/a> (27.8B, 16GB)<\/li>\n<li><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/hemmingway-1-open-27b-model-specialized-for-human-like-writing\/\">Hemmingway-1 Text Generation Model: 12GB+ VRAM, GGUF Builds<\/a> (26.9B, 12GB)<\/li>\n<\/ul>\n<p><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/models-by-task-en\/#task-text\">See all text generation models \u2192<\/a><\/p>\n<p><!-- \/lmw:same-task --><\/p>\n<p><!-- lmw:next-steps --><\/p>\n<h2>What to Read Next<\/h2>\n<ul>\n<li><strong>Find models by VRAM<\/strong> (This model does not fit even in 80GB; its smallest build needs about 690GB) \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/vram-guide-en\/\">VRAM quick reference, including models over 80GB<\/a><\/li>\n<li><strong>Explore the same model family<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/model-naiveai-naive-n0-5-en\/\">Naive-N0.5 family overview (1 articles, 2 converted builds)<\/a><\/li>\n<li><strong>Formats this model is available in<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/format-fp8-en\/\">FP8 format guide and models<\/a><\/li>\n<li><strong>How to read Terminal-Bench, SWE-bench<\/strong> \u2192 <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/glossary-benchmarks-en\/\">Benchmark glossary<\/a><\/li>\n<\/ul>\n<p><!-- \/lmw:next-steps --><\/p>\n<h2>Sources<\/h2>\n<ul>\n<li><a href=\"https:\/\/huggingface.co\/NaiveAI\/Naive-N0.5-Flash\">https:\/\/huggingface.co\/NaiveAI\/Naive-N0.5-Flash<\/a><\/li>\n<li><a href=\"https:\/\/naive.ai\/en\/research\/\">https:\/\/naive.ai\/en\/research\/<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/XiaomiMiMo\/MiMo-V2.5-Base\">https:\/\/huggingface.co\/XiaomiMiMo\/MiMo-V2.5-Base<\/a><\/li>\n<li><a href=\"https:\/\/huggingface.co\/Baekpica\/Naive-N0.5-Flash-Mixed-Quant-GGUF\">https:\/\/huggingface.co\/Baekpica\/Naive-N0.5-Flash-Mixed-Quant-GGUF<\/a><\/li>\n<\/ul>\n<p><!-- lmw:updates --><\/p>\n<h2>Update History<\/h2>\n<ul>\n<li>2026-10-05: Added converted builds to \u201cQuantized and Converted Variants\u201d: NaiveAI\/Naive-N0.5-Flash-FP8-Draft, NaiveAI\/Naive-N0.5-Flash-FP8<\/li>\n<li>2026-10-06: Added our own measurements: Japanese token efficiency.<\/li>\n<\/ul>\n<p><!-- \/lmw:updates --><\/p>\n","protected":false},"excerpt":{"rendered":"<p>NaiveAI has released Naive-N0.5-Flash, a 309B MoE model with 15.5B active parameters, a 1M context window, and MIT license.<\/p>\n","protected":false},"author":1,"featured_media":9967,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[310],"tags":[3005,3007,165,3009,3011,1547,1555],"class_list":["post-9968","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-new-models","tag-mimo-v2-5-en","tag-mimo-v2-6-en","tag-moe-en","tag-naive-n0-5-flash-en","tag-naiveai-en","tag-verified","tag--en"],"lang":"en","translations":{"en":9968,"ja":9966},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9968","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=9968"}],"version-history":[{"count":4,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9968\/revisions"}],"predecessor-version":[{"id":10171,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/posts\/9968\/revisions\/10171"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media\/9967"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=9968"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/categories?post=9968"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/tags?post=9968"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}