{"id":6735,"date":"2026-09-28T17:30:34","date_gmt":"2026-09-28T08:30:34","guid":{"rendered":"https:\/\/localmodelwatch.tsuchitsuchi.com\/observations-en\/"},"modified":"2026-09-29T01:30:02","modified_gmt":"2026-09-28T16:30:02","slug":"observations-en","status":"publish","type":"page","link":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/observations-en\/","title":{"rendered":"Our Measurements: Japanese Token Efficiency, CPU Runs, GGUF Internals and Engine Support"},"content":{"rendered":"<p>This page collects what Local Model Watch measures and records itself, rather than what model cards say. We read the tokenizer and the GGUF header of each model we cover, run small models on our own CPU-only server, and keep daily records of Hugging Face and of each inference engine&#8217;s model registry. Everything below is computed by code from those records; no language model writes or grades these figures. Each model&#8217;s article shows the same measurements in its \u201cOur Own Measurements\u201d section, including the verbatim answers to our five Japanese questions (temperature 0, up to 1024 tokens).<\/p>\n<h2>Japanese Token Efficiency<\/h2>\n<p>We count how many tokens each model&#8217;s tokenizer needs for a fixed text we wrote ourselves (876 Japanese characters across six genres) and for its English translation. Fewer tokens mean more Japanese fits in the same context length, and faster generation per character. Models that share a tokenizer are listed on one row; \u201creference\u201d marks well-known tokenizers we measure for comparison.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>#<\/th>\n<th>Tokens per 1,000 Japanese chars<\/th>\n<th>Ratio to English<\/th>\n<th>Vocabulary<\/th>\n<th>Models using this tokenizer<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>1<\/td>\n<td>497<\/td>\n<td>0.85\u00d7<\/td>\n<td>99,574<\/td>\n<td>LLM-jp-3 (reference)<\/td>\n<\/tr>\n<tr>\n<td>2<\/td>\n<td>546<\/td>\n<td>0.99\u00d7<\/td>\n<td>248,070\u301c248,077<\/td>\n<td>Qwopus3.8-27B-Flash-GGUF, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-n25-mini-released\/\">Nex-N2.5-mini<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-agi-announces-nex-n25-pro-agent-model\/\">Nex-N2.5-Pro<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/edge0-35b-a3b-preview-sparse-moe\/\">Edge0-35B-A3B-preview<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B-GGUF<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">vectionlabs_Salience-27B-R6-GGUF<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/18\/ternary-bonsai-2-27b-gguf\/\">Ternary-Bonsai-2-27B-gguf<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-distill-qwen-9b-gguf\/\">MiMo-V2.6-Distill-Qwen-9B-GGUF<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/tev1-4b-experimental-released\/\">Tev1-4B-experimental<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/tev1-08b-experimental-2\/\">Tev1-0.8B-experimental<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/apple-lensvlm-9b-2\/\">LensVLM-9B<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/hemmingway-1-open-27b-model-specialized-for-human-like-writing\/\">Hemmingway-1<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/orcasaq2-27b-qwen3-8-27b-quantized\/\">OrcaSAQ-2-27B<\/a><\/td>\n<\/tr>\n<tr>\n<td>3<\/td>\n<td>564<\/td>\n<td>1.03\u00d7<\/td>\n<td>262,144\u301c262,145<\/td>\n<td>Gemma 3 (reference), <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/pantheon-reasoning-26b-gguf\/\">Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/\">TheDrummer_Orion-26B-A4B-v1.1-GGUF<\/a><\/td>\n<\/tr>\n<tr>\n<td>4<\/td>\n<td>565<\/td>\n<td>1.04\u00d7<\/td>\n<td>250,624<\/td>\n<td>K2-Horizon-MoVA-36B-A4B-GGUF, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-375b-a23b-nvfp4-released-2\/\">K2-Horizon-375B-A23B-NVFP4<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/23\/ifm-k2-horizon-32b-nvfp4-released\/\">K2-Horizon-32B-NVFP4<\/a><\/td>\n<\/tr>\n<tr>\n<td>5<\/td>\n<td>642<\/td>\n<td>1.15\u00d7<\/td>\n<td>125,017<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/liquidai-lfm2-5-vl-dspark-released\/\">LFM2.5-VL-3B-DSpark<\/a><\/td>\n<\/tr>\n<tr>\n<td>6<\/td>\n<td>666<\/td>\n<td>1.21\u00d7<\/td>\n<td>130,560<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">MiniCPM5-2B<\/a><\/td>\n<\/tr>\n<tr>\n<td>7<\/td>\n<td>688<\/td>\n<td>1.26\u00d7<\/td>\n<td>151,665\u301c151,675<\/td>\n<td>Qwen3 (reference), <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/tencent-simple-attention-sparsification-qwen3\/\">Simple-Attention-Sparsification<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/17\/qwen-25-1b-rlcd-mlx-constrained-decoding\/\">Qwen-2.5-1B-RLCD<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/22\/mimo-v2-6-flash-rl-gguf-released\/\">MiMo-V2.6-Flash-RL-GGUF<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/27\/mimo-v2-6-pro-rl-released\/\">MiMo-V2.6-Pro-RL<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/xiaomi-mimo-v2-6-mopd-models\/\">MiMo-V2.6-Flash-MOPD<\/a><\/td>\n<\/tr>\n<tr>\n<td>8<\/td>\n<td>705<\/td>\n<td>1.29\u00d7<\/td>\n<td>129,280<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/deepseek-v4-1-flash-released\/\">DeepSeek-V4.1-Flash<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/10\/nvidia-releases-deepseek-v4-pro-nvfp4\/\">DeepSeek-V4-Pro-0813-nvfp4-DSpark<\/a>, <a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/deepseek-v4-pro-0813-released\/\">DeepSeek-V4-Pro-0813<\/a><\/td>\n<\/tr>\n<tr>\n<td>9<\/td>\n<td>744<\/td>\n<td>1.36\u00d7<\/td>\n<td>128,256<\/td>\n<td>Llama 3.2 (reference)<\/td>\n<\/tr>\n<tr>\n<td>10<\/td>\n<td>795<\/td>\n<td>1.45\u00d7<\/td>\n<td>200,019<\/td>\n<td>gpt-oss (reference)<\/td>\n<\/tr>\n<tr>\n<td>11<\/td>\n<td>1,119<\/td>\n<td>2.05\u00d7<\/td>\n<td>100,278<\/td>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/eleutherai-olmo-3-7b-reward-hacking-models\/\">olmo3-7b-sdf-sft-clean150<\/a><\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Running Models on a CPU Only<\/h2>\n<p>Measured on our own server: Neoverse-N1, 3 threads, no GPU, with the official llama.cpp builds (b11223). Speeds come from llama-bench (512-token prompt, 128-token generation); peak memory is measured while answering our Japanese questions with a 4,096-token context and includes the memory-mapped model file. Each model&#8217;s article shows its answers verbatim.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Quant<\/th>\n<th>File<\/th>\n<th>Prompt<\/th>\n<th>Generation<\/th>\n<th>In Japanese<\/th>\n<th>Peak memory<\/th>\n<th>Measured<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">MiniCPM5-2B<\/a><\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>1.45GB<\/td>\n<td>34.4 tok\/s<\/td>\n<td>13.9 tok\/s<\/td>\n<td>20.9 chars\/s<\/td>\n<td>2.9GB<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<tr>\n<td>Qwen3-4B (reference)<\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>2.33GB<\/td>\n<td>18.9 tok\/s<\/td>\n<td>6.8 tok\/s<\/td>\n<td>\u2014<\/td>\n<td>5.7GB<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/eleutherai-olmo-3-7b-reward-hacking-models\/\">olmo3-7b-sdf-sft-clean150<\/a><\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>4.16GB<\/td>\n<td>11.3 tok\/s<\/td>\n<td>4.9 tok\/s<\/td>\n<td>4.4 chars\/s<\/td>\n<td>12.4GB<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<tr>\n<td>Qwen3-8B (reference)<\/td>\n<td><code>Q4_K_M<\/code><\/td>\n<td>4.68GB<\/td>\n<td>10.4 tok\/s<\/td>\n<td>4.6 tok\/s<\/td>\n<td>\u2014<\/td>\n<td>9.5GB<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Quality Loss by Quantization<\/h2>\n<p>For each model we download its quantizations one at a time and compare them with the F16 (or Q8_0) file on a fixed Japanese text \u2014 the opening of Natsume Soseki&#8217;s <em>Botchan<\/em> (public domain, from Aozora Bunko) \u2014 using llama.cpp&#8217;s KL-divergence mode (context 512 tokens). Each cell is how often the most likely next token matches the baseline. <em>Botchan<\/em> is famous and may be in the training data, so compare quantizations of the same model rather than models with each other.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>Baseline<\/th>\n<th>Q2_K<\/th>\n<th>Q3_K_M<\/th>\n<th>IQ4_XS<\/th>\n<th>Q4_K_M<\/th>\n<th>Q5_K_M<\/th>\n<th>Q6_K<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>Qwen3-4B (reference)<\/td>\n<td><code>Q8_0<\/code><\/td>\n<td>54.1%<\/td>\n<td>72.7%<\/td>\n<td>83.0%<\/td>\n<td>83.9%<\/td>\n<td>90.3%<\/td>\n<td>91.7%<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">MiniCPM5-2B<\/a><\/td>\n<td><code>Q8_0<\/code><\/td>\n<td>37.4%<\/td>\n<td>63.6%<\/td>\n<td>76.9%<\/td>\n<td>81.5%<\/td>\n<td>89.1%<\/td>\n<td>90.8%<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>Inside the GGUF Files<\/h2>\n<p>We read only the header of each GGUF file (not the weights) with HTTP range requests. Of 18 files, 12 record that they were quantized with an importance matrix (imatrix), and 18 ship a chat template that mentions tool calls. The average bits per weight is the file&#8217;s data size divided by the number of weights.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>File<\/th>\n<th>Avg bits<\/th>\n<th>Main types<\/th>\n<th>imatrix<\/th>\n<th>Context<\/th>\n<th>Tools in template<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/hemmingway-1-open-27b-model-specialized-for-human-like-writing\/\">Hemmingway-1<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.10<\/td>\n<td>Q4_K 75% \/ Q6_K 18%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/apple-lensvlm-9b-2\/\">LensVLM-9B<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.21<\/td>\n<td>Q4_K 72% \/ Q6_K 21%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/xing4-0-29b-a4b-china-telecom\/\">Xing4.0-29B-A4B<\/a><\/td>\n<td><code>IQ4_NL<\/code> (XingChen-AGI)<\/td>\n<td>5.15<\/td>\n<td>IQ4_NL 93% \/ BF16 5%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/deepseek-v4-pro-0813-released\/\">DeepSeek-V4-Pro-0813<\/a><\/td>\n<td><code>Q4_K_M<\/code> (DevQuasar)<\/td>\n<td>4.84<\/td>\n<td>Q4_K 84% \/ Q6_K 16%<\/td>\n<td>no<\/td>\n<td>1,048,576<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B<\/a><\/td>\n<td><code>Q4_K_M<\/code> (unsloth)<\/td>\n<td>4.82<\/td>\n<td>IQ4_XS 33% \/ Q4_K 27%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/plamo-3-610m-fin-instruct-2\/\">plamo-3-610m-fin-instruct<\/a><\/td>\n<td><code>Q4_K_M<\/code> (mradermacher)<\/td>\n<td>5.13<\/td>\n<td>Q4_K 70% \/ Q6_K 30%<\/td>\n<td>no<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/tev1-4b-experimental-released\/\">Tev1-4B-experimental<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.15<\/td>\n<td>Q4_K 65% \/ Q5_K 15%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/16\/eleutherai-olmo-3-7b-reward-hacking-models\/\">olmo3-7b-sdf-sft-clean150<\/a><\/td>\n<td><code>Q4_K_M<\/code> (EleutherAI)<\/td>\n<td>4.90<\/td>\n<td>Q4_K 81% \/ Q6_K 19%<\/td>\n<td>no<\/td>\n<td>65,536<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/salience-27b-r6-gguf-2\/\">vectionlabs_Salience-27B-R6-GGUF<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.10<\/td>\n<td>Q4_K 75% \/ Q6_K 18%<\/td>\n<td>yes<\/td>\n<td>1,048,576<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/15\/intern-s2-397b-gguf-released\/\">Intern-S2-397B-GGUF<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.05<\/td>\n<td>Q4_K 77% \/ Q6_K 17%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/14\/orion-26b-a4b-v1-1-gguf-quantizations\/\">TheDrummer_Orion-26B-A4B-v1.1-GGUF<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.69<\/td>\n<td>Q4_K 52% \/ Q8_0 20%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/12\/signal-3-8-27b-gguf-overview\/\">Signal-3.8-27B-GGUF<\/a><\/td>\n<td><code>Q4_K_M<\/code> (agentionai)<\/td>\n<td>4.97<\/td>\n<td>Q4_K 79% \/ Q6_K 20%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/11\/pantheon-reasoning-26b-gguf\/\">Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.77<\/td>\n<td>Q4_K 51% \/ Q8_0 22%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-agi-announces-nex-n25-pro-agent-model\/\">Nex-N2.5-Pro<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.06<\/td>\n<td>Q4_K 78% \/ Q6_K 17%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-n25-mini-released\/\">Nex-N2.5-mini<\/a><\/td>\n<td><code>Q4_K_M<\/code> (bartowski)<\/td>\n<td>5.15<\/td>\n<td>Q4_K 76% \/ Q6_K 12%<\/td>\n<td>yes<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">MiniCPM5-2B<\/a><\/td>\n<td><code>Q4_K_M<\/code> (openbmb)<\/td>\n<td>4.95<\/td>\n<td>Q4_K 78% \/ Q6_K 22%<\/td>\n<td>no<\/td>\n<td>131,072<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td>K2-Horizon-MoVA-36B-A4B-GGUF<\/td>\n<td><code>Q4_K_M<\/code> (IFM)<\/td>\n<td>4.78<\/td>\n<td>Q4_K 87% \/ Q6_K 13%<\/td>\n<td>no<\/td>\n<td>524,288<\/td>\n<td>yes<\/td>\n<\/tr>\n<tr>\n<td>Qwopus3.8-27B-Flash-GGUF<\/td>\n<td><code>Q4_K_M<\/code> (Jackrong)<\/td>\n<td>4.92<\/td>\n<td>Q4_K 80% \/ Q6_K 20%<\/td>\n<td>no<\/td>\n<td>262,144<\/td>\n<td>yes<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>How Soon GGUF Versions Appear<\/h2>\n<p>The time from the creation of the original model&#8217;s Hugging Face repository to the creation of its first GGUF repository, for the models we covered (repository creation time, not public release: repositories are often created privately before release, and a negative value means the GGUF repository was created first, e.g. with early access). Median over 11 models: <strong>15.2 hours<\/strong>.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Who made the first GGUF<\/th>\n<th>Models<\/th>\n<th>Median hours<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>bartowski<\/td>\n<td>4<\/td>\n<td>32.0<\/td>\n<\/tr>\n<tr>\n<td>mradermacher<\/td>\n<td>2<\/td>\n<td>23.5<\/td>\n<\/tr>\n<tr>\n<td>The original publisher<\/td>\n<td>2<\/td>\n<td>-13.4<\/td>\n<\/tr>\n<tr>\n<td>unsloth<\/td>\n<td>2<\/td>\n<td>101.3<\/td>\n<\/tr>\n<tr>\n<td>LiquidAI<\/td>\n<td>1<\/td>\n<td>0.0<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Model<\/th>\n<th>First GGUF<\/th>\n<th>Hours<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/hemmingway-1-open-27b-model-specialized-for-human-like-writing\/\">Hemmingway-1<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/Altworld_Hemmingway-1-GGUF\">bartowski\/Altworld_Hemmingway-1-GGUF<\/a><\/td>\n<td>3.4<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/apple-lensvlm-9b-2\/\">LensVLM-9B<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/LensVLM-9B-GGUF\">bartowski\/LensVLM-9B-GGUF<\/a><\/td>\n<td>44.4<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/28\/xing4-0-29b-a4b-china-telecom\/\">Xing4.0-29B-A4B<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/XingChen-AGI\/Xing4.0-29B-A4B-GGUF\">XingChen-AGI\/Xing4.0-29B-A4B-GGUF<\/a><\/td>\n<td>6.7<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/deepseek-v4-pro-0813-released\/\">DeepSeek-V4-Pro-0813<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/DeepSeek-V4-Pro-0813-GGUF\">unsloth\/DeepSeek-V4-Pro-0813-GGUF<\/a><\/td>\n<td>10.6<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/26\/qwen3-8-27b-overview-specs-performance\/\">Qwen3.8-27B<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/unsloth\/Qwen3.8-27B-GGUF\">unsloth\/Qwen3.8-27B-GGUF<\/a><\/td>\n<td>192.1<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/25\/liquidai-lfm2-5-vl-dspark-released\/\">LFM2.5-VL-3B-DSpark<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/LiquidAI\/LFM2.5-VL-3B-DSpark-GGUF\">LiquidAI\/LFM2.5-VL-3B-DSpark-GGUF<\/a><\/td>\n<td>0.0<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/plamo-3-610m-fin-instruct-2\/\">plamo-3-610m-fin-instruct<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/mradermacher\/plamo-3-610m-fin-instruct-GGUF\">mradermacher\/plamo-3-610m-fin-instruct-GGUF<\/a><\/td>\n<td>15.2<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/24\/tev1-4b-experimental-released\/\">Tev1-4B-experimental<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/togethercomputer_Tev1-4B-experimental-GGUF\">bartowski\/togethercomputer_Tev1-4B-experimental-GGUF<\/a><\/td>\n<td>19.6<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-agi-announces-nex-n25-pro-agent-model\/\">Nex-N2.5-Pro<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/bartowski\/Nex-N2.5-Pro-GGUF\">bartowski\/Nex-N2.5-Pro-GGUF<\/a><\/td>\n<td>101.1<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/09\/nex-n25-mini-released\/\">Nex-N2.5-mini<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/mradermacher\/Nex-N2.5-mini-GGUF\">mradermacher\/Nex-N2.5-mini-GGUF<\/a><\/td>\n<td>31.8<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/2026\/09\/08\/minicpm5-2b-released\/\">MiniCPM5-2B<\/a><\/td>\n<td><a href=\"https:\/\/huggingface.co\/openbmb\/MiniCPM5-2B-GGUF\">openbmb\/MiniCPM5-2B-GGUF<\/a><\/td>\n<td>-33.5<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>How Soon Inference Engines Register New Architectures<\/h2>\n<p>Every day we fetch each engine&#8217;s model registry from its source code and note the first day each architecture name appears (tracking since 2026-09-18). For the models we covered, we compare that day with the day the model&#8217;s repository was created. \u201cNot registered\u201d means the name is not in the registry today, not that the model cannot run.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<p class=\"lmw-table-hint\" style=\"margin:0 0 4px;font-size:0.85em;opacity:0.7;\">\u2192 Scroll horizontally to see all columns<\/p>\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Engine<\/th>\n<th>Registered when the model appeared<\/th>\n<th>Registered later<\/th>\n<th>Not registered<\/th>\n<th>Already registered before tracking began<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td>llama.cpp<\/td>\n<td>9<\/td>\n<td>0<\/td>\n<td>4<\/td>\n<td>8<\/td>\n<\/tr>\n<tr>\n<td>vLLM<\/td>\n<td>10<\/td>\n<td>0<\/td>\n<td>2<\/td>\n<td>9<\/td>\n<\/tr>\n<tr>\n<td>MLX (mlx-lm)<\/td>\n<td>6<\/td>\n<td>0<\/td>\n<td>9<\/td>\n<td>6<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n<h2>How Often Inference Engines Release<\/h2>\n<p>Releases recorded from each engine&#8217;s GitHub repository in the 30 days up to 2026-09-28.<\/p>\n<div class=\"lmw-table-scroll\" tabindex=\"0\" style=\"overflow-x:auto;-webkit-overflow-scrolling:touch;max-width:100%;\">\n<table style=\"width:max-content;min-width:100%;border-collapse:collapse;\">\n<thead>\n<tr>\n<th>Project<\/th>\n<th>Releases<\/th>\n<th>Median days between releases<\/th>\n<th>Latest<\/th>\n<\/tr>\n<\/thead>\n<tbody>\n<tr>\n<td><a href=\"https:\/\/github.com\/unslothai\/unsloth\/releases\">Unsloth<\/a><\/td>\n<td>13<\/td>\n<td>1.1<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/turboderp-org\/exllamav3\/releases\">ExLlamaV3<\/a><\/td>\n<td>9<\/td>\n<td>3.3<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/ggml-org\/ggml\/releases\">ggml<\/a><\/td>\n<td>6<\/td>\n<td>0.8<\/td>\n<td>2026-09-25<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/ollama\/ollama\/releases\">Ollama<\/a><\/td>\n<td>6<\/td>\n<td>4.0<\/td>\n<td>2026-09-23<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/Comfy-Org\/ComfyUI\/releases\">ComfyUI<\/a><\/td>\n<td>3<\/td>\n<td>5.7<\/td>\n<td>2026-09-21<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/LostRuins\/koboldcpp\/releases\">KoboldCpp<\/a><\/td>\n<td>3<\/td>\n<td>5.4<\/td>\n<td>2026-09-26<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/ggml-org\/llama.cpp\/releases\">llama.cpp<\/a><\/td>\n<td>3<\/td>\n<td>9.5<\/td>\n<td>2026-09-24<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/open-webui\/open-webui\/releases\">Open WebUI<\/a><\/td>\n<td>3<\/td>\n<td>10.8<\/td>\n<td>2026-09-22<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/invoke-ai\/InvokeAI\/releases\">InvokeAI<\/a><\/td>\n<td>2<\/td>\n<td>21.1<\/td>\n<td>2026-09-28<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/sgl-project\/sglang\/releases\">SGLang<\/a><\/td>\n<td>2<\/td>\n<td>13.8<\/td>\n<td>2026-09-19<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/vllm-project\/vllm\/releases\">vLLM<\/a><\/td>\n<td>2<\/td>\n<td>12.9<\/td>\n<td>2026-09-22<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/containers\/ramalama\/releases\">RamaLama<\/a><\/td>\n<td>1<\/td>\n<td>\u2014<\/td>\n<td>2026-09-25<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/ggml-org\/whisper.cpp\/releases\">whisper.cpp<\/a><\/td>\n<td>1<\/td>\n<td>\u2014<\/td>\n<td>2026-09-11<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/mozilla-ai\/llamafile\/releases\">llamafile<\/a><\/td>\n<td>1<\/td>\n<td>\u2014<\/td>\n<td>2026-09-16<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/mudler\/LocalAI\/releases\">LocalAI<\/a><\/td>\n<td>1<\/td>\n<td>\u2014<\/td>\n<td>2026-09-18<\/td>\n<\/tr>\n<tr>\n<td><a href=\"https:\/\/github.com\/rsxdalv\/tts-webui\/releases\">TTS WebUI<\/a><\/td>\n<td>1<\/td>\n<td>\u2014<\/td>\n<td>2026-09-01<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<\/div>\n","protected":false},"excerpt":{"rendered":"<p>This page collects what Local Model Watch measures and records itself, rather than what model cards say. We read the tokenizer and the GGUF header of each model we cover, run small models on our own CPU-only server, and keep daily records of Hugging Face and of each inference engine&#8217;s model registry. Everything below is [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"parent":0,"menu_order":0,"comment_status":"closed","ping_status":"closed","template":"","meta":{"footnotes":""},"class_list":["post-6735","page","type-page","status-publish","hentry"],"lang":"en","translations":{"en":6735,"ja":6734},"_links":{"self":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/6735","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages"}],"about":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/types\/page"}],"author":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/comments?post=6735"}],"version-history":[{"count":3,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/6735\/revisions"}],"predecessor-version":[{"id":6867,"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/pages\/6735\/revisions\/6867"}],"wp:attachment":[{"href":"https:\/\/localmodelwatch.tsuchitsuchi.com\/en\/wp-json\/wp\/v2\/media?parent=6735"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}