Our Measurements: Japanese Token Efficiency, CPU Runs, GGUF Internals and Engine Support

September 29, 2026

This page collects what Local Model Watch measures and records itself, rather than what model cards say. We read the tokenizer and the GGUF header of each model we cover, run small models on our own CPU-only server, and keep daily records of Hugging Face and of each inference engine’s model registry. Everything below is computed by code from those records; no language model writes or grades these figures. Each model’s article shows the same measurements in its “Our Own Measurements” section, including the verbatim answers to our five Japanese questions (temperature 0, up to 1024 tokens).

Japanese Token Efficiency

We count how many tokens each model’s tokenizer needs for a fixed text we wrote ourselves (876 Japanese characters across six genres) and for its English translation. Fewer tokens mean more Japanese fits in the same context length, and faster generation per character. Models that share a tokenizer are listed on one row; “reference” marks well-known tokenizers we measure for comparison.

→ Scroll horizontally to see all columns

# Tokens per 1,000 Japanese chars Ratio to English Vocabulary Models using this tokenizer
1 497 0.85× 99,574 LLM-jp-3 (reference)
2 546 0.99× 248,070〜248,077 Qwopus3.8-27B-Flash-GGUF, Nex-N2.5-mini, Nex-N2.5-Pro, Edge0-35B-A3B-preview, Signal-3.8-27B-GGUF, Intern-S2-397B-GGUF, vectionlabs_Salience-27B-R6-GGUF, Ternary-Bonsai-2-27B-gguf, MiMo-V2.6-Distill-Qwen-9B-GGUF, Tev1-4B-experimental, Tev1-0.8B-experimental, Qwen3.8-27B, LensVLM-9B, Hemmingway-1, OrcaSAQ-2-27B
3 564 1.03× 262,144〜262,145 Gemma 3 (reference), Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF, TheDrummer_Orion-26B-A4B-v1.1-GGUF
4 565 1.04× 250,624 K2-Horizon-MoVA-36B-A4B-GGUF, K2-Horizon-375B-A23B-NVFP4, K2-Horizon-32B-NVFP4
5 642 1.15× 125,017 LFM2.5-VL-3B-DSpark
6 666 1.21× 130,560 MiniCPM5-2B
7 688 1.26× 151,665〜151,675 Qwen3 (reference), Simple-Attention-Sparsification, Qwen-2.5-1B-RLCD, MiMo-V2.6-Flash-RL-GGUF, MiMo-V2.6-Pro-RL, MiMo-V2.6-Flash-MOPD
8 705 1.29× 129,280 DeepSeek-V4.1-Flash, DeepSeek-V4-Pro-0813-nvfp4-DSpark, DeepSeek-V4-Pro-0813
9 744 1.36× 128,256 Llama 3.2 (reference)
10 795 1.45× 200,019 gpt-oss (reference)
11 1,119 2.05× 100,278 olmo3-7b-sdf-sft-clean150

Running Models on a CPU Only

Measured on our own server: Neoverse-N1, 3 threads, no GPU, with the official llama.cpp builds (b11223). Speeds come from llama-bench (512-token prompt, 128-token generation); peak memory is measured while answering our Japanese questions with a 4,096-token context and includes the memory-mapped model file. Each model’s article shows its answers verbatim.

→ Scroll horizontally to see all columns

Model Quant File Prompt Generation In Japanese Peak memory Measured
MiniCPM5-2B Q4_K_M 1.45GB 34.4 tok/s 13.9 tok/s 20.9 chars/s 2.9GB 2026-09-28
Qwen3-4B (reference) Q4_K_M 2.33GB 18.9 tok/s 6.8 tok/s — 5.7GB 2026-09-28
olmo3-7b-sdf-sft-clean150 Q4_K_M 4.16GB 11.3 tok/s 4.9 tok/s 4.4 chars/s 12.4GB 2026-09-28
Qwen3-8B (reference) Q4_K_M 4.68GB 10.4 tok/s 4.6 tok/s — 9.5GB 2026-09-28

Quality Loss by Quantization

For each model we download its quantizations one at a time and compare them with the F16 (or Q8_0) file on a fixed Japanese text — the opening of Natsume Soseki’s Botchan (public domain, from Aozora Bunko) — using llama.cpp’s KL-divergence mode (context 512 tokens). Each cell is how often the most likely next token matches the baseline. Botchan is famous and may be in the training data, so compare quantizations of the same model rather than models with each other.

→ Scroll horizontally to see all columns

Model Baseline Q2_K Q3_K_M IQ4_XS Q4_K_M Q5_K_M Q6_K
Qwen3-4B (reference) Q8_0 54.1% 72.7% 83.0% 83.9% 90.3% 91.7%
MiniCPM5-2B Q8_0 37.4% 63.6% 76.9% 81.5% 89.1% 90.8%

Inside the GGUF Files

We read only the header of each GGUF file (not the weights) with HTTP range requests. Of 18 files, 12 record that they were quantized with an importance matrix (imatrix), and 18 ship a chat template that mentions tool calls. The average bits per weight is the file’s data size divided by the number of weights.

→ Scroll horizontally to see all columns

Model File Avg bits Main types imatrix Context Tools in template
Hemmingway-1 Q4_K_M (bartowski) 5.10 Q4_K 75% / Q6_K 18% yes 262,144 yes
LensVLM-9B Q4_K_M (bartowski) 5.21 Q4_K 72% / Q6_K 21% yes 262,144 yes
Xing4.0-29B-A4B IQ4_NL (XingChen-AGI) 5.15 IQ4_NL 93% / BF16 5% yes 262,144 yes
DeepSeek-V4-Pro-0813 Q4_K_M (DevQuasar) 4.84 Q4_K 84% / Q6_K 16% no 1,048,576 yes
Qwen3.8-27B Q4_K_M (unsloth) 4.82 IQ4_XS 33% / Q4_K 27% yes 262,144 yes
plamo-3-610m-fin-instruct Q4_K_M (mradermacher) 5.13 Q4_K 70% / Q6_K 30% no 262,144 yes
Tev1-4B-experimental Q4_K_M (bartowski) 5.15 Q4_K 65% / Q5_K 15% yes 262,144 yes
olmo3-7b-sdf-sft-clean150 Q4_K_M (EleutherAI) 4.90 Q4_K 81% / Q6_K 19% no 65,536 yes
vectionlabs_Salience-27B-R6-GGUF Q4_K_M (bartowski) 5.10 Q4_K 75% / Q6_K 18% yes 1,048,576 yes
Intern-S2-397B-GGUF Q4_K_M (bartowski) 5.05 Q4_K 77% / Q6_K 17% yes 262,144 yes
TheDrummer_Orion-26B-A4B-v1.1-GGUF Q4_K_M (bartowski) 5.69 Q4_K 52% / Q8_0 20% yes 262,144 yes
Signal-3.8-27B-GGUF Q4_K_M (agentionai) 4.97 Q4_K 79% / Q6_K 20% yes 262,144 yes
Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF Q4_K_M (bartowski) 5.77 Q4_K 51% / Q8_0 22% yes 262,144 yes
Nex-N2.5-Pro Q4_K_M (bartowski) 5.06 Q4_K 78% / Q6_K 17% yes 262,144 yes
Nex-N2.5-mini Q4_K_M (bartowski) 5.15 Q4_K 76% / Q6_K 12% yes 262,144 yes
MiniCPM5-2B Q4_K_M (openbmb) 4.95 Q4_K 78% / Q6_K 22% no 131,072 yes
K2-Horizon-MoVA-36B-A4B-GGUF Q4_K_M (IFM) 4.78 Q4_K 87% / Q6_K 13% no 524,288 yes
Qwopus3.8-27B-Flash-GGUF Q4_K_M (Jackrong) 4.92 Q4_K 80% / Q6_K 20% no 262,144 yes

How Soon GGUF Versions Appear

The time from the creation of the original model’s Hugging Face repository to the creation of its first GGUF repository, for the models we covered (repository creation time, not public release: repositories are often created privately before release, and a negative value means the GGUF repository was created first, e.g. with early access). Median over 11 models: 15.2 hours.

Who made the first GGUF Models Median hours
bartowski 4 32.0
mradermacher 2 23.5
The original publisher 2 -13.4
unsloth 2 101.3
LiquidAI 1 0.0

How Soon Inference Engines Register New Architectures

Every day we fetch each engine’s model registry from its source code and note the first day each architecture name appears (tracking since 2026-09-18). For the models we covered, we compare that day with the day the model’s repository was created. “Not registered” means the name is not in the registry today, not that the model cannot run.

→ Scroll horizontally to see all columns

Engine Registered when the model appeared Registered later Not registered Already registered before tracking began
llama.cpp 9 0 4 8
vLLM 10 0 2 9
MLX (mlx-lm) 6 0 9 6

How Often Inference Engines Release

Releases recorded from each engine’s GitHub repository in the 30 days up to 2026-09-28.

Project Releases Median days between releases Latest
Unsloth 13 1.1 2026-09-28
ExLlamaV3 9 3.3 2026-09-28
ggml 6 0.8 2026-09-25
Ollama 6 4.0 2026-09-23
ComfyUI 3 5.7 2026-09-21
KoboldCpp 3 5.4 2026-09-26
llama.cpp 3 9.5 2026-09-24
Open WebUI 3 10.8 2026-09-22
InvokeAI 2 21.1 2026-09-28
SGLang 2 13.8 2026-09-19
vLLM 2 12.9 2026-09-22
RamaLama 1 — 2026-09-25
whisper.cpp 1 — 2026-09-11
llamafile 1 — 2026-09-16
LocalAI 1 — 2026-09-18
TTS WebUI 1 — 2026-09-01