Our Measurements: Japanese Token Efficiency, CPU Runs, GGUF Internals and Engine Support
This page collects what Local Model Watch measures and records itself, rather than what model cards say. We read the tokenizer and the GGUF header of each model we cover, run small models on our own CPU-only server, and keep daily records of Hugging Face and of each inference engine’s model registry. Everything below is computed by code from those records; no language model writes or grades these figures. Each model’s article shows the same measurements in its “Our Own Measurements” section, including the verbatim answers to our five Japanese questions (temperature 0, up to 1024 tokens).
Japanese Token Efficiency
We count how many tokens each model’s tokenizer needs for a fixed text we wrote ourselves (876 Japanese characters across six genres) and for its English translation. Fewer tokens mean more Japanese fits in the same context length, and faster generation per character. Models that share a tokenizer are listed on one row; “reference” marks well-known tokenizers we measure for comparison.
→ Scroll horizontally to see all columns
Running Models on a CPU Only
Measured on our own server: Neoverse-N1, 3 threads, no GPU, with the official llama.cpp builds (b11223). Speeds come from llama-bench (512-token prompt, 128-token generation); peak memory is measured while answering our Japanese questions with a 4,096-token context and includes the memory-mapped model file. Each model’s article shows its answers verbatim.
→ Scroll horizontally to see all columns
| Model | Quant | File | Prompt | Generation | In Japanese | Peak memory | Measured |
|---|---|---|---|---|---|---|---|
| MiniCPM5-2B | Q4_K_M |
1.45GB | 34.4 tok/s | 13.9 tok/s | 20.9 chars/s | 2.9GB | 2026-09-28 |
| Qwen3-4B (reference) | Q4_K_M |
2.33GB | 18.9 tok/s | 6.8 tok/s | — | 5.7GB | 2026-09-28 |
| olmo3-7b-sdf-sft-clean150 | Q4_K_M |
4.16GB | 11.3 tok/s | 4.9 tok/s | 4.4 chars/s | 12.4GB | 2026-09-28 |
| Qwen3-8B (reference) | Q4_K_M |
4.68GB | 10.4 tok/s | 4.6 tok/s | — | 9.5GB | 2026-09-28 |
Quality Loss by Quantization
For each model we download its quantizations one at a time and compare them with the F16 (or Q8_0) file on a fixed Japanese text — the opening of Natsume Soseki’s Botchan (public domain, from Aozora Bunko) — using llama.cpp’s KL-divergence mode (context 512 tokens). Each cell is how often the most likely next token matches the baseline. Botchan is famous and may be in the training data, so compare quantizations of the same model rather than models with each other.
→ Scroll horizontally to see all columns
| Model | Baseline | Q2_K | Q3_K_M | IQ4_XS | Q4_K_M | Q5_K_M | Q6_K |
|---|---|---|---|---|---|---|---|
| Qwen3-4B (reference) | Q8_0 |
54.1% | 72.7% | 83.0% | 83.9% | 90.3% | 91.7% |
| MiniCPM5-2B | Q8_0 |
37.4% | 63.6% | 76.9% | 81.5% | 89.1% | 90.8% |
Inside the GGUF Files
We read only the header of each GGUF file (not the weights) with HTTP range requests. Of 18 files, 12 record that they were quantized with an importance matrix (imatrix), and 18 ship a chat template that mentions tool calls. The average bits per weight is the file’s data size divided by the number of weights.
→ Scroll horizontally to see all columns
| Model | File | Avg bits | Main types | imatrix | Context | Tools in template |
|---|---|---|---|---|---|---|
| Hemmingway-1 | Q4_K_M (bartowski) |
5.10 | Q4_K 75% / Q6_K 18% | yes | 262,144 | yes |
| LensVLM-9B | Q4_K_M (bartowski) |
5.21 | Q4_K 72% / Q6_K 21% | yes | 262,144 | yes |
| Xing4.0-29B-A4B | IQ4_NL (XingChen-AGI) |
5.15 | IQ4_NL 93% / BF16 5% | yes | 262,144 | yes |
| DeepSeek-V4-Pro-0813 | Q4_K_M (DevQuasar) |
4.84 | Q4_K 84% / Q6_K 16% | no | 1,048,576 | yes |
| Qwen3.8-27B | Q4_K_M (unsloth) |
4.82 | IQ4_XS 33% / Q4_K 27% | yes | 262,144 | yes |
| plamo-3-610m-fin-instruct | Q4_K_M (mradermacher) |
5.13 | Q4_K 70% / Q6_K 30% | no | 262,144 | yes |
| Tev1-4B-experimental | Q4_K_M (bartowski) |
5.15 | Q4_K 65% / Q5_K 15% | yes | 262,144 | yes |
| olmo3-7b-sdf-sft-clean150 | Q4_K_M (EleutherAI) |
4.90 | Q4_K 81% / Q6_K 19% | no | 65,536 | yes |
| vectionlabs_Salience-27B-R6-GGUF | Q4_K_M (bartowski) |
5.10 | Q4_K 75% / Q6_K 18% | yes | 1,048,576 | yes |
| Intern-S2-397B-GGUF | Q4_K_M (bartowski) |
5.05 | Q4_K 77% / Q6_K 17% | yes | 262,144 | yes |
| TheDrummer_Orion-26B-A4B-v1.1-GGUF | Q4_K_M (bartowski) |
5.69 | Q4_K 52% / Q8_0 20% | yes | 262,144 | yes |
| Signal-3.8-27B-GGUF | Q4_K_M (agentionai) |
4.97 | Q4_K 79% / Q6_K 20% | yes | 262,144 | yes |
| Gryphe_Pantheon-Reasoning-26B-A4B-1.1-V2-GGUF | Q4_K_M (bartowski) |
5.77 | Q4_K 51% / Q8_0 22% | yes | 262,144 | yes |
| Nex-N2.5-Pro | Q4_K_M (bartowski) |
5.06 | Q4_K 78% / Q6_K 17% | yes | 262,144 | yes |
| Nex-N2.5-mini | Q4_K_M (bartowski) |
5.15 | Q4_K 76% / Q6_K 12% | yes | 262,144 | yes |
| MiniCPM5-2B | Q4_K_M (openbmb) |
4.95 | Q4_K 78% / Q6_K 22% | no | 131,072 | yes |
| K2-Horizon-MoVA-36B-A4B-GGUF | Q4_K_M (IFM) |
4.78 | Q4_K 87% / Q6_K 13% | no | 524,288 | yes |
| Qwopus3.8-27B-Flash-GGUF | Q4_K_M (Jackrong) |
4.92 | Q4_K 80% / Q6_K 20% | no | 262,144 | yes |
How Soon GGUF Versions Appear
The time from the creation of the original model’s Hugging Face repository to the creation of its first GGUF repository, for the models we covered (repository creation time, not public release: repositories are often created privately before release, and a negative value means the GGUF repository was created first, e.g. with early access). Median over 11 models: 15.2 hours.
| Who made the first GGUF | Models | Median hours |
|---|---|---|
| bartowski | 4 | 32.0 |
| mradermacher | 2 | 23.5 |
| The original publisher | 2 | -13.4 |
| unsloth | 2 | 101.3 |
| LiquidAI | 1 | 0.0 |
How Soon Inference Engines Register New Architectures
Every day we fetch each engine’s model registry from its source code and note the first day each architecture name appears (tracking since 2026-09-18). For the models we covered, we compare that day with the day the model’s repository was created. “Not registered” means the name is not in the registry today, not that the model cannot run.
→ Scroll horizontally to see all columns
| Engine | Registered when the model appeared | Registered later | Not registered | Already registered before tracking began |
|---|---|---|---|---|
| llama.cpp | 9 | 0 | 4 | 8 |
| vLLM | 10 | 0 | 2 | 9 |
| MLX (mlx-lm) | 6 | 0 | 9 | 6 |
How Often Inference Engines Release
Releases recorded from each engine’s GitHub repository in the 30 days up to 2026-09-28.
| Project | Releases | Median days between releases | Latest |
|---|---|---|---|
| Unsloth | 13 | 1.1 | 2026-09-28 |
| ExLlamaV3 | 9 | 3.3 | 2026-09-28 |
| ggml | 6 | 0.8 | 2026-09-25 |
| Ollama | 6 | 4.0 | 2026-09-23 |
| ComfyUI | 3 | 5.7 | 2026-09-21 |
| KoboldCpp | 3 | 5.4 | 2026-09-26 |
| llama.cpp | 3 | 9.5 | 2026-09-24 |
| Open WebUI | 3 | 10.8 | 2026-09-22 |
| InvokeAI | 2 | 21.1 | 2026-09-28 |
| SGLang | 2 | 13.8 | 2026-09-19 |
| vLLM | 2 | 12.9 | 2026-09-22 |
| RamaLama | 1 | — | 2026-09-25 |
| whisper.cpp | 1 | — | 2026-09-11 |
| llamafile | 1 | — | 2026-09-16 |
| LocalAI | 1 | — | 2026-09-18 |
| TTS WebUI | 1 | — | 2026-09-01 |