OpenBMB Releases MiniCPM5-2B: A SOTA 2B On-Device Model

At a Glance
| Item | Value |
|---|---|
| Repository | openbmb/MiniCPM5-2B |
| Published | 2026-09-06 |
| License | apache-2.0 |
| Formats | safetensors |
| Paper | arXiv:2506.07900, arXiv:2602.09003 |
| Source type | Primary source (the publisher itself) |
Values determined by this site’s code at collection time. Dates are JST.
Overview
OpenBMB has released “MiniCPM5-2B", a new 2B-class on-device model. Following the “MiniCPM5-1B" in the same series, it adopts the standard LlamaForCausalLM architecture and is designed to deliver superior performance in code reasoning, mathematical reasoning, long-context understanding, tool use, and agent tasks. According to the model card, it achieves SOTA compared to open-source models of the same size.
Specifications
- Number of parameters: 2,516,756,480
- Non-embedding parameters: 1,981,982,720
- Number of layers: 42
- Attention heads (GQA): 16 for Q, 2 for KV
- Context length: 131,072
- License: apache-2.0
Performance
From the evaluation result comparison table published in the model card, some comparison targets have been narrowed down (MiniCPM5-2B, LFM2.5-2.6B, Qwen3.5-2B, Gemma-4-E2B-it, Qwen3.5-4B, granite-4.2-3B), and key figures are extracted and shown below.
→ Scroll horizontally to see all columns
| MiniCPM5-2B | 2B-class Models / LFM2.5-2.6B | 2B-class Models / Qwen3.5-2B | 2B-class Models / Gemma-4-E2B-it | 4B-class Models / Qwen3.5-4B | 4B-class Models / granite-4.2-3B | |
|---|---|---|---|---|---|---|
| Average | 53.9 | 33.2 | 28.0 | 24.6 | 51.1 | 42.7 |
| Code Reasoning | ||||||
| LiveCodeBench v6 | 69.1 | 42.1 | 20.2 | 42.9 | 56.4 | 58.9 |
| LCB-Pro 25Q2 (Easy) | 68.0 | 30.9 | 10.3 | 27.1 | 58.3 | 54.6 |
| LCB-Pro 25Q2 (Medium) | 17.5 | 0.0 | 0.0 | 0.0 | 7.0 | 5.3 |
| OJBench | 32.5 | 11.2 | 2.6 | 11.6 | 24.8 | 21.8 |
| SciCode (wbg) | 26.3 † | 14.2 † | 2.8 † | 20.9 † | 16.1 † | 24.9 † |
| Math Reasoning | ||||||
| AIME 2025 | 86.5 | 41.9 | 29.6 | 31.7 | 78.8 | 79.4 |
| AIME 2026 | 86.5 | 45.2 | 29.0 | 39.8 | 82.7 | 83.5 |
| HMMT Feb 2026 | 63.8 | 33.7 | 20.5 | 17.8 | 64.0 | 60.8 |
| MATH-500 | 94.6 | 89.6 | 85.8 | 85.4 | 99.0 | 97.0 |
| Instruction Following | ||||||
| IFBench | 66.3 | 59.0 | 46.0 | 25.7 | 59.0 | 73.0 |
| IFEval | 86.7 | 93.4 | 77.5 | 31.4 | 90.2 | 93.7 |
| Multi-IF | 71.8 | 76.8 | 57.1 | 40.3 | 73.6 | 75.9 |
| General Knowledge | ||||||
| MMLU-Pro | 70.8 | 65.2 | 64.3 | 56.0 | 78.0 | 65.8 |
| MMLU-Redux | 84.7 | 80.0 | 80.0 | 71.8 | 88.7 | 78.9 |
| HLE | 8.9 † | 6.2 † | 2.6 † | 4.8 † | 9.9 † | 6.6 † |
| GPQA-Diamond | 70.2 † | 55.8 † | 45.6 † | 43.3 † | 77.1 † | 55.9 † |
| SuperGPQA | 40.8 | 26.2 | 38.6 | 30.3 | 52.8 | 39.9 |
| Long Context | ||||||
| AA-LCR | 59.0 † | 5.3 † | 28.7 † | 17.0 † | 61.0 † | 24.3 † |
| NoLiMa | 68.1 | 0.7 | 17.1 | 3.9 | 43.5 | 5.1 |
| LongBenchPro | 44.8 | 23.7 | 8.2 | 42.2 | 58.4 | 34.8 |
| LongBench v2 | 43.7 | 30.3 | 24.9 | 33.2 | 47.3 | 36.0 |
| Tool Use | ||||||
| τ³-Bench Banking | 20.8 † | 7.2 † | 2.1 | 3.9 | 6.8 † | 5.6 † |
| τ²-Bench Telecom | 97.1 | 90.4 | 69.0 † | 20.8 † | 92.1 † | 40.9 |
| BFCL v4 | 66.6 | 61.1 | 43.6 | 36.6 | 56.8 | 52.2 |
| Coding Agent | ||||||
| SWE-bench Verified | 46.4 | 6.0 | 5.0 | 2.0 | 33.6 | 36.8 |
| SWE-bench Pro | 14.4 | 0.6 | 0.8 | 0.0 | 28.2 | 12.3 |
| Terminal-Bench v2.1 | 8.6 † | 4.5 † | 3.0 † | 0.4 † | 25.8 † | 13.9 † |
| Search Agent | ||||||
| BrowseComp-ZH | 43.5 | 9.8 | 18.2 | 4.7 | 39.6 | 21.1 |
| BrowseComp Top100 | 39.7 | 13.7 | 19.3 | 6.0 | 33.3 | 19.0 |
| GAIA Text-103 | 88.7 | 49.5 | 47.9 | 30.1 | 78.6 | 57.3 |
| General Agent | ||||||
| GDPval-AA v2 | 19.6 † | 4.5 | 0.0 | 0.0 | 11.7 | 0.0 † |
| Claw-Gym | 59.2 | 19.3 | 25.5 | 31.3 | 51.6 | 60.0 |
| WildClaw | 23.9 | 10.2 | 9.2 | 8.9 | 17.0 | 20.0 |
| QwenClaw | 42.9 | 19.3 | 18.2 | 14.5 | 37.1 | 36.4 |
According to measurements by the publisher, MiniCPM5-2B recorded an average score of 53.9, showing performance that exceeds 2B-class models of the same scale and some 4B-class models. In particular, it reaches high figures in coding and mathematics-related metrics such as LiveCodeBench v6 and AIME 2025. On the other hand, results show it falling behind some other models in specific instruction-following items such as IFEval. It demonstrates superiority in long-context, tool use, and agent-related tasks (such as GAIA Text-103).
Strengths and Use Cases
It is designed for local assistants, coding agents, tool-using workflows, and reasoning scenarios. It maintains a compact deployment footprint while featuring native long-context support (131,072 tokens).
Hardware Requirements
Estimated requirements (calculated by Local Model Watch) — 2.5B parameters
| Your VRAM | Quantization | File size | Est. memory needed |
|---|---|---|---|
| 4GB (laptop iGPU / phone class) | Q8_0 | 2.5GB | 3.0GB |
| 8GB (RTX 4060 / 3060 Ti, etc.) | F16 | 4.7GB | 5.6GB |
Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-09-18): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.
Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. File sizes are measured from the converted build openbmb/MiniCPM5-2B-GGUF. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.
Recent Models in the Same Size Class
Models with up to 4B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.
→ Scroll horizontally to see all columns
| Model | Parameters | Smallest VRAM tier | License | Our article |
|---|---|---|---|---|
| harshatheg/Qwen-2.5-1B-RLCD | 1.5B | 8GB | apache-2.0 | Fast Structured Generation on Apple Silicon with MLX and Qwen (2026-09-16) |
| tencent/Simple-Attention-Sparsification | 4.0B | 12GB | — | Tencent Releases Simple Attention Sparsification for Qwen3 (2026-09-14) |
How to Get It
It supports diverse backends and formats such as Transformers, vLLM, SGLang, llama.cpp, Ollama, LM Studio, MLX, ArcLight, vLLM Ascend, and LiteRT-LM. An example command for serving with SGLang is as follows:
pip install "sglang[srt]>=0.5.16"
python -m sglang.launch_server --model-path openbmb/MiniCPM5-2B --port 30000
Quantized and Converted Variants
→ Scroll horizontally to see all columns
| Added | Publisher | Format | Repository | Smallest VRAM tier (build, est. memory) |
|---|---|---|---|---|
| 2026-09-18 | openbmb | GGUF | openbmb/MiniCPM5-2B-GGUF | Q8_0 3.0GB (fits in 4GB VRAM) |
| 2026-09-18 | openbmb | MLX | openbmb/MiniCPM5-2B-MLX | MLX 1.6GB (fits in 4GB VRAM) |
| 2026-09-18 | bartowski | GGUF | bartowski/MiniCPM5-2B-GGUF | Q8_0 3.0GB (fits in 4GB VRAM) |
| 2026-09-18 | openbmb | GPTQ | openbmb/MiniCPM5-2B-GPTQ | GPTQ 2.3GB (fits in 4GB VRAM) |
| 2026-09-18 | mlx-community | MLX | mlx-community/MiniCPM5-2B-8bit | MLX 8bit 3.0GB (fits in 4GB VRAM) |
In addition, 43 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model’s publisher or established quantization maintainers.
This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files.
Related Articles
Sources
- https://huggingface.co/openbmb/MiniCPM5-2B
- https://arxiv.org/pdf/2602.09003
- https://ultradata.openbmb.cn/
- https://huggingface.co/datasets/openbmb/UltraX-Preview
- https://huggingface.co/datasets/openbmb/UltraData-Code
- https://huggingface.co/datasets/openbmb/UltraData-SFT-Agent-2609
- https://huggingface.co/datasets/openbmb/UltraData-RL-2609
- https://huggingface.co/datasets/openbmb/Ultra-FineWeb
- https://huggingface.co/datasets/openbmb/Ultra-FineWeb-L3
- https://huggingface.co/datasets/openbmb/UltraData-Math
- https://huggingface.co/datasets/openbmb/UltraData-SFT-2605
- https://flagos.io/
- https://github.com/flagos-ai/FlagGems
- https://github.com/flagos-ai/vllm-plugin-FL
- https://panhaoxuan.notion.site/justrl-ii-scaling-small-llms-to-128k-reasoning-with-a-critic
- https://huggingface.co/openbmb/MiniCPM5-1B
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-SFT
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-Midtrain
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-Base
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-GGUF
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-MLX
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-GPTQ
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-DSpark
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-2B-DSpark-GGUF
- https://huggingface.co/litert-community/MiniCPM5-2B
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-SFT
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-Base
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-GGUF
- https://www.modelscope.cn/models/OpenBMB/MiniCPM5-1B-MLX
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/transformers.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-transformers/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/vllm.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-vllm/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/sglang.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-sglang/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/llama_cpp.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-llama-cpp/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/ollama.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-ollama/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/lmstudio.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-lmstudio/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/mlx.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-mlx/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/arclight.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-arclight/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/vllm_ascend.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-vllm-ascend/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/deployment/litert.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-deploy-litert/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/finetune/trl.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-finetune-trl/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/finetune/llamafactory.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-finetune-llamafactory/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/finetune/ms_swift.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-finetune-ms-swift/SKILL.md
- https://github.com/OpenBMB/MiniCPM/blob/main/docs/finetune/unsloth.md
- https://github.com/OpenBMB/MiniCPM/blob/main/skills/minicpm5-finetune-unsloth/SKILL.md
- https://www.modelscope.cn/models/FlagRelease/MiniCPM5-2B-nvidia-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-nvidia-FlagOS
- https://www.modelscope.cn/models/FlagRelease/MiniCPM5-2B-hygon-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-hygon-FlagOS
- https://www.modelscope.cn/models/FlagRelease/MiniCPM5-2B-metax-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-metax-FlagOS
- https://www.modelscope.cn/models/FlagRelease/MiniCPM5-2B-iluvatar-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-iluvatar-FlagOS
- https://www.modelscope.cn/models/FlagRelease/MiniCPM5-2B-zhenwu-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-zhenwu-FlagOS
- https://www.modelscope.cn/models/FlagRelease/MiniCPM5-2B-mthreads-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-mthreads-FlagOS
- https://www.modelscope.cn/models/FlagRelease/MiniCPM5-2B-kunlunxin-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-kunlunxin-FlagOS
- https://modelscope.cn/models/FlagRelease/MiniCPM5-2B-ascend-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-ascend-FlagOS
- https://modelscope.cn/models/FlagRelease/MiniCPM5-2B-Armv9-FlagOS
- https://huggingface.co/FlagRelease/MiniCPM5-2B-Armv9-FlagOS

