OpenBMB Releases MiniCPM5-2B: A SOTA 2B On-Device Model

September 18, 2026

OpenBMB Releases MiniCPM5-2B, SOTA On-Device Model in 2B Class

At a Glance

Item Value
Repository openbmb/MiniCPM5-2B
Published 2026-09-06
License apache-2.0
Formats safetensors
Paper arXiv:2506.07900, arXiv:2602.09003
Source type Primary source (the publisher itself)

Values determined by this site’s code at collection time. Dates are JST.

Overview

OpenBMB has released “MiniCPM5-2B", a new 2B-class on-device model. Following the “MiniCPM5-1B" in the same series, it adopts the standard LlamaForCausalLM architecture and is designed to deliver superior performance in code reasoning, mathematical reasoning, long-context understanding, tool use, and agent tasks. According to the model card, it achieves SOTA compared to open-source models of the same size.

Specifications

  • Number of parameters: 2,516,756,480
  • Non-embedding parameters: 1,981,982,720
  • Number of layers: 42
  • Attention heads (GQA): 16 for Q, 2 for KV
  • Context length: 131,072
  • License: apache-2.0

Performance

From the evaluation result comparison table published in the model card, some comparison targets have been narrowed down (MiniCPM5-2B, LFM2.5-2.6B, Qwen3.5-2B, Gemma-4-E2B-it, Qwen3.5-4B, granite-4.2-3B), and key figures are extracted and shown below.

→ Scroll horizontally to see all columns

MiniCPM5-2B 2B-class Models / LFM2.5-2.6B 2B-class Models / Qwen3.5-2B 2B-class Models / Gemma-4-E2B-it 4B-class Models / Qwen3.5-4B 4B-class Models / granite-4.2-3B
Average 53.9 33.2 28.0 24.6 51.1 42.7
Code Reasoning
LiveCodeBench v6 69.1 42.1 20.2 42.9 56.4 58.9
LCB-Pro 25Q2 (Easy) 68.0 30.9 10.3 27.1 58.3 54.6
LCB-Pro 25Q2 (Medium) 17.5 0.0 0.0 0.0 7.0 5.3
OJBench 32.5 11.2 2.6 11.6 24.8 21.8
SciCode (wbg) 26.3 † 14.2 † 2.8 † 20.9 † 16.1 † 24.9 †
Math Reasoning
AIME 2025 86.5 41.9 29.6 31.7 78.8 79.4
AIME 2026 86.5 45.2 29.0 39.8 82.7 83.5
HMMT Feb 2026 63.8 33.7 20.5 17.8 64.0 60.8
MATH-500 94.6 89.6 85.8 85.4 99.0 97.0
Instruction Following
IFBench 66.3 59.0 46.0 25.7 59.0 73.0
IFEval 86.7 93.4 77.5 31.4 90.2 93.7
Multi-IF 71.8 76.8 57.1 40.3 73.6 75.9
General Knowledge
MMLU-Pro 70.8 65.2 64.3 56.0 78.0 65.8
MMLU-Redux 84.7 80.0 80.0 71.8 88.7 78.9
HLE 8.9 † 6.2 † 2.6 † 4.8 † 9.9 † 6.6 †
GPQA-Diamond 70.2 † 55.8 † 45.6 † 43.3 † 77.1 † 55.9 †
SuperGPQA 40.8 26.2 38.6 30.3 52.8 39.9
Long Context
AA-LCR 59.0 † 5.3 † 28.7 † 17.0 † 61.0 † 24.3 †
NoLiMa 68.1 0.7 17.1 3.9 43.5 5.1
LongBenchPro 44.8 23.7 8.2 42.2 58.4 34.8
LongBench v2 43.7 30.3 24.9 33.2 47.3 36.0
Tool Use
τ³-Bench Banking 20.8 † 7.2 † 2.1 3.9 6.8 † 5.6 †
τ²-Bench Telecom 97.1 90.4 69.0 † 20.8 † 92.1 † 40.9
BFCL v4 66.6 61.1 43.6 36.6 56.8 52.2
Coding Agent
SWE-bench Verified 46.4 6.0 5.0 2.0 33.6 36.8
SWE-bench Pro 14.4 0.6 0.8 0.0 28.2 12.3
Terminal-Bench v2.1 8.6 † 4.5 † 3.0 † 0.4 † 25.8 † 13.9 †
Search Agent
BrowseComp-ZH 43.5 9.8 18.2 4.7 39.6 21.1
BrowseComp Top100 39.7 13.7 19.3 6.0 33.3 19.0
GAIA Text-103 88.7 49.5 47.9 30.1 78.6 57.3
General Agent
GDPval-AA v2 19.6 † 4.5 0.0 0.0 11.7 0.0 †
Claw-Gym 59.2 19.3 25.5 31.3 51.6 60.0
WildClaw 23.9 10.2 9.2 8.9 17.0 20.0
QwenClaw 42.9 19.3 18.2 14.5 37.1 36.4

According to measurements by the publisher, MiniCPM5-2B recorded an average score of 53.9, showing performance that exceeds 2B-class models of the same scale and some 4B-class models. In particular, it reaches high figures in coding and mathematics-related metrics such as LiveCodeBench v6 and AIME 2025. On the other hand, results show it falling behind some other models in specific instruction-following items such as IFEval. It demonstrates superiority in long-context, tool use, and agent-related tasks (such as GAIA Text-103).

Strengths and Use Cases

It is designed for local assistants, coding agents, tool-using workflows, and reasoning scenarios. It maintains a compact deployment footprint while featuring native long-context support (131,072 tokens).

Hardware Requirements

Estimated requirements (calculated by Local Model Watch) — 2.5B parameters

Your VRAM Quantization File size Est. memory needed
4GB (laptop iGPU / phone class) Q8_0 2.5GB 3.0GB
8GB (RTX 4060 / 3060 Ti, etc.) F16 4.7GB 5.6GB

Inference engine support (architecture name matched against each project’s own model registry in its source code, checked 2026-09-18): llama.cpp: registered, vLLM: registered, MLX (mlx-lm): registered.

Memory estimates add a 20% runtime overhead (KV cache, etc.) to the actual size of the distributed files. Actual usage varies with context length, batch size and inference engine. These figures are computed by this site from file sizes, not published by the model’s authors. File sizes are measured from the converted build openbmb/MiniCPM5-2B-GGUF. Compare with other models in our VRAM quick reference. What the quantization names mean: glossary.

Recent Models in the Same Size Class

Models with up to 4B parameters that Local Model Watch covered recently, listed by code from our article log for comparison. VRAM tiers are this site’s estimates; licenses are as stated on the model cards.

→ Scroll horizontally to see all columns

Model Parameters Smallest VRAM tier License Our article
harshatheg/Qwen-2.5-1B-RLCD 1.5B 8GB apache-2.0 Fast Structured Generation on Apple Silicon with MLX and Qwen (2026-09-16)
tencent/Simple-Attention-Sparsification 4.0B 12GB Tencent Releases Simple Attention Sparsification for Qwen3 (2026-09-14)

How to Get It

It supports diverse backends and formats such as Transformers, vLLM, SGLang, llama.cpp, Ollama, LM Studio, MLX, ArcLight, vLLM Ascend, and LiteRT-LM. An example command for serving with SGLang is as follows:

pip install "sglang[srt]>=0.5.16"
python -m sglang.launch_server --model-path openbmb/MiniCPM5-2B --port 30000

Quantized and Converted Variants

→ Scroll horizontally to see all columns

Added Publisher Format Repository Smallest VRAM tier (build, est. memory)
2026-09-18 openbmb GGUF openbmb/MiniCPM5-2B-GGUF Q8_0 3.0GB (fits in 4GB VRAM)
2026-09-18 openbmb MLX openbmb/MiniCPM5-2B-MLX MLX 1.6GB (fits in 4GB VRAM)
2026-09-18 bartowski GGUF bartowski/MiniCPM5-2B-GGUF Q8_0 3.0GB (fits in 4GB VRAM)
2026-09-18 openbmb GPTQ openbmb/MiniCPM5-2B-GPTQ GPTQ 2.3GB (fits in 4GB VRAM)
2026-09-18 mlx-community MLX mlx-community/MiniCPM5-2B-8bit MLX 8bit 3.0GB (fits in 4GB VRAM)

In addition, 43 converted build(s) from other uploaders exist on Hugging Face; this site lists only builds from the model’s publisher or established quantization maintainers.

This section is appended automatically by Local Model Watch when a converted build of this model appears after publication. Memory figures are estimated from the size of the distributed files.

Related Articles

Sources