MiniCPM5-2B Reality Check: 46.4 SWE-bench Verified and 14.4 Pro Are Internal Runs; AA v4.3 Is Estimated
OpenBMB’s September 7 MiniCPM5-2B reports 46.4 SWE-bench Verified and 14.4 SWE-bench Pro, but both are internal reproductions. Artificial Analysis independently measured 15 on v4.2; its current v4.3 score of 14 is explicitly estimated.
What OpenBMB actually released
OpenBMB released MiniCPM5-2B on September 7, 2026 as a compact, dense, text-only reasoning model aimed at local assistants, coding agents, tool-use workflows and resource-constrained deployment. The final model is a standard LlamaForCausalLM with 2,516,756,480 total parameters, 1,981,982,720 non-embedding parameters, 42 layers, grouped-query attention with 16 query heads and 2 key/value heads, and a 131,072-token context window. The weights are published under Apache 2.0.
This identity matters because the broader MiniCPM family includes multimodal systems, but MiniCPM5-2B itself is text-only. Image capability from other MiniCPM checkpoints should not be silently attributed to this release.
OpenBMB publishes the final BF16 RL+OPD checkpoint plus SFT, mid-training and base checkpoints. It also provides GGUF builds for llama.cpp/Ollama/LM Studio, MLX 4-bit for Apple Silicon, GPTQ 4-bit, and a DSpark draft model for speculative decoding. The model card documents Transformers, vLLM, SGLang, llama.cpp, Ollama, LM Studio, MLX, ArcLight and Ascend deployment paths.
Primary source:
The 53.9 average is a vendor-selected comparison, not a universal ranking
OpenBMB reports an average score of 53.9 for MiniCPM5-2B across the benchmark table it selected, above the 51.1 average shown for Qwen3.5-4B in the same table. That is a useful first-party summary, but it should not be generalized into “MiniCPM5-2B beats every 4B model.”
The table itself shows why. MiniCPM5-2B has several strong rows, including 69.1 on LiveCodeBench v6, 86.5 on AIME 2025, 86.5 on AIME 2026, 66.6 on BFCL v4 and 97.1 on τ²-Bench Telecom. Yet Qwen3.5-4B is higher on several difficult rows, including GPQA-Diamond, SWE-bench Pro and Terminal-Bench v2.1.
OpenBMB explicitly says the table compares a chosen set of 2B-class models and lists larger models “for reference.” The correct reading is therefore: MiniCPM5-2B is unusually competitive for its size inside OpenBMB’s published comparison set, not that one vendor average establishes an across-the-board capability ranking.
SWE-bench Verified: 46.4, but internally reproduced
OpenBMB reports 46.4 on SWE-bench Verified for MiniCPM5-2B. The same table gives Qwen3.5-4B 33.6.
The provenance is important. OpenBMB marks scores imported from Artificial Analysis with a dagger and states that all other scores are reproduced internally. The MiniCPM5-2B SWE-bench Verified row has no dagger, so 46.4 is an OpenBMB/internal result rather than an independent Artificial Analysis measurement.
In the public model card reviewed for this article, OpenBMB does not pin the exact SWE-bench Verified dataset revision, task count, agent scaffold, harness version, number of retries, internet policy, raw trajectories or per-task results for that row. Those omissions do not make 46.4 false, but they limit reproducibility. A third-party rerun should therefore pin those details before treating the score as independently confirmed.
SWE-bench Pro: 14.4 is separate, and it tells a different story
OpenBMB separately reports 14.4 on SWE-bench Pro. That result is also un-daggered and therefore internally reproduced under the provenance note in the model card.
It must not be merged with the 46.4 Verified score. SWE-bench Verified and SWE-bench Pro are different evaluation suites with different task populations and difficulty profiles. The model card’s own comparison makes the distinction especially useful: MiniCPM5-2B is ahead of Qwen3.5-4B on its reported Verified row (46.4 vs 33.6) but far behind it on the reported Pro row (14.4 vs 28.2).
That split is more informative than the average headline. It suggests that the model can look very strong on one software-engineering suite while still having a substantial gap on another. Until a version-pinned independent reproduction exists, the safest description is promising vendor coding-agent evidence with a clear SWE-bench Pro weakness in the same vendor table.
Terminal-Bench is independently sourced, but it is the older v2.1
The model card lists 8.6 on Terminal-Bench v2.1 for MiniCPM5-2B and marks it with the dagger used for official Artificial Analysis results. Artificial Analysis independently reported the result rounded to 9% in its September 7 release analysis.
This is materially weaker than Qwen3.5-4B’s 25.8 in OpenBMB’s table and shows that MiniCPM5-2B’s strengths are not uniform across long-horizon terminal work.
There is also a benchmark-version issue. Artificial Analysis moved its Intelligence Index from Terminal-Bench v2.1 to Terminal-Bench v4.0 on September 7. Its v4.3 methodology runs 66 Terminal-Bench 4.0 tasks three times and reports average pass@1. MiniCPM5-2B’s current Artificial Analysis model page does not show a completed independent v4.3 evaluation yet. So the old 8.6/9% v2.1 result should not be silently relabeled as Terminal-Bench 4.0 performance.
Independent sources:
- Artificial Analysis release analysis for MiniCPM5-2B
- Artificial Analysis Intelligence Index v4.3 methodology
Artificial Analysis measured 15 on v4.2; the current v4.3 number is an estimate
Artificial Analysis independently measured 15 on Intelligence Index v4.2 for MiniCPM5-2B. At release, that was the highest measured score Artificial Analysis reported for an open-weight model below 4B total parameters. The same evaluation reported GDPval-AA v2 Elo 831, AA-Briefcase Elo 438, roughly 21% on τ³-Banking, 26% on SciCode, 9% on Humanity’s Last Exam, 59% on AA-LCR v1.1, and roughly 9% on Terminal-Bench v2.1.
Artificial Analysis also reported that MiniCPM5-2B used about 19,000 output tokens per Intelligence Index task, including about 11,000 reasoning tokens. That is a useful efficiency signal within that benchmark setup, but it is not wall-clock latency and should not be converted into tokens-per-second or local GPU throughput.
The current Artificial Analysis model page now shows 14 on Intelligence Index v4.3, but it explicitly labels this number “Estimate (independent evaluation forthcoming).” That distinction matters. The page currently shows output speed as N/A and cost per Intelligence Index task as N/A, so there is no completed v4.3 speed/cost measurement to publish as if it were measured.
Artificial Analysis also warns against treating version changes as model regressions. It previously recorded MiniCPM5-2B at 23 on v4.1.1 and 15 on v4.2, while explaining that the evaluation mix and weights changed and the two scores are not directly comparable. v4.3 changes the mix again: it replaces τ³-Banking with AutomationBench-AA and Terminal-Bench v2.1 with Terminal-Bench v4.0. The current 14 estimate therefore should not be read as evidence that the model “lost one point” after release.
Current model page:
What v4.3 changes
Artificial Analysis v4.3 keeps category weights at Agents 30%, Coding 20%, General 30% and Scientific Reasoning 20%, but changes two important component evaluations.
Terminal-Bench v4.0 replaces v2.1. Artificial Analysis says it runs all 66 tasks three times and reports average pass@1.
AutomationBench-AA replaces τ³-Banking. It uses a private held-out set of 657 tasks from Zapier’s v1.0.6 benchmark across Finance, HR, Marketing, Operations, Sales and Support. Artificial Analysis scores the share of objectives completed, with a guardrail violation reducing a task’s score to zero. Because AutomationBench-AA is held out, evaluations with private questions or answers now account for 45% of the v4.3 index weight.
These changes are exactly why MiniCPM5-2B’s measured v4.2 score and estimated v4.3 score should be labeled with their benchmark versions.
Training openness is unusually substantial for a small release
OpenBMB is not only publishing final weights. The model card links the base, mid-training and SFT checkpoints and describes a three-stage training recipe covering base training, mid-training and post-training.
OpenBMB says post-training starts with 400B tokens of deep-thinking SFT, followed by specialized RL teachers for mathematics, code, agentic tasks, writing and other domains, then On-Policy Distillation. It says OPD merges capabilities from 16 expert models, including five agentic experts.
OpenBMB also publishes or links the underlying UltraData resources. The release notes describe 500K agent training samples in UltraData-SFT-Agent-2609 and 80K+ RL samples in UltraData-RL-2609.
Training/data sources:
These are first-party descriptions of the training program. They improve auditability, but they do not independently validate the benchmark outcomes.
Price, latency and practical local access
MiniCPM5-2B is open-weight and designed for self-hosting, so there is no single universal first-party per-token tariff comparable with a proprietary API price. The current Artificial Analysis page displays zero-dollar token fields for the open-weight model while simultaneously showing N/A for cost per Intelligence Index task. Those fields should not be interpreted as zero operating cost: hardware, electricity, hosting, quantization and runtime choices still determine deployment cost.
The official release provides several practical deployment options, including BF16, GGUF, MLX 4-bit, GPTQ 4-bit and DSpark-assisted decoding. That makes it unusually accessible for local experimentation.
However, I did not accept a controlled, version-pinned independent latency or throughput benchmark for the exact final checkpoint in this bounded review. A claim such as “X tokens per second on a 4090/Mac” would require the exact quantization, context length, batch/concurrency, prompt length, output length, backend, software version and hardware. Until that evidence is pinned, latency remains workload- and hardware-dependent rather than a universal model property.
Multimodal capability: do not inherit it from the MiniCPM family
The current MiniCPM5-2B model is explicitly text input and text output. Artificial Analysis likewise lists it as text-only. It does not process images.
This is worth saying because MiniCPM is a broad family with well-known vision-language releases. A model-family reputation is not evidence for the exact checkpoint. For MiniCPM5-2B, multimodal benchmark scores are therefore not reported/not applicable, rather than silently borrowed from another MiniCPM model.
Public feedback is early and self-selected
The September 7 r/LocalLLaMA release thread shows immediate interest in the model’s small footprint. One commenter proposed placing it between speech recognition and text-to-speech in a local voice pipeline; another expressed disappointment that this exact checkpoint has no vision input. The thread also contains general enthusiasm around the parameter-to-capability tradeoff.
That discussion is useful for identifying likely real-world use cases, but it is not controlled evidence. Participants choose whether to post, hardware and prompts are not standardized, and early reactions can overrepresent enthusiasts.
I did not accept a stable, directly attributable OpenBMB X status containing additional reproducible benchmark or deployment evidence in this review. No X quote or “community consensus” is invented.
Practical verdict
MiniCPM5-2B is a notable small open-weight release because it combines a 2.52B-parameter dense architecture, 131K context, unusually broad local deployment support and a substantial published training-data/checkpoint trail.
Its benchmark evidence needs careful labeling:
- SWE-bench Verified 46.4: OpenBMB/internal reproduction; exact harness/revision/retries/raw trajectories not pinned in the public model card.
- SWE-bench Pro 14.4: separate OpenBMB/internal reproduction; do not merge it with Verified.
- Terminal-Bench v2.1 8.6/9%: sourced from Artificial Analysis, but it is the older v2.1 benchmark.
- Artificial Analysis Intelligence Index v4.2 15: independently measured at release.
- Artificial Analysis Intelligence Index v4.3 14: currently an estimate, with the independent evaluation explicitly still forthcoming.
The most valuable next evidence is a third-party, version-pinned full SWE-bench Verified and SWE-bench Pro reproduction for the exact final MiniCPM5-2B checkpoint, including harness, suite revision, task count, retries, internet policy, raw trajectories and cost/runtime. A completed Artificial Analysis v4.3 run plus reproducible local latency tests on named hardware would then make its capability-per-dollar and capability-per-watt story much easier to judge.
This article is built from the source material below. Open the originals for full context and the latest updates.