Key Features

The model natively supports 131,072 tokens of context.
The published count is 2,516,756,480 parameters including embeddings.
It uses standard LlamaForCausalLM for compatibility with mainstream runtimes.
SGLang supports its XML-style tool calls through the minicpm5 parser.
The release lists GGUF formats for llama.cpp, Ollama, and LM Studio.
MLX and four-bit variants support local Apple Silicon inference.
Base, mid-training, SFT, and final post-trained checkpoints are listed.
The release includes UltraData resources for pretraining, coding, agents, and reinforcement learning.

The model uses the standard LlamaForCausalLM architecture with 42 layers and grouped-query attention. It has roughly 2.52 billion total parameters, including embeddings, and a native 131,072-token context. Released formats and inference recipes support several common server, desktop, and Apple Silicon runtimes.


MiniCPM5-2B is useful for applications needing local execution and an inspectable training ecosystem. The release includes checkpoints at different training stages and associated UltraData datasets. Tool-calling deployments are best started from the documented SGLang configuration and its dedicated parser.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!