The model uses the standard LlamaForCausalLM architecture with 42 layers and grouped-query attention. It has roughly 2.52 billion total parameters, including embeddings, and a native 131,072-token context. Released formats and inference recipes support several common server, desktop, and Apple Silicon runtimes.
MiniCPM5-2B is useful for applications needing local execution and an inspectable training ecosystem. The release includes checkpoints at different training stages and associated UltraData datasets. Tool-calling deployments are best started from the documented SGLang configuration and its dedicated parser.

