The backbone contains 770B total parameters and activates 49B per token across 78 layers. Gated DeepSeek Sparse Attention and cross-layer index reuse reduce attention overhead, while identity Hyper-Connections expand information flow. A native multi-token prediction layer supports speculative decoding alongside a one-million-token context.
The model is useful for research and infrastructure teams evaluating large open-weight agents. It can work across code and office artifacts, and the release includes serving, fine-tuning, and quantization documentation. Tencent notes that this early version can reason too long and over-verify its work.

