Key Features

The backbone has 770B total parameters with 49B activated per token.
The backbone contains 78 layers, with MoE in all but the first.
The model supports a one-million-token context window.
Gated DSA combines sparse attention with cross-layer IndexCache reuse.
A native multi-token prediction layer is included for decoding acceleration.
The release covers engineering, office analysis, game development, and scientific research.
The model card includes vLLM, SGLang, fine-tuning, and quantization guidance.
The preview may spend too long reasoning and repeatedly verifying its work.

The backbone contains 770B total parameters and activates 49B per token across 78 layers. Gated DeepSeek Sparse Attention and cross-layer index reuse reduce attention overhead, while identity Hyper-Connections expand information flow. A native multi-token prediction layer supports speculative decoding alongside a one-million-token context.


The model is useful for research and infrastructure teams evaluating large open-weight agents. It can work across code and office artifacts, and the release includes serving, fine-tuning, and quantization documentation. Tencent notes that this early version can reason too long and over-verify its work.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!