The architecture combines Gated DeltaNet with Qwen Sparse Attention, four-branch gated residual streams, and context-indexed n-gram embeddings. A 125B-parameter main model is supplemented by 51B embedding parameters, with 6B activated per token. Native context is 262,144 tokens and can extend to one million using YaRN.
The release is useful for researchers exploring capacity and serving efficiency in long-context multimodal models. Public checkpoints are available on Hugging Face and ModelScope. A production version is hosted as Qwen3.8-Flash with a default one-million-token context and built-in tools, under separate usage-based pricing.

