Qwen3.8-Flash-Next

NEW

Key Features

The main model has 125B parameters plus 51B n-gram embedding parameters.
Approximately 6B parameters are active per token.
Gated DeltaNet is combined with Qwen Sparse Attention.
Gated Residual expands information flow into four dynamically controlled branches.
Local-context lookup tables add capacity with relatively little extra computation.
The open model supports 262,144 tokens natively.
YaRN extends the supported context to one million tokens.
QwenCloud serves the production version as Qwen3.8-Flash.

The architecture combines Gated DeltaNet with Qwen Sparse Attention, four-branch gated residual streams, and context-indexed n-gram embeddings. A 125B-parameter main model is supplemented by 51B embedding parameters, with 6B activated per token. Native context is 262,144 tokens and can extend to one million using YaRN.


The release is useful for researchers exploring capacity and serving efficiency in long-context multimodal models. Public checkpoints are available on Hugging Face and ModelScope. A production version is hosted as Qwen3.8-Flash with a default one-million-token context and built-in tools, under separate usage-based pricing.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!