Qwen3.8-2.4T-A95B

NEW

Key Features

2.4T total parameters.
95B activated parameters per token.
Mixture-of-experts architecture.
512 routed experts.
262K native context window.
Context extensible to about 1.01M tokens.
Strong coding and agentic task focus.
vLLM, SGLang, and quantized-runtime support.

The model uses 512 routed experts with ten routed experts plus one shared expert active for each token. It has a native 262K-token context window extensible to about 1.01 million tokens, and it uses gated linear attention and gated attention blocks to manage large-scale inference.


The Hugging Face release includes instructions for vLLM, SGLang, Docker Model Runner, and quantized local runtimes. It is useful for organizations with substantial inference infrastructure that need high-capacity open model weights for long-context agent workloads.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!