Ling-3.0-flash-base

NEW

Key Features

124B total-parameter base model.
5.1B non-embedding parameters activated.
Highly sparse 1/64 MoE architecture.
512 routed experts.
Eight routed plus one shared active expert.
Hybrid KDA and Gated MLA attention.
Warmup-Stable and Merge training.
MIT-licensed open checkpoint.

The architecture uses 512 routed experts, eight routed experts plus one shared expert per token, and native hybrid linear attention combining KDA with Gated MLA. Its Warmup-Stable and Merge training approach replaces conventional learning-rate decay with weighted checkpoint merging, supporting continual pretraining and dynamic data expansion.


This checkpoint is intentionally a base model rather than a ready-made chat assistant. It is useful for researchers and teams building domain-adapted models, testing long-context MoE systems, exploring continued pretraining, or scaling strategies from Ling-3.0-tiny-base to the larger flash model.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!