The architecture uses 512 routed experts, eight routed experts plus one shared expert per token, and native hybrid linear attention combining KDA with Gated MLA. Its Warmup-Stable and Merge training approach replaces conventional learning-rate decay with weighted checkpoint merging, supporting continual pretraining and dynamic data expansion.
This checkpoint is intentionally a base model rather than a ready-made chat assistant. It is useful for researchers and teams building domain-adapted models, testing long-context MoE systems, exploring continued pretraining, or scaling strategies from Ling-3.0-tiny-base to the larger flash model.

