The model uses a single-stream design with a 40-layer backbone, a sparse middle made of multi-head mixture-of-experts layers, and a smaller activated expert set. The training system focuses on communication, kernels, routing stability, optimization, and memory efficiency.
MAGI-2 is intended for researchers investigating large-scale generative video, physical plausibility, long-term consistency, and audio-video synchronization. The preview is explicitly a research milestone, so its results should be treated as evidence about scaling direction rather than a finished commercial product.

