The model is designed to preserve expressive intent across songs of up to five minutes, render instruments with physical realism, and produce vocals that sound performed. Its pipeline uses fine-grained temporal descriptions, a global-local Hybrid-LM, residual vector quantization, hidden-state fusion, flow matching, and a Flow-VAE.
MiniMax presents Music 3.0 as open weights with production-oriented quality and versatility. It is useful for creators, game and film teams, prototyping, music education, and developers who need controllable full-song generation rather than isolated loops.

