Key Features

Complete song generation from concepts and lyrics.
Composition, arrangement, performance, and production in one pass.
Songs up to five minutes long.
Expressive intent preserved across long structures.
Realistic instrument rendering.
Performed-sounding vocal generation.
Global-local Hybrid-LM architecture.
Open-weight production-ready release.

The model is designed to preserve expressive intent across songs of up to five minutes, render instruments with physical realism, and produce vocals that sound performed. Its pipeline uses fine-grained temporal descriptions, a global-local Hybrid-LM, residual vector quantization, hidden-state fusion, flow matching, and a Flow-VAE.


MiniMax presents Music 3.0 as open weights with production-oriented quality and versatility. It is useful for creators, game and film teams, prototyping, music education, and developers who need controllable full-song generation rather than isolated loops.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!