Its approximately 3.59-billion-parameter architecture connects symbolic scores, semantic music tokens, and acoustic representations. Autoregressive and non-autoregressive experts share attention while retaining separate transformations. Companion MERT2 music representations and SheetSage2 audio-to-score transcription connect recorded music with the same editable composition workflow.
The system supports original music, covers, and targeted edits, with examples across languages and genres. Public research comparisons include differing candidate-selection budgets, so reported rankings should be interpreted with their evaluation settings. The released repository supplies inference guidance and an agent-oriented music workflow.

