The underlying MuScriptor model frames multi-instrument transcription as language modeling over a piano-roll token sequence. It conditions on mel-spectrogram chunks and autoregressively predicts pitch, timing, and instrument tokens, with training that combines large synthetic MIDI pretraining, real audio fine-tuning, and reinforcement-style post-training on curated high-quality transcriptions.
The tool is useful for musicians, producers, composers, educators, and researchers who need to recover musical structure from dense recordings rather than isolated stems. Mirelo and Kyutai open-sourced the research model and inference code while also offering a free product experience in Mirelo Studio with an improved production version.


