Key Features

Unified speech, music, effects, and ambience generation.
LLM-driven autoregressive generation.
Per-token conditional flow matching.
Variable-length output with a learned stop head.
768-dimensional semantic-acoustic latents.
25 Hz latent generation at 16 kHz output.
Audio-text alignment pretraining.
Nine-language support with emotion control.

The system combines a pretrained large language model and audio tokenizer with per-token conditional flow matching. It generates variable-length 16 kHz audio autoregressively using 768-dimensional semantic-acoustic latents at 25 Hz and a learned stop head.


Audio-text alignment pretraining maps audio latents into the language model token space before generation training. The result is a research system aimed at multimodal sound design, interactive environments, synthetic data, and multilingual expressive audio generation.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!