The 1.5-billion-parameter model combines a multimodal language model, a 50-Hz audio VAE, and a hybrid rectified-flow Transformer. Semantic instructions and optional reference audio condition generation. Training progresses from generation to joint generation and editing, followed by preference optimization and reward-based refinement.
AuK is suited to researchers and developers building controllable speech workflows. Its distilled AuK-Flash variant performs four-step inference without classifier-free guidance, offering a faster alternative under matched conditions. The project describes source-code and weight releases and provides examples spanning the supported task families.

