Key Features

Natural-language instructions specify generation or edits with optional reference audio.
Content editing is one of the five supported speech task families.
AuK includes enhancement and separation tasks within its common interface.
Paralinguistic editing adjusts speech characteristics beyond the spoken words.
Acoustic editing is supported alongside content and vocal-delivery changes.
The foundational AuK model contains approximately 1.5B parameters.
AuK-Flash is a distilled four-step model that does not require inference-time CFG.
Training uses approximately 3.03B instruction-audio instances across five task families.

The 1.5-billion-parameter model combines a multimodal language model, a 50-Hz audio VAE, and a hybrid rectified-flow Transformer. Semantic instructions and optional reference audio condition generation. Training progresses from generation to joint generation and editing, followed by preference optimization and reward-based refinement.


AuK is suited to researchers and developers building controllable speech workflows. Its distilled AuK-Flash variant performs four-step inference without classifier-free guidance, offering a faster alternative under matched conditions. The project describes source-code and weight releases and provides examples spanning the supported task families.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!