Key Features

Zero-shot neural text-to-speech.
Transformer Text-to-Semantic module.
Non-autoregressive Semantic-to-Mel module.
Emotion and voice characteristic replication.
25 Hz semantic codec compression.
Zipformer S2M architecture.
Chinese, English, Japanese, Spanish, and Arabic.
Cross-lingual emotion transfer.

The update compresses the semantic codec from 50 Hz to 25 Hz, reducing sequence length and inference cost, and replaces the earlier S2M backbone with a more efficient Zipformer architecture. New cross-lingual strategies use boundary-aware alignment, token-level concatenation, and instruction-guided generation.


IndexTTS 2.5 supports Chinese, English, Japanese, Spanish, and Arabic, including emotion transfer without target-language emotional training data. It is useful for narration, localization, character voices, accessibility, and research into zero-shot emotional TTS.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!