The update compresses the semantic codec from 50 Hz to 25 Hz, reducing sequence length and inference cost, and replaces the earlier S2M backbone with a more efficient Zipformer architecture. New cross-lingual strategies use boundary-aware alignment, token-level concatenation, and instruction-guided generation.
IndexTTS 2.5 supports Chinese, English, Japanese, Spanish, and Arabic, including emotion transfer without target-language emotional training data. It is useful for narration, localization, character voices, accessibility, and research into zero-shot emotional TTS.

