The release offers separate real-time and recorded-audio endpoints. The Live API provides bidirectional streaming with sub-second latency, while the Interactions API adds word timestamps and speaker attribution for recordings. Automatic language detection covers more than 85 languages, with regional and accent variation included.
Developers can access the model through Google AI Studio and the Gemini Enterprise Agent Platform, while related capabilities appear in Google consumer tools. Recorded-audio speaker identification supports up to three speakers, with larger groups experimental. API usage is classified as Paid and depends on current platform pricing.

