Gemini 3.5 Transcribe

NEWHOT

Key Features

The gemini-3.5-transcribe-live endpoint provides bidirectional streaming through the Live API.
The gemini-3.5-transcribe endpoint handles recordings through the Interactions API.
Smart transcription can clean disfluencies and interpret spoken self-corrections.
Custom vocabulary helps recognition of jargon and unusual spellings.
The announcement describes automatic detection and transcription of more than 85 languages.
Recorded-audio processing supports word-level timestamps.
Recorded audio supports attribution for up to three speakers; larger groups are experimental.
The streaming interface targets sub-second latency for interactive voice applications.

The release offers separate real-time and recorded-audio endpoints. The Live API provides bidirectional streaming with sub-second latency, while the Interactions API adds word timestamps and speaker attribution for recordings. Automatic language detection covers more than 85 languages, with regional and accent variation included.


Developers can access the model through Google AI Studio and the Gemini Enterprise Agent Platform, while related capabilities appear in Google consumer tools. Recorded-audio speaker identification supports up to three speakers, with larger groups experimental. API usage is classified as Paid and depends on current platform pricing.

Get more likes & reach the top of search results by adding this button on your site!

Embed button preview - Light theme
Embed button preview - Dark theme
TurboType Banner

Subscribe to the AI Search Newsletter

Get top updates in AI to your inbox every weekend. It's free!