Google Launches Gemini 3.5 Transcribe with 85+ Language Support
Google launched Gemini 3.5 Transcribe, a new speech-to-text model that converts raw audio directly into accurate, polished, and formatted text for both real-time and pre-recorded audio.
Details
- Smart transcription: automatically handles speaker self-corrections, strips filler words, and auto-formats the resulting text rather than producing a raw verbatim dump
- Multi-speaker identification: attributes speech to up to three distinct speakers in a recording, complete with timestamps
- Function calling and custom vocabulary: can delegate complex downstream tasks to other Gemini models, and adapts to specialized jargon or unusual spellings
- Language coverage: automatically detects and transcribes over 85 languages
- Benchmark results: word error rate of 4.0% in streaming mode and 2.6% for pre-recorded audio, a 70% improvement in transcription latency over the previous model, and 5.50%/5.04% WER (streaming/non-streaming) on the FLEURS benchmark
- Availability: public preview for developers via Google AI Studio and Google Antigravity, public preview for enterprises through the Gemini Enterprise Agent Platform, and rollout to general users through the Gemini app on macOS and the Rambler app on Android in select countries and languages, with Chrome support coming soon
What happened next
Gemini 3.5 Transcribe positions Google to compete more directly in real-time transcription and voice-agent tooling, an area where latency and multi-speaker accuracy matter as much as raw language coverage; broader rollout to Chrome and additional Rambler markets is still pending.