Google Launches Gemini 3.8 Flash and Flash-Lite TTS With 2,000+ Voices
Google launched two new text-to-speech models: Gemini 3.8 Flash TTS for expressive character design, and Gemini 3.8 Flash-Lite TTS for cost-efficient, high-volume use, describing them as its most expressive audio generation models yet.
Details
- Voice creation and library: custom voices can be generated from natural-language prompts across 100+ languages, alongside a library of 2,000+ production-ready voices with regional varieties
- Voice replication: voices can be recreated from a 30-second audio sample with consent verification
- Performance control: supports line-by-line delivery with stage directions and natural script cues, long-form generation spanning hours of continuous audio, two-speaker scene staging, and vocal bursts like laughs and sighs with backchanneling
- Benchmarks: Flash TTS scored 71.4 overall on Hume AI’s Voice Design Benchmark and led on accent modeling at 60.8, taking the #1 and #2 spots on Hume AI’s Overall Quality Index; both models ranked at the top in blind preference tests across languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi
- Availability: live now via the Gemini API and Google AI Studio for developers, coming soon to Gemini Enterprise, and rolling out to general users in Gemini Notebook and Google Vids
What happened next
The launch lands the same week as Gemini’s new Connected Apps rollout and study-notebook expansion, continuing a broad push across Google’s Gemini lineup — voice, integrations, and education — within a single week.