Google introduced two new text-to-speech models on Wednesday — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — calling them its most expressive audio generation models yet. The models are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids, according to Google's announcement from Leland Rechis, group product manager, and Alan Cowen, director of research science on the Gemini audio team.
The launch moves voice generation beyond static voice presets into what Google describes as a creative studio: voices designed from a text prompt, performances directed line by line, and safety scaffolding built in from the start. For more on this and other releases, see our AI model news.
Two Models, Two Jobs
The pair splits the workload in familiar Google fashion. Gemini 3.8 Flash TTS is built for deep creative direction and character design — creating entirely new voices from scratch via natural language prompts for gaming, immersive audiobooks, podcasts and interactive media, with granular control over acting cues, pacing, dialect shifts and backchanneling. Gemini 3.8 Flash-Lite TTS is the volume play: optimized for high-volume dubbing, audio content creation and expressive voice agents at scale, with fine-grained control over tone, pacing and nuance at lower cost.
Google's positioning frames the pair as transforming voice generation "from static presets into a dynamic creative studio." The company points to audiobooks, games and podcasts as natural destinations for the flagship model, while the Lite variant targets the unglamorous but enormous market for dubbing and always-on voice agents, where cost per generated minute matters more than star power. Both models are integrated into Google products — Gemini Notebook and Google Vids among them — signaling that the technology will surface in consumer-facing tools rather than remaining an API-only play.
A Full Vocal Studio, From 30 Voices to Infinite
Google says developers can now scale from 30 original voices to what it calls an infinite library. The generative voice design system lets users create bespoke voices by specifying role, accent and voice characteristics across more than 100 languages and dialects — Google's demos range from a high-energy DJ voice from Melbourne to a "super-tinny, monotone robot" and a Japanese dragon.
Alongside generation, the models ship with a library of more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English.
Voice replication is included with guardrails: users can recreate a consistent vocal profile from just a 30-second audio sample — of their own voice, or one they have the rights to use — backed by built-in consent verification, SynthID watermarking and C2PA credentials. Google is also adding save-and-scale functionality to keep custom voices consistent across projects, with voice remixing — fine-tuning timbre, pitch, pace and accent via prompts — listed as coming soon.
Directing the Performance, Line by Line
Both models support precise per-line direction. Developers can write their own stage directions or let the model steer delivery from natural script cues — the difference between a calm customer service agent and a whispered suspense scene. Scripted vocal bursts such as
Two features target longer-form and multi-voice work. Long-form generation maintains voice quality and pacing across hours of continuous audio with minimal speaker drift — aimed at podcasts and audiobooks — while native two-speaker scene staging directs multi-turn conversations from a single script with natural turn-taking and clearly separated voices.
Benchmarks: No. 1 on Hume AI
Google is leading the announcement with third-party numbers. Gemini 3.8 Flash TTS took the #1 overall spot on Hume AI's Voice Design Benchmark with a score of 71.4, and leads in accent modeling at 60.8. On Hume AI's Overall Quality Index, the two new models claim the #1 and #2 spots respectively.
Benchmark claims from the vendor should always be read with some caution, but Hume AI — an independent voice AI research company — is a reasonable referee for speech quality, and the scores give Google a concrete talking point against rivals in voice generation.
The Audio Arms Race Continues
The new models join a fast-growing Gemini audio family that already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live and 3.8 Live Extended Thinking — the real-time voice models Google rolled out earlier this month. The cadence is no accident: expressive, reliable speech synthesis is becoming core infrastructure for voice agents, customer support, accessibility tools and media production, and Google wants Gemini to be the default layer for all of it.
For developers, the immediate takeaway is practical rather than futuristic: custom-branded voices, multilingual dubbing and hours-long narration are now a prompt away inside the Gemini API and AI Studio — with watermarks attached.
The economics will decide how much of that potential gets used. Google has not emphasized pricing in the launch materials, but the existence of a Flash-Lite tier implies a deliberate low-cost option for bulk workloads, mirroring the tiering strategy it applies across the Gemini model family. Enterprises on Gemini Enterprise get the same capabilities behind their existing agreements, which is where Google expects the bulk of high-volume audio production to land.
What remains to be seen is how quickly creative industries adopt synthesized narration. Watermarking and consent verification lower the legal and ethical barriers, and the benchmark numbers suggest quality is no longer the bottleneck. If Google's claims hold up in practice, the gap between a recorded human voiceover and a generated one is now narrow enough that cost and turnaround — not authenticity — will make the decision for most projects.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →