Google introduced two new text-to-speech models on Wednesday — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — calling them its most expressive audio generation models yet. The models are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids, according to Google's announcement from Leland Rechis, group product manager, and Alan Cowen, director of research science on the Gemini audio team.
The launch moves voice generation beyond static voice presets into what Google describes as a creative studio: voices designed from a text prompt, performances directed line by line, and safety scaffolding built in from the start. For more on this and other releases, see our AI model news.
Two Models, Two Jobs
The pair splits the workload in familiar Google fashion. Gemini 3.8 Flash TTS is built for deep creative direction and character design — creating entirely new voices from scratch via natural language prompts for gaming, immersive audiobooks, podcasts and interactive media, with granular control over acting cues, pacing, dialect shifts and backchanneling. Gemini 3.8 Flash-Lite TTS is the volume play: optimized for high-volume dubbing, audio content creation and expressive voice agents at scale, with fine-grained control over tone, pacing and nuance at lower cost.
Google's positioning frames the pair as transforming voice generation "from static presets into a dynamic creative studio." The company points to audiobooks, games and podcasts as natural destinations for the flagship model, while the Lite variant targets the unglamorous but enormous market for dubbing and always-on voice agents, where cost per generated minute matters more than star power. Both models are integrated into Google products — Gemini Notebook and Google Vids among them — signaling that the technology will surface in consumer-facing tools rather than remaining an API-only play.
A Full Vocal Studio, From 30 Voices to Infinite
Google says developers can now scale from 30 original voices to what it calls an infinite library. The generative voice design system lets users create bespoke voices by specifying role, accent and voice characteristics across more than 100 languages and dialects — Google's demos range from a high-energy DJ voice from Melbourne to a "super-tinny, monotone robot" and a Japanese dragon.
Alongside generation, the models ship with a library of more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French and Scots English.
Voice replication is included with guardrails: users can recreate a consistent vocal profile from just a 30-second audio sample — of their own voice, or one they have the rights to use — backed by built-in consent verification, SynthID watermarking and C2PA credentials. Google is also adding save-and-scale functionality to keep custom voices consistent across projects, with voice remixing — fine-tuning timbre, pitch, pace and accent via prompts — listed as coming soon.
Directing the Performance, Line by Line
Both models support precise per-line direction. Developers can write their own stage directions or let the model steer delivery from natural script cues — the difference between a calm customer service agent and a whispered suspense scene. Scripted vocal bursts such as
Two features target longer-form and multi-voice work. Long-form generation maintains voice quality and pacing across hours of continuous audio with minimal speaker drift — aimed at podcasts and audiobooks — while native two-speaker scene staging directs multi-turn conversations from a single script with natural turn-taking and clearly separated voices.
Benchmarks: No. 1 on Hume AI
Google is leading the announcement with third-party numbers. Gemini 3.8 Flash TTS took the #1 overall spot on Hume AI's Voice Design Benchmark with a score of 71.4, and leads in accent modeling at 60.8. On Hume AI's Overall Quality Index, the two new models claim the #1 and #2 spots respectively.
Madai ya viwango kutoka kwa muuzaji yanapaswa kusomwa kila wakati kwa tahadhari, lakini Hume AI - kampuni huru ya utafiti ya AI ya sauti - ni mwamuzi anayefaa wa ubora wa usemi, na alama hizo zinaipa Google nafasi ya kuzungumza dhidi ya wapinzani katika utengenezaji wa sauti.
Mashindano ya Silaha za Sauti Yanaendelea
Wanamitindo hao wapya wanajiunga na familia ya sauti ya Gemini inayokua kwa kasi ambayo tayari inajumuisha 3.5 Tafsiri Papo Hapo, 3.5 Nukuu, 3.8 Moja kwa Moja na 3.8 Fikra Zilizopanuliwa Papo Hapo - miundo ya sauti ya wakati halisi ambayo Google ilizinduliwa mapema mwezi huu. Mwanguko huo haukutokea kwa bahati mbaya: usanisi wa usemi unaoeleweka na unaotegemewa unakuwa miundombinu kuu kwa mawakala wa sauti, usaidizi kwa wateja, zana za ufikivu na utengenezaji wa maudhui, na Google inataka Gemini iwe safu chaguomsingi kwa yote hayo.
Kwa wasanidi programu, uchukuaji wa papo hapo ni wa vitendo badala ya wa siku zijazo: sauti zenye chapa maalum, uandikaji wa lugha nyingi na masimulizi ya saa moja sasa ni haraka ndani ya Gemini API na Studio ya AI - na alama za maji zimeambatishwa.
Uchumi utaamua ni kiasi gani cha uwezo huo kitatumika. Google haijasisitiza bei katika nyenzo za uzinduzi, lakini kuwepo kwa kiwango cha Flash-Lite kunamaanisha chaguo la kimakusudi la gharama ya chini kwa mzigo mwingi wa kazi, kuakisi mkakati wa kupanga unaotumika katika familia ya mfano wa Gemini. Enterprises kwenye Gemini Enterprise hupata uwezo sawa nyuma ya makubaliano yao yaliyopo, ambapo Google inatarajia sehemu kubwa ya uzalishaji wa sauti ya juu kutua.
Kinachobakia kuonekana ni jinsi tasnia za ubunifu zinavyochukua masimulizi yaliyounganishwa haraka. Uthibitishaji wa alama za maji na idhini hupunguza vizuizi vya kisheria na kimaadili, na nambari za benchmark zinapendekeza ubora sio kizuizi tena. Ikiwa madai ya Google yatadumu kivitendo, pengo kati ya sauti iliyorekodiwa ya mwanadamu na ile iliyotolewa sasa ni finyu kiasi kwamba gharama na mabadiliko - sio uhalisi - itafanya uamuzi kwa miradi mingi.
---
Kaa Mbele ya AIPata habari za hivi punde za AI, uchanganuzi na mafanikio - yote katika sehemu moja.
Soma habari zaidi za AI →