Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, a pair of live dialogue models the company describes as its most advanced yet for natural voice conversation. The models began rolling out on September 15 across the Gemini API, Google AI Studio, Search Live, Gemini Live and Google Workspace, with enterprise access in private preview through Gemini Enterprise.
The announcement came from Tom Ouyang, a principal engineer, and Malini Jaganathan, a member of technical staff, writing on behalf of the Gemini Audio Team. It extends the Gemini 3.8 family that debuted earlier this month with the Flash and Flash Cyber models, and it marks Google's most aggressive push yet into real-time, multimodal voice assistance — a market where OpenAI's voice mode and a wave of speech startups are competing for the same daily-use case.
The launch fits into a remarkably crowded stretch for Google's model line; coverage of the latest AI developments shows the company has now shipped repeatedly in the span of a few weeks as it fights rivals on pricing, coding and now voice.
Built for Real-Time Conversation
What distinguishes the Live models from a standard chat model with a microphone attached is latency and state. Google's Live API documentation describes a low-latency interface that processes continuous streams of audio, images and text over a stateful WebSocket connection, accepting raw 16kHz PCM audio, JPEG images at up to one frame per second, and text, and returning spoken audio in real time.
The models process visual input in near real time, letting the assistant see what the user sees — Google's demonstration videos show Gemini 3.8 Live guiding an employee onboarding session using visual context, playing chess in near real time, and building complete business plans and custom marketing toolkits through natural speech. The models can also detect and transition between 97 supported languages automatically, mid-conversation, without the user switching settings.
Perhaps most significant for practical use: the models execute tools and API calls in the background while the conversation continues, acknowledging requests and keeping the dialogue going while tasks finish. That turns the assistant from a strictly turn-based respondent into something closer to an agent that happens to talk.
The Numbers Google Is Touting
Google reports that Gemini 3.8 Live Extended Thinking captured the top overall position on Artificial Analysis' Speech-to-Speech Quality Index with a score of 82.6. On agentic task completion, the company cites 68.6% on the τ-Voice benchmark and 35.1% on Sierra's τ-Voice-banking evaluation, alongside a 97.7% score on Big Bench Audio.
Those are Google's own reported figures, and independent verification will take time — but the claim to the top of Artificial Analysis' chart, if it holds, would put Google ahead of the speech-optimized offerings from OpenAI and dedicated voice AI companies on overall conversation quality. Google also says Extended Thinking maintains a highly competitive price point, an implicit nod to the price pressure that has defined the 3.8 generation.
All audio generated by the models is watermarked with SynthID, Google's imperceptible watermark, so AI-generated speech remains detectable — a compliance and misinformation safeguard that is becoming standard across Google's generative products.
Where You Can Use It
Availability differs between the two models. Gemini 3.8 Live is available to everyone inside Search Live, Google's real-time visual search mode. Extended Thinking is available in Gemini Live, in Docs for Google AI Pro and Ultra subscribers through Workspace, and in Gmail and Keep for all Google users.
Developers get the models through the Gemini API and Google AI Studio, and Google says the Live API is already used by platforms including Agora, Fishjam, LiveKit, Pipecat, Vercel and Vision Agents to build voice-driven interfaces on top of Google's real-time infrastructure. Salesforce, Genspark and Lumeris are named as launch partners for the new models.
A Crowded Voice AI Market
Voice is becoming the interface battleground of 2026 because it is where AI finally escapes the keyboard. A model that can see, listen, switch languages and run tools while talking is competing not with other chatbots but with the phone call, the customer support line and the personal assistant.
Google's advantage is distribution: Search, Android, Workspace and Chrome give the Live models an installed base no rival can match. The company is betting that being present in every corner of a user's day — now with its best voice models — will matter more than any single benchmark win. Rivals will get their chance to respond, but Google has made clear that voice, once a demo feature, is now a first-class product line.
For developers, the message is similarly direct. By publishing exact input formats, documenting background tool execution and seeding the ecosystem with Live API platforms, Google is trying to make its stack the default substrate for anyone building spoken interfaces — from customer support bots to wearable assistants. Whether Gemini 3.8 Live's quality claims hold up under independent testing, the availability play is already in motion.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →