Google has launched Gemini 3.5 Transcribe, a new speech-to-text model built for real-time transcription that recognizes more than 85 languages automatically — and polishes output along the way, stripping filler words and correcting verbal stumbles without being asked.

The launch positions Google's speech stack as a direct challenger in the fast-growing transcription market, where accuracy and latency increasingly decide which services developers adopt for calls, meetings, captions, and voice interfaces. For more context on this story, see our ongoing AI industry coverage.

What the Model Can Do

According to Google's announcement, coverage of which was published by The Decoder, Gemini 3.5 Transcribe goes well beyond converting spoken words into text:

  • Automatic language detection across 85-plus languages, with no configuration required
  • Filler word removal, cleaning up "um," "uh," and similar disfluencies
  • Correction of slips of the tongue, producing readable prose rather than verbatim noise
  • Autonomous text formatting, including punctuation and structure

The company reports a word error rate of 4.0 percent for streaming audio and 2.6 percent for recorded audio, alongside latency reductions of roughly 70 percent compared with its predecessor, Chirp 3 — a substantial margin in live captioning scenarios where delay breaks conversations.

Two Interfaces for Two Jobs

Developers get access through two distinct application programming interfaces:

Live API — `gemini-3.5-transcribe-live`

Handles real-time streaming with very low latency, targeting live captions, voice agents, and interactive assistants where response time matters most.

Interactions API — `gemini-3.5-transcribe`

Processes recorded audio and returns transcripts with speaker attribution and timestamps — the workflow typical for meetings, interviews, podcasts, and media post-production.

Beyond transcription, the model supports function calling, meaning it can delegate follow-up tasks to other models in the Gemini family — for example triggering an image generation request or a web search based on something said during a session. That turns a transcript engine into a controllable interface layer for voice-driven applications.

Already Shipping Across Google Products

The model is live in Google AI Studio and on the Gemini Enterprise Agent Platform for developers today. Google has also wired it into consumer surfaces: Gboard on Android uses it via an internal component called Rambler for dictation, the Gemini app on macOS applies it to voice input, and Chrome integration is slated to follow.

That distribution matters. Speech recognition earns its keep at scale, and embedding the model in keyboards, desktop apps, and browsers puts it in front of billions of input interactions daily — while feeding the evaluation loops needed to keep error rates falling across accents, dialects, and noisy environments.

Competitive Stakes in Speech AI

Transcription has quietly become a competitive battleground among frontier labs and specialist vendors alike. Accurate, low-latency speech understanding is the entry ticket for agentic products that act on spoken commands, enterprise meeting intelligence, accessibility features, and multilingual customer service automation.

By pairing aggressive latency improvements with self-correcting output, Google is betting that developers prefer transcripts that read like edited text rather than raw phonetic capture — trading strict verbatim fidelity for usability. Teams needing exact records, such as legal or medical transcribers, may still want verbatim modes; for everyone else, the practical difference between hearing someone and understanding them continues to narrow.

With Gemini 3.5 Transcribe now generally available in AI Studio and enterprise tooling, the first meaningful test will be adoption metrics from voice-agent developers — a market segment growing quickly enough that even incremental accuracy gains translate into sizable switching decisions.

Why Latency Is the Real Battleground

Error rates get the headlines, but latency quietly determines which products are viable. A transcription engine feeding a live caption overlay needs to keep pace with conversation, or viewers disengage. A voice agent negotiating a customer support call cannot pause awkwardly while its speech layer catches up. Google's reported 70 percent reduction in delay relative to Chirp 3 targets precisely that gap — shortening the distance between a speaker finishing a sentence and downstream systems acting on it.

For agentic applications this compounds: every turn of a spoken dialogue involves recognition, reasoning, and response. Shaving milliseconds from the recognition stage gives builders headroom to spend on more capable — and slower — reasoning models without breaching conversational timing budgets.

The Verbatim Trade-Off

One design choice deserves scrutiny. By default, Gemini 3.5 Transcribe curates output: filler words vanish, stumbles get corrected, and formatting is applied automatically. For most business uses — minutes, captions, searchable archives — that is exactly what users want.

But certain workflows depend on verbatim records. Legal depositions, compliance reviews, and journalistic fact-checking sometimes require knowing precisely what was said, hesitations included. Organizations in regulated industries will want clarity on how much of the smoothing is optional and configurable before migrating sensitive pipelines onto the new model. Google's documentation and API parameters will effectively set where the verbatim-versus-polished line sits for the industry.

A Multilingual Footprint as Moat

Coverage of 85-plus languages with automatic detection is not just a feature checklist item. Multilingual reach concentrates value in markets where competitors historically lag: regional accents, code-switching between languages mid-sentence, and low-resource languages all punish models trained narrowly on English-heavy corpora.

For global enterprises running support centers across dozens of countries, a single model handling language detection transparently reduces integration complexity — one pipeline instead of many per-market configurations. Combined with distribution through Gboard's billions of Android installations, Google is positioning transcribe-quality multilingualism as infrastructure rather than a premium capability.

What to Watch

Three signals will show whether Gemini 3.5 Transcribe changes market share or merely raises expectations: developer migration in third-party benchmarks over coming months; how quickly rivals match the latency numbers rather than the accuracy claims; and whether the function-calling hooks spawn a wave of voice-first agents built natively on Gemini. The launch is live now, and pricing tiers through AI Studio should make experimentation cheap enough that switching costs fall to nearly nothing.

---

Stay Ahead of AI

Get the latest AI news, analysis, and breakthroughs — all in one place.

Read more AI news →