Google Ships Gemini 3.5 Transcribe Speech-to-Text
Google introduced Gemini 3.5 Transcribe on 26 August 2026 as its most precise speech-to-text model. Artificial Analysis measured a 2.6% non-streaming word error rate and a 70% faster time to final transcription versus Chirp 3.
PromptCrates Editorial
Staff Writer

Google introduced Gemini 3.5 Transcribe on 26 August 2026 as its most precise speech-to-text model. Artificial Analysis measured a 2.6% non-streaming word error rate and said time to final transcription is 70% faster than Chirp 3.
What Gemini 3.5 Transcribe actually ships today
This is a speech-to-text model, not the overdue Gemini 3.5 Pro launch. Google's blog, signed by Diego Melendo Casado and Luke Leonhard of the Gemini Audio team, called it the company's most precise speech-to-text system yet and said it converts raw audio into polished, formatted text instead of leaving disfluencies for a later cleanup pass.
The Verge's Jess Weatherbed filed at 17:00 UTC on 26 August 2026. Google had also mentioned 3.5 Live and 3.5 Live Experimental updates that would build on Gemini's voice-chat speech recognition. After that story published, Google told The Verge those additional models are not launching yet and did not give a new date. Only Gemini 3.5 Transcribe is the 26 August ship. Do not write a Live launch into this changelog.
Two APIs carry the model. Real-time streaming runs through the Live API as gemini-3.5-transcribe-live and is specified for sub-second latency on interactive voice apps. Pre-recorded audio — meetings, call logs, uploaded files — runs through the Interactions API as gemini-3.5-transcribe and adds speaker attribution plus word-level timestamps. Those IDs are the ones to pin. Do not invent a third SKU.
Smart transcription is the product claim. Google said the model handles self-corrections such as "let's meet Tuesday—no, Wednesday," removes filler words ("ums" and "ahs"), auto-formats text, and accepts a custom vocabulary so jargon and unusual spellings survive. It automatically detects more than 85 languages and attributes speech in pre-recorded audio for up to three speakers. Support for more than three speakers is experimental. Do not write a four-speaker GA date.
If you already track Z.ai's Ox Alpha open-weight GLM, keep weights and speech-to-text on two lines. Ox Alpha is an open checkpoint. Gemini 3.5 Transcribe is a Google API and a set of consumer dictation surfaces.
The 2.6% word-error file versus Chirp 3
Artificial Analysis is the third-party yardstick Google cited. Average word error rate is 4.0% in streaming and 2.6% in non-streaming use. Time to final transcription improves 70% versus Chirp 3, Google's previous transcription model. On the FLEURS multilingual benchmark, across a set of top languages and locales, the model posted 5.50% streaming and 5.04% non-streaming word error, which Google said improves on Chirp 3. Those four percentages plus the 70% latency claim are the only scoreboard numbers in the blog.
Google also said the model holds up in noisy, real-world audio and captures alphanumeric entities such as postal codes and order IDs. That is a qualitative claim. It is not a second WER table. Do not invent a noise-condition score that Artificial Analysis did not print.
Function calling sits next to transcription in the Gemini macOS app. The blog said the model can delegate image generation and file analysis to other Gemini models. That path is currently available in the Gemini macOS app. It is not a general API promise in the 26 August post.
Developer platforms named as Live API partners are Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. Customer quotes in the blog come from Vivo, Intellitek Health, and Lingopal, who pointed at latency, accuracy, and language coverage. Treat those as named early users, not as a published customer count.
If your coding stack already points at OpenAI's GPT-5.6 Sol, Terra, and Luna in AWS Kiro, keep Kiro as a coding-agent surface. This announcement is a speech-to-text model and two Gemini API IDs.
Where Gemini 3.5 Transcribe is live, and where it is not
The availability matrix is three rows. Developers get a public preview in the Gemini API via Google AI Studio and Google Antigravity. Enterprises get a public preview on Gemini Enterprise Agent Platform, with Gemini Enterprise for Customer Experience coming soon. Consumers get the model in the Gemini app on macOS in English and in Rambler on Android in select countries and languages. Chrome talk-to-type is coming soon. Do not write a global consumer GA.
The surface list is specific. On Gboard on Android, Rambler turns spoken thoughts into formatted text and lets a user edit, correct spelling, and change writing style by voice. On Google Antigravity, the model can pair screen context and chat history, with permission, so file names and agent notes transcribe more accurately. In Google AI Studio Build mode, developers can vibe-code apps by voice. In the Gemini macOS app, voice commands can summarize local files, move text across apps, or generate images at the cursor by calling other Gemini models in the background.
The Verge's consumer line matches the blog: English for all macOS Gemini app users, Rambler on Android in select countries and languages, developer preview in AI Studio and Antigravity, Chrome later. Select countries is the Android limit. Do not invent a country list.
If you already watch Instinct's $250 million Series B, keep the personal-agent raise and this dictation model apart. Instinct is a private-beta life-organization agent. Gemini 3.5 Transcribe is a speech-to-text engine Google is putting under Gboard, Antigravity, and the Gemini API.
The useful pin is: on 26 August 2026 Google introduced Gemini 3.5 Transcribe; streaming WER is 4.0% and non-streaming WER is 2.6%; time to final transcription is 70% faster than Chirp 3; FLEURS is 5.50% / 5.04%; the IDs are gemini-3.5-transcribe-live and gemini-3.5-transcribe; more than 85 languages; up to three speakers on recordings; public preview in AI Studio, Antigravity, and Gemini Enterprise Agent Platform; macOS Gemini in English; Rambler on Android in select countries; Chrome coming soon; 3.5 Live is not in this launch.
Sources
- Intelligent transcription with Gemini 3.5 Transcribe — Google, 26 August 2026
- Google's new AI transcription edits out your 'ums' and 'ahs' — The Verge, 26 August 2026


