AI Voice & Audio

Gemini 3.5 Transcribe: Google's New Speech Model

Google launched Gemini 3.5 Transcribe on 26 August 2026, cleaning up filler words and speaker self-corrections automatically across more than 85 languages.

Editorial noteThis news analysis is based on the linked primary sources. Performance and product claims are attributed to the announcing vendor unless the article explicitly says they were independently tested.

Google introduced Gemini 3.5 Transcribe on 26 August 2026, a speech-to-text model built around the idea that raw transcription is rarely what anyone actually wants. Rather than reproducing every word as spoken, the model removes filler words such as "ums" and "ahs", automatically formats the output, and resolves speaker self-corrections — Google's own example being "let's meet Tuesday—no, Wednesday", which the model is designed to render as the corrected intent rather than the literal utterance [1].

Google says the model automatically detects and transcribes more than 85 languages, supports custom vocabulary for specialist jargon and unusual spellings, and identifies up to three speakers with timestamps, with support beyond three described as experimental [1]. It is available now in public preview through the Gemini API via Google AI Studio, Google Antigravity and the Gemini Enterprise Agent Platform, and reaches consumers through the Rambler feature in Gboard on Android and the Gemini app on macOS, with Chrome listed as coming soon [1].

The accuracy figures, and where they come from

Google cites word error rate (WER) measurements it attributes to Artificial Analysis, an independent benchmarking organisation: 4.0% WER for streaming transcription and 2.6% for non-streaming [1]. On the FLEURS multilingual benchmark, Google reports 5.50% WER streaming and 5.04% non-streaming [1]. The company also claims a 70% improvement in time-to-final-transcription compared with Chirp 3, its previous speech model [1].

This is a slightly stronger evidential position than most model announcements we have covered this month, because the headline accuracy figures are attributed to a third-party evaluator rather than to Google's own internal testing. The caveat is that Google has selected which third-party figures to publish and in which framing, and Next AI Compare has not independently reproduced them. The 70% latency improvement, notably, is a comparison against Google's own previous model rather than against any competitor.

Why "structured" matters more than "accurate"

Speech-to-text accuracy has been good enough for several years that WER alone is no longer the deciding factor for most business use. The remaining friction is that an accurate transcript of a person talking is still an awkward document: full of restarts, hesitations, corrections and no paragraph structure. Someone still has to clean it up before it becomes meeting notes, a summary or a draft.

Gemini 3.5 Transcribe targets that cleanup step rather than the transcription step. Whether that is genuinely useful depends on the use case, and the trade-off runs in both directions — a model that silently discards a self-correction is making an editorial judgement about intent, which is helpful for meeting notes and potentially problematic for a legal, medical or journalistic record where the literal utterance is what matters. Google's announcement does not indicate whether the cleanup behaviour can be disabled to obtain a verbatim transcript.

An unusual addition: function calling inside a transcription model

The most architecturally interesting claim is that the model can delegate tasks to other Gemini models mid-transcription, with Google citing image generation and file analysis as examples [1]. This blurs the line between a transcription model and an assistant: a spoken instruction could, in principle, trigger an action rather than simply being written down. Google has not detailed how the model distinguishes an instruction from speech that merely sounds like one, which is the obvious failure mode for such a feature.

How this sits against the alternatives

The transcription market is not short of credible options, and this release does not obviously displace them. ElevenLabs has been actively developing its Scribe line, adding realtime entity detection to Scribe v2 in a changelog entry dated 3 August 2026, and shipped a Dubbing v2 API on 10 August 2026 covering more than 90 languages while preserving speaker voice and pacing [2]. Deepgram remains a widely used API-first choice for developers building realtime speech recognition. OpenAI maintains its own transcription models within its audio lineup.

Cross-vendor WER comparisons are genuinely difficult to make responsibly, because results depend heavily on the test dataset, audio conditions and whether streaming or batch processing is measured — which is why we are not reproducing competitor figures alongside Google's here. What can be said is that Google's differentiation in this release is less about raw accuracy than about post-processing and distribution: it is shipping directly into Gboard and Chrome, which is a reach advantage no independent transcription vendor can match.

Why this matters

Transcription is becoming an input layer for other AI work rather than a standalone product. Meeting recordings feed summarisation, voice notes feed drafting, and support calls feed CRM records. A model that outputs clean, formatted, speaker-attributed text reduces the work needed before that downstream step. Google shipping it directly into Gboard and Chrome rather than only as an API also means a large number of people will use it without ever choosing it.

Who should care

Anyone dictating regularly on Android already has access through Gboard's Rambler feature and will encounter the cleanup behaviour by default. Developers building meeting-notes, voice-note or call-analysis products should evaluate the API preview, particularly the custom vocabulary support if their domain involves specialist terminology. Teams handling multilingual audio should note the 85-language automatic detection, which removes the need to specify a language in advance. Anyone whose use case requires a verbatim record — legal, compliance, medical, journalistic — should establish whether the cleanup can be turned off before adopting it.

Practical implications for buyers and users

The three-speaker limit is the most consequential practical constraint for business use, since a great many meetings involve four or more participants and Google describes support beyond three as experimental. Test with your actual meeting sizes rather than a two-person sample. Because everything on the developer side is in public preview, terms, behaviour and availability may change before general release, so this is a stage for evaluation rather than production commitment. The absence of published pricing in the announcement means cost per hour of audio cannot yet be compared against ElevenLabs, Deepgram or OpenAI, which is the comparison most buyers will actually need.

Limitations, availability and unresolved questions

Google has not published pricing for Gemini 3.5 Transcribe in the announcement. All three developer access routes are public preview rather than general availability, and no general-availability date is given. Chrome support is stated as forthcoming without a date. There is no published guidance on whether verbatim, uncleaned output is available as an option, nor on how audio data is retained or used — a material question for anyone transcribing confidential meetings. Google has also not detailed accuracy variation across the 85-plus supported languages, and aggregate FLEURS figures can conceal significant per-language differences.

Verdict

Gemini 3.5 Transcribe is a well-targeted release that competes on output usability rather than headline accuracy, and the decision to cite third-party evaluation figures rather than purely internal ones is more credible than the industry norm. Its distribution through Gboard and Chrome is arguably a bigger competitive advantage than the model itself. The three-speaker ceiling limits its immediate usefulness for meeting transcription specifically, and the open question of whether verbatim output is available should be resolved before anyone adopts it for a record-keeping use case. Worth testing now in preview; too early to standardise a workflow on.

The current AI Voice & Audio shortlist

Where this sits in the wider market: our current shortlist for AI Voice & Audio, what each tool is best at and the main caution to check before committing.

ToolBest forCurrent positionImportant caution
ElevenLabs
Best expressive speech
Narration, dubbing, voice design and agentsEleven v3 remains the flagship synthesis model, alongside Scribe v2 for transcription in 90+ languages with speaker diarisation, Eleven Music v2, and Flash v2.5 for latency-sensitive work.Obtain clear consent for cloned voices and define retention rules.
OpenAI Realtime
Agent infrastructure
Low-latency multimodal assistantsOpenAI’s current realtime family is designed for audio-in/audio-out applications with tool use, supported by separate transcription and audio models.Realtime quality depends on turn detection, network conditions and tool latency.
Gemini Audio
Live translation
Google-based multimodal and translation workflowsGoogle’s 2026 model cards include live audio, TTS and translation-oriented Gemini releases for conversational applications.Check supported languages and regional availability for the exact model.
Deepgram
Developer pick
Realtime speech recognition and voice APIsDeepgram remains a practical API-first option for developers building transcription and conversational voice systems.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
Hume AI
Expressive agents
Emotion-aware conversational interfacesHume focuses on expressive voice and conversational systems where delivery and interaction style matter.Emotion-related claims should be evaluated carefully in sensitive use cases.
Adobe Firefly
Commercially safe audio
Music, voiceover and sound effects inside one creative workflowAdobe made Generate Music, Generate Speech and Generate Sound Effects generally available on 20 August 2026, with ElevenLabs offered as an alternative speech engine. Adobe describes the generated music as universally licensed and safe for commercial use.Read Adobe’s actual licensing terms before using generated audio in paid campaigns, and check whether the ElevenLabs option carries the same assurances.
Descript
Editing workflow
Podcasts, interviews and transcript-led editingDescript combines transcription, voice tools and text-based audio/video editing in one creator workflow.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
Murf
Business voiceovers
Training, presentations and corporate narrationMurf packages synthetic voice into a straightforward studio for business production teams.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
PlayHT
Long-form option
Voice libraries and API-based narrationPlayHT remains an option for teams comparing large voice libraries and programmatic generation.Plans, limits and model availability change frequently; confirm the current vendor page before purchasing.
Suno
Music creation
Rapid song concepts and music ideationSuno serves music-generation workflows rather than speech production, with fast ideation from natural-language prompts.Review commercial-use terms and avoid imitating living artists.
Udio
Music alternative
Music exploration and arrangement ideasUdio is another specialist music-generation environment for experimenting with composition and style.Rights and platform terms deserve the same attention as output quality.

Sources and verification notes

Primary product documentation checked for this update: