Google Launches Gemini 3.5 Transcribe for Faster Speech

Picture of Wird-e- Ali

Wird-e- Ali

Google Launches Gemini 3.5 Transcribe for Faster Speech

Google has launched Gemini 3.5 Transcribe, a specialised speech-to-text model designed to convert spoken audio into accurate, formatted text while better handling natural conversations.

The new model is aimed at both real-time voice applications and recorded audio. Google says Gemini 3.5 Transcribe delivers lower error rates and improves time to final transcription by 70% compared with its previous Chirp 3 model.

Gemini 3.5 Transcribe is available through two developer interfaces: the Live API for real-time audio and the Interactions API for recorded audio such as meetings and calls.

Real-time and recorded transcription

The Live API provides continuous, two-way audio streaming with sub-second latency. Google is positioning it for real-time voice agents, interactive applications and live captioning.

The Interactions API is designed for pre-recorded audio. It supports features such as speaker attribution, word-level timestamps and custom vocabulary, making it more suitable for meetings, calls and other longer recordings.

The two interfaces offer different capabilities, limits and pricing, giving developers the option to choose the system that best matches their application.

Smarter transcription removes filler words

One of Gemini 3.5 Transcribe’s key features is its smart transcription capability.

The system can remove filler words such as “um” and “uh”, clean up repeated phrases and false starts, and understand corrections made while a person is speaking.

For example, if someone says, “Let’s meet Tuesday — no, Wednesday,” the system can reflect the correction directly in the final transcript rather than simply reproducing the speaker’s hesitation.

Developers can also provide custom vocabulary to improve recognition of technical terms, acronyms, brand names, proper nouns and specialised jargon.

Google’s documentation allows up to 1,000 custom vocabulary terms, although it recommends using up to 100 terms for the best results.

Supports more than 85 languages

Gemini 3.5 Transcribe supports more than 85 languages and can automatically detect the language being spoken.

The model can also handle regional accents and dialects and switch between languages during a conversation without requiring developers to manually configure each language.

It is designed to recognise specialised terminology as well as alphanumeric information such as postal codes and order IDs, including in noisy real-world environments.

For recorded audio, the model can identify different speakers and assign labels to their speech.

Verbatim and smart transcription options

Developers can choose between verbatim and smart transcription modes.

Verbatim transcription preserves fillers, repetitions, false starts and other details of natural speech. Smart transcription instead focuses on producing a cleaner and more readable transcript by removing disfluencies and resolving spoken corrections.

The choice depends on how the transcript will be used. Verbatim mode may be more suitable when an exact record of what was said is required, while smart transcription can be useful for meetings, notes and general content creation.

Google notes that smart transcription cannot be combined with speaker diarization or word-level timestamps.

Live API comes with time limits

The Live API is designed for low-latency voice applications and provides interim transcription while a person is still speaking before delivering the final version.

However, continuous Live API sessions are limited to 10 minutes.

Speaker diarization and word-level timestamps are also unavailable through the Live API, making the Interactions API more appropriate for certain recorded-audio applications.

The Interactions API can handle recordings of up to one hour under standard conditions. When speaker diarization or word-level timestamps are enabled, the maximum duration is reduced to 30 minutes.

Google reports improved accuracy

Google says Gemini 3.5 Transcribe significantly improves speech recognition compared with Chirp 3.

According to measurements cited by Google, the model achieved an average Word Error Rate of 4% for streaming transcription and 2.6% for non-streaming transcription.

On the multilingual FLEURS benchmark, Google reported Word Error Rates of 5.50% for streaming and 5.04% for non-streaming use across selected languages and locales.

Google also reported a 70% improvement in time to final transcription compared with Chirp 3.

Gemini 3.5 Transcribe reaches Google products

The technology is already being integrated into several Google products and services.

On Android, the model powers the new Rambler feature in Gboard, which converts spoken thoughts into formatted text and allows users to edit content, correct spelling and change writing styles using voice commands.

Google is also using the model in Antigravity, where screen context and chat history can improve recognition of file names, active documents and other relevant information when users give permission.

Google AI Studio allows developers to use the technology to build voice-driven applications, while the Gemini macOS app can use natural speech for formatted text and voice commands involving screen context.

Google also says voice dictation is coming to Chrome, allowing users to speak into web fields to create replies, posts and prompts.

Available through Google’s developer platforms

Gemini 3.5 Transcribe is available to developers through Google AI Studio and the Gemini API, while enterprise customers can access it through the Gemini Enterprise Agent Platform.

The model is currently available in public preview and is offered as an API-based managed service rather than an open-weight or self-hosted model.

Potential applications include real-time voice agents, live captioning, meeting transcription, contact centres, clinical documentation, media localisation, legal and insurance intake, dictation and voice-controlled software.

Google said the Live API is also supported by several developer platforms, including LiveKit, Pipecat, Agora, Fishjam, Vercel, Vision Agents and others.

With multilingual support, custom vocabulary, speaker identification and faster processing, Google is positioning Gemini 3.5 Transcribe as a major upgrade for developers and businesses building voice-based applications.

Slso read: Google Confirms Android Show Return Ahead of Google I/O 2026

Related News

Type to Search