Gemini 3.5 Transcribe Sets 2.6% WER Baseline and Enables Voice-Triggered Workflows

Gemini 3.5 Transcribe hits 2.6% WER, slashes final latency by 70%, and adds function calling to voice AI.
Audio waveform transforming to text with WER and latency markers, cyan calibration tool, deep blue cloud panels.
Waveform to text with WER calibration and cloud panels. By Andres SEO Expert.

Key Takeaways

  • Gemini 3.5 Transcribe hits a 2.6% word error rate non-streaming
  • Dual APIs enable sub-second streaming or recorded-file transcription
  • Function calling turns voice into agentic workflows across 85+ languages

A 2.6% Word Error Rate Redraws the Voice AI Baseline

Google DeepMind’s Gemini Audio team broke the news on today: Gemini 3.5 Transcribe is now available in public preview through the Gemini API.

The speech-to-text model converts raw audio directly into clean, formatted text, with an average non-streaming word error rate of 2.6% as measured by Artificial Analysis.

That precision arrives alongside two separate developer entry points: the Live API for real-time streaming and the Interactions API for pre-recorded audio.

Inside the Dual-API Architecture for Real-Time and Recorded Audio

According to Google DeepMind, Gemini 3.5 Transcribe ships as two model identifiers with distinct operational profiles.

  • gemini-3.5-transcribe-live: Streams audio in both directions with latency below one second for interactive voice applications.
  • gemini-3.5-transcribe: Processes pre-recorded files such as meetings and call logs, attaching speaker labels and word-level timing data.

That split matters because streaming latency and non-streaming accuracy require different engineering trade-offs.

The streaming profile posts a 4.0% WER, while non-streaming use-cases reach 2.6% WER.

Final transcription arrives 70% sooner than on the previous Chirp 3 model, according to Artificial Analysis.

On the FLEURS multilingual benchmark, streaming WER reaches 5.50% and non-streaming WER reaches 5.04%, also improving over Chirp 3.

Smart transcription beyond recognition

Gemini 3.5 Transcribe handles self-corrections, strips filler words, and auto-formats spoken content into readable text.

A user who says one day and then corrects themselves mid-sentence gets a clean final instruction with the intended date.

Function calling turns audio into action

Function calls let the model hand off complex tasks to other Gemini models, including image generation and file analysis.

That capability is currently available in the Gemini macOS app, where voice commands can pair with screen context to summarize local files or generate images at the cursor.

Custom vocabulary and multilingual reach

Custom vocabulary support helps the model recognize specialized jargon, postal codes, order IDs, and unique spellings submitted by developers.

It automatically detects and transcribes more than 85 languages, including regional accents and dialect variation.

For pre-recorded audio, multi-speaker identification attributes up to three speakers with timestamps; support for more than three speakers remains experimental.

Beyond the API into Google surfaces

The company is rolling the model into Gboard on Android through the Rambler feature, Google Antigravity, Google AI Studio Build mode, and the Gemini app on macOS.

Chrome support will soon enable dictated input in any web form field.

Public preview access is live through Google AI Studio and Google Antigravity on the Gemini API.

Enterprise access runs through the Gemini Enterprise Agent Platform, while support for Gemini Enterprise for Customer Experience is on the near-term roadmap.

Why 70% Faster Final Transcription Reshapes the Voice AI Stack

Developer platforms such as Agora, LiveKit, Pipecat, Vercel, LangChain, Fishjam, and Vision Agents are already positioned to build on the Gemini Live API.

Companies including Vivo, Intellitek Health, and Lingopal have highlighted latency, accuracy, and language coverage as early reasons for interest.

The public preview arrives during a period of intense focus on model verification and real-time infrastructure.

Recent moves to open model usage data to independent researchers reinforce the broader demand for benchmark claims that hold outside curated demos.

Separate real-time infrastructure research has demonstrated a 39x faster failover path for large language model serving, a signal that voice stacks will increasingly treat latency recovery as a core resilience metric.

For the voice AI market, those signals collide with Gemini 3.5 Transcribe’s function-calling layer.

The model is not simply transcribing words; it is becoming an entry point for agentic workflows where spoken commands trigger file analysis, image generation, and cross-app actions.

Voice Interfaces Now Ship With Developer-Grade Precision

The Gemini Audio team just turned speech-to-text from a commodity utility into a developer-facing orchestration layer, and the API preview makes that layer available now. For teams building voice-driven AI products that need to scale, programmatic SEO AI automation is how Andres SEO Expert approaches it — contact us.

Frequently Asked Questions

What is Gemini 3.5 Transcribe?

Gemini 3.5 Transcribe is a speech-to-text model from Google DeepMind that converts raw audio into clean, formatted text. It offers a non-streaming word error rate of 2.6% and a streaming WER of 4.0%, with final transcription arriving 70% faster than Chirp 3.

What is the word error rate of Gemini 3.5 Transcribe?

The non-streaming WER is 2.6%, while the streaming profile posts a 4.0% WER, as measured by Artificial Analysis. On the FLEURS multilingual benchmark, streaming WER reaches 5.50% and non-streaming WER reaches 5.04%.

What APIs are available for Gemini 3.5 Transcribe?

There are two model identifiers: gemini-3.5-transcribe-live for real-time streaming with sub-second latency, and gemini-3.5-transcribe for pre-recorded audio, which adds speaker labels and word-level timing data.

How does Gemini 3.5 Transcribe compare to Chirp 3?

Final transcription arrives 70% sooner than on the previous Chirp 3 model, according to Artificial Analysis, with improved WER on both standard and multilingual benchmarks.

How many languages does Gemini 3.5 Transcribe support?

It automatically detects and transcribes more than 85 languages, including regional accents and dialect variation, making it suitable for global voice AI applications.

Can Gemini 3.5 Transcribe identify different speakers?

Yes, for pre-recorded audio, it supports multi-speaker identification for up to three speakers with timestamps. Support for more than three speakers remains experimental.

Prev Next

Subscribe to My Newsletter

Subscribe to my email newsletter to get the latest posts delivered right to your email. Pure inspiration, zero spam.
You agree to the Terms of Use and Privacy Policy