Google Announces Gemini 3.5 Transcribe Featuring Advanced Multilingual and Real-Time Capabilities

Gemini 3.5 Transcribe supports a 96,000-token context window, enabling processing of extended audio sessions without losing track of earlier content.
Google differentiates real-time transcription pathways: live audio can be handled via a dedicated Live API or the Cloud Speech-to-Text API for latency-sensitive use, instead of routing live Gemini audio through the general model.
Developers have access to two Gemini 3.5 Transcribe APIs: the Live API for real-time streaming (gemini-3.5-transcribe-live) and the Interactions API for pre-recorded audio (gemini-3.5-transcribe), with the ability to edit transcripts via voice.
Gemini 3.5 Transcribe adds voice-driven editing, allowing users to modify text by speaking during or after transcription.
Language support for Gemini 3.5 Transcribe includes coverage of up to 85 languages, expanding multilingual transcription capabilities.
Google has unveiled Gemini 3.5 Transcribe, its most advanced speech-to-text model yet, capable of converting raw audio into formatted text with unprecedented accuracy Seeking Alpha. The new model handles up to 85 languages, speaker attribution, word-level timestamps, and can clean up disfluencies while organizing messy spoken words into polished, structured text UK News Hour.
Gemini 3.5 Transcribe powers new features across Google's ecosystem, including Rambler in Gboard on Android, the Gemini app on macOS, and incoming Chrome dictation Android Pure. Developers can access the model through two separate APIs: one for real-time streaming audio and another for pre-recorded files, with optional translation, summarization, and even voice-driven editing capabilities Google announcements.
Unlike traditional speech recognition, Gemini 3.5 Transcribe converts raw audio into structured text AOL. The model understands context deeply enough to clean up filler words, false starts, and broken sentences. It recognizes specialized terminology across industries and handles diverse accents and noisy environments more reliably than earlier models Yahoo Tech.
Google offers the Live API (gemini-3.5-transcribe-live) for real-time streaming audio, designed for low-latency applications. The Interactions API (gemini-3.5-transcribe) handles pre-recorded files Android Pure. This split approach lets developers choose the right tool based on whether they need instant results or can accept slightly more processing time for accuracy.
Both APIs include speaker attribution, so transcripts identify who said what. Word-level timestamps mark exactly when each word was spoken. Users can edit transcripts using voice commands during or after transcription, adding another layer of control Google announcements.
Gemini 3.5 Transcribe supports a massive 96,000-token context window, allowing the model to process extended audio sessions without losing track of earlier content. This means longer conversations stay coherent, and the model maintains better understanding of context across minutes of speech, not just seconds Key Point Summary. Emotion detection and stronger contextual handling benefit directly from this expanded window.
The model integrates through Google AI Studio, the Gemini API, and the Gemini Enterprise Agent Platform UK News Hour. Businesses can deploy Gemini 3.5 Transcribe on-premise where compliance requires it, while monetization and API-first options let companies embed transcription directly into their workflows Summary. Chrome dictation, Android's Rambler, and macOS Gemini app all tap the same underlying model Android Pure.
Publishers
12
Articles
5
Reach
17