Google Ships Gemini 3.5 Transcribe: Auto-Detects 85+ Languages and Strips Out Your "Ums" and Self-Corrections
Google released the speech-to-text model Gemini 3.5 Transcribe on August 27, replacing Chirp 3. It automatically detects and transcribes more than 85 languages, handles background noise, filler words, self-corrections and specialized terminology, and formats the output; for pre-recorded audio it attributes speech with timestamps for up to three speakers (more than three is experimental). Google cites Artificial Analysis measurements of 4.0% average word error rate streaming and 2.6% non-streaming, with time to final transcription down 70% versus Chirp 3 — vendor-supplied figures with no independent reproduction. It is in public preview for developers through the Gemini API and AI Studio.











