Editorial illustration for Google Launches 3.5 Transcribe AI Model to Edit 'Ums' from Audio
Google's Gemini 3.5 Transcribe Removes Filler Words
Google Launches 3.5 Transcribe AI Model to Edit 'Ums' from Audio
Google added a new model to its Gemini Audio lineup on Tuesday: Gemini 3.5 Transcribe, a transcription tool built to clean up speech-to-text output as it works, stripping out "um" and "uh" and reformatting text without waiting for a human editor to do it later. The company says it detects specialized jargon automatically, handles more than 85 languages, and can attribute dialogue across up to three speakers in pre-recorded audio, complete with word-level timestamps. Users can also feed it custom vocabulary lists so niche terms and unusual spellings get transcribed correctly the first time.
Google is pitching 3.5 Transcribe as a real jump from its previous model, Chirp 3, particularly on multilingual accuracy and word error rates. The release lands alongside updates to 3.5 Live and a 3.5 Live Experimental version, both aimed at the voice recognition system behind Gemini's chat mode.
The timing is a little awkward. Google still hasn't shipped Gemini 3.5 Pro, the flagship model it promised for June. Instead, the company is rolling out audio tools while that larger release sits unfinished, and Google's own framing of the announcement hints at why that gap is starting to draw attention.
Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um” and “uh.”
Why this matters
For developers building on Gemini, this is a reminder that Google ships peripheral models faster than flagship ones. Gemini 3.5 Live Translate and now 3.5 Transcribe have arrived while 3.5 Pro, promised for June, still hasn't landed. If you're planning a product roadmap around Gemini's core reasoning model, that gap should factor into your timeline, not just your excitement about audio features.
The transcription upgrade itself is real and useful: 85-plus languages, jargon detection, cleaned-up filler words, and Google's own claim of a "major advancement" over Chirp 3 in word error rates. That's worth testing if you handle meeting notes, call transcripts, or multilingual voice data. But treat vendor benchmarks skeptically until you run your own comparisons against Chirp 3 or Whisper on your actual audio, not curated demo clips.
The bigger pattern here is Google iterating on Gemini's audio stack in public while the headline model stays in limbo. Watch whether 3.5 Pro shows up before Q3 ends, because that delay says more about Google's priorities than any transcription feature does.
Common Questions Answered
What are the key features of Google's Gemini 3.5 Transcribe model?
Gemini 3.5 Transcribe automatically removes filler words like "um" and "uh" from speech-to-text output while also reformatting text in real-time. The model supports more than 85 languages, detects specialized jargon automatically, can attribute dialogue across up to three speakers in pre-recorded audio, and provides word-level timestamps for precise editing.
How does Gemini 3.5 Transcribe improve upon the previous Chirp 3 model?
According to Google, Gemini 3.5 Transcribe represents a major advancement from Chirp 3, particularly in multilingual performance and reducing wording error rates. The new model also introduces the ability to edit naturally with just your voice and automatically format text without requiring human intervention after transcription.
Why is the release timeline of Gemini 3.5 Transcribe significant for developers?
Gemini 3.5 Transcribe's arrival demonstrates that Google is shipping peripheral audio models faster than its flagship Gemini 3.5 Pro model, which was promised for June but has not yet launched. Developers planning product roadmaps around Gemini's core reasoning capabilities should account for this gap in timeline expectations, not just the excitement around new audio features.
Can Gemini 3.5 Transcribe handle multiple speakers in audio files?
Yes, Gemini 3.5 Transcribe can attribute dialogue across up to three speakers in pre-recorded audio files. The model also provides word-level timestamps for each speaker's contribution, enabling precise tracking of who said what and when throughout the recording.
Further Reading
- Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text - Ars Technica
- Google rolls out Gemini 3.5 Transcribe to improve real-time dialogue and speech recognition - Android Authority
- Google launches Gemini 3.5 Transcribe with sub-second streaming and 85-language support - daily.dev
- Google unveils Gemini 3.5 Transcribe speech-to-text model - Investing.com
- Intelligent transcription with Gemini 3.5 Transcribe - Google Blog