New Delhi: Google has introduced Gemini 3.5 Transcribe, a new speech-to-text AI model aimed at turning spoken audio into cleaner and more accurate written text. The model can automatically detect and transcribe more than 85 languages, including different regional accents and dialects.
The company announced the model on August 26, saying it is being rolled out across the Gemini API, Google AI Studio and several Google products. Developers can use it for live voice apps, captions, recorded meetings and call transcripts. Some of its features are already appearing in Google’s consumer products, including the Gemini app on macOS and Rambler on Android.
What Gemini 3.5 Transcribe can do
Speech-to-text sounds fairly simple until you actually try dictating something in a noisy room. People pause, correct themselves, change sentences halfway through and throw in plenty of “ums” and “ahs”. Gemini 3.5 Transcribe has been built to deal with these common problems.
Google says its “smart transcription” feature can understand corrections such as, “let’s meet Tuesday, no, Wednesday”, and use the corrected version in the final text. It can remove filler words and format the transcript without requiring a separate cleanup step.
The model supports custom vocabulary too. This could help people working with unusual product names, technical terms, order numbers or industry-specific words that normal transcription tools often get wrong.
For recorded audio, Gemini 3.5 Transcribe can identify and label up to three speakers with timestamps. Support for more than three speakers remains experimental.
Google claims lower transcription errors
Google cited measurements from Artificial Analysis, which found an average Word Error Rate of 4.0 per cent for streaming transcription and 2.6 per cent for non-streaming use.
On Google’s FLEURS multilingual benchmark, the model recorded a 5.50 per cent error rate in streaming mode and 5.04 per cent for non-streaming transcription.
Google says the new model improves the time taken to produce a final transcript by 70 per cent compared with its earlier Chirp 3 transcription model.
Developers get two versions. gemini-3.5-transcribe-live handles real-time audio through the Live API with sub-second latency. gemini-3.5-transcribe is meant for recorded audio through the Interactions API.
Where users will see it
Google is already putting the model into several products. Rambler on Android can convert spoken thoughts into formatted text and remove filler words. Users can use voice commands to correct spelling or change the writing style.
The Gemini app on macOS goes a little further. Google says voice commands can work with screen context and call other Gemini models for tasks such as analysing local files or generating images.
Chrome is next on Google’s list. The company says a talk-to-type feature is coming soon, letting users dictate text inside web fields.
As of now, Gemini 3.5 Transcribe is available in public preview through the Gemini API, Google AI Studio and the Gemini Enterprise Agent Platform.









