Google has introduced Gemini 3.5 Transcribe, a multilingual speech-to-text model that the company says outperforms competing offerings from OpenAI and Deepgram.
Google has unveiled Gemini 3.5 Transcribe, a new speech-to-text model designed to convert spoken language into written text. The company says the system supports more than 85 languages and can automatically detect which language is being spoken, a capability that includes cases where speakers switch or mix languages within the same conversation.
The tool can distinguish between up to eight different speakers in a single audio recording and automatically adds timestamps to transcripts, making it easier to follow who said what and when. It also filters out common filler words and drawn-out vocalizations such as "uh" to produce cleaner, more readable text output.
According to Google's stated test results, Gemini 3.5 Transcribe outperforms competing speech-recognition offerings from OpenAI, Deepgram, and other providers on accuracy. Google claims the model is currently the best speech-to-text solution available, positioning it as a new benchmark in the increasingly competitive market for voice-to-text services.
https://t.me/business_ua/19447