
Google announced the release of Gemini 3.5 Transcribe, a new addition to its Gemini Audio family designed to enhance transcription capabilities across multiple languages and use cases. The model represents what Google characterizes as a significant improvement over its previous transcription system, particularly in handling multiple languages and reducing transcription errors.
The transcription service offers several key features for users. It can automatically detect and preserve specialized vocabulary and jargon through customized vocabulary inputs, reducing the need for manual corrections. The tool also removes common filler words such as “um” and “uh” from transcriptions while automatically formatting text. For multi-speaker scenarios, the model can attribute speech to up to three different speakers in pre-recorded audio and provide precise word-level timestamps for reference.
Google indicated that voice-based editing capabilities allow users to make corrections through spoken commands rather than manual text input. The company also supports communication in more than 85 languages, significantly expanding accessibility for international users.
The rollout of Gemini 3.5 Transcribe began today in English for macOS Gemini app users and through the Rambler dictation feature on Android in select countries and languages. Developers can access the tool through public preview via the Gemini API, AI Studio, and Antigravity. Chrome browser support is expected to arrive in a future update. Google initially indicated that companion models called Gemini 3.5 Live and 3.5 Live Experimental would launch concurrently, though the company later clarified that only 3.5 Transcribe was launching at this time.
Article Attribution | Read More at Article Source
Article summary produced by Claude AI