Google is rolling out Gemini 3.1 Flash TTS, its latest text-to-speech AI model, promising more natural and expressive synthesized voices. The update brings granular control over vocal performance, aiming to empower developers and enterprises building next-generation audio applications.
First detailed by Deepmind, the model achieves an impressive Elo score of 1,211 on the Artificial Analysis TTS leaderboard, indicating a strong human preference for its output quality.
Enhanced Control with Audio Tags
A key innovation is the introduction of audio tags. These allow users to embed natural language commands directly into text inputs to precisely direct vocal style, pacing, and delivery. This feature places developers in the "director's chair," enabling detailed scene direction and speaker-specific instructions.
Users can configure audio profiles for distinct characters and apply "Director's Notes" for pace, tone, and accent adjustments. Inline tags offer further mid-sentence expression changes.
The precise parameters can be exported as Gemini API code for consistent voice application across projects.
