Google's Gemini 3.1 Flash TTS adds expressive AI voice

Google's Gemini 3.1 Flash TTS introduces advanced audio tags for expressive AI speech, supporting over 70 languages with SynthID watermarking.

Illustration of sound waves emanating from a Google Gemini AI interface.
Google's Gemini 3.1 Flash TTS enhances AI-generated speech with greater control and expressiveness.· Deepmind

Google is rolling out Gemini 3.1 Flash TTS, its latest text-to-speech AI model, promising more natural and expressive synthesized voices. The update brings granular control over vocal performance, aiming to empower developers and enterprises building next-generation audio applications.

First detailed by Deepmind, the model achieves an impressive Elo score of 1,211 on the Artificial Analysis TTS leaderboard, indicating a strong human preference for its output quality.

Enhanced Control with Audio Tags

A key innovation is the introduction of audio tags. These allow users to embed natural language commands directly into text inputs to precisely direct vocal style, pacing, and delivery. This feature places developers in the "director's chair," enabling detailed scene direction and speaker-specific instructions.

Users can configure audio profiles for distinct characters and apply "Director's Notes" for pace, tone, and accent adjustments. Inline tags offer further mid-sentence expression changes.

The precise parameters can be exported as Gemini API code for consistent voice application across projects.

Developers can begin experimenting with these advanced controls in Google AI Studio.

Global Scale and Security

Gemini 3.1 Flash TTS supports over 70 languages, facilitating localized and expressive speech experiences worldwide. To combat misinformation, all audio generated by the model is watermarked using SynthID, an imperceptible digital signature that reliably identifies AI-generated content.

The model is available in preview via the Gemini API and Google AI Studio for developers, and on Vertex AI for enterprises. Google Vids will also integrate the technology for Workspace users.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.