Voxtral TTS

Voxtral TTSVoxtral TTS
Voxtral TTS

Voxtral TTS

AI-powered text-to-speech and voice cloning with zero-shot capabilities.

2026Active
Rate

About

Voxtral TTS is an AI text-to-speech and voice cloning tool powered by Mistral AI's open-source 4B model. It enables users to clone any voice from just 2-3 seconds of audio, generating studio-quality AI speech with natural intonation, rhythm, and emotional rendering. Designed for real-time applications, Voxtral TTS supports 9 languages and cross-lingual voice cloning, offering an open-source solution for voice generation.
Comments
(6)
4 positive1 mixed1 negative
Reddit
r/MistralAIu/BalterBlackMay 3, 2026Negative

Mistral advertises it’s Voxtral Model with the ability to speak German. Yet iI didn’t find any video of a demonstration. Anyone got an example?

View on Reddit
Reddit
r/MistralAIu/SelectionCalm70Apr 22, 2026Mixed

MISTRAL-FM is a never ending radio station where two AI hosts, Camille and Hugo, actually talk between the songs. Everything they say is generated live by Mistral AI no scripts, no recordings, just fresh banter every few tracks. Fair warnin…

View on Reddit
Reddit
r/MistralAIu/robotrossartMar 31, 2026Positive💎

Voice (TTS): Mistral Voxtral (the voice is incredibly crisp).

View on Reddit
Reddit
r/LocalLLaMAu/Nunki08Mar 26, 2026Positive

Mistral AI to release Voxtral TTS, a 3-billion-parameter text-to-speech model with open weights that the company says outperformed ElevenLabs Flash v2.5 in human preference tests. The model runs on about 3 GB of RAM, achieves 90-millisecond…

View on Reddit
Reddit
r/MistralAIu/Nunki08Mar 26, 2026Positive

Mistral AI to release Voxtral TTS, a 3-billion-parameter text-to-speech model with open weights that the company says outperformed ElevenLabs Flash v2.5 in human preference tests. The model runs on about 3 GB of RAM, achieves 90-millisecond…

View on Reddit
Reddit
r/StableDiffusionu/fruesomeMar 26, 2026Positive

Voxtral TTS: open-weight model for natural, expressive, and ultra-fast text-to-speech. Highlights. Realistic, emotionally expressive speech in 9 popular languages with support for diverse dialects. Very low latency for time-to-first-audio.

View on Reddit

Some comments are pulled from public discussions around the web (look for the source icon). Quotes are excerpts; click through to read the full thread.

Frequently asked

What does Voxtral TTS do?

Voxtral TTS is an AI text-to-speech and voice cloning tool powered by Mistral AI's open-source 4B model. It enables users to clone any voice from just 2-3 seconds of audio, generating studio-quality AI speech with natural intonation, rhythm, and emotional rendering. Designed for real-time applications, Voxtral TTS supports 9 languages and cross-lingual voice cloning, offering an open-source solution for voice generation.

When was Voxtral TTS founded?

Voxtral TTS was founded in 2026.

What industry does Voxtral TTS operate in?

Voxtral TTS operates in Voice Agent, Voice AI, Voice Agents, Voice Cloning, Voice Synthesis, Text-to-Speech.