The race to make AI agents sound less like robots and more like humans just got a new front-runner. AI startup Cartesia has unveiled Sonic-3, a text-to-speech (TTS) model it claims is the fastest and most emotionally expressive on the market, capable of generating laughter and a full range of emotions in real-time conversations.
For anyone who has suffered through a laggy, monotone call with an automated agent, Cartesia’s claims are significant. The company reports an end-to-end latency of just 190 milliseconds, well below the typical threshold for human conversational response. This speed, combined with the ability to generate non-speech sounds like laughter, aims to eliminate the uncanny, stilted nature of most current voice AI. In demos, the voice can sound palpably excited or even "devastatingly sad," a far cry from the neutral tone of typical assistants.