Editorial illustration for Cartesia's Sonic-3.6 TTS Model Adds Natural Pauses and Hinglish Code-Switching
Sonic-3.6 TTS Adds Natural Pauses, Hinglish Code-Switching
Cartesia put out Sonic-3.6 this week, the latest version of its real-time text-to-speech model and a follow-up to Sonic-3.5 from three months back. The headline feature is naturalness, specifically the model's handling of pauses and code-switching between languages like Hinglish, where a speaker mixes Hindi and English mid-sentence. That's a hard problem for TTS systems built around clean, single-language training data, and it's the kind of thing that either sounds right or doesn't when you hear it.
What sets this release apart from a typical vendor announcement is that the naturalness claim didn't have to rely on Cartesia's own word. Artificial Analysis runs two independent speech leaderboards, and Sonic-3.6 now sits at #1 on both. One of those boards, the Controlled Voice arena, strips away the usual advantage of having a bigger voice catalog by cloning every competing model onto the same eight reference voices. That setup isolates the underlying synthesis engine itself, which is where the real comparison happens.
Cartesia has released Sonic-3.6, the newest version of its real-time text-to-speech model. It arrives roughly three months after Sonic-3.5. The new change is naturalness, and this one is independently checkable.
Why this matters
For anyone building voice products, the Controlled Voice score is the number worth watching, not the flashy overall Elo. By cloning every model onto the same eight reference voices, Artificial Analysis strips away the advantage of a company's proprietary voice library and tests the synthesis engine itself. Sonic-3.6 topping that board, at 1,123 Elo, is a harder claim to dismiss than winning on home-turf voices. Add Hinglish code-switching and filler-word naturalness, and Cartesia is chasing a real gap: most TTS still sounds foreign to how multilingual users actually speak, pauses and mid-sentence language switches included.
The pricing detail matters just as much for founders doing unit economics. At $49 per 1M characters, Cartesia sits at half of ElevenLabs' Eleven v3 rate, which changes the calculus for anyone running high-volume voice agents or dubbing pipelines. We'd still want to hear these demos outside a curated launch reel; benchmark Elo and real-world latency under load are different animals. Next thing to watch: whether Sonic-3.6 holds its Controlled Voice lead once more labs submit models cloned onto those same eight voices.
Common Questions Answered
What are the main improvements in Cartesia's Sonic-3.6 compared to Sonic-3.5?
Sonic-3.6 focuses on naturalness as its headline feature, with significant improvements in handling pauses and code-switching between languages like Hinglish. The model can now better manage mixed-language speech where speakers blend Hindi and English mid-sentence, which has historically been a challenging problem for text-to-speech systems trained on single-language data.
What is Hinglish code-switching and why is it difficult for TTS models?
Hinglish code-switching refers to when speakers mix Hindi and English languages within the same sentence or utterance. This is challenging for traditional text-to-speech systems because they are typically built around clean, single-language training data, making it difficult to synthesize naturally sounding speech that seamlessly transitions between two languages.
What does the Controlled Voice score measure and why does it matter?
The Controlled Voice score, measured by Artificial Analysis, tests the synthesis engine itself by cloning every model onto the same eight reference voices, which strips away the advantage of a company's proprietary voice library. This metric is more meaningful than overall Elo scores for voice product builders because it provides an apples-to-apples comparison of the actual synthesis quality rather than relying on each company's custom voices.
What Elo rating did Sonic-3.6 achieve on the Controlled Voice benchmark?
Sonic-3.6 topped the Controlled Voice benchmark with a score of 1,123 Elo, which represents a significant achievement in demonstrating the model's synthesis engine capabilities independent of proprietary voice libraries. This ranking is considered a harder claim to dismiss than winning on home-turf voices because it reflects the core technology performance.
How does Sonic-3.6 handle filler words and natural pauses in speech?
Sonic-3.6 has improved its naturalness in handling filler words and pauses, which are essential elements of natural-sounding human speech. This improvement, combined with its code-switching capabilities, makes the model better equipped to generate speech that sounds more authentic and human-like across different linguistic contexts.
Further Reading
- Cartesia Docs: Sonic 3.6 (Beta) - Cartesia Docs
- Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas - MarkTechPost
- Cartesia Sonic-3.6: #1 TTS on Artificial Analysis - ExplainX
- Artificial Analysis Speech Arena leaderboard update featuring Cartesia Sonic 3.6 - Artificial Analysis
- Introducing Sonic-3.6: our most lifelike TTS yet, and another step change in naturalness across 44 languages - Cartesia