Skip to main content
Gradium AI's new text-to-speech (TTS) model achieves an 81% pass rate, demonstrating advanced AI voice synthesis.

Editorial illustration for Gradium AI’s New TTS Model Scores 81% Pass Rate

Gradium AI's TTS Model Hits 81% Accuracy Rate

4 min read

A voice agent can hold a fine conversation and still fail the one moment that matters: reading back an order number, a phone digit, an email address. That's the gap Gradium AI is targeting with a new text-to-speech model it made the default across its API and Studio on August 31, 2026. No migration required. Existing voices, including custom clones, keep working as before.

The company built its case on a 500-sentence hard-case evaluation set, open-sourced on Hugging Face under a CC BY 4.0 license, covering five languages: English, German, French, Spanish, and Portuguese. Ten criteria structure the test, seven atomic ones (spelling, acronyms, alphanumeric strings, dates, regular numbers, large and floating numbers, email) and three composite ones (Orders, IT Ticket, Claims) that chain several of those into a single realistic agent turn. Scoring came from independent native-speaker raters working under strict pass rules, loudness-normalized audio, randomized order, and a 40-comparison cap with a mandatory break to guard against fatigue. Against that bar, and against rival models, Gradium is reporting its results.

Gradium AI has released a new text-to-speech model and made it the default across its API and Studio. The company reports an 81.0% human-rated pass rate on a 500-sentence hard-case set spanning five languages, ahead of Cartesia Sonic 3.6 at 75.1% and ElevenLabs v3 Conversational at 65.4%. Time to first audio is 216 ms at P50 on Coval, 170 ms faster than the model it replaces.

Why this matters

An 81% pass rate still means one in five hard sentences comes out wrong, and those are the ones with an order number or a callback digit attached. That gap is the whole ballgame for anyone building voice agents that handle real transactions, not demo scripts. Gradium's benchmark methodology, loudness-normalized audio, randomized order, raters capped at 40 comparisons, is more rigorous than most vendor self-reports we see, but it's still Gradium grading Gradium against competitors on Gradium's chosen test set.

We'd want to see Cartesia and ElevenLabs run the same 500-sentence set independently before treating 75.1% versus 65.4% as settled fact. The 216ms time-to-first-audio number matters more for founders shipping call-center or IVR products, where latency compounds with every turn. For researchers, the real story is that "hard-case" benchmarks spanning digits, emails, and callback numbers are becoming the standard yardstick, because word-error-rate on clean speech stopped telling anyone anything useful a while ago.

Watch whether Fish Audio and others publish rebuttal numbers.

Common Questions Answered

What specific problem does Gradium AI's new TTS model address in voice agents?

Gradium AI's new text-to-speech model targets the critical gap where voice agents fail to accurately read back order numbers, phone digits, and email addresses during transactions. While conversational ability may be strong, these specific moments are where traditional TTS models often falter, which is essential for real-world transactional voice agents.

How does Gradium AI's 81% pass rate compare to competing TTS models?

Gradium AI's new model achieved an 81.0% human-rated pass rate on a 500-sentence hard-case evaluation set, outperforming Cartesia Sonic 3.6 at 75.1% and ElevenLabs v3 Conversational at 65.4%. This benchmark was conducted across five languages using a rigorous methodology including loudness-normalized audio and randomized comparison order.

What is the time to first audio performance improvement in Gradium AI's new model?

The new Gradium AI TTS model achieves a time to first audio of 216 milliseconds at P50 on Coval, which represents a 170 millisecond improvement over the model it replaces. This faster response time enhances the user experience in real-time voice agent interactions.

Did existing Gradium AI users need to migrate to the new TTS model?

No migration was required when Gradium AI made the new TTS model the default across its API and Studio on August 31, 2026. Existing voices, including custom clones, continue to work as before, ensuring seamless adoption for current users.

Why is an 81% pass rate still significant despite meaning one in five hard sentences fails?

The 81% pass rate matters because the failing sentences are specifically those containing order numbers or callback digits—the critical moments in real transactions where accuracy is non-negotiable. For voice agents handling actual transactions rather than demo scripts, this performance gap represents the difference between successful and failed customer interactions.

LIVE06:22Gradium AI’s New TTS Model Scores 81% Pass Rate