Editorial illustration for Open TTS Leaderboard Adds Speaker Similarity Column, Ranks Top Multilingual Models
TTS Leaderboard Adds Speaker Similarity, Ranks Models
Hugging Face counted more than 8,000 text-to-speech models on its Hub as of September 30, 2026. That number alone explains why ranking them has turned into a real problem. Arena-style leaderboards, where listeners pick between two synthesized clips to generate Elo scores, have become the default reference for the TTS community, following the same Bradley-Terry approach used elsewhere. But these arenas were built for a slower era of model releases, and they're showing strain.
The gap shows up clearest in who gets ranked at all. On Artificial Analysis, only 16 of 92 listed models are open-weights, with Voice Arena showing a similar imbalance toward commercial, API-based systems. Getting an API model onto a leaderboard takes little more than a key. Hosting and serving an open model takes infrastructure, and open-source authors have less incentive to chase a listing than commercial providers do.
There's a second problem beyond coverage: consistency. Human raters drift. An arena has no real mechanism to guarantee the same voters are applying the same standard of "better" across sessions, let alone across months of new releases.
While human preference is the ultimate decider, arenas cannot scale to keep up with the pace of TTS releases. This may partly explain why open-source models are underrepresented on arena-style leaderboards: as of Sep 30, 2026, only 16 of the 92 models on Artificial Analysis are open-weights, with a similar skew on Voice Arena.
Why this matters
For anyone building voice products, the SIM column matters more than another leaderboard refresh. Word error rate tells you if a model transcribes cleanly, not whether it sounds like the person you fed it. With k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 now ranked on speaker similarity alongside the new Pareto plots for SIM versus inference speed and model size, developers finally get a way to weigh voice cloning fidelity against latency and compute cost in one view, rather than guessing from marketing claims or cherry-picked demos.
With over 8,000 TTS models on the Hugging Face Hub as of September 2026, that kind of sorting is overdue. We're skeptical that any single benchmark captures everything MOS or MUSHRA testing would, since human listening panels still catch nuance automated metrics miss. But standardizing comparison across this many open releases beats the current mess of scattered claims.
Worth watching: whether smaller labs start optimizing specifically for SIM scores, and whether that improves real-world voice cloning or just gamifies the leaderboard.
Common Questions Answered
Why did Hugging Face add a Speaker Similarity column to the Open TTS Leaderboard?
The Speaker Similarity column was added because word error rate alone doesn't indicate whether a text-to-speech model sounds like the original speaker, which is crucial for voice cloning applications. For developers building voice products, speaker similarity metrics provide essential information about voice cloning fidelity that traditional evaluation methods fail to capture.
What scalability problem do arena-style TTS leaderboards face according to the article?
Arena-style leaderboards, which use human preference voting to generate Elo scores, cannot scale to keep up with the rapid pace of text-to-speech model releases. As of September 30, 2026, Hugging Face counted more than 8,000 TTS models on its Hub, making it impossible for traditional arenas to evaluate all models in a timely manner.
Why are open-source models underrepresented on existing arena-style TTS leaderboards?
Open-source models are underrepresented because arena-style leaderboards built for slower model release cycles cannot scale to evaluate the growing number of open-weights models being released. As of September 30, 2026, only 16 of the 92 models on Artificial Analysis were open-weights, demonstrating this significant skew away from open-source options.
What new evaluation metrics does the Open TTS Leaderboard provide alongside speaker similarity?
The Open TTS Leaderboard now includes Pareto plots that allow developers to weigh speaker similarity against inference speed and model size. This enables builders of voice products to make informed trade-offs between voice cloning fidelity, latency performance, and computational requirements when selecting models.
Which multilingual TTS models are highlighted as top performers on the updated leaderboard?
The article highlights k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 as top-ranked models now evaluated on speaker similarity metrics. These models represent the current state-of-the-art in multilingual text-to-speech with strong performance across the new evaluation criteria.
Further Reading
- IndexTTS 2.5 Technical Report - arXiv
- Qwen-Audio-3.0-TTS: Freely Controllable and Highly Expressive Speech Synthesis - arXiv
- Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot Voice Transfer - arXiv
- Zero-shot Cross-lingual Voice Transfer for TTS - arXiv
- MiniMax-Speech: A high-performance speech synthesis model supporting 32 languages - arXiv