Skip to main content
Alibaba Qwen Audio 3.0 TTS Plus ranking above Google Gemini and Sonic on a digital leaderboard.

Editorial illustration for Alibaba's Qwen Audio 3.0 TTS Plus Leads Rankings Over Gemini, Sonic

Alibaba's Qwen Audio 3.0 TTS Tops Speech Rankings

4 min read

Alibaba's Qwen-Audio-3.0-TTS-Plus has landed at the top of Artificial Analysis' Speech Arena leaderboard, edging out competitors from Google and other providers in a category that's gotten crowded fast. The model posted an Elo score of 1,236, narrowly beating Simba 3.2 at 1,234. Google's Gemini 3.1 Flash TTS trailed at 1,214, with Sonic 3.5 rounding out the top four at 1,207.

Alibaba split the release into two versions aimed at different jobs. Flash handles real-time conversation with latency around 300 milliseconds, while Plus is built for output quality rather than speed. The model covers 16 languages, reaching into territory competitors often skip, like Tagalog, Malay, Thai, and Vietnamese, plus a handful of Chinese dialects. Users can shape delivery through plain-language instructions or drop in tags like "[angry]" or "[giggles]" to push specific emotional beats into the audio.

Pricing runs $27.60 per million characters through Alibaba Cloud Model Studio, and the company has published a set of audio samples for anyone who wants to hear the output directly. Where the model stands relative to rivals on raw generation speed, and what that tradeoff means for developers picking a provider, is where the numbers get interesting.

Alibaba's new text-to-speech model, Qwen-Audio-3.0-TTS-Plus, leads Artificial Analysis' Speech Arena leaderboard for provider voices. With an Elo score of 1,236, it sits just ahead of Simba 3.2 (1,234). Gemini 3.1 Flash TTS (1,214) and Sonic 3.5 (1,207) follow behind.

Why this matters

The Elo gap between Qwen-Audio-3.0-TTS-Plus (1,236) and Simba 3.2 (1,234) is two points. That's noise, not dominance, and it tells us the TTS leaderboard has become crowded enough that "leads the rankings" is a fragile headline, good for a press cycle, not much else. What's more useful here is the language coverage: Tagalog, Malay, Thai, Vietnamese support puts Alibaba ahead of Gemini and Sonic for teams building products outside the usual English/Mandarin/Spanish default.

If you're shipping voice products for Southeast Asian markets, that's the real signal, not the Elo score. The Flash/Plus split also matters practically: 300ms latency for Flash means real conversational agents become viable, while Plus is clearly optimized for narration, dubbing, or anything where output quality outweighs response time. Anyone benchmarking TTS vendors right now should treat Artificial Analysis' Speech Arena as a rough proxy, not gospel.

Ask for the underlying voice samples and check language coverage against your actual user base before you commit engineering time to switching providers over a two-point Elo difference.

Common Questions Answered

What Elo score did Alibaba's Qwen-Audio-3.0-TTS-Plus achieve on the Artificial Analysis Speech Arena leaderboard?

Alibaba's Qwen-Audio-3.0-TTS-Plus achieved an Elo score of 1,236 on the Artificial Analysis Speech Arena leaderboard, placing it at the top of the rankings. This score narrowly edges out its closest competitor, Simba 3.2, which scored 1,234, demonstrating a very competitive landscape in text-to-speech models.

How does Qwen-Audio-3.0-TTS-Plus compare to Google's Gemini 3.1 Flash TTS in the rankings?

Qwen-Audio-3.0-TTS-Plus leads Google's Gemini 3.1 Flash TTS by 22 Elo points, with a score of 1,236 compared to Gemini's 1,214. This places Gemini in third position on the Speech Arena leaderboard, behind both Qwen and Simba 3.2.

What language coverage advantage does Qwen-Audio-3.0-TTS-Plus offer over competitors?

Qwen-Audio-3.0-TTS-Plus provides support for languages including Tagalog, Malay, Thai, and Vietnamese, which puts Alibaba ahead of competitors like Gemini and Sonic for teams building products outside the typical English, Mandarin, and Spanish defaults. This expanded language coverage is particularly valuable for developers targeting Southeast Asian and other emerging markets.

Why does the article suggest the Elo gap between top competitors represents 'noise' rather than dominance?

The article points out that the two-point Elo difference between Qwen-Audio-3.0-TTS-Plus (1,236) and Simba 3.2 (1,234) is minimal enough to be considered statistically insignificant, indicating the text-to-speech leaderboard has become crowded with highly competitive models. This narrow margin suggests that claiming definitive leadership is more of a temporary press headline than a meaningful technical advantage.

What are the two different versions Alibaba released for Qwen-Audio-3.0-TTS-Plus?

Alibaba split the Qwen-Audio-3.0-TTS-Plus release into two versions designed for different purposes: Flash, which handles real-time conversation with low latency requirements, and Plus, which is optimized for other use cases. This dual-version approach allows developers to choose the model variant that best fits their specific application needs.

Further Reading

LIVE19:09Report: US Weighs Ban on Chinese AI Models Amid IP Theft Concerns