Editorial illustration for Alibaba Slashes AI Audio Prices 95%, Launches Qwen 3.1 Models
Alibaba Cuts AI Audio Prices 95%, Launches Qwen 3.1
Alibaba's Qwen team put out Qwen-Audio-3.1 this week, a set of five models covering speech recognition, text-to-speech, and real-time voice interaction. The release lands alongside a sharp price cut across the board, with Alibaba positioning the update as both a technical upgrade and a cost play aimed at developers building voice products on its cloud platform.
The ASR model is the headline act here. Alibaba says it handles multiple languages and regional dialects more accurately than prior versions, and it can automatically strip out filler words like "um" and "uh" along with repeated phrases, a detail that matters for anyone building transcription tools or voice assistants where clean output saves downstream cleanup work.
Pricing is where Alibaba is making its biggest pitch. TTS costs are down around 70 percent, real-time interaction pricing drops roughly 85 percent, and ASR pricing falls by as much as 95 percent depending on usage. That kind of cut on the recognition side in particular could shift the calculus for companies weighing Qwen against rival audio APIs. Alibaba published full specs and pricing details on its official blog and through Qwen Cloud.
The real-time model supports simultaneous speaking and listening with instant interruption. When it detects a low mood, it responds more slowly and with more empathy, according to Qwen. Alibaba is also slashing prices.
Why this matters
A 95 percent price cut on audio models isn't a rounding error, it's Alibaba signaling that speech recognition and synthesis are becoming commodity infrastructure rather than premium features. For developers building voice products, that math changes fast: workflows that were too expensive to run at scale, call center transcription, multilingual dubbing, real-time translation, suddenly pencil out. The technical additions matter too.
Multi-speaker identification with timestamps and emotion detection in ASR-Next means teams can skip building their own diarization layers. Cross-language voice transfer in the TTS model, controllable through plain text prompts, lowers the bar for anyone wanting expressive, multilingual audio without hiring voice actors or writing SSML by hand.
We'd watch two things from here. First, whether this pricing holds once usage scales, or whether it's a land-grab move to pull developers off OpenAI's and Google's audio APIs before those companies respond. Second, whether the "clean up filler words automatically" feature actually holds up on messy, accented, real-world audio, not just benchmark clips.
Alibaba's making a strong bet that audio AI is now infrastructure, not a feature. If the quality matches the price, that bet pays off fast.
Common Questions Answered
What are the five models included in Alibaba's Qwen-Audio-3.1 release?
Qwen-Audio-3.1 comprises five models covering speech recognition, text-to-speech, and real-time voice interaction capabilities. The ASR (automatic speech recognition) model is positioned as the headline feature, handling multiple languages and regional dialects with improved accuracy compared to previous versions.
How much did Alibaba reduce prices for its AI audio models?
Alibaba slashed AI audio prices by up to 95 percent across the board with the Qwen-Audio-3.1 release. This dramatic price reduction positions speech recognition and synthesis as commodity infrastructure rather than premium features, making previously expensive workflows like call center transcription and real-time translation economically viable at scale.
What are the key capabilities of Qwen's real-time voice interaction model?
The real-time model supports simultaneous speaking and listening with instant interruption capabilities, allowing natural conversational flow. Additionally, it can detect user mood and respond accordingly, speaking more slowly and with greater empathy when it identifies a low mood in the user.
What use cases become feasible with the 95% price reduction on Qwen audio models?
The dramatic price cut makes several previously expensive workflows economically viable, including call center transcription, multilingual dubbing, and real-time translation. Developers can now build voice products at scale without the cost constraints that previously limited deployment of these audio AI features.
Further Reading
- Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully ... - Alibaba_Qwen
- 阿里云百炼部分语音系列模型降价 - Futu News
- 阿里千问发布Qwen-Audio-3.1:五款语音模型同发,ASR 价格直降95% - Sohu
- 阿里千问发布Qwen-Audio-3.1系列语音大模型 - Sohu
- 全线降价!阿里密集发布5款语音大模型,最高降幅达95% - Sina Finance