Skip to main content
Alibaba Qwen-Audio-3.1 AI model, two neural networks for voice decisions, advanced audio processing.

Editorial illustration for Alibaba's Qwen-Audio-3.1 Uses Two Models for Voice Decisions

Alibaba's Qwen-Audio-3.1 Dual Models for Voice Agents

• 4 min read

Alibaba's Qwen team put out five audio models this week, and the one drawing attention is Qwen-Audio-3.1-Realtime, a full-duplex speech system built for voice agents that need to call tools mid-conversation. The release comes with a price cut across the board: roughly 85% off Realtime, 70% off text-to-speech, and up to 95% off automatic speech recognition. None of it ships as open weights. Instead, Qwen is running it as a managed API, with qwen-audio-3.1-realtime-plus live on QwenCloud over WebSocket.

The model page spells out the specs: 262K tokens of context, 245K max input, 16K max output, capped by default at 60 requests and 100K tokens per minute. Audio input runs $6.4 per million tokens, text input $0.8, and combined text-and-audio output $24, with text output itself free. Function calling, web search, structured outputs, context caching and fine-tuning are all built in. A second model, Qwen-Audio-3.1-ASR-Flash-Filetrans, handles offline transcription of long audio, with support for hot words, speaker separation and Chinese dialects.

What makes the Realtime model different isn't the price tag. It's how Qwen split the job of holding a conversation into two separate systems working in tandem.

Alibaba’s Qwen team has released Qwen-Audio-3.1, a 5-model audio stack spanning ASR, TTS and realtime interaction. The main model is Qwen-Audio-3.1-Realtime, a full-duplex speech model built for voice agents that call tools. Qwen also cut prices: about 85% on Realtime, about 70% on TTS and up to 95% on ASR.

Why this matters

Splitting the decision model from the speech renderer is a practical fix for the thing that makes voice agents feel broken: they don't know when to shut up. By handing turn-taking to a dedicated model rather than baking it into one giant end-to-end system, Qwen is treating "when to speak" as its own engineering problem, separate from "what to say." That's a useful distinction for anyone building voice agents that need to call tools mid-conversation without stepping on the user.

The catch is that this only ships as a managed API on QwenCloud, no open weights, so builders get the behavior but not the internals. Combined with steep price cuts (85% on Realtime, up to 95% on ASR), Alibaba looks to be racing OpenAI and Google on cost and latency for voice infrastructure rather than on openness. For founders, that's a cheaper backend to build on.

For researchers, it's another capable system you can only study from the outside. Worth watching whether Qwen eventually open-sources any piece of this stack, and how the two-model approach holds up under real multi-turn tool-calling load rather than demo conditions.

Common Questions Answered

What is the main architectural difference in Qwen-Audio-3.1-Realtime compared to traditional voice agents?

Qwen-Audio-3.1-Realtime uses two separate models instead of one end-to-end system: a dedicated decision model for turn-taking and a speech renderer for audio output. This split architecture solves a key problem in voice agents where they don't know when to stop speaking, treating 'when to speak' as a separate engineering challenge from 'what to say.'

What pricing reductions did Alibaba announce for the Qwen-Audio-3.1 models?

Alibaba cut prices dramatically across its audio model stack, with approximately 85% off Qwen-Audio-3.1-Realtime, about 70% off text-to-speech, and up to 95% off automatic speech recognition. These price reductions make the audio models significantly more accessible for developers building voice applications.

How does Qwen-Audio-3.1-Realtime handle tool calling during conversations?

Qwen-Audio-3.1-Realtime is a full-duplex speech system specifically built for voice agents that need to call tools mid-conversation without interrupting the user's experience. The dedicated decision model manages turn-taking to ensure the agent knows when to speak and when to listen while executing tool calls.

Why did Alibaba release Qwen-Audio-3.1 as a managed API rather than open weights?

Alibaba chose to deploy Qwen-Audio-3.1 as a managed API service rather than releasing it as open weights, with the main model qwen-audio-3.1-realtime-plus available on QwenCloud over WebSocket. This approach allows Alibaba to maintain control over the service while providing developers with easy access through their infrastructure.

What audio capabilities are included in Alibaba's five-model Qwen-Audio-3.1 stack?

The Qwen-Audio-3.1 stack spans three main audio capabilities: automatic speech recognition (ASR) for converting speech to text, text-to-speech (TTS) for generating audio from text, and real-time interaction through the Qwen-Audio-3.1-Realtime model. Together, these five models provide a comprehensive audio processing solution for voice agent applications.

LIVE08:29Anthropic warns its AI could end humanity as investors profit