Editorial illustration for Qwen AI Challenges Gemini Flash on Price and Performance
Qwen AI Challenges Gemini Flash on Price and Performance
Alibaba's Qwen team put a price tag on its new multimodal model this week that looks aimed squarely at Google. Qwen3.8-Omni-Flash, released as the company's first multimodal system built for AI agents, handles audio and video together, draws its own conclusions, and can act on them without a human nudging each step. Think editing a vlog, translating a short clip, or summarizing a two-hour movie, all inside a single one-million-token context window.
The bigger story is what Qwen charges for it. API access runs $0.15 per million input tokens and $0.47 per million output tokens, a fraction of what Google asks for Gemini 3.8 Flash. Qwen says its model gets close to matching Gemini on audio-video benchmarks despite that gap.
The model ships through Qwen Studio, Qwen Cloud, and the API, backed by open-source Qwen-MM-Plugins for video editing, speaker recognition, and PDF video notes that plug into agent tools like Claude Code and Gemini CLI. A companion tool, Qwen-Live Harness, adds real-time camera and microphone interaction. Here's how Qwen's pricing stacks up against Google's.
API pricing sits at $0.15 per million input tokens and $0.47 per million output tokens. Qwen estimates audio input at under $0.01 per hour, while 720p video with audio at one frame per second runs about $0.20, not counting response costs. For comparison, Gemini 3.8 Flash charges $0.75 for input and $3.75 for output per million tokens at its introductory rate, with prices set to double on January 1, 2027.
Why this matters
For teams building agents that watch video and listen to audio at the same time, the math here is hard to ignore. Alibaba is pricing Qwen3.8-Omni-Flash at $0.15 per million input tokens and $0.47 per million output tokens, with 720p video at one frame per second running about $0.20 an hour. That's the kind of number that changes what's economical to build, not just what's technically possible. A one-million-token context window means longer videos and longer meetings can go in whole, without the usual chunking workarounds.
The catch is that "comes close to matching Gemini 3.8 Flash" is Qwen's own framing, not an independent benchmark, so anyone weighing this for production should test on their own audio-video workloads before trusting the comparison. Still, the fact that a Chinese lab is undercutting Google on price while claiming near-parity on multimodal tasks tells us the Flash-tier price war is real and Google no longer sets that market alone. Worth watching: whether Google responds with its own price cuts, and whether third-party benchmarks back up Qwen's claims once developers start running real vlog-editing and translation workloads against both.
Further Reading
- Qwen3.8-Omni-Flash: Omni Senses. Agentic Delivery. - Qwen
- Alibaba Launches Qwen3.8-Omni-Flash: Native Multimodal, Million ... - AIbase
- Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use - MarkTechPost
- Qwen3.8-Omni-Flash — Alibaba's 1M-context omni model - AI TL;DR
- Alibaba's Qwen3.8-Omni-Flash Cuts Video AI Costs by 89 ... - AlphaSignal