Editorial illustration for New AI Model Claims Lead Over OpenAI on Software Engineering Test
Alibaba's Qwen3.8 Outperforms OpenAI on Code Tests
New AI Model Claims Lead Over OpenAI on Software Engineering Test
Alibaba's Qwen team put out a new flagship model overnight, and the numbers it's claiming are aimed squarely at OpenAI and Anthropic's home turf. Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model built for agentic computer use, the kind of long-horizon task where an AI has to navigate software, write code, and reproduce research without constant hand-holding. On OSWorld-Verified, a benchmark that tracks how well models operate a computer autonomously, Qwen says its new system scored 86.1, ahead of both GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0. Alibaba also reports the top score on PaperBench, which tests a model's ability to replicate published research.
There's a second piece to this launch that matters as much as the scores. Alibaba plans to release Qwen3.8-Max's weights next week alongside a smaller Qwen3.8-27B model, something no previous Max-tier Qwen release has done. Whether that openness actually holds up, and under what license, is still unknown. Moonshot's recent Kimi K3 release showed that "open" doesn't always mean unrestricted.
Most notably, Qwen reports that Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark measuring how well ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0), while also posting the highest reported score on PaperBench and leading or remaining highly competitive across software engineering, research reproduction, multimodal reasoning, and visual web development benchmarks.
Why this matters
For developers picking a model to wire into an agentic coding pipeline, Qwen3.8-Max's numbers are worth logging, not worshipping. Alibaba is claiming wins on agentic computer-use tasks while conceding, in its own disclosures, that OpenAI's model still tops SWE-Pro and Anthropic's Opus 4.8 leads on other software benchmarks and Agents' Last Exam. That's a more honest framing than most launch posts offer, and it's also a signal: no single lab is running the table anymore on the benchmarks that actually predict whether an agent can ship working code unsupervised.
For founders building on these models, the practical move is to treat Qwen3.8-Max as a strong option for specific workloads, not a wholesale replacement for GPT-5.6 or Claude. For researchers, the interesting question is why performance is fragmenting by task type rather than converging, since a 2.4-trillion-parameter MoE model besting rivals on one axis while losing on others suggests architecture and training choices matter more than raw scale right now. Watch for independent benchmark reruns before trusting Alibaba's own numbers.
Common Questions Answered
How does Qwen3.8-Max perform on the OSWorld-Verified benchmark compared to competitors?
Qwen3.8-Max scores 86.1 on the OSWorld-Verified benchmark, which measures autonomous computer operation capabilities. This score places it ahead of GPT-5.6 Sol Max (83.2) and Fable 5 (85.0), demonstrating Alibaba's claim to leadership in agentic computer-use tasks.
What type of model architecture is Qwen3.8-Max and what is it designed for?
Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model built specifically for agentic computer use. It is designed to handle long-horizon tasks where AI must navigate software, write code, and reproduce research with minimal human intervention.
In which software engineering benchmarks does Qwen3.8-Max claim the highest scores?
Qwen reports that Qwen3.8-Max posts the highest reported score on PaperBench and leads or remains highly competitive across software engineering, research reproduction, multimodal reasoning, and visual web development benchmarks. However, the company acknowledges that OpenAI's model still tops SWE-Pro and Anthropic's Opus 4.8 leads on other software benchmarks.
Why does Alibaba's framing of Qwen3.8-Max's capabilities stand out compared to typical model launches?
Alibaba provides a more honest assessment by acknowledging areas where competitors still lead, such as OpenAI's performance on SWE-Pro and Anthropic's Opus 4.8 on other benchmarks. This transparent approach contrasts with typical launch announcements that often overstate capabilities without conceding competitive weaknesses.
What does Qwen3.8-Max's competitive performance indicate about the current state of AI model development?
Qwen3.8-Max's strong showing suggests that no single AI lab is dominating across all benchmarks and capabilities anymore. Different models now excel in different areas, indicating a more distributed landscape of AI advancement rather than one clear leader.
Further Reading
- Alibaba says newest Qwen AI model is second only to Anthropic’s Claude Fable 5 - South China Morning Post
- Alibaba Rallies After Unveiling New Qwen AI Model - Yahoo Finance
- Qwen3.7: The Agent Frontier - Qwen / Alibaba
- Alibaba unveils Qwen3.5 as China's chatbot race shifts to agentic AI - CNBC
- Alibaba Unveils Cutting-Edge AI Coding Model Qwen3-Coder - Alibaba Cloud Press Room