Editorial illustration for Alibaba's Qwen3.8-Max Writes 7,600 Lines of Code in Five Days
Alibaba's Qwen3.8-Max Codes 7,600 Lines in 5 Days
Alibaba's Qwen team put out Qwen3.8-Max this week, a 2.4-trillion-parameter model with 95 billion active parameters per query, and the pitch is different from the usual chatbot upgrade. Instead of chasing better one-shot answers, the team built it to grind through jobs that take days, not seconds. Think reproducing a research paper's results and then improving on them, or running a simulated e-commerce operation from end to end without a human checking in every few minutes.
The model builds on the Qwen3.5 architecture and follows a preview version Alibaba floated back in mid-July, offered through the Token Plan, Qoder, and QoderWork at a tenth of standard pricing. That early version already carried the 2.4 trillion parameter count and got ranked just behind Fable 5, though Alibaba held back benchmark numbers at the time. Qwen3.8-Max also marks a first for the Qwen-Max line: it's the first model in that class where Alibaba plans to release the actual weights, expected next week.
To back up the claims about long-horizon performance, the team ran the model through a set of coding tests meant to show what it can do left alone for extended stretches.
Alibaba's new flagship model Qwen3.8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips autonomously.
Why this matters Five days of unattended compute producing 7,600 lines of code and 33 training runs is the kind of number that should make anyone building AI agents sit up, but it's also the kind of number that deserves a second look before we start rewriting research workflows around it. Alibaba's own team ran the benchmarks, on a paper it apparently chose, with success criteria it presumably defined. That's not a knock on Qwen3.8-Max so much as a reminder that "beat the paper's method" needs independent replication before it becomes a hiring decision for some lab's next hire.
For developers, the real signal isn't the parameter count, it's the shift toward multi-day autonomous runs as a benchmark category at all. If open-weight models can plausibly run a simulated business or extend a paper's results without a human checking in every hour, procurement conversations change, and so does the calculus on what tasks you still assign to a person. Founders should watch for third-party reproductions of these exact tasks.
Researchers should ask which 18 ideas failed before the four that worked, because that ratio matters more than the highlight reel.
Common Questions Answered
What are the key specifications of Alibaba's Qwen3.8-Max model?
Qwen3.8-Max is a 2.4-trillion-parameter model with 95 billion active parameters per query. Unlike typical chatbot upgrades focused on one-shot answers, this model is specifically designed to handle complex tasks that take days to complete without human intervention.
What types of long-horizon tasks can Qwen3.8-Max perform autonomously?
Qwen3.8-Max can handle complex tasks such as reproducing research paper results and improving upon them, designing chips autonomously, and running simulated end-to-end e-commerce operations. The model demonstrated its capability by producing 7,600 lines of code and completing 33 training runs over a five-day period without human oversight.
How does Qwen3.8-Max differ from traditional chatbot model upgrades?
Rather than chasing better one-shot answers like typical chatbot upgrades, Qwen3.8-Max is built to grind through jobs that take days or weeks to complete. This focus on long-horizon task completion represents a different approach to AI model development, targeting AI agents and research workflows rather than immediate response quality.
What important considerations should be kept in mind when evaluating Qwen3.8-Max's benchmark results?
The benchmarks were conducted by Alibaba's own team on a paper they selected, with success criteria they defined, which means the results warrant careful scrutiny before restructuring research workflows around them. This highlights the importance of independent verification and understanding the specific conditions under which the model was tested.
Further Reading
- Alibaba unveils latest AI model with enhanced coding, reasoning capabilities - Xinhua
- Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot's Kimi K3 Open-Weight Launch - MarkTechPost
- Qwen3.8 Benchmarks: A Coding-Agent Evaluation Guide - NXCode
- Qwen3.8-Max Review: I Tested Alibaba's 2.4T Model - Thomas Wiegold
- Alibaba's Qwen Unveils Preview of Flagship AI Model - Yahoo Finance