Skip to main content
Alibaba Qwen3.8-Max AI model, 2.4 trillion parameters, represented by a complex neural network graphic.

Editorial illustration for Alibaba's Qwen3.8-Max Model Hits 2.4 Trillion Parameters

Alibaba's Qwen3.8-Max Reaches 2.4T Parameters

3 min read

Alibaba's Qwen team put out Qwen3.8-Max this week, a mixture-of-experts model with 2.4 trillion total parameters, the largest in the Qwen lineup so far. It takes text, image and video as input and returns text, and it's live now through Alibaba's hosted API, which is compatible with both OpenAI and DashScope endpoints. Open weights for Qwen3.8-Max are due next week, alongside a smaller checkpoint, Qwen3.8-27B, also headed for open release.

The scale here raises an obvious question: who can actually run this thing outside a rented API. Alibaba hasn't disclosed the activated-parameter count for the MoE architecture, which means the real serving cost, and whether this fits on anything short of a multi-node datacenter cluster, is still unclear. The benchmark numbers Alibaba has published tell a more specific story about where the model actually improved versus its predecessor, particularly on coding and agentic tasks rather than raw reasoning scores. Some of those comparisons come with caveats worth flagging before taking the headline figures at face value.

The open weights are a different matter. At 2.4T total parameters, the checkpoint is a multi-node datacenter artifact. Alibaba has not disclosed the activated-parameter count.

Why this matters

For teams evaluating Qwen against GPT and Gemini-class models, the two-checkpoint strategy is the story here, not the parameter count. Alibaba is giving enterprises a 2.4-trillion-parameter flagship for benchmark bragging rights while shipping Qwen3.8-27B as the model people will actually deploy on their own racks. That split matters more than most model announcements because it addresses the real bottleneck for founders and researchers outside hyperscaler budgets: GPU access.

If Qwen3.8-27B performs anywhere close to its bigger sibling on the four named use cases (software engineering, legal and financial review, media and e-commerce, design), it becomes a serious open-weights option for companies that can't run trillion-parameter MoE models but still want multimodal input. We'd treat the "most capable Qwen yet" framing with normal skepticism until independent benchmarks land alongside next week's weight release. Watch whether Qwen3.8-27B's on-premise performance holds up under real workloads, not just the vendor's own framing, since that's what will determine if this becomes a genuine alternative to closed frontier models for cost-conscious builders.

Common Questions Answered

What are the key specifications of Alibaba's Qwen3.8-Max model?

Qwen3.8-Max is a mixture-of-experts model with 2.4 trillion total parameters, making it the largest in the Qwen lineup to date. The model accepts text, image, and video as input while returning text output, and is currently available through Alibaba's hosted API with compatibility for both OpenAI and DashScope endpoints.

Why is Alibaba releasing both Qwen3.8-Max and Qwen3.8-27B as separate checkpoints?

Alibaba's two-checkpoint strategy addresses the real GPU access bottleneck for enterprises and researchers outside hyperscaler budgets. The company is providing the 2.4-trillion-parameter Qwen3.8-Max as a flagship model for benchmark performance, while Qwen3.8-27B serves as the practical model that teams can actually deploy on their own infrastructure.

When will open weights for Qwen3.8-Max become available?

Open weights for Qwen3.8-Max are scheduled for release next week, alongside the smaller Qwen3.8-27B checkpoint. Both models will be made available for open release to the community.

What is the significance of the activated-parameter count not being disclosed for Qwen3.8-Max?

Alibaba has not disclosed the activated-parameter count for the 2.4 trillion parameter model, which is notable because at this scale the checkpoint represents a multi-node datacenter artifact. The undisclosed activated-parameter count means the actual number of parameters used during inference may be significantly lower than the total parameter count, which is typical for mixture-of-experts models.

LIVE11:39OpenAI's Astra Solves 10 Long-Standing Math Problems