Skip to main content
Qwen3-Max Thinking, Gemini 3 Pro, and GPT-5.2 AI models on a leaderboard, illustrating AI exam performance. [arxiv.org](https

Editorial illustration for Qwen3-Max Thinking Beats Gemini 3 Pro, GPT-5.2 on Humanity's Last Exam

Qwen3-Max Beats Gemini Pro in Ultimate AI Reasoning Test

Qwen3-Max Thinking Beats Gemini 3 Pro, GPT-5.2 on Humanity's Last Exam

Updated: 3 min read

Alibaba's latest model just scored higher on a major benchmark than Google's and OpenAI's best. It didn't win by being bigger. It won by being thrifty.

Qwen3-Max Thinking now tops Gemini 3 Pro and GPT-5.2 on Humanity's Last Exam, a final boss for AI reasoning. The usual story is more data, more parameters. This one is about using less.

The model has a trick. Instead of barreling down a single path, it runs a tight internal feedback loop. It critiques its own half-formed thoughts, spots dead-ends early, and funnels its computational energy into the parts of a problem that remain fuzzy.

The technique is called "test-time scaling." It trades raw compute for a more deliberate, almost cautious, form of intelligence.

The Architecture: "Test-Time Scaling" Redefined The core innovation driving Qwen3-Max-Thinking is a departure from standard inference methods. While most models generate tokens linearly, Qwen3 utilizes a "heavy mode" driven by a technique known as "Test-time scaling." In simple terms, this technique allows the model to trade compute for intelligence. But unlike naive "best-of-N" sampling--where a model might generate 100 answers and pick the best one -- Qwen3-Max-Thinking employs an experience-cumulative, multi-round strategy.

When the model encounters a complex query, it doesn't just guess; it engages in iterative self-reflection. It uses a proprietary "take-experience" mechanism to distill insights from previous reasoning steps. This allows the model to: Identify Dead Ends: Recognize when a line of reasoning is failing without needing to fully traverse it.

Focus Compute: Redirect processing power toward "unresolved uncertainties" rather than re-deriving known conclusions. By avoiding redundant reasoning, the model integrates richer historical context into the same window. The Qwen team reports that this method drove massive performance jumps without exploding token costs: GPQA (PhD-level science): Scores improved from 90.3 to 92.8.

That GPQA jump, from 90.3 to 92.8, is significant. It's the kind of gain you'd expect from a whole new generation of hardware, not a smarter way to use the old one. The benchmark victory matters.

What matters more is the shift it represents. Intelligence is being redefined not as sheer output but as efficient allocation. The model stops when it's stuck.

It doubles down only on confusion.

This feels less like engineering and more like a crude cognitive strategy. It's learning to waste less time. The frontier is no longer just about asking models harder questions.

It's about watching them develop an internal economy of thought, a reluctance to commit until they've checked their own work. The next challenge is figuring out what to do with a machine that hesitates.

Common Questions Answered

What is the key innovation of Qwen3-Max-Thinking's 'Test-Time Scaling' approach?

Qwen3-Max-Thinking introduces a novel approach to model inference that allows trading computational resources for enhanced intelligence. Unlike traditional linear token generation, this technique enables the model to dynamically switch to a 'heavy mode' for more complex reasoning tasks, potentially improving performance on challenging benchmarks.

How does Qwen3-Max-Thinking perform on the 'Humanity's Last Exam' benchmark?

The model reportedly outperformed both Gemini 3 Pro and GPT-5.2 on this challenging benchmark, which tests reasoning, mathematical, and commonsense capabilities. However, the result is based on a single, highly stylized test, and broader performance verification remains pending.

What makes the Qwen3-Max model unique in the current AI landscape?

Qwen3-Max stands out with its massive 1T parameter scale and 36T tokens of pre-training data, utilizing an advanced Mixture of Experts (MoE) architecture. The model introduces a groundbreaking thinking mode that allows for dynamic computational resource allocation, enabling more sophisticated reasoning across different types of tasks.

LIVE02:29Brain Waves Could Guide AI on When to Learn, Neuroscientist Says