Editorial illustration for Gemini 3 Deep Think Boosts Reasoning with Mathematical and Algorithmic Rigor
Gemini 3 Deep Think: AI Reasoning Breakthrough
Gemini 3 Deep Think Boosts Reasoning with Mathematical and Algorithmic Rigor
The hardest tests we have for human intelligence are now being passed by a machine.
Gemini 3 Deep Think just posted a score of 48.4% on Humanity’s Last Exam, which is meant to stump the most advanced models. It did that without any tools. It scored 84.6% on ARC-AGI-2.
Its competitive programming rank, an Elo of 3455 on Codeforces, is staggering. It performed at a gold-medal level on the 2025 International Math Olympiad.
This isn't just about math. The same system now hits gold-medal standards on the written sections of the 2025 International Physics and Chemistry Olympiads. The claim is that a single approach, built on mathematical and algorithmic rigor, can conquer these wildly different fields.
Today, we’re releasing a major upgrade to Gemini 3 Deep Think, our specialized reasoning mode, built to push the frontier of intelligence and solve modern challenges across science, research, and engineering.
Benchmarks are just numbers. The shift here is a kind of proof. If one method can produce gold-medal results in high school math, competitive coding, university-level physics, and theoretical chemistry, then the method itself is the story.
The machine isn't just retrieving answers. It is applying a consistent, verifiable process to problems that require deep, structured thought. We are watching a tool learn how to think in a way we recognize as rigorous.
That changes what the tool is for.
Common Questions Answered
What makes Gemini 3 Deep Think different from previous AI models in reasoning capabilities?
Gemini 3 Deep Think introduces Advanced Parallel Reasoning, which explores multiple hypothesis paths simultaneously instead of following a single chain of thought. This 'System 2' approach allows the model to pause, explore multiple hypotheses, and critically critique its own logic before generating an output, marking a significant departure from traditional 'System 1' language models that simply predict the next token.
How did Gemini 3 Deep Think perform on challenging academic benchmarks?
Gemini 3 Deep Think demonstrated exceptional performance on rigorous benchmarks like Humanity's Last Exam, achieving 41.0% accuracy without using external tools, and ARC-AGI-2, where it scored an unprecedented 45.1% with code execution. These results build on previous achievements, including gold-medal level performances at the International Mathematical Olympiad and International Collegiate Programming Contest World Finals.
What is the key cognitive architecture difference between System 1 and System 2 reasoning in AI?
System 1 reasoning, typical of standard Large Language Models, is fast, automatic, and impulsive - like a student quickly blurting out an answer during a pop quiz. In contrast, System 2 reasoning, exemplified by Gemini 3 Deep Think, is slow, deliberate, and analytical - similar to a mathematician carefully working through a complex proof, allocating computational time to thinking and exploring multiple logical paths before generating an output.
Further Reading
- A new era of intelligence with Gemini 3 — Google Blog
- Gemini Deep Think: Redefining the Future of Scientific Research — Google DeepMind Blog
- Gemini 3 Deep Think: Enable Advanced AI Reasoning Today — Gend.co
- Gemini 3 - Google DeepMind — Google DeepMind