Skip to main content
Gemini 3 Deep Think logo, a stylized "G" with radiating lines, symbolizing advanced AI reasoning and problem-solving. [blog.g

Editorial illustration for Gemini 3 Deep Think Boosts Reasoning with Mathematical and Algorithmic Rigor

Gemini 3 Deep Think: AI Reasoning Breakthrough

Gemini 3 Deep Think Boosts Reasoning with Mathematical and Algorithmic Rigor

Updated: 3 min read

The hardest tests we have for human intelligence are now being passed by a machine.

Gemini 3 Deep Think just posted a score of 48.4% on Humanity’s Last Exam, which is meant to stump the most advanced models. It did that without any tools. It scored 84.6% on ARC-AGI-2.

Its competitive programming rank, an Elo of 3455 on Codeforces, is staggering. It performed at a gold-medal level on the 2025 International Math Olympiad.

This isn't just about math. The same system now hits gold-medal standards on the written sections of the 2025 International Physics and Chemistry Olympiads. The claim is that a single approach, built on mathematical and algorithmic rigor, can conquer these wildly different fields.

Elevating reasoning with mathematical and algorithmic rigor Last year, we showed that specialized versions of Deep Think could successfully navigate some of the toughest challenges in reasoning, achieving gold-medal standards at math and programming world championships. More recently, Deep Think has enabled specialized agents to conduct research-level mathematics exploration. The updated Deep Think mode continues to push the frontiers of intelligence, reaching new heights across the most rigorous academic benchmarks, including: - Setting a new standard (48.4%, without tools) on Humanity's Last Exam, a benchmark designed to test the limits of modern frontier models - Achieving an unprecedented 84.6% on ARC-AGI-2, verified by the ARC Prize Foundation - Attaining a staggering Elo of 3455 on Codeforces, a benchmark consisting of competitive programming challenges - Reaching gold-medal level performance on the International Math Olympiad 2025 Navigating complex scientific domains Beyond mathematics and competitive coding, Gemini 3 Deep Think now also excels across broad scientific domains such as chemistry and physics. Our updated Deep Think mode demonstrates gold medal-level results on the written sections of the 2025 International Physics Olympiad and Chemistry Olympiad.

Benchmarks are just numbers. The shift here is a kind of proof. If one method can produce gold-medal results in high school math, competitive coding, university-level physics, and theoretical chemistry, then the method itself is the story.

The machine isn't just retrieving answers. It is applying a consistent, verifiable process to problems that require deep, structured thought. We are watching a tool learn how to think in a way we recognize as rigorous.

That changes what the tool is for.

Common Questions Answered

What makes Gemini 3 Deep Think different from previous AI models in reasoning capabilities?

Gemini 3 Deep Think introduces Advanced Parallel Reasoning, which explores multiple hypothesis paths simultaneously instead of following a single chain of thought. This 'System 2' approach allows the model to pause, explore multiple hypotheses, and critically critique its own logic before generating an output, marking a significant departure from traditional 'System 1' language models that simply predict the next token.

How did Gemini 3 Deep Think perform on challenging academic benchmarks?

Gemini 3 Deep Think demonstrated exceptional performance on rigorous benchmarks like Humanity's Last Exam, achieving 41.0% accuracy without using external tools, and ARC-AGI-2, where it scored an unprecedented 45.1% with code execution. These results build on previous achievements, including gold-medal level performances at the International Mathematical Olympiad and International Collegiate Programming Contest World Finals.

What is the key cognitive architecture difference between System 1 and System 2 reasoning in AI?

System 1 reasoning, typical of standard Large Language Models, is fast, automatic, and impulsive - like a student quickly blurting out an answer during a pop quiz. In contrast, System 2 reasoning, exemplified by Gemini 3 Deep Think, is slow, deliberate, and analytical - similar to a mathematician carefully working through a complex proof, allocating computational time to thinking and exploring multiple logical paths before generating an output.

LIVE00:05NVIDIA Toolkit Accelerates OpenFold3 Co-Folding Workflow