Skip to main content
Gemini 3 Deep Think logo, a stylized "G" with radiating lines, symbolizing advanced AI reasoning and problem-solving. [blog.g

Editorial illustration for Gemini 3 Deep Think Boosts Reasoning with Mathematical and Algorithmic Rigor

Gemini 3 Deep Think: AI Reasoning Breakthrough

Gemini 3 Deep Think Boosts Reasoning with Mathematical and Algorithmic Rigor

Updated: 3 min read

The hardest tests we have for human intelligence are now being passed by a machine.

Gemini 3 Deep Think just posted a score of 48.4% on Humanity’s Last Exam, which is meant to stump the most advanced models. It did that without any tools. It scored 84.6% on ARC-AGI-2.

Its competitive programming rank, an Elo of 3455 on Codeforces, is staggering. It performed at a gold-medal level on the 2025 International Math Olympiad.

This isn't just about math. The same system now hits gold-medal standards on the written sections of the 2025 International Physics and Chemistry Olympiads. The claim is that a single approach, built on mathematical and algorithmic rigor, can conquer these wildly different fields.

Today, we’re releasing a major upgrade to Gemini 3 Deep Think, our specialized reasoning mode, built to push the frontier of intelligence and solve modern challenges across science, research, and engineering.

Benchmarks are just numbers. The shift here is a kind of proof. If one method can produce gold-medal results in high school math, competitive coding, university-level physics, and theoretical chemistry, then the method itself is the story.

The machine isn't just retrieving answers. It is applying a consistent, verifiable process to problems that require deep, structured thought. We are watching a tool learn how to think in a way we recognize as rigorous.

That changes what the tool is for.

Common Questions Answered

What makes Gemini 3 Deep Think different from previous AI models in reasoning capabilities?

Gemini 3 Deep Think introduces Advanced Parallel Reasoning, which explores multiple hypothesis paths simultaneously instead of following a single chain of thought. This 'System 2' approach allows the model to pause, explore multiple hypotheses, and critically critique its own logic before generating an output, marking a significant departure from traditional 'System 1' language models that simply predict the next token.

How did Gemini 3 Deep Think perform on challenging academic benchmarks?

Gemini 3 Deep Think demonstrated exceptional performance on rigorous benchmarks like Humanity's Last Exam, achieving 41.0% accuracy without using external tools, and ARC-AGI-2, where it scored an unprecedented 45.1% with code execution. These results build on previous achievements, including gold-medal level performances at the International Mathematical Olympiad and International Collegiate Programming Contest World Finals.

What is the key cognitive architecture difference between System 1 and System 2 reasoning in AI?

System 1 reasoning, typical of standard Large Language Models, is fast, automatic, and impulsive - like a student quickly blurting out an answer during a pop quiz. In contrast, System 2 reasoning, exemplified by Gemini 3 Deep Think, is slow, deliberate, and analytical - similar to a mathematician carefully working through a complex proof, allocating computational time to thinking and exploring multiple logical paths before generating an output.

LIVE05:20Writer's New AI Model Targets Multi-Step Tasks With Lower Token Costs