Skip to main content
A focused analyst in a sleek office watches a monitor displaying Gemini AI search results, HLE and DeepSearchQA charts.

Editorial illustration for Google's Gemini Deep Research Agent Tops Academic Benchmark Tests

Gemini AI Shatters Academic Research Benchmarks Decisively

Gemini Deep Research agent posts top results on HLE, DeepSearchQA, leads BrowseComp

Updated: 3 min read

46.4% on Humanity’s Last Exam. 66.1% on DeepSearchQA. And 59.2% on BrowseComp, our best ever.

The new Gemini Deep Research agent doesn’t just inch ahead; it clears the bar. These aren’t incremental gains. They represent a leap in how machines tackle the kind of sprawling, multi-step research that confounds simpler models.

The secret? An agent built to generate thorough reports at a fraction of the cost. We’re also open-sourcing DeepSearchQA, a benchmark designed to measure exactly what matters: the messy, real-world complexity of deep web research.

This isn’t just a lab result. Deep Research is becoming more useful, more intelligent, and soon it will land inside Google Search, NotebookLM, Google Finance, and an upgraded Gemini App. The era of shallow searches is over.

Today, we are releasing a significantly more powerful Gemini Deep Research agent, available via the Interactions API.

This isn’t just a leap on a leaderboard. Deep Research now delivers expert-grade synthesis at a fraction of the cost, and it lands where people actually work: Search, NotebookLM, Finance, the Gemini app. Humanity’s Last Exam tested the fringe of knowledge.

DeepSearchQA tests the messy, multi-step grunt work that defines real research. By open-sourcing that benchmark, we ensure the race stays honest and the bar keeps rising. The agent leads both.

What was once a novelty is becoming a utility. That shift is what matters.

Common Questions Answered

What benchmark tests did the Gemini Deep Research agent excel in?

The Gemini Deep Research agent achieved state-of-the-art results on the Humanity's Last Exam (HLE) with a 46.4% score and performed exceptionally well on DeepSearchQA with a 66.1% performance. These benchmark tests demonstrate the agent's advanced capabilities in complex information retrieval and analytical tasks.

In which Google products will the Gemini Deep Research agent be integrated?

Google plans to integrate the Gemini Deep Research agent across multiple products, including Google Search, NotebookLM, Google Finance, and the Gemini App. This widespread integration suggests a strategic approach to enhancing AI-powered research and information synthesis across different platforms.

How does the Gemini Deep Research agent improve upon previous AI research capabilities?

The Gemini Deep Research agent represents a significant advancement in AI's ability to handle nuanced research tasks that previously challenged artificial intelligence systems. It can generate well-researched reports at a lower cost and demonstrates superior comprehension and analytical skills across complex information retrieval challenges.

LIVE15:01Nvidia and Microsoft form open AI security alliance, exclude OpenAI