Skip to main content
Scientists in a lab gather around a monitor displaying GPT-5.2 test results, as one points to a falling performance graph.

Editorial illustration for GPT-5 Struggles with Research Tasks Despite Strong Test Performance

GPT-5 Falters in Research Tasks Despite High Test Scores

GPT-5.2 leads FrontierScience test, but falters on real research tasks

Updated: 3 min read

A new benchmark lands, and GPT-5.2 sits at the top. OpenAI's FrontierScience test is designed to be grueling: each problem demands three to five hours of work, scored on a ten-point rubric, with the model grading itself. On the Olympiad set, GPT-5.2 hits 77 percent, a commanding lead over Gemini 3 Pro's 76 percent.

But here’s the rub. On the Research tasks, the messy, open-ended problems that mimic real scientific work, GPT-5.2 manages only 25 percent. It ties with the older GPT-5, while the newer GPT-5.1 unexpectedly lags at 19 percent.

Claude Opus 4.5 and Grok 4 trail further. The gap is stark: progress on expert-level questions is real, but when the rubber meets the road on actual research, even the best models still stumble.

OpenAI says existing scientific benchmarks are running out of headroom. When the company released GPQA—a "Google-proof" multiple-choice test for PhD-level science questions—in November 2023, GPT-4 scored 39 percent.

The numbers tell a clear story: GPT-5.2 can blaze through Olympiad-level puzzles, but when research demands open-ended inquiry, its edge dulls. A 25 percent score on genuine research tasks is an improvement over yesterday’s models, yet it’s still a gaping chasm away from the kind of autonomous scientific thinking that truly advances knowledge. Performance scales with compute, sure, but brute force has diminishing returns on problems that require creativity, context, and even serendipity.

The o3 model’s odd regression at higher reasoning intensity hints at something deeper: more thinking doesn’t always mean better thinking. OpenAI celebrates progress, and they should. But the real test isn’t how fast a model can solve a known problem.

It’s whether it can find the problem worth solving. On that front, the frontier remains stubbornly human.

Common Questions Answered

How did GPT-5.2 perform on different types of scientific benchmarks?

GPT-5.2 demonstrated impressive performance on Olympiad tasks, scoring 77 percent, but struggled significantly with more complex research challenges, only achieving a 25 percent success rate. The model's performance varied dramatically depending on the task complexity and reasoning intensity.

What makes research tasks challenging for AI models like GPT-5.2?

Research tasks are not about finding a single correct answer, but require sophisticated evaluation across a ten-point rubric that tests nuanced problem-solving capabilities. OpenAI suggests these tasks should take three to five hours to solve, highlighting the depth and complexity beyond standard testing metrics.

How does OpenAI evaluate the reasoning capabilities of GPT-5.2?

OpenAI tests reasoning models at different intensity levels, including 'high' and 'xhigh' reasoning modes, with tests conducted without browsing capabilities to assess the model's intrinsic problem-solving skills. The evaluation process goes beyond simple test scores to examine the model's ability to handle complex, open-ended scientific investigations.

LIVE05:20Writer's New AI Model Targets Multi-Step Tasks With Lower Token Costs