Editorial illustration for GPT-5 Struggles with Research Tasks Despite Strong Test Performance
GPT-5 Falters in Research Tasks Despite High Test Scores
GPT-5.2 leads FrontierScience test, but falters on real research tasks
A new benchmark lands, and GPT-5.2 sits at the top. OpenAI's FrontierScience test is designed to be grueling: each problem demands three to five hours of work, scored on a ten-point rubric, with the model grading itself. On the Olympiad set, GPT-5.2 hits 77 percent, a commanding lead over Gemini 3 Pro's 76 percent.
But here’s the rub. On the Research tasks, the messy, open-ended problems that mimic real scientific work, GPT-5.2 manages only 25 percent. It ties with the older GPT-5, while the newer GPT-5.1 unexpectedly lags at 19 percent.
Claude Opus 4.5 and Grok 4 trail further. The gap is stark: progress on expert-level questions is real, but when the rubber meets the road on actual research, even the best models still stumble.
The numbers tell a clear story: GPT-5.2 can blaze through Olympiad-level puzzles, but when research demands open-ended inquiry, its edge dulls. A 25 percent score on genuine research tasks is an improvement over yesterday’s models, yet it’s still a gaping chasm away from the kind of autonomous scientific thinking that truly advances knowledge. Performance scales with compute, sure, but brute force has diminishing returns on problems that require creativity, context, and even serendipity.
The o3 model’s odd regression at higher reasoning intensity hints at something deeper: more thinking doesn’t always mean better thinking. OpenAI celebrates progress, and they should. But the real test isn’t how fast a model can solve a known problem.
It’s whether it can find the problem worth solving. On that front, the frontier remains stubbornly human.
Common Questions Answered
How did GPT-5.2 perform on different types of scientific benchmarks?
GPT-5.2 demonstrated impressive performance on Olympiad tasks, scoring 77 percent, but struggled significantly with more complex research challenges, only achieving a 25 percent success rate. The model's performance varied dramatically depending on the task complexity and reasoning intensity.
What makes research tasks challenging for AI models like GPT-5.2?
Research tasks are not about finding a single correct answer, but require sophisticated evaluation across a ten-point rubric that tests nuanced problem-solving capabilities. OpenAI suggests these tasks should take three to five hours to solve, highlighting the depth and complexity beyond standard testing metrics.
How does OpenAI evaluate the reasoning capabilities of GPT-5.2?
OpenAI tests reasoning models at different intensity levels, including 'high' and 'xhigh' reasoning modes, with tests conducted without browsing capabilities to assess the model's intrinsic problem-solving skills. The evaluation process goes beyond simple test scores to examine the model's ability to handle complex, open-ended scientific investigations.
Further Reading
- GPT-5.2 tops OpenAI's new FrontierScience test but struggles with real research problems — The Decoder
- Evaluating AI's ability to perform scientific research tasks — OpenAI
- frontierscience: evaluating AI's ability to perform scientific research tasks — ArXiv / OpenAI (technical paper)
- AI Is Getting Better at Science. OpenAI Is Testing How Far It Can Go — TIME
- OpenAI Unleashes FrontierScience for AI-Fueled Scientific Reasoning — eWEEK