Editorial illustration for GPT-5.2 Outperforms Professionals on 70.9% of Tasks, OpenAI Analysis Reveals
GPT-5.2 Beats Pros on 70.9% of Tasks, OpenAI Study Shows
Analysis overhauls AI Index; GPT-5.2 beats professionals on 70.9% of tasks
The AI industry has stopped bragging about test scores and started clocking in. A major research group has scrapped its old benchmarks for measuring intelligence. They now look at whether models can do a job, not pass a test.
The results from this new AI Intelligence Index are blunt. According to OpenAI, GPT-5.2 now beats or ties top professionals on 70.9% of well-specified tasks across 44 occupations. Companies like Notion, Box, Shopify, Harvey, and Zoom call its performance state-of-the-art.
The question is no longer about solving a math puzzle. It is about completing the work.
On the original GDPval evaluation, GPT-5.2 beat or tied top industry professionals on 70.9% of well-specified tasks, according to OpenAI. The company claims GPT-5.2 "outperforms industry professionals at well-specified knowledge work tasks spanning 44 occupations," with companies including Notion, Box, Shopify, Harvey, and Zoom observing "state-of-the-art long-horizon reasoning and tool-calling performance." The emphasis on economically measurable output is a philosophical shift in how the industry thinks about AI capability. Rather than asking whether a model can pass a bar exam or solve competition math problems -- achievements that generate headlines but don't necessarily translate to workplace productivity -- the new benchmarks ask whether AI can actually do jobs. Graduate-level physics problems expose the limits of today's most advanced AI models While GDPval-AA measures practical productivity, another new evaluation called CritPT reveals just how far AI systems remain from true scientific reasoning.
This pivot matters. The industry is finally valuing economic output over academic achievement. But a new evaluation called CritPT draws a hard line.
It shows these models still fail at graduate-level physics. They cannot reason scientifically.
So we have a split. AI can now reliably perform many defined jobs. It remains hopeless at the open-ended inquiry that drives real discovery.
The tool is getting very good. The scientist is not inside it.
Common Questions Answered
How did GPT-5.2 perform across professional tasks in OpenAI's evaluation?
According to OpenAI's research, GPT-5.2 beat or tied top industry professionals on 70.9% of well-specified tasks. The system demonstrated exceptional performance across 44 different occupations, showcasing advanced long-horizon reasoning and tool-calling capabilities.
Which major tech companies have observed GPT-5.2's performance?
Companies including Notion, Box, Shopify, Harvey, and Zoom have directly observed GPT-5.2's state-of-the-art performance. These organizations have noted the AI's remarkable ability to handle complex knowledge work tasks with unprecedented efficiency.
What makes GPT-5.2's performance significant for workplace dynamics?
GPT-5.2's ability to outperform professionals on nearly 71% of tasks suggests a potential transformative shift in how knowledge work is conducted. The research challenges traditional assumptions about professional competence and indicates that AI could fundamentally reshape workplace productivity and task execution.
Further Reading
- GPT-5.2 Review: Benchmarks (AIME 100%), Visual AI ... — Vertu
- GPT-5.2: Pricing, Context Window, Benchmarks, and More — LLM Stats
- GPT-5.2 Benchmarks — Vellum AI
- How GPT-5.2 stacks up against Gemini 3.0 and Claude Opus 4.5 — RDWorld Online