Skip to main content
OpenAI unveils GPT-5.5 breakthrough, achieving 82.7% on Terminal-Bench 2.0 and 84.9% on GDPval, showcasing advanced AI perfor

Editorial illustration for OpenAI launches GPT-5.5, hits 82.7% on Terminal-Bench 2.0, 84.9% on GDPval

GPT-5.5 Crushes Benchmarks with 82.7% Agentic AI Score

OpenAI launches GPT-5.5, hits 82.7% on Terminal-Bench 2.0, 84.9% on GDPval

Updated: 3 min read

The numbers are stark. 82.7% on Terminal-Bench 2.0. 84.9% on GDPval.

But the real story of OpenAI’s GPT-5.5 isn’t found in a single score. It’s buried in a benchmark called Expert-SWE, where the median human task takes twenty hours. That is not a quick fix.

That is a full day’s work for a senior engineer, a sprawling refactor, a deep-dive debugging session, a feature build that touches a dozen files. For the first time, a model is being measured against that kind of endurance. And early testers report something more telling than a percentage: GPT-5.5 grasps the *shape* of a software system.

It sees the architecture, not just the syntax. It understands why a failure is happening, where the fix lives, and, critically, what else will break when you apply it. This is not a faster autocomplete.

This is a model that has been fully retrained, not fine-tuned, to act as an agent. The scores on Terminal-Bench and GDPval are impressive, but they are the headline. The real story is in the long, messy, multi-session work that defines professional engineering.

For long-horizon coding specifically, OpenAI also reports results on Expert-SWE, an internal benchmark measuring tasks with a median estimated human completion time of 20 hours. This benchmark is significant because it reflects the kind of extended, multi-session engineering work -- large refactors, feature builds, debugging deep in a codebase -- that agentic tools are increasingly being asked to handle autonomously. Developers who tested the system early said GPT-5.5 has a better understanding of the "shape" of a software system, and can better understand why something is failing, where the fix is needed, and what else in the codebase would be affected.

GPT-5.5 doesn’t just score higher, it understands the architecture of failure. That elusive “shape” of a system, the ability to trace a bug’s root through tangled dependencies and predict the collateral damage of a fix, is what separates incremental improvement from genuine agentic autonomy. The numbers on Terminal-Bench and GDPval are impressive, but the deeper signal lives in Expert-SWE: a model that can shoulder 20-hour engineering tasks without losing context, without needing a human to hold its hand through every refactor. This is the moment a tool stops being a faster copilot and starts becoming a true teammate, one that sees the whole blueprint, not just the next line.

Common Questions Answered

How does GPT-5.5 perform on the Expert-SWE benchmark for long-horizon coding tasks?

GPT-5.5 is evaluated on Expert-SWE, an internal benchmark that measures tasks with a median estimated human completion time of 20 hours. This benchmark is crucial as it assesses the model's ability to handle complex, extended engineering work like large refactors, feature builds, and deep codebase debugging.

What makes GPT-5.5 different from previous OpenAI language models?

GPT-5.5 is the first fully retrained base model since GPT-4.5, achieving impressive scores of 82.7% on Terminal-Bench 2.0 and 84.9% on GDPval. The model is designed to be an agentic assistant that understands underlying goals, rather than simply following a rigid checklist of instructions.

What are the key challenges in evaluating GPT-5.5's capabilities?

While GPT-5.5 shows promising benchmark results, the evaluation is limited by internal benchmarks and specific task families. The true test of the model lies in its ability to perform prolonged, iterative work across complex coding environments and interdependent modules.

LIVE04:46Naval Postgraduate School Activates NVIDIA AI Supercomputer for In-House Training