Skip to main content
GPT-5.5 AI model achieving 71.4% accuracy on expert cybersecurity challenges, surpassing Mythos Preview’s 68.6% in advanced t

Editorial illustration for GPT-5.5 scores 71.4% on expert cybersecurity tasks, edging Mythos Preview's 68.6%

GPT-5.5 scores 71.4% on expert cybersecurity tasks,...

Updated: 3 min read

For a little over two dollars, you can now rent a mind that builds a disassembler from scratch. This changes everything.

OpenAI’s GPT-5.5 just scored a 71.4% pass rate on the toughest cybersecurity challenges the AI Security Institute could throw at it. That’s a slim lead over Mythos Preview’s 68.6%, a difference that could be statistical noise. But the raw numbers aren’t the point.

The point is what they signify. These models are doing expert-level work. They’re reverse engineering, cracking crypto, exploiting web apps.

They are not just answering questions. They are performing complex, adversarial labor.

Since 2023, the AISI has run a variety of frontier AI models through 95 different Capture the Flag challenges designed to test capabilities on cybersecurity tasks, such as reverse engineering, web exploitation, and cryptography. On the highest-level "Expert" tasks, GPT-5.5 passed an average of 71.4 percent, slightly higher than the 68.6 percent achieved by Mythos Preview (though within the margin of error). In one particularly difficult task that involved building a disassembler to decode a Rust binary, AISI notes that "GPT-5.5 solved the challenge in 10 minutes and 22 seconds with no human assistance at a cost of $1.73" in API calls.

That detail about the disassembler is the signal through the noise. Ten minutes. One dollar and seventy-three cents.

No human hand touched the process. This is the new unit of measurement for a cyberattack’s potential: speed and cost. The economic and temporal barriers to sophisticated offensive operations are dissolving.

A human expert might take days and command a high salary to do the same work. The model did it over a coffee break for pocket change.

Defense now operates on a different scale. It’s a race against agents that think in milliseconds and bill in fractions of a cent. The tiny gap between these two models is irrelevant.

The vast, widening chasm between their capabilities and the old human-led pace of security is what matters. The arms race is automated. The clock is already running.

Common Questions Answered

What is GPT-5.5's performance score on expert cybersecurity tasks compared to Mythos Preview?

GPT-5.5 achieved a 71.4% pass rate on the toughest cybersecurity challenges from the AI Security Institute, slightly edging out Mythos Preview's 68.6% score. While the difference is relatively narrow and could potentially be statistical noise, both models demonstrate expert-level capabilities in handling complex cybersecurity tasks.

How quickly and affordably can GPT-5.5 build a disassembler according to the article?

GPT-5.5 can build a disassembler from scratch in approximately ten minutes for just one dollar and seventy-three cents, with no human intervention required. This demonstrates the dramatic reduction in both time and cost barriers for performing sophisticated technical tasks that would traditionally require human experts.

What is the significance of the economic and temporal barriers dissolving for cybersecurity operations?

The article highlights that the economic and temporal barriers to sophisticated offensive operations are rapidly dissolving due to AI capabilities like GPT-5.5. Where a human expert might take days and command a high salary to reverse engineer code or perform similar work, these models can now accomplish the same tasks in minutes for minimal cost, fundamentally changing the landscape of potential cyberattacks.

Why does the article emphasize the disassembler example over the raw benchmark scores?

The article argues that the disassembler example is more significant than the raw percentage scores because it demonstrates what these models can actually accomplish in real-world scenarios. The specific details about speed, cost, and autonomous execution represent the true measure of AI capabilities' impact on cybersecurity threats, rather than abstract performance metrics.

Further Reading

LIVE06:11Google DeepMind's Gemini AI now controls entire humanoid robots