Skip to main content
Security researcher at a computer, demonstrating AI guardrail bypass for offensive security work.

Editorial illustration for Security researcher says AI guardrails don't impede his offensive work

AI Guardrails Don't Stop Offensive Security Work

4 min read

Anthropic's Mythos and Fable models spent part of this summer under U.S. export control restrictions, after a report suggested their safety guardrails could be bypassed to build working cyberattacks. Fable 5 came back to general availability on July 1. Mythos 5 returned only for vetted American organizations, folded into a government review process that Anthropic itself has leaned into by marketing the model as something close to a doomsday tool, one that requires careful screening and tight restrictions before anyone gets near it.

That posture isn't limited to Mythos. Anthropic runs a Cyber Verification Program for its other models, and OpenAI has its own version, Trusted Access for Cyber, both designed to let vetted researchers work with fewer restrictions than the general public gets. The intent is straightforward: keep malicious hackers from turning frontier models into attack tools. But the people who spend their careers hunting for vulnerabilities before criminals do say the vetting process and the guardrails built around it are getting in the way of the actual work, not just the misuse the companies are trying to prevent.

Chris Anley, the chief scientist at security consulting giant NCC Group, said that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability worth fixing. But if a guardrail prompts the model to refuse to answer the question outright, the guardrail hurts defenders, he said.

Why this matters

Cali's workflow points to a gap between how AI companies imagine misuse and how research actually happens. He isn't asking a model to write an exploit; he's using it to speed up reverse engineering and tool-building, tasks that sit well inside any reasonable guardrail policy but still get lumped into the same anxious conversations about "AI-assisted hacking." That mismatch matters for us because it suggests export controls and vetting programs, like the ones now attached to Anthropic's Mythos and Fable, are reacting to worst-case scenarios rather than how skilled researchers actually work day to day. If the guardrails are calibrated around what a model could theoretically do rather than what practitioners are actually doing with it, we end up with policy that's loud but imprecise.

For founders building on these models and researchers relying on them, the lesson is to watch where restrictions actually bite versus where they just generate headlines. Cali's experience is one data point, not a verdict, but it's a useful check against the assumption that every guardrail debate maps cleanly onto real offensive security work.

Common Questions Answered

Why were Anthropic's Mythos and Fable models placed under U.S. export control restrictions?

A report suggested that the safety guardrails on these models could be bypassed to build working cyberattacks, prompting regulatory concerns about their potential misuse. This led to export restrictions being imposed on both models during the summer, with Fable 5 eventually returning to general availability on July 1 and Mythos 5 being restricted to vetted American organizations.

According to Chris Anley from NCC Group, how do AI guardrails impact offensive cybersecurity research?

Anley argues that asking an AI model to exploit a bug is a crucial step in confirming whether a vulnerability is real and worth fixing. However, when guardrails cause the model to refuse to answer such questions outright, they actually hurt defensive security researchers who need this capability to validate vulnerabilities.

What is the gap between how AI companies view misuse and how security researchers actually use these models?

Security researchers like Cali use AI models for reverse engineering and tool-building tasks that fall well within reasonable guardrail policies, yet these legitimate uses get conflated with concerns about AI-assisted hacking. This mismatch suggests that export controls and vetting programs may be overly broad and could inadvertently impede legitimate defensive security work.

What restrictions were placed on Mythos 5 after the export control review?

Mythos 5 was restricted to vetted American organizations only and was folded into a government review process. Anthropic has marketed the model as something close to a doomsday tool that requires careful screening and tight restrictions on access.

Further Reading

LIVE03:50Security researcher says AI guardrails don't impede his offensive work