Editorial illustration for Opus 5 Hits Zero Percent Attack Rate Against AI Browser Prompt Injections
Opus 5 Blocks All Browser Prompt Injection Attacks
Opus 5 Hits Zero Percent Attack Rate Against AI Browser Prompt Injections
Anthropic ran 129 test scenarios pitting Opus 5 against browser-based prompt injection attacks, the trick where an attacker hides instructions in a webpage's text to hijack an AI agent's behavior. The attack success rate came back at zero percent, according to the model's system card. That number matters because prompt injection has been the industry's stubborn problem. OpenAI said as much in December, admitting the flaw may never get fully solved.
Anthropic's own numbers show progress even without special conditions. In a general test run by security firm Gray Swan, the attack success rate after 15 attempts fell from 5.5 percent on Opus 4.8 to 2.0 percent on Opus 5. But the headline zero percent figure comes with a catch: it only holds when Auto Mode is switched on, the setting available in products like Claude Cowork.
Turn it off, and Opus 5's rate climbs to 3.7 percent, worse than Sonnet 5's 0.93 percent under the same conditions. That gap raises the question of what Auto Mode is actually doing, and why the model alone isn't enough to close the hole.
Prompt injection, where an attacker slips past an AI model's instructions through manipulated inputs like hidden text on a webpage, fails against Opus 5 in almost every case. For browser agents, the attack success rate hit zero percent across 129 test scenarios, per the system card. That's a big deal given that OpenAI admitted in December that prompt injection may never be fully solved.
Why this matters Zero percent across 129 scenarios sounds like a solved problem, but it's Anthropic grading its own homework. The system card comes from the same company selling Opus 5, and "nearly immune" in the write-up doesn't quite match the "zero percent" headline number, worth watching for. That gap is exactly where developers and founders building browser agents should stay skeptical rather than declare victory.
Gray Swan's third-party test result got cut off before the actual figure showed up, which matters more than Anthropic's internal numbers do. If an outside red team using different attack patterns can't replicate zero percent, the real story is narrower than "solved." Still, the contrast with OpenAI's December admission that prompt injection might never fully close is notable. It suggests architecture choices in how an agent parses page content versus instructions can meaningfully shrink the attack surface.
For anyone shipping agents that click, scroll, and read arbitrary web pages, the next thing to watch isn't Anthropic's benchmark. It's whether Gray Swan, or another outside lab, publishes a full number that holds up under adversarial pressure Anthropic didn't design the test for.
Common Questions Answered
What attack success rate did Opus 5 achieve against browser-based prompt injection attacks?
Opus 5 achieved a zero percent attack success rate across 129 test scenarios involving browser-based prompt injection attacks, according to Anthropic's system card. This represents a significant milestone since prompt injection has been a persistent industry problem that OpenAI admitted in December may never be fully solved.
How do attackers use prompt injection to hijack AI agents?
Attackers hide malicious instructions in a webpage's text to manipulate an AI agent's behavior through prompt injection attacks. These hidden instructions attempt to override the model's original instructions and cause the AI to perform unintended actions.
Why should developers remain skeptical about Opus 5's zero percent prompt injection result?
Anthropic is grading its own homework since the system card comes from the same company selling Opus 5, and the write-up uses language like 'nearly immune' which doesn't perfectly match the 'zero percent' headline. Developers should wait for independent third-party test results before declaring the prompt injection problem fully solved.
What did OpenAI say about the future of prompt injection vulnerabilities?
OpenAI admitted in December that prompt injection may never be fully solved, highlighting how stubborn this security problem has been across the AI industry. This context makes Anthropic's claimed zero percent success rate particularly noteworthy as a potential breakthrough.
Further Reading
- Anthropic published the prompt injection failure rates that enterprise ... - VentureBeat
- How to Harden Claude Code Against Prompt Injection - Pete Builds
- Claude Computer Use and Prompt Injection Resistance: The ... - AI Tech News
- Prompt Injection Now Has a Number: 31.5% Agent Hijack - Rogue AI
- Anthropic's AI Browser Agent Got Hijacked 31.5% of the Time - YouTube