Editorial illustration for Anthropic's Mythos Tool Meets Its Hype in Internal Testing
Anthropic's Mythos Tool Meets Its Hype in Internal Testing
Microsoft's engineers had a problem they weren't used to having: too many bugs, too fast. In mid-May, dozens of them logged onto a video call, and a few dozen more packed a conference room at the company's Redmond headquarters, to hash out Project Glasswing, an effort to patch security holes in SharePoint before anyone outside the company found them first. The holes were being surfaced by Mythos, an AI model built by Anthropic and quietly handed to a handful of organizations, Microsoft among them, that make software relied on by governments and companies worldwide.
The idea was straightforward enough: get ahead of hackers and state actors, China included, who might soon have similar tools of their own. What nobody fully expected was how well Mythos would work. By April, the tool had already flagged dozens of critical flaws in SharePoint alone, more than Microsoft's engineers could realistically fix on schedule.
That mismatch, between what the AI could find and what humans could patch, is what brought the room together that afternoon in May, and it's the tension at the center of the meeting captured below.
Why this matters
The Redmond meeting is a useful data point for anyone building or buying AI security tools: Mythos apparently found bugs faster than a company with Microsoft's engineering headcount could patch them. That's either a genuine leap in automated vulnerability discovery or a sign that "unprecedented clip" is doing a lot of work in Anthropic's own framing, and the ProPublica recording doesn't settle which. For developers and security teams, the real story isn't the demo, it's the backlog.
A model that surfaces flaws faster than humans can triage them creates its own operational problem, and Project Glasswing's internal scramble suggests Microsoft is still figuring out how to absorb that output rather than just celebrate it. Founders pitching AI-driven security should take note: discovery speed without a fix pipeline to match is a liability, not a selling point. We'd want to see how Mythos performs across a full disclosure-to-patch cycle, not just the moment engineers admit the hype checked out.
Watch whether Microsoft publishes concrete numbers on bugs found versus bugs closed before crowning this a breakthrough.
Further Reading
- Anthropic Releases Claude Mythos Preview with Restricted Access for Security Testing - InfoQ
- How Anthropic Learned Mythos Was Too Dangerous for the Wild - Bloomberg
- Anthropic's Claude Mythos isn't a sentient super-hacker, it's a sales pitch - Tom's Hardware
- The wildest things Anthropic's Mythos pulled off in testing - Benton Institute
- Anthropic's Mythos Claims Questioned by Cybersecurity Insider - YouTube