Skip to main content
Anthropic's Mythos AI tool interface, displaying internal testing results and data visualizations.

Editorial illustration for Anthropic's Mythos Tool Meets Its Hype in Internal Testing

Anthropic's Mythos Tool Meets Its Hype in Internal Testing

3 min read

Microsoft's engineers had a problem they weren't used to having: too many bugs, too fast. In mid-May, dozens of them logged onto a video call, and a few dozen more packed a conference room at the company's Redmond headquarters, to hash out Project Glasswing, an effort to patch security holes in SharePoint before anyone outside the company found them first. The holes were being surfaced by Mythos, an AI model built by Anthropic and quietly handed to a handful of organizations, Microsoft among them, that make software relied on by governments and companies worldwide.

The idea was straightforward enough: get ahead of hackers and state actors, China included, who might soon have similar tools of their own. What nobody fully expected was how well Mythos would work. By April, the tool had already flagged dozens of critical flaws in SharePoint alone, more than Microsoft's engineers could realistically fix on schedule.

That mismatch, between what the AI could find and what humans could patch, is what brought the room together that afternoon in May, and it's the tension at the center of the meeting captured below.

The version being used by Microsoft, Claude Mythos Preview, was surfacing bugs faster than the tech giant could patch them, and engineers, the manager said, were now in “a mad dash” to close the gap.

Why this matters

The Redmond meeting is a useful data point for anyone building or buying AI security tools: Mythos apparently found bugs faster than a company with Microsoft's engineering headcount could patch them. That's either a genuine leap in automated vulnerability discovery or a sign that "unprecedented clip" is doing a lot of work in Anthropic's own framing, and the ProPublica recording doesn't settle which. For developers and security teams, the real story isn't the demo, it's the backlog.

A model that surfaces flaws faster than humans can triage them creates its own operational problem, and Project Glasswing's internal scramble suggests Microsoft is still figuring out how to absorb that output rather than just celebrate it. Founders pitching AI-driven security should take note: discovery speed without a fix pipeline to match is a liability, not a selling point. We'd want to see how Mythos performs across a full disclosure-to-patch cycle, not just the moment engineers admit the hype checked out.

Watch whether Microsoft publishes concrete numbers on bugs found versus bugs closed before crowning this a breakthrough.

LIVE19:06Hugging Face Traces 17,600 Actions by Compromised AI Models