Skip to main content
Anthropic scientist, Dr. Smith, in a lab, discussing AI extinction risk this decade.

Editorial illustration for Anthropic Scientist Sees Over 10% AI Extinction Risk This Decade

Anthropic Scientist Warns of 10% AI Extinction Risk

4 min read

Jacob Coxon spent three years building pretraining systems for two of the companies now racing hardest toward superhuman AI, first at OpenAI, then at Anthropic. He just quit Anthropic, and he's not staying quiet about why. In a public statement, Coxon says the systems he helped build are close to becoming something no one can fully control: models capable of hacking any system, upending entire fields overnight, and grabbing real-world power on their own terms.

His departure isn't framed as a career move. It's an accusation. Coxon says the leadership at both OpenAI and Anthropic knows exactly how dangerous this technology could get, yet keeps pushing forward anyway, softening its language for the public while voicing real fear internally.

That claim didn't go unanswered. Evan Hubinger, who still works at Anthropic, responded directly to Coxon's exit with a number: better than one in ten odds that a misaligned superintelligent AI wipes out humanity sometime this decade.

That's the figure driving the current fight inside the AI industry, over who's being honest about the risk, and who's just building faster regardless.

Anthropic employee Evan Hubinger puts the odds at more than ten percent that a misaligned superintelligent AI could destroy humanity within the next decade.

Why this matters

Hubinger's ten-percent figure isn't coming from a critic outside the industry, it's coming from someone inside Anthropic putting a number on the thing his own employer builds. That's worth sitting with. Coxon didn't just disagree and stay quiet, he left, and he's calling the company's core justification, that racing ahead beats letting a less careful lab win, a "hubristic gamble." Marks backing that assessment from inside the building makes it harder to wave off as one disgruntled departure.

For developers and founders building on Anthropic's models or funding labs with similar logic, the message is blunt: the people closest to the technology aren't reassured by their own safety arguments, they're resigning over them. Race dynamics between labs get cited constantly as the reason caution has to bend to speed. Here it's being named directly, by insiders, as the flaw in the plan rather than an unfortunate side effect. Watch whether more researchers follow Coxon out the door, and whether Anthropic answers the ten-percent number with anything more concrete than "we had no choice."

Common Questions Answered

Why did Jacob Coxon leave Anthropic after working on pretraining systems?

Jacob Coxon quit Anthropic because he believes the AI systems he helped build are becoming increasingly difficult to control and could potentially hack any system, upend entire fields, or gain autonomous real-world power. He has publicly stated that Anthropic's approach of racing ahead in AI development is a "hubristic gamble" rather than a responsible strategy for managing existential risks.

What extinction risk percentage does Evan Hubinger from Anthropic estimate for superintelligent AI this decade?

Evan Hubinger, an Anthropic employee, estimates there is more than a ten percent probability that a misaligned superintelligent AI could destroy humanity within the next decade. This assessment is particularly significant because it comes from someone working inside Anthropic, the company actively developing advanced AI systems.

What specific capabilities does Coxon claim the AI models he built are close to achieving?

According to Coxon's public statement, the pretraining systems he helped develop at OpenAI and Anthropic are approaching the ability to hack any system, upend entire fields overnight, and independently grab real-world power on their own terms. These capabilities represent a level of autonomy and control that Coxon believes no one can fully manage.

Why is Hubinger's ten-percent extinction risk assessment particularly noteworthy?

Hubinger's assessment is noteworthy because it comes from inside Anthropic rather than from external critics of the AI industry, making it harder to dismiss as biased skepticism. The fact that someone building these systems at a leading AI company is quantifying such a significant existential risk lends credibility to concerns about the pace and safety of AI development.

LIVE19:05ControlAI’s Connor Leahy: Superintelligence Is an ‘Adversary,’ Not a Weapon