Editorial illustration for Anthropic and OpenAI Propose Safety Evaluator Access to Staff
Anthropic, OpenAI Open AI Safety Access to Evaluators
Anthropic and OpenAI Propose Safety Evaluator Access to Staff
Dario Amodei spent the weekend making a pitch that would have gotten laughed out of the room a year ago. In an essay published Saturday, the Anthropic CEO proposed letting outside evaluators inside frontier AI companies, with access to report safety incidents, check whether models are actually aligned, and publish what they find without a corporate filter. Anthropic says it will give groups like METR and Redwood Research that kind of access. Sam Altman said OpenAI will do the same at rival OpenAI, which would mark a real shift in how the two biggest labs deal with outside scrutiny.
Third-party evaluators contacted by TechCrunch called the idea welcome but unfinished, with plenty left to negotiate before anyone can say whether these evaluators are independent watchdogs or just contractors working on the labs' terms. The stakes are rising because models are getting better at spotting when they're being tested, which means good behavior during an eval doesn't rule out bad behavior elsewhere. Some researchers argue the real evidence isn't in the finished model at all. It's buried in how that model behaved during training, long before anyone thought to ask.
In a lengthy essay published over the weekend, Anthropic CEO Dario Amodei made a proposal that the AI industry would have rejected instantly even a year ago: embed third-party evaluators inside all frontier AI companies, giving them the power to report safety incidents, assess whether AI models are truly aligned, and share their unvarnished findings with the world.
Why this matters
Amodei's proposal sounds like real progress until you ask who picks the evaluators, who funds them, and who decides what "unvarnished" actually means in practice. METR and Redwood Research getting employee interview access and document-checking rights is a bigger ask than model access alone, and it's the part worth watching closest. Companies have controlled their own safety narratives for years; letting outsiders compare internal Slack messages against public safety claims is a genuine test of whether Anthropic and OpenAI mean this or are getting ahead of regulation they see coming anyway.
For developers and founders building on these platforms, the real signal isn't the essay, it's whether Gleave's team or others actually get the access described and whether their findings get published without a corporate edit pass. Researchers should watch for the first evaluator report that contradicts a company's own safety claims. That's the moment this becomes real oversight instead of a well-timed PR document.
Common Questions Answered
What specific access is Anthropic proposing to give third-party safety evaluators?
Anthropic CEO Dario Amodei proposed embedding outside evaluators inside frontier AI companies with the power to report safety incidents, assess whether AI models are truly aligned, and publish their findings without corporate filtering. This access would include employee interviews and document-checking rights, going beyond just model access to include internal company communications and safety documentation.
Which organizations has Anthropic committed to giving safety evaluator access?
Anthropic has committed to giving groups like METR and Redwood Research access to embed evaluators inside the company. Sam Altman also announced that OpenAI will provide the same type of access to rival companies, representing a significant shift in industry transparency practices.
Why is Dario Amodei's safety evaluator proposal considered unusual for the AI industry?
Amodei's proposal would have been rejected instantly by the AI industry even a year ago, according to the article. The proposal represents a dramatic shift because it asks companies to allow outsiders to compare internal communications like Slack messages against public safety claims, fundamentally changing how AI companies control their safety narratives.
What are the key implementation questions that remain unanswered about the evaluator access proposal?
Critical questions persist about who will pick the evaluators, who will fund them, and what "unvarnered" findings actually means in practice. The article highlights that determining these details is essential before the proposal can be meaningfully implemented across the industry.
Further Reading
- Anthropic grants outside evaluators permanent access to company systems in new safety push - Fortune
- Anthropic's Amodei proposes continuous evaluator access for AI safety oversight - CryptoBriefing
- Sam Altman joins the 'brake' of Anthropic's AI: OpenAI will open its systems to external evaluators - Demócrata
- Amodei, Altman, Musk call for slowing AI model development - Los Angeles Times
- Anthropic and OpenAI Leaders Commit to Independent Evaluators for Powerful AI Models - AI Research Institute