Skip to main content
AI lab safety audit: a red "X" over a checklist, symbolizing unfulfilled basic safety controls.

Editorial illustration for Nonprofit Finds No AI Labs Fully Apply Basic Safety Controls

AI Labs Fail Safety Controls, Guidelight Scorecard Reveals

Nonprofit Finds No AI Labs Fully Apply Basic Safety Controls

4 min read

Anthropic got a C+. So did OpenAI. Nobody else came close.

That's the grading from Guidelight, a nonprofit that just published its first scorecard on how well AI companies actually police their own systems while building them. The group looked at five labs, Anthropic, OpenAI, Google, xAI, and Meta, using only what those companies have already made public: system cards, safety reports, blog posts. No inside access, no interviews, just a check against six practices that Guidelight considers baseline for controlling powerful AI internally.

The list covers things like logging what an AI system does behind the scenes, requiring review before it takes risky actions, having a way to shut it down fast if something goes wrong, and having an actual plan for what happens if a model turns out to be misaligned. Guidelight was founded by Page Hedley and Steven Adler, both of whom used to work on safety at OpenAI, so the group knows roughly what these internal processes are supposed to look like.

The results split the five companies into rough tiers, with wide gaps between the top and the bottom.

No AI company fully applies basic control measures to its own internal AI systems. That's the takeaway from the first assessment by the nonprofit Guidelight.

Why this matters

Guidelight's findings land at an awkward moment for an industry that keeps asking regulators and the public to trust its internal judgment on frontier risk. Anthropic, OpenAI, Google, xAI, and Meta all publish safety frameworks, but none of them, according to this review of their own public disclosures, fully logs internal AI activity or gates risky actions through a working review mechanism. That's a gap between what labs say in system cards and what they actually run day to day.

For developers and researchers, the lesson isn't that these companies are reckless. It's that "we have a safety framework" and "we apply it to our own systems" are two different claims, and right now only the first one is well documented. Founders building on top of these platforms should treat vendor safety marketing with the same skepticism they'd apply to any other unverified compliance claim. And if a five-person nonprofit can spot these gaps just by reading blog posts and system cards, that's a low bar the labs themselves should be able to clear before anyone else checks their homework.

Common Questions Answered

What grades did Anthropic and OpenAI receive in Guidelight's AI safety scorecard?

Both Anthropic and OpenAI received a C+ grade from Guidelight's first safety scorecard assessment. This was the highest score among the five AI labs evaluated, indicating that even the leading companies fell short of fully applying basic safety controls to their internal systems.

Which AI companies were included in Guidelight's safety assessment?

Guidelight evaluated five major AI labs: Anthropic, OpenAI, Google, xAI, and Meta. The nonprofit assessed these companies based solely on their publicly available materials including system cards, safety reports, and blog posts, without requiring inside access or conducting interviews.

What critical gap did Guidelight find between AI labs' public safety frameworks and their actual practices?

According to Guidelight's review, no AI company fully logs internal AI activity or implements working review mechanisms to gate risky actions, despite publishing safety frameworks in their system cards. This represents a significant disconnect between what these labs publicly claim about their safety practices and what they actually execute in their day-to-day operations.

Why is Guidelight's finding particularly problematic for the AI industry?

The findings are significant because AI companies have been asking regulators and the public to trust their internal judgment on frontier risks, yet Guidelight's assessment shows they are not fully applying basic control measures to their own systems. This gap between industry claims and demonstrated practices undermines the credibility of their self-regulation efforts during a critical period of AI development.

LIVE16:12GLM-5.3 Jumps 246 Points on Key Benchmark, Ranks Second to Claude Opus