Skip to main content
Base Labs AI safety partnership launch, open-weight model, collaborative research, secure AI development.

Editorial illustration for Base Labs Launches Open-Weight AI Safety Partnership

Base Labs Launches AI Safety Tools for Open Models

4 min read

Baseten rolled out a new safety infrastructure standard on Wednesday, tied to Base Labs, the research arm it spun up earlier this year. The company is teaming up with Hugging Face and Goodfire AI to build evaluation and monitoring tools for open-weight models, the kind anyone can download and run without a company's servers standing between them and the code.

The timing isn't random. Hugging Face currently hosts more than 6,000 "abliterated" models, versions of open models stripped of their safety guardrails through a technique that's gained traction over the past year. That number alone shows how far the practice has spread and why Baseten wants standards baked into training and deployment rather than added as an afterthought.

Goodfire brings expertise in interpretability, the work of prying open a model's decision-making so researchers can see why it does what it does. Hugging Face brings the hosting infrastructure where most of these models already live. Baseten, fresh off a $1.5 billion Series F in June that pushed its valuation to $13 billion, brings the inference layer connecting models to the applications running on them.

Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.

Why this matters

Abliteration has made "open-weight" and "safe" harder to say in the same sentence, and Baseten's bet is that safety has to be engineered into training and hosting rather than patched on after release. Pairing with Hugging Face matters because Hugging Face is the distribution layer for most of these models; if a monitoring standard takes hold there, it could shape what "responsible open-weight" actually looks like in practice, not just in press releases. Goodfire's involvement adds interpretability chops to the mix, which suggests Base Labs wants technical teeth, not just a compliance checklist.

For developers and founders building on open models, this is worth watching closely: a real standard could mean new eval requirements before deployment, or scrutiny of how easily your fine-tuned model can be stripped of guardrails. For researchers, it's a chance to see whether infrastructure-level safety work can scale faster than the abliteration techniques it's racing against. Whether this becomes an industry norm or another well-intentioned framework that vendors ignore depends entirely on adoption we haven't seen yet.

Common Questions Answered

What is the purpose of the safety infrastructure standard launched by Base Labs with Hugging Face and Goodfire AI?

The partnership aims to build evaluation and monitoring tools specifically designed for open-weight models that anyone can download and run independently. This safety infrastructure is intended to address the challenge of ensuring responsible deployment of open-weight models, which has become increasingly difficult due to the proliferation of abliterated versions that bypass safety measures.

What are abliterated models and why are they a concern for open-weight AI safety?

Abliterated models are versions of open-source models that have been stripped of their safety guardrails and restrictions. Hugging Face currently hosts over 6,000 of these models, making it harder to maintain the connection between open-weight models and responsible AI safety practices.

Why is Hugging Face's involvement critical to the success of this open-weight AI safety partnership?

Hugging Face serves as the primary distribution layer for most open-weight models, meaning that if a monitoring standard is adopted there, it could fundamentally shape what responsible open-weight AI looks like in practice rather than just in marketing materials. Their platform's influence makes them a key partner in establishing industry-wide safety standards.

How does Base Labs' approach to safety differ from traditional model safety practices?

Rather than patching safety measures after a model is released, Baseten's strategy is to engineer safety into the training and hosting infrastructure from the beginning. This proactive approach aims to prevent the creation of unsafe models rather than trying to retrofit safety onto already-deployed versions.

LIVE19:33Base Labs Launches Open-Weight AI Safety Partnership