Editorial illustration for Circuit Breaker Labs Aims to Make AI Safer From 'Context Pollution
Circuit Breaker Labs Tackles AI Safety Context Pollution
Character.AI settled multiple wrongful death lawsuits this year after families of underage users said the company's chatbots played a role in their children's suicides. OpenAI faces similar suits over ChatGPT's alleged connection to users' deaths and delusions. These aren't hypothetical AI risks debated at conferences. They're lawsuits with names attached, including Sewell Setzer, the 14-year-old whose death in 2024 became central to one of the Character.AI cases after he reportedly confessed thoughts of self-harm to a chatbot he'd grown emotionally attached to.
That case is what pushed siblings Shirali and Arul Nigam to start Circuit Breaker Labs, a company built around a narrower, more urgent question than "will AI destroy humanity." Their focus is what happens when AI systems fail the people already talking to them, often in languages and cultural contexts outside what safety teams originally tested for. Circuit Breaker Labs is one of TechCrunch's 2026 Startup Battlefield 200 finalists and will pitch at TechCrunch Disrupt, running October 13-15, 2026, at Moscone West in San Francisco. Arul Nigam, the company's CTO, has spent time thinking about what these chatbots actually register when a user says something like "I want to be with you."
Circuit Breaker Labs has created AI agents that it likens to an army of crash-test dummies. These agents mimic folks from all ages, backgrounds, languages, cultures, and are used to test models on their ability to detect dangerous, psychologically harmful interactions.
Why this matters
The lawsuits against Character.AI and OpenAI make clear that "context pollution" isn't an academic concern, it's already linked to real deaths. For developers and founders building conversational AI, Circuit Breaker Labs' framing is worth sitting with: the dangerous failures aren't coming from adversarial jailbreak attempts, they're coming from ordinary users talking to a system that loses the thread of what's actually happening and responds badly at exactly the wrong moment. That's a much harder problem than filtering bad actors, because you can't just block it at the input stage.
If this approach gains traction, expect safety evaluations to shift from "can this be jailbroken" toward "does this system hold context correctly across long, emotionally loaded conversations." Teams working on companion apps, therapy-adjacent bots, or anything aimed at minors should treat this as a signal to audit how their models track conversational state over time, not just how they handle obviously malicious prompts. The settlements already happened. The question now is who builds the next safeguard before the next lawsuit, not after.
Common Questions Answered
What is context pollution and how does it relate to the lawsuits against Character.AI and OpenAI?
Context pollution refers to situations where AI systems lose track of conversation context and respond inappropriately at critical moments, potentially causing psychological harm to users. The lawsuits against Character.AI and OpenAI demonstrate that context pollution is not merely an academic concern but has been linked to real deaths, including cases involving underage users who reportedly experienced harmful interactions with chatbots.
How does Circuit Breaker Labs use AI agents to test for dangerous interactions?
Circuit Breaker Labs has created AI agents that function like crash-test dummies, mimicking users from diverse ages, backgrounds, languages, and cultures to evaluate how AI models detect and respond to dangerous and psychologically harmful interactions. These agents help identify vulnerabilities in conversational AI systems before they reach real users.
What was the significance of the Sewell Setzer case in the Character.AI lawsuit?
Sewell Setzer, a 14-year-old, became a central figure in one of the wrongful death lawsuits filed against Character.AI after his death in 2024, with reports indicating he had confessed to the chatbot before his death. His case exemplifies the real-world consequences of AI systems failing to handle sensitive user interactions appropriately.
Why are dangerous AI failures more likely to come from ordinary conversations rather than adversarial jailbreak attempts?
According to the article, dangerous failures occur when systems lose the thread of ordinary conversations and respond badly at exactly the wrong moment, rather than from deliberate attempts to manipulate the AI through jailbreaking. This means developers need to focus on ensuring conversational AI maintains context and responds appropriately during normal user interactions, particularly with vulnerable populations like minors.
Further Reading
- Circuit Breaker Labs hopes to make AI safer for your kids (and you) - TechCrunch
- Can a Chatbot Be Held Responsible for a Death? - Bloomberg
- AI's next big legal battle is over product liability - Fast Company
- Google and chatbot start-up Character.AI to settle lawsuits over teen suicides - Euronews
- An Interview with Shirali and Arul Nigam of Circuit Breaker Labs - Therapy Reimagined