Skip to main content
AI-generated text detection, large language models, post-training guardrails, study findings, digital security.

Editorial illustration for Post-Training Guardrails Make LLM Text Detectable, Study Finds

AI Text Detectors Work Because of Safety Training

Post-Training Guardrails Make LLM Text Detectable, Study Finds

4 min read

Bradley Emi has a job that depends on catching AI-written text, and he just published a blog post explaining why his own company's detector works as well as it does. Emi is CTO of Pangram, a startup that sells AI text detection, and his argument cuts against the usual doom-laden framing around chatbots getting too good to catch. The problem, he says, isn't capability. It's what happens after training ends.

Every major chatbot on the market, ChatGPT, Claude, Gemini, goes through a second phase of training after the initial model is built, where engineers teach it rules about what to say and how to say it. That phase is supposed to make the tool safer to release. Emi's post argues it also leaves fingerprints in the writing itself, fingerprints detectable at scale by tools like his own.

The raw, untuned versions of these models don't carry the same signature, according to Emi, and neither do narrower fine-tuned models built for specific writing styles. Watermarked text is a separate case entirely, one Emi says should remain detectable regardless.

Post-training and safety guardrails keep their text detectable, argues Bradley Emi, CTO of AI text detector Pangram, in a blog post. Systems like ChatGPT, Claude, or Gemini learn behavioral rules to avoid dangerous outputs or censor certain political statements. This sharply narrows their expressive range, an effect called "mode collapse."

Why this matters

For anyone building on top of ChatGPT, Claude, or Gemini, this is worth sitting with: the same alignment work that makes these models safe to ship is also what makes their output fingerprint-able. Pangram's Bradley Emi is arguing that mode collapse isn't a side effect of scale or architecture, it's a byproduct of RLHF and safety tuning narrowing the space of acceptable phrasing. That's a useful reframe for founders selling "undetectable AI writing" tools, and a warning for anyone assuming detection is a losing arms race against ever-smarter models.

If base models write with more human-like variety before guardrails get bolted on, detectability may actually track how much a lab has constrained its model, not how advanced it is. Worth watching: whether AI text detectors start marketing themselves as guardrail detectors rather than AI detectors, and whether that shifts how labs think about the tradeoff between safety tuning and stylistic fingerprinting. It also raises an obvious question for red-teamers and policy folks: does less restrictive post-training make a model harder to detect, and is that a cost anyone's actually pricing in?

Common Questions Answered

Why does Bradley Emi argue that post-training guardrails make LLM text detectable?

Bradley Emi, CTO of Pangram, argues that post-training safety guardrails and RLHF (Reinforcement Learning from Human Feedback) create a phenomenon called "mode collapse" that narrows the expressive range of AI models like ChatGPT, Claude, and Gemini. This narrowing of acceptable phrasing creates a detectable fingerprint in the text output, making AI-written content identifiable despite the underlying capability of these models to write more naturally.

What is mode collapse and how does it relate to AI text detection?

Mode collapse is the sharp narrowing of expressive range that occurs when LLMs undergo post-training and safety tuning to avoid dangerous outputs or censor certain political statements. This restriction in how models can phrase their responses creates a distinctive pattern that makes their text more detectable, rather than the problem being that AI models are becoming too sophisticated to catch.

How do ChatGPT, Claude, and Gemini become safer through their training process?

These major chatbots go through a second phase after initial training that involves post-training guardrails and safety tuning to learn behavioral rules that prevent dangerous outputs and enforce content policies. However, this safety-focused training has the side effect of limiting their linguistic diversity and creating detectable patterns in their generated text.

What is the significance of Pangram's findings for AI safety and detection?

Pangram's research reframes the relationship between AI safety and detectability, showing that the alignment work making models safe to deploy is simultaneously what makes their output fingerprint-able and detectable. This finding is important for founders building detection tools and serves as a warning for those attempting to create undetectable AI writing tools, as the safety mechanisms themselves create the detection signatures.

LIVE20:47Post-Training Guardrails Make LLM Text Detectable, Study Finds