Skip to main content
AI model secrets revealed: a glowing brain graphic with interconnected nodes, symbolizing complex neural networks.

Editorial illustration for New Method Could Let Anyone Distill AI Models' Secrets

New Method Lets Anyone Extract AI Model Secrets

New Method Could Let Anyone Distill AI Models' Secrets

4 min read

Ask a frontier AI model to solve a hard math problem and it will show its work: pages of reasoning tokens laying out the steps before it lands on an answer. Companies like OpenAI, Anthropic and Google have treated that scratchpad as proprietary, something users can glimpse but not fully extract. A team of computer scientists just found a way around that.

Researchers from the University of Tübingen, the Max Planck Institute, MATS Research, and the security firm Snyk built a method for pulling the hidden reasoning traces out of models accessed through an API, the standard way developers query systems like Claude Opus 4.8 and GPT 5.6 Sol. The technique worked across every major frontier provider they tested. Once they had those traces in hand, they compared them against the outputs of open-weight models built elsewhere, including Moonshot AI's Kimi K3, looking for signs that one model's reasoning had been used to train another.

What they found raises pointed questions about how some competing labs, particularly in China, may have built their systems, and exposes a security hole with consequences well beyond bragging rights over benchmarks.

“All major frontier model providers we tested share this vulnerability,” says Alexander Panfilov, a computer scientist at University of Tübingen in Germany who was involved with the work. “It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks.”

Why this matters

For anyone building on top of closed models, this changes the threat model around what "hidden" reasoning actually protects. Panfilov's method suggests the chain-of-thought tokens providers like OpenAI keep behind the curtain aren't as sealed off as advertised, and that has real consequences for licensing terms, API pricing, and the whole premise that reasoning traces can be monetized separately from outputs. We'd resist the urge to treat this as proof that specific Chinese labs distilled specific US models; the researchers themselves stop short of that claim.

But the more interesting story for developers and founders is upstream: if reasoning can be extracted this way, the "moat" around frontier reasoning capability is thinner than boardrooms have been assuming. Expect closed-model providers to respond with tighter output filtering, more aggressive rate limiting on token-level access, or contractual language explicitly banning this kind of extraction. Worth watching whether OpenAI or Anthropic issue any technical response, and whether this method gets replicated against other frontier systems in the next few months.

Common Questions Answered

What vulnerability did researchers discover in frontier AI models' reasoning tokens?

Researchers from the University of Tübingen, Max Planck Institute, MATS Research, and Snyk discovered a method to extract the chain-of-thought reasoning that companies like OpenAI, Anthropic, and Google had kept proprietary. According to Alexander Panfilov, all major frontier model providers tested share this vulnerability, which enables large-scale reasoning distillation attacks and can lead to personal information leakage.

How does this method threaten the monetization of AI reasoning traces?

The discovered vulnerability undermines the premise that reasoning traces can be monetized separately from model outputs, since the chain-of-thought tokens are not as sealed off as providers advertised. This has real consequences for licensing terms, API pricing, and the business model that depends on keeping reasoning scratchpads proprietary and hidden from users.

What are chain-of-thought tokens and why did companies keep them proprietary?

Chain-of-thought tokens are the reasoning steps that frontier AI models generate when solving complex problems, showing their work before arriving at an answer. Companies treated these reasoning scratchpads as proprietary intellectual property that users could glimpse but not fully extract, allowing them to potentially monetize reasoning separately from final outputs.

Which institutions collaborated on discovering this AI model vulnerability?

The research team included computer scientists from the University of Tübingen in Germany, the Max Planck Institute, MATS Research, and the security firm Snyk. Alexander Panfilov from the University of Tübingen was one of the key researchers involved in developing the method to extract hidden reasoning from frontier AI models.

LIVE15:12Nvidia's Switchyard Router Cuts AI Task Costs to a Third