Editorial illustration for Experts: Kimi K3's Gains Not From Costly, Slow Frontier Model API
Kimi K3 Built by Copying Anthropic Model, Experts Say
Experts: Kimi K3's Gains Not From Costly, Slow Frontier Model API
Michael Kratsios, the White House's top science advisor, posted on X this week that Moonshot AI built Kimi K3 by copying Anthropic's Fable model, using chips barred from export to China in the process. "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable," Kratsios wrote. He didn't say where the allegation came from, and Moonshot hasn't answered questions about how it trained the model, which is currently the largest open-weight LLM available.
Treasury Secretary Scott Bessent made a similar claim recently, saying officials are "finding watermarks of our U.S. large language models on many of the Chinese models, and that that's unacceptable." Nobody at Treasury has explained what those watermarks actually look like. The accusations land as Washington debates whether to restrict Chinese open-weight models altogether, a prospect that has already unsettled AI labs and researchers who rely on them.
The timeline is what's drawing scrutiny. Fable, Anthropic's model, only went public on July 1. Kimi K3 arrived two weeks later. Researchers who study distillation, the practice of querying a rival model to reverse-engineer its behavior, say that window is too tight to explain what Moonshot pulled off.
“I’ve been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning],” Nathan Lambert, an AI researcher at the Allen Institute for AI, said in a podcast released yesterday.
Why this matters Kratsios's accusation lands at a moment when Washington is weighing whether to ban Chinese open-weight models outright, and the technical case for "covert distillation" is thinner than the political rhetoric suggests. If querying Anthropic's Fable through its API is as slow and costly as described, Moonshot's team would have needed enormous budgets and patience to build Kimi K3 that way, not exactly the profile of a covert operation. For developers and founders building on open-weight models, the real story might be more mundane: firms train on whatever public data and prior model outputs are available, American or otherwise, and calling that theft sets a precedent that could boomerang on U.S.
labs too. Researchers should watch how the White House defines "distillation" in any coming policy, because a vague standard could justify restricting access to models that compete well on price and performance. The chip provenance question is separate and verifiable.
The plagiarism claim, so far, is not. Until someone produces output logs or training traces tying Kimi K3 directly to Fable's API, this reads less like forensic evidence and more like a talking point built for a trade fight.
Common Questions Answered
What allegation did Michael Kratsios make about how Moonshot AI built Kimi K3?
Michael Kratsios, the White House's top science advisor, alleged that Moonshot AI built Kimi K3 by copying Anthropic's Fable model using chips that are barred from export to China. He characterized this as large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research.
Why might covert distillation of Anthropic's Fable model be impractical for building Kimi K3?
According to the article, querying Anthropic's Fable through its API is both slow and costly, which would require enormous budgets and patience to build a model like Kimi K3 covertly. This operational profile doesn't align with what a covert operation would realistically look like, suggesting the technical case for distillation is weaker than the political rhetoric suggests.
What is Nathan Lambert's perspective on the effectiveness of distillation for Chinese AI models?
Nathan Lambert, an AI researcher at the Allen Institute for AI, believes that distillation has become less impactful over time as Chinese models get closer to the frontier and training regimes shift toward reinforcement learning. This suggests that distillation may not be as viable a strategy for advancing Chinese AI capabilities as some critics claim.
What political context surrounds Kratsios's accusation about Kimi K3?
Kratsios's accusation comes at a time when Washington is weighing whether to ban Chinese open-weight models outright. The timing suggests the allegation is part of broader policy discussions about restricting Chinese AI development and protecting American technological advantages.
Further Reading
- Moonshot Ships Kimi K3: 2.8T Parameters, 1M Context, Frontier-Scale Open-Weight Model - How AI Works
- Kimi K3: Moonshot AI's 2.8T Open-Weight Model - Eigent AI
- Kimi K3: The Complete Guide to Moonshot AI's 2.8T Model - Amplifi Labs
- Kimi K3: Moonshot AI Unveils a 2.8-Trillion-Parameter Open-Weight Model - WE0 AI
- Kimi K3: Benchmarks, Pricing & Review — Moonshot's 2.8T Frontier Model - The AI Rankings