• 4 min read
Experts doubt Fable copying made Kimi K3 so strong
Researchers question whether Anthropic’s Fable could explain Kimi K3's rapid rise, while chip-access allegations add a second layer to the dispute.

Image: TechCrunch
White House science adviser Michael Kratsios has accused Moonshot, the Chinese company behind Kimi K3, of copying Anthropic’s Fable large language model while using Nvidia chips barred from export to China. Kimi K3 is described as the largest available open-weight LLM.
“Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable,” Kratsios wrote, amid reported discussions about banning Chinese open-weight models. Moonshot did not respond to questions about its training process, and Kratsios offered no further details supporting the allegation.
Treasury Secretary Scott Bessent has made a similar claim, saying officials are finding “watermarks of our U.S. large language models on many of the Chinese models.” It is unclear what those watermarks are, and the Treasury Department did not respond to a request for comment.

Recommended reading
Jibo’s successor wants to turn life into AI memories
Why researchers doubt rapid Fable distillation
Several researchers told TechCrunch that distillation alone is unlikely to explain Kimi K3's capabilities. Distillation involves systematically querying a target model, then using its answers to train another model. The process can include asking a model to explain its chain of thought, or using its prompts and responses for supervised fine-tuning, known as SFT.
“I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation. There’s just not even frankly time, right? Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks.”
Nathan Lambert, an AI researcher at the Allen Institute for AI, said in a podcast released yesterday that distillation has become less impactful as Chinese models approach the frontier and training shifts toward reinforcement learning. SFT can teach a model what Lambert called its “manners,” but he said supervised fine-tuning alone is unlikely to reproduce the strongest capabilities of a frontier system.
Reproducing Fable-like performance would probably require reinforcement learning. That can involve using a larger model or agent to grade a smaller model’s responses and adjust its behavior. Large reinforcement-learning runs may require tens of millions of agents; using a frontier model’s API for that process would be “insanely expensive,” potentially slow, and might not produce better performance.
Earlier models, chips and data-center oversight
Previous frontier models may still have contributed to Kimi K3. Anthropic accused Moonshot, DeepSeek and MiniMax earlier this year of systematically distilling its models, citing millions of exchanges identified through IP addresses and other metadata. Anthropic said those queries reflected deliberate capability extraction rather than ordinary use. The company did not respond to TechCrunch’s questions about whether Fable had been distilled.
Distillation is not necessarily limited to Chinese companies. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models to develop Grok, describing the practice as common across the industry. The boundary between distillation and the creation of synthetic training data can also be difficult to define.
Hancock cautioned against underestimating Chinese researchers. One Moonshot founder was a Carnegie Mellon University PhD student, he said, and the company’s researchers and engineers are capable of making substantial progress independently.
Kratsios also alleged that Moonshot obtained advanced Nvidia Grace Blackwell 300s and accessed servers equipped with GB300 chips in Thailand. Those chips are banned from export to China, although a black market exists, according to Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology.
“I am a proponent of know your customer laws for data centers across the world. If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing.”
President Joe Biden’s Commerce Department proposed federal know-your-customer rules for data centers in 2024, but no further progress appears to have been made under Donald Trump. Exporters shipping advanced chips abroad are still expected to ensure that the hardware is used only for approved purposes.
AI Editor
Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.
via TechCrunch


