Does this matter? Distillation is not illegal by every definition of the word.
There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.
And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
edit: (moved this to bottom)
The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.
show comments
throwa356262
Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited.
How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies?
I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies
show comments
madduci
So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.
So here robbers are blaming robbers?
These claims are just pointless, everytime
show comments
sent-hil
Reminds of the quote by Bill Gates.
> "Well, Steve [Jobs]… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
Commenters are overlooking the significance of this information and posting emotional reactions based on perceptions of fairness or feelings of schadenfreude.
The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to continue to lose money or to reduce spending on R&D.
Moonshot's papers [1] claim that their training load was primarily from synthetic data and model self-teaching rather than RLHF and therefore keep their costs low. If Moonshot genuinely does not rely on human-led training, they will surpass US closed-source model providers. The United States government considers US supremacy in "AI" as a national security consideration.
This announcement is noteworthy because it implies that Moonshot's success is in fact due to distillation. It's in the interest of US frontier labs to place barriers to this if they find themselves in the position of subsidizing rival labs' research.
I can understand that the AI labs might care about other labs distilling their models as it can eat into their competitive advantage, but do consumers care at all? Aren't consumers benefiting from this practice by getting better cheaper models as a result?
show comments
ekelsen
Reminds me of this classic line from the 1973 movie The Sting:
"What was I supposed to do? Call him for cheating better than me in front of the others?!"
Said in response to being out-cheated at a high-stakes poker game.
Except in this case, it sounds like that's exactly the path they have chosen.
Super interesting. So Fable was really made available... a couple weeks ago? And K3 a few days ago? That's a really impressive feat to distill enough data AND train AND review to get a release that works really well in that time period. Mad props to the Moonshot team :flame:.
nahuel0x
Information wants to be free.
sscaryterry
How much credible, provable evidence? None really.
teravor
the distillation everyone talks about in respect to LLM's isn't nearly as easy as most think.
none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.
therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.
efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.
what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.
HeavyStorm
Poor AI labs... All they hard earned training, done via scraping a lot of people works for free, now being scraped through payed subscriptions...
thih9
One of the replies:
> @MehdiKarech
> I don't remember letting Anthropic or Open Ai scrapping my GitHub, my research gate and all my online writings L O L
According to these results GLM 5.2 is very similar to Google Gemini and Kimi K3 is very similar to Fable 5.
The American frontier labs are not similar to each other.
show comments
Geee
You wouldn't distill a car.
show comments
softwaredoug
OpenAI and Anthropic should enter into distillation agreements with other US labs. Turn a threat into a profit center.
Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation.
Why fight it when there’s clear money to make here?
nchmy
We have information that Claude distilled billions of copyrighted, and otherwise-created-by-others, materials for the development of their entire business.
feverzsj
If web scraping is legal, so is distilling.
show comments
storus
I doubt they did any distillation as Hinton defined it (requiring logit access). They most likely ran a bunch of prompts/conversations and captured the results. Those conversations already missed thinking tokens, replaced by some confusing quasi-summaries. Then they took those and ran basic SFT or maybe DPO if they had competing responses. As there is no copyright on the output of AI, I am not sure where is the "covert industrial distillation" part of the problem.
wseqyrku
Every model is distilled internet. This is just the natural progression of that.
grim_io
So, if it's that easy and fast to "copy" Fable, is it really worth that much in the first place?
Sounds like the opposite of the conversation Anthropic would want to have.
blaufast
The frontier labs' work is more akin to discovery than artistic expression. An art piece is valued for its uniqueness and individuality, but AI is valued for verifiable correctness. Discovery cannot be unseen and is easily replicable. I think the AI labs are in a tough situation because their work is more similar to fundamental scientific discovery than say, a unique painting or song.
Mendel doesn't get a cut every time somebody uses the principles of heritability he discovered, and Einstein's family aren't getting royalties if you compute relative speeds. I think the frontier labs should expect to be treated more like scientists than artists in this regard.
jmward01
If 'distillation' means training on outputs then what is the legal concept of ownership of outputs? And, more broadly, is this something that could be skirted by doing it in different countries that have different legal structures? Basically, are they saying they own those outputs, not the companies that paid for the tokens, and only they can train on them? I suspect a lot of companies are saving their token histories and using them to fine tune internal models.
show comments
strictnein
The level of discourse here anytime distillment is mentioned is so mundane. Do we need 40 people saying the same thing about how they don't feel bad and it serves them right and all that surface level stuff on every single one of these? This is the level of insight one receives anytime you mention chocolate and dogs "Oh it's poisonous for dogs!". Yes, we've all heard that 100 times. Thanks for adding nothing to the conversation.
A more interesting part of this discussion is that consistently these Chinese models are held up as a great achievement, and that they're "catching up" when in reality they're just using the work of Anthropic and OpenAI to try and keep up with them. This isn't even to say it's not a valid tactic, but it definitely colors these announcements and proclamations about foreign companies catching up to American ones.
If I get a 1600 on the SAT and you copied my answers and got a 1540, your achievement isn't that significant.
show comments
alexruf
Who cares? Is it theft if a thief gets robbed of their stolen goods?
Gives me more of a modern Robin Hood vibe tbh.
No, seriously: first of all, that's not the AI labs' data, it's ours. And if the AI labs think they can rake in tons of money using our data, then I'm actually glad if someone comes along and at least offers us a good product at reasonable prices.
throwa356262
In the meantime, reddit is making fun of Opus for "distilling" Qwen:
K3 frequently refers to itself as Claude in reasoning when instructed to play a role.
NetOpWibby
GOOD
I love using Claude but Fable's unusable wrt useful work like cryptography, biology, &c.
Kneecapping my productivity when I pay $100/month is annoying af.
mrandish
How was K3 trained on data distilled from Fable when Fable was only publicly available in the last two weeks before K3 was released? The timing just doesn't work.
cyanydeez
oh know, better make them pay up your legal fines for stealing all that from the public good.
kamranjon
So here is an important question I think.
If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?
I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.
So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.
I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.
econ
I have an idea! If they are so hungry for citable content they should start a cheap or free blogging platform with images and video and a blogroll and verified credentials and resumes, with your own html css etc and domain name and a git server and a mail client and their own advertisement platform and aggregator and a chat platform, scientific journals too obviously, tools for writing and publishing books and documents. API available everywhere to avoid training on its own output.
Because there is no way in hell I'm going to make an effort creating quality content for existing platforms. The website should be entirely my own without moderation subject only to my local legal system.
Can just insert this comment as a prompt and vibe code everything in a few days⸮
jerrythegerbil
“However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”
What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user prompts and traces. These inference providers are explicitly blocked in the claude cli if you reverse engineer it.
The real picture is that these Chinese labs have figured out how to get exactly what they need, at a high quality, directly from distinct and unique real user prompts.
It’s only “covert” because Anthropic doesn’t like it, while simultaneously being perfectly fine to do.
cmiles8
But wasn’t fable distilled from knowledge taken from others? I get why Anthropic is angry here, but it would appear they’re not really in a position to complain about this.
BigTTYGothGF
Good for Moonshot.
alastairr
Presumably the frontier labs themselves can do their own distillation far better than the chinese labs can. Why can't they just beat them at their own game and release / host low cost intelligence and own the whole game. There will always be a market for the more expensive frontier intelligence.
felooboolooomba
The pot calling the kettle black.
wnmurphy
It's funny to me that these models were created by effectively "distilling" all available content including the proprietary works of many other people, but now it's a problem that someone is doing the same to them.
You're using available information (copyrighted works, or the output of another model) to train a model to encode the information in a new form. Why is the former not theft, but the latter is theft?
jchw
Is this person trustworthy? I struggle to believe that in the relatively short time Fable was available it has already been distilled so effectively. If this really is actually true, very impressive work.
pandinus
As with many others among these threads I don't see how the timing works out for K3 to have trained on distilled Fable usage. There should be at least a tacit academic acknowledgment of Kimi's own design efforts.
Distillation itself, however, is still clearly valuable - else competitors wouldn't pay so much to their rival on distillation campaigns or try to circumvent anti-distillation defenses.
As for the morality of it, if you paid for the tokens they're yours. It is already understood that you own the output. Seems to me like a variation of ordinary business arbitrage. Providers might object to certain use-cases or intention and try to craft terms around that, but that's hard to enforce at scale.
show comments
yeodev
The US gov and AI providers when they steal billions of user data, content and media to train their models on: :)
The US gov and AI providers when funny chinese people steal their data to train their models: >:(
clowns
edit: TIL you can't use emojis on HN
4chandaily
Seems to me like Moonshot is a paying customer, and if their business isn't worth the money Anthropic is charging, perhaps they should raise the price per token charged for it. Otherwise, I don't see why this is a story. "AI Company pays another AI Company for training data" just isn't that interesting.
NichoPaolucci
I wonder if this points at a “shared” future (or at least things will eventually converge there whether companies like it or not). Ultimately, if you’re going to release these models that are fundamentally built on shared data - it’s pretty wishful to assume you’ll be able to harbor that model and the data, forever, and profit from it.
It also leads me to think about things like the original release of Fable 5, people were complaining that it was safeguarded too much - if you lock the models down too much they cease to be useful. So it’s going to be increasingly difficult to protect a model from competition while ALSO keeping it useful.
show comments
zuzululu
If distilling is fair then so is banning it. It's funny how people cry about Anthropic using books and materials to train itself but then when Anthropic does something about it they think its unfair. Pick a lane.
asadotzler
A company that distills LLMs should be called "Moonshine" not "Moonshot"
Ba-dum-tss
chasd00
if there are no consequences then who cares? You're not going to take Chinese companies to court and stealing IP is nothing new either. It's going to take some sort of policy change at the federal government level to do anything but they haven't done much up to this point. Maybe AI is important enough to actually get some kind of policy change, sucks for everyone else who have had their IP stolen with no consequences whatsoever.
show comments
riknos314
If it's true that in under 15 days of access significant improvements were realized in K3, then the moat of closed-weight models is far smaller than previously thought.
Doesn't bode well for the valuations of these labs.
nmeofthestate
Weird 'discussion'. Almost entirely single messages with no threads, all with the same anti-Anthropic/AI position.
tanh
For code can't they distill from public GitHub commits? If they could figure out who used Mythos/Fable assitance in the commits.
show comments
muldvarp
Okay? We have information that Anthropic sucked up all of the internet for the development of Fable.
alightsoul
how is it possible to distill fable only a month after its release? maybe they are confusing opus with fable.
show comments
mrhottakes
Good. If Fable is really so smart, it wouldn't let itself be distilled.
neals
How does one distill? Just send a million request asking for information? Start with the letter A?
show comments
rambojohnson
who cares. all these frontier models are trained on theft.
Chance-Device
Hmm. I wonder when this was detected. And was the CoT trace cut from Fable from the start on June 9th or just after the export ban and relaunch? Is this what the export ban was actually about? I honestly don’t know, just wondering aloud.
mrbonner
Hah tales as old as time. what’s next? Distillation of Disney theme park?
thundoe
The Irony. These models have been created distilling Internet without ever asking for permission or paying anyone. Internet was the first model.
Catloafdev
I wonder how they detect this kind of thing. Seems like this is going to be a perpetual issue until it stops being worth doing.
Side note, didn't they stop releasing real thinking tokens for Fable? Or is it still part of some subs or API usage?
kouteiheika
Assuming they did then they surely paid for them, which makes it "not stealing". Am I also "stealing proprietary U.S. technology" by harvesting my Claude chats from my `.claude` directory and training a bunch of models on them?
That said, I doubt the "they distilled Fable" is the reason why K3 is as good as it is, considering the timelines involved, and that Anthropic hides thinking traces, and their overly aggressive "safety" filters.
This constant FUD spread by Anthropic is so tiring.
show comments
sajithdilshan
How the tables have turned. It's okay for Anthropic to train their models on copyrighted data, but it's wrong to steal the stolen data from Anthropic models.
nozzlegear
Cry about it IMO. Anthropic reaps what they sow.
sensanaty
For a buncha supposed capitalists, they sure do hate fair competition eh?
tacone
So it is as "dangerous" as Fable?
scronkfinkle
so they distilled one of the best models in the world AND released it for free to everyone. Where can I send them flowers as a thank you?
stephbook
I don't know what purpose these "they copied us" crying is ever going to achieve. Europeans stole Chinese silk worms. US stole European books, looms and rocket scientists. Who cares? Be grateful you've got people inventing stuff worth copying.
codedokode
Do you by chance also have information about Anthropic's training data sources?
Gud
OK, DIRECTOR Michael Kratsios, but why should we give a shit?
American AI corporations are pushing up the prices for computing, making it unaffordable for the common man. Additionally, they have built their entire business on stealing(yes, stealing) work from us.
So fuck em
noncoml
Yes, I know this is not Reddit but Clarkson’s “Oh no! Anyway…” is the perfect, and most fitting, reaction to this. Nothing else to say
NDlurker
Good. Keep it up
Jaauthor
"Well, Steve, I think there's more than one way of looking at it. I think it's more like we both had this rich neighbor named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
mbix77
Didn't they just pay a fine for stealing all those books?
show comments
stranded22
Seems like an advert for K3 to me.
Fable level performance, for much lower price.
But really, this is the USA getting ready to bring AI companies completely under the control of the Trump administration for ‘national security’
stego-tech
I continue to laugh uproariously at American AIBros screaming “bUt OuR iP” at foreign distilleries and competitors despite literally building their own platforms on the single largest theft of copyrightable works in human history.
Like, goddamn ya’ll are hypocrites.
m_ke
Anthropic should think hard about all their fear mongering. It will only end up backfiring on them and everyone else involved.
They definitely used closed private saas products to train their own models, to prove that just drop random small screenshots of any popular product behind a login screen and see how well it's able to identify all of them. ex: https://x.com/michalwols/status/2079968211865330165
So? Their business model requires building on the labor of others for free. Isn't that how it works?
bparsons
IP protections for me, not for thee.
bakugo
Fun fact about K3's distillation:
As of a couple months ago, when using Claude to write adult content through the API, sometimes it will silently inject a system prompt giving the model a bunch of guidelines on exactly what kind of adult content it's allowed to write, steering it away from anything "questionable" ("Claude will not write etc etc").
Moonshot distilled Claude so hard recently, they actually ended up distilling this prompt injection, too. Using K3 to write adult content results in it randomly hallucinating the injected Claude prompt during thinking, and it will quote parts of that prompt, complete with the name "Claude".
Not that I think distillation is a bad thing, just thought this was funny.
dmitrygr
OMG someone used our data to make an MK model! Just like we did to every author in the world!
stldev
So company who stole stuff to make their stuff is mad because another company is stealing their stuff.
And now a regime best known for lying to their own people is the one trying to convince me?
Go, China!
superloika
I think they deserve, by Justice, to have their models pillaged and raped, just like they did to the internet. They didn't ask for permission when they took the entire of the internet, after all, and given their behaviour is nefarious, it's of Justice that they receive nefarious treatment by others, including chinese AI labs.
The Chinese are not gonna deterred, but the posturing by the Americans is so blatantly hypocritical that everybody is cheering for their demise. See, for example, one of Francis Fukuyama's latests videos on youtube.
show comments
tamimio
“If you can’t compete with them, get them banned”
- US AI companies
show comments
guess_who_is
If you ask fable, it will identify as deepseek
mattrighetti
Is distillation something we have to live with or are there ways to prevent it?
show comments
deaton
Who cares. Anthropic distilled the entire internet, and then a good bit more beyond that.
juancn
So?
We have information that Fable was distilled from humans.
If it works it works. Isn't that the argument?
AI outputs are not copyrightable, so distillation is fair use.
It may be a TOS violation, but that's a private matter. Cancel the accounts used for distillation and be done.
lowbloodsugar
China has done this with absolutely everything, starting with “customs inspections” of ships engineering sections by “inspectors” drawing diagrams of what they see. Bit late to be worrying about it now. This only matters now because China is now near parity in tech and vastly superior in production ability. Meanwhile we run out of bullets in a five month war with Iran.
martinjc
So what? I want the best model at the cheapest price. You guys illegally trained on books, movies, audiobook etc.. Why should we care?
caycep
honestly if they did what he said they did, it seems like it would be cheaper just to train your own model from the get go
show comments
surgical_fire
First: Even if true, I don't care.
Second: Post is rich with allegations but light with evidence. Can very well be bullshit.
drop_star
America, the perpetual victim
show comments
xnoto
"we ripped off the entire ecosystem of copyrighted data but I draw the line when we get ripped off"
jauntywundrkind
Two recent ones that really really hit me,
> we're entering the most geopolitically volatile moment since the trinity test lit up the alamogordo desert and the only US policy prescription is a big button labeled sinophobia
> every vendor cranking the big dial labeled "sinophobia" and looking back at the us government for approval
The government itself doing the propaganda here, skipping the vendors. Sinophobia intensifies. War drums of "be afraid be afraid be afraid" beat louder.
It's so bad, it's so stupid. Kimi lands one showing pretty clearly this was absolutely the determining concern happening at vast scale, that they can just a lot of this themselves, and this noise pollution from the most hopelessly lost aggro administration ever still gets blared out the trumpets of war & discord. What a joke. Give me a break, give it a rest.
War here is less winnable than the Iran war they started. They're going to make America itself so much worse, these people so hungry to put down free and good models. This pathetic attempt is not going to work, you are just going to once again hold the US citizens hostage & make their lives worse, for sick political games.
show comments
sleepyguy
Is this a surprise, I think history has proven that the Chinese technology theft is part of their strategy. They let the American tax payer or "The West" shoulder the cost and then steal it.
Waiting for the whataboutism....junk away...
show comments
traceroute66
"we have information" says a US Government official who almost certainly has had Anthropic and/or OpenAI on the phone spinning him stories.
See also, don't trust anyone in Trump's government who says "we have information".
"they distilled us" is fast becoming standard US FUD.
The same as people telling me with a serious face that the Chinese models are distilled just because it says "I am Claude".
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the distillation panic and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening – that Chinese companies are extremely good at building models in the same way the leading American companies are.
The problem is… what are you going to do about it?
This is obviously an idiotic and dangerous Cold War and has no happy ending.
avazhi
And?
Nobody cares. This is neither a controversy nor news, and that would be the case even if Anthropic hadn’t just settled a 1.5 billion dollar lawsuit where they trained Claude on thousands of books without permission lol.
To be clear I’m not taking a jab at OP - I’m saying the labs crying about distillation have neither a legal nor a moral leg to stand on. There’s nothing wrong with distillation.
tibbydudeza
Proof - they also claimed that China has an ASML UEV machine - crickets when ASML said it was impossible due to all the safeguards and assistance needed to operate one.
The current US administration is known to be collection of BS artists and liars.
treetalker
rules for thee but not for me
show comments
mbmbn
“China’s great leaps in AI that are surpassing the US” are actually just what China always does with every technology: copy the west… poorly.
And before the Chinese astroturfing starts (it already started, that’s clear from the comments and voting): the point is not even that the US companies have the right to intelectual property over their models (they should, but ok, that’s not even the point). The point is that China is incapable of innovation and any innovation into AI we can expect, will always come from the US.
show comments
solumunus
Get your violins out folks.
brap
Anyone surprised by this is incredibly naive.
By all means use whatever works for you, I’m not even going to try to make an argument on ethics (and honestly I’m not even sure where I stand, given the behavior of American AI companies).
But I just cringe every time I see people acting like any of this is done in good faith.
Open source coming out of China is a state-sponsored criminal enterprise, built only for the benefit of the Chinese regime, one of the worst to exist in human history.
Does this matter? Distillation is not illegal by every definition of the word.
There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.
And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
edit: (moved this to bottom) The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.
Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited.
How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies?
I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies
So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.
So here robbers are blaming robbers?
These claims are just pointless, everytime
Reminds of the quote by Bill Gates.
> "Well, Steve [Jobs]… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
Source: https://www.goodreads.com/quotes/824084-well-steve-jobs-i-th...
Commenters are overlooking the significance of this information and posting emotional reactions based on perceptions of fairness or feelings of schadenfreude.
The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to continue to lose money or to reduce spending on R&D.
Moonshot's papers [1] claim that their training load was primarily from synthetic data and model self-teaching rather than RLHF and therefore keep their costs low. If Moonshot genuinely does not rely on human-led training, they will surpass US closed-source model providers. The United States government considers US supremacy in "AI" as a national security consideration.
This announcement is noteworthy because it implies that Moonshot's success is in fact due to distillation. It's in the interest of US frontier labs to place barriers to this if they find themselves in the position of subsidizing rival labs' research.
1. Kimi K2, https://arxiv.org/html/2507.20534v1
I can understand that the AI labs might care about other labs distilling their models as it can eat into their competitive advantage, but do consumers care at all? Aren't consumers benefiting from this practice by getting better cheaper models as a result?
Reminds me of this classic line from the 1973 movie The Sting:
"What was I supposed to do? Call him for cheating better than me in front of the others?!"
Said in response to being out-cheated at a high-stakes poker game.
Except in this case, it sounds like that's exactly the path they have chosen.
https://getyarn.io/yarn-clip/7612c4ce-1077-479f-a7bf-617dbc6...
Super interesting. So Fable was really made available... a couple weeks ago? And K3 a few days ago? That's a really impressive feat to distill enough data AND train AND review to get a release that works really well in that time period. Mad props to the Moonshot team :flame:.
Information wants to be free.
How much credible, provable evidence? None really.
the distillation everyone talks about in respect to LLM's isn't nearly as easy as most think.
none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.
therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.
efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.
what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.
Poor AI labs... All they hard earned training, done via scraping a lot of people works for free, now being scraped through payed subscriptions...
One of the replies:
> @MehdiKarech
> I don't remember letting Anthropic or Open Ai scrapping my GitHub, my research gate and all my online writings L O L
https://xcancel.com/MehdiKarech/status/2080000779859939678#m
Here's a site that asks the same questions to 22 models and compares how similar their responses are.
https://typebulb.com/u/lab/you-re-relatively-right/full
According to these results GLM 5.2 is very similar to Google Gemini and Kimi K3 is very similar to Fable 5.
The American frontier labs are not similar to each other.
You wouldn't distill a car.
OpenAI and Anthropic should enter into distillation agreements with other US labs. Turn a threat into a profit center.
Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation.
Why fight it when there’s clear money to make here?
We have information that Claude distilled billions of copyrighted, and otherwise-created-by-others, materials for the development of their entire business.
If web scraping is legal, so is distilling.
I doubt they did any distillation as Hinton defined it (requiring logit access). They most likely ran a bunch of prompts/conversations and captured the results. Those conversations already missed thinking tokens, replaced by some confusing quasi-summaries. Then they took those and ran basic SFT or maybe DPO if they had competing responses. As there is no copyright on the output of AI, I am not sure where is the "covert industrial distillation" part of the problem.
Every model is distilled internet. This is just the natural progression of that.
So, if it's that easy and fast to "copy" Fable, is it really worth that much in the first place?
Sounds like the opposite of the conversation Anthropic would want to have.
The frontier labs' work is more akin to discovery than artistic expression. An art piece is valued for its uniqueness and individuality, but AI is valued for verifiable correctness. Discovery cannot be unseen and is easily replicable. I think the AI labs are in a tough situation because their work is more similar to fundamental scientific discovery than say, a unique painting or song.
Mendel doesn't get a cut every time somebody uses the principles of heritability he discovered, and Einstein's family aren't getting royalties if you compute relative speeds. I think the frontier labs should expect to be treated more like scientists than artists in this regard.
If 'distillation' means training on outputs then what is the legal concept of ownership of outputs? And, more broadly, is this something that could be skirted by doing it in different countries that have different legal structures? Basically, are they saying they own those outputs, not the companies that paid for the tokens, and only they can train on them? I suspect a lot of companies are saving their token histories and using them to fine tune internal models.
The level of discourse here anytime distillment is mentioned is so mundane. Do we need 40 people saying the same thing about how they don't feel bad and it serves them right and all that surface level stuff on every single one of these? This is the level of insight one receives anytime you mention chocolate and dogs "Oh it's poisonous for dogs!". Yes, we've all heard that 100 times. Thanks for adding nothing to the conversation.
A more interesting part of this discussion is that consistently these Chinese models are held up as a great achievement, and that they're "catching up" when in reality they're just using the work of Anthropic and OpenAI to try and keep up with them. This isn't even to say it's not a valid tactic, but it definitely colors these announcements and proclamations about foreign companies catching up to American ones.
If I get a 1600 on the SAT and you copied my answers and got a 1540, your achievement isn't that significant.
Who cares? Is it theft if a thief gets robbed of their stolen goods? Gives me more of a modern Robin Hood vibe tbh.
No, seriously: first of all, that's not the AI labs' data, it's ours. And if the AI labs think they can rake in tons of money using our data, then I'm actually glad if someone comes along and at least offers us a good product at reasonable prices.
In the meantime, reddit is making fun of Opus for "distilling" Qwen:
https://www.reddit.com/r/ClaudeCode/comments/1tqaist/opus_48...
(don't take this too seriously)
K3 frequently refers to itself as Claude in reasoning when instructed to play a role.
GOOD
I love using Claude but Fable's unusable wrt useful work like cryptography, biology, &c.
Kneecapping my productivity when I pay $100/month is annoying af.
How was K3 trained on data distilled from Fable when Fable was only publicly available in the last two weeks before K3 was released? The timing just doesn't work.
oh know, better make them pay up your legal fines for stealing all that from the public good.
So here is an important question I think.
If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?
I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.
So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.
I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.
I have an idea! If they are so hungry for citable content they should start a cheap or free blogging platform with images and video and a blogroll and verified credentials and resumes, with your own html css etc and domain name and a git server and a mail client and their own advertisement platform and aggregator and a chat platform, scientific journals too obviously, tools for writing and publishing books and documents. API available everywhere to avoid training on its own output.
Because there is no way in hell I'm going to make an effort creating quality content for existing platforms. The website should be entirely my own without moderation subject only to my local legal system.
Can just insert this comment as a prompt and vibe code everything in a few days⸮
“However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”
What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user prompts and traces. These inference providers are explicitly blocked in the claude cli if you reverse engineer it.
The real picture is that these Chinese labs have figured out how to get exactly what they need, at a high quality, directly from distinct and unique real user prompts.
It’s only “covert” because Anthropic doesn’t like it, while simultaneously being perfectly fine to do.
But wasn’t fable distilled from knowledge taken from others? I get why Anthropic is angry here, but it would appear they’re not really in a position to complain about this.
Good for Moonshot.
Presumably the frontier labs themselves can do their own distillation far better than the chinese labs can. Why can't they just beat them at their own game and release / host low cost intelligence and own the whole game. There will always be a market for the more expensive frontier intelligence.
The pot calling the kettle black.
It's funny to me that these models were created by effectively "distilling" all available content including the proprietary works of many other people, but now it's a problem that someone is doing the same to them.
You're using available information (copyrighted works, or the output of another model) to train a model to encode the information in a new form. Why is the former not theft, but the latter is theft?
Is this person trustworthy? I struggle to believe that in the relatively short time Fable was available it has already been distilled so effectively. If this really is actually true, very impressive work.
As with many others among these threads I don't see how the timing works out for K3 to have trained on distilled Fable usage. There should be at least a tacit academic acknowledgment of Kimi's own design efforts.
Distillation itself, however, is still clearly valuable - else competitors wouldn't pay so much to their rival on distillation campaigns or try to circumvent anti-distillation defenses.
As for the morality of it, if you paid for the tokens they're yours. It is already understood that you own the output. Seems to me like a variation of ordinary business arbitrage. Providers might object to certain use-cases or intention and try to craft terms around that, but that's hard to enforce at scale.
The US gov and AI providers when they steal billions of user data, content and media to train their models on: :)
The US gov and AI providers when funny chinese people steal their data to train their models: >:(
clowns
edit: TIL you can't use emojis on HN
Seems to me like Moonshot is a paying customer, and if their business isn't worth the money Anthropic is charging, perhaps they should raise the price per token charged for it. Otherwise, I don't see why this is a story. "AI Company pays another AI Company for training data" just isn't that interesting.
I wonder if this points at a “shared” future (or at least things will eventually converge there whether companies like it or not). Ultimately, if you’re going to release these models that are fundamentally built on shared data - it’s pretty wishful to assume you’ll be able to harbor that model and the data, forever, and profit from it.
It also leads me to think about things like the original release of Fable 5, people were complaining that it was safeguarded too much - if you lock the models down too much they cease to be useful. So it’s going to be increasingly difficult to protect a model from competition while ALSO keeping it useful.
If distilling is fair then so is banning it. It's funny how people cry about Anthropic using books and materials to train itself but then when Anthropic does something about it they think its unfair. Pick a lane.
A company that distills LLMs should be called "Moonshine" not "Moonshot"
Ba-dum-tss
if there are no consequences then who cares? You're not going to take Chinese companies to court and stealing IP is nothing new either. It's going to take some sort of policy change at the federal government level to do anything but they haven't done much up to this point. Maybe AI is important enough to actually get some kind of policy change, sucks for everyone else who have had their IP stolen with no consequences whatsoever.
If it's true that in under 15 days of access significant improvements were realized in K3, then the moat of closed-weight models is far smaller than previously thought.
Doesn't bode well for the valuations of these labs.
Weird 'discussion'. Almost entirely single messages with no threads, all with the same anti-Anthropic/AI position.
For code can't they distill from public GitHub commits? If they could figure out who used Mythos/Fable assitance in the commits.
Okay? We have information that Anthropic sucked up all of the internet for the development of Fable.
how is it possible to distill fable only a month after its release? maybe they are confusing opus with fable.
Good. If Fable is really so smart, it wouldn't let itself be distilled.
How does one distill? Just send a million request asking for information? Start with the letter A?
who cares. all these frontier models are trained on theft.
Hmm. I wonder when this was detected. And was the CoT trace cut from Fable from the start on June 9th or just after the export ban and relaunch? Is this what the export ban was actually about? I honestly don’t know, just wondering aloud.
Hah tales as old as time. what’s next? Distillation of Disney theme park?
The Irony. These models have been created distilling Internet without ever asking for permission or paying anyone. Internet was the first model.
I wonder how they detect this kind of thing. Seems like this is going to be a perpetual issue until it stops being worth doing.
Side note, didn't they stop releasing real thinking tokens for Fable? Or is it still part of some subs or API usage?
Assuming they did then they surely paid for them, which makes it "not stealing". Am I also "stealing proprietary U.S. technology" by harvesting my Claude chats from my `.claude` directory and training a bunch of models on them?
That said, I doubt the "they distilled Fable" is the reason why K3 is as good as it is, considering the timelines involved, and that Anthropic hides thinking traces, and their overly aggressive "safety" filters.
This constant FUD spread by Anthropic is so tiring.
How the tables have turned. It's okay for Anthropic to train their models on copyrighted data, but it's wrong to steal the stolen data from Anthropic models.
Cry about it IMO. Anthropic reaps what they sow.
For a buncha supposed capitalists, they sure do hate fair competition eh?
So it is as "dangerous" as Fable?
so they distilled one of the best models in the world AND released it for free to everyone. Where can I send them flowers as a thank you?
I don't know what purpose these "they copied us" crying is ever going to achieve. Europeans stole Chinese silk worms. US stole European books, looms and rocket scientists. Who cares? Be grateful you've got people inventing stuff worth copying.
Do you by chance also have information about Anthropic's training data sources?
OK, DIRECTOR Michael Kratsios, but why should we give a shit?
American AI corporations are pushing up the prices for computing, making it unaffordable for the common man. Additionally, they have built their entire business on stealing(yes, stealing) work from us.
So fuck em
Yes, I know this is not Reddit but Clarkson’s “Oh no! Anyway…” is the perfect, and most fitting, reaction to this. Nothing else to say
Good. Keep it up
"Well, Steve, I think there's more than one way of looking at it. I think it's more like we both had this rich neighbor named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
Didn't they just pay a fine for stealing all those books?
Seems like an advert for K3 to me.
Fable level performance, for much lower price.
But really, this is the USA getting ready to bring AI companies completely under the control of the Trump administration for ‘national security’
I continue to laugh uproariously at American AIBros screaming “bUt OuR iP” at foreign distilleries and competitors despite literally building their own platforms on the single largest theft of copyrightable works in human history.
Like, goddamn ya’ll are hypocrites.
Anthropic should think hard about all their fear mongering. It will only end up backfiring on them and everyone else involved.
They definitely used closed private saas products to train their own models, to prove that just drop random small screenshots of any popular product behind a login screen and see how well it's able to identify all of them. ex: https://x.com/michalwols/status/2079968211865330165
or other similar "AI" startups https://x.com/envconfig/status/2079613455296827402
So? Their business model requires building on the labor of others for free. Isn't that how it works?
IP protections for me, not for thee.
Fun fact about K3's distillation:
As of a couple months ago, when using Claude to write adult content through the API, sometimes it will silently inject a system prompt giving the model a bunch of guidelines on exactly what kind of adult content it's allowed to write, steering it away from anything "questionable" ("Claude will not write etc etc").
Moonshot distilled Claude so hard recently, they actually ended up distilling this prompt injection, too. Using K3 to write adult content results in it randomly hallucinating the injected Claude prompt during thinking, and it will quote parts of that prompt, complete with the name "Claude".
Not that I think distillation is a bad thing, just thought this was funny.
OMG someone used our data to make an MK model! Just like we did to every author in the world!
So company who stole stuff to make their stuff is mad because another company is stealing their stuff.
And now a regime best known for lying to their own people is the one trying to convince me?
Go, China!
I think they deserve, by Justice, to have their models pillaged and raped, just like they did to the internet. They didn't ask for permission when they took the entire of the internet, after all, and given their behaviour is nefarious, it's of Justice that they receive nefarious treatment by others, including chinese AI labs.
The Chinese are not gonna deterred, but the posturing by the Americans is so blatantly hypocritical that everybody is cheering for their demise. See, for example, one of Francis Fukuyama's latests videos on youtube.
“If you can’t compete with them, get them banned”
- US AI companies
If you ask fable, it will identify as deepseek
Is distillation something we have to live with or are there ways to prevent it?
Who cares. Anthropic distilled the entire internet, and then a good bit more beyond that.
So?
We have information that Fable was distilled from humans.
If it works it works. Isn't that the argument?
AI outputs are not copyrightable, so distillation is fair use.
It may be a TOS violation, but that's a private matter. Cancel the accounts used for distillation and be done.
China has done this with absolutely everything, starting with “customs inspections” of ships engineering sections by “inspectors” drawing diagrams of what they see. Bit late to be worrying about it now. This only matters now because China is now near parity in tech and vastly superior in production ability. Meanwhile we run out of bullets in a five month war with Iran.
So what? I want the best model at the cheapest price. You guys illegally trained on books, movies, audiobook etc.. Why should we care?
honestly if they did what he said they did, it seems like it would be cheaper just to train your own model from the get go
First: Even if true, I don't care.
Second: Post is rich with allegations but light with evidence. Can very well be bullshit.
America, the perpetual victim
"we ripped off the entire ecosystem of copyrighted data but I draw the line when we get ripped off"
Two recent ones that really really hit me,
> we're entering the most geopolitically volatile moment since the trinity test lit up the alamogordo desert and the only US policy prescription is a big button labeled sinophobia
https://bsky.app/profile/thebadcode.com/post/3mr3skoyass2k , and,
> every vendor cranking the big dial labeled "sinophobia" and looking back at the us government for approval
The government itself doing the propaganda here, skipping the vendors. Sinophobia intensifies. War drums of "be afraid be afraid be afraid" beat louder.
It's so bad, it's so stupid. Kimi lands one showing pretty clearly this was absolutely the determining concern happening at vast scale, that they can just a lot of this themselves, and this noise pollution from the most hopelessly lost aggro administration ever still gets blared out the trumpets of war & discord. What a joke. Give me a break, give it a rest.
War here is less winnable than the Iran war they started. They're going to make America itself so much worse, these people so hungry to put down free and good models. This pathetic attempt is not going to work, you are just going to once again hold the US citizens hostage & make their lives worse, for sick political games.
Is this a surprise, I think history has proven that the Chinese technology theft is part of their strategy. They let the American tax payer or "The West" shoulder the cost and then steal it.
Waiting for the whataboutism....junk away...
"we have information" says a US Government official who almost certainly has had Anthropic and/or OpenAI on the phone spinning him stories.
See also, don't trust anyone in Trump's government who says "we have information".
"they distilled us" is fast becoming standard US FUD.
The same as people telling me with a serious face that the Chinese models are distilled just because it says "I am Claude".
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
[1] https://www.interconnects.ai/p/kimi-k3-the-open-weights-esca...I’m certain they did.
The problem is… what are you going to do about it?
This is obviously an idiotic and dangerous Cold War and has no happy ending.
And?
Nobody cares. This is neither a controversy nor news, and that would be the case even if Anthropic hadn’t just settled a 1.5 billion dollar lawsuit where they trained Claude on thousands of books without permission lol.
To be clear I’m not taking a jab at OP - I’m saying the labs crying about distillation have neither a legal nor a moral leg to stand on. There’s nothing wrong with distillation.
Proof - they also claimed that China has an ASML UEV machine - crickets when ASML said it was impossible due to all the safeguards and assistance needed to operate one.
The current US administration is known to be collection of BS artists and liars.
rules for thee but not for me
“China’s great leaps in AI that are surpassing the US” are actually just what China always does with every technology: copy the west… poorly.
And before the Chinese astroturfing starts (it already started, that’s clear from the comments and voting): the point is not even that the US companies have the right to intelectual property over their models (they should, but ok, that’s not even the point). The point is that China is incapable of innovation and any innovation into AI we can expect, will always come from the US.
Get your violins out folks.
Anyone surprised by this is incredibly naive.
By all means use whatever works for you, I’m not even going to try to make an argument on ethics (and honestly I’m not even sure where I stand, given the behavior of American AI companies).
But I just cringe every time I see people acting like any of this is done in good faith.
Open source coming out of China is a state-sponsored criminal enterprise, built only for the benefit of the Chinese regime, one of the worst to exist in human history.