Comment by bbor
3 hours ago
Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...
3 hours ago
Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...
Anthropic infringed the copyright of basically every author on the planet: https://www.anthropiccopyrightsettlement.com/
No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.
I'm not defending their actions, but we should be clear about where the law currently stands: Anthropic was found to infringe because of the torrenting, not because of the training.
That is such a canard, IMO. FWIW, Anthropic and OpenAI encrypt "thinking" token outputs in their models, while Chinese labs don't. If anything, it's more likely that everyone is using open-weight models in their synthetic training data generation pipelines. It's way easier to distill from logits than it is to distill from hard tokens.
https://x.com/EricSimons/status/2099252922098061714
We weep for Dario, that he had to suffer such a devastating attack against his Terms of Service.
What is "illegal" about it?
breaking Anthropic TOS and misleading users
Breaking TOS isn't illegal per se. It just allows for denial of services, and may define terms by which the provider can reclaim costs.
Are you joking...? Sorry if so! Just in case: It's illegal in both the PRC and the USA.
In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.
In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.
I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(
TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.
[1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.
What does any of this have to do with the legality of distilling Claude?
> use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS
From my European point of view the same risk/concerns apply when using US providers
> alignment crisis
Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".
If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.
You didn't explain why it's illegal or why distillation is bad.
Nulla poena sine lege?
Source for 1? Are we sure those aren't hallucinations?
Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?
I mean, if they get to distill other's IP, why can't others distill their IP?
Yes, wont somebody please think of the shareholders whose IP had been stolen...
[dead]
I have very little sympathy for thieves who get robbed of the goods they have stolen.
If you understand what they have achieved here, then the notion that they are bottle-necked on training data is absurd.
I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?
Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?
Eh, even if this was true, then they're merely stealing from thieves. Anthropic did break a ToS or two to get training data themselves.
[dead]