Comment by bbor

4 hours ago

Well, other than the infrastructure they got from illegally routing millions of paying customers' requests through Anthropic's Opus 4.8 in a distillation attack...

Anthropic infringed the copyright of basically every author on the planet: https://www.anthropiccopyrightsettlement.com/

No real reason to respect any terms they might want to impose. Besides, if you want to break TOS, just have an agent do it; "everyone" running these things agrees there's no corporate or moral liability for what your AI does.

  • I'm not defending their actions, but we should be clear about where the law currently stands: Anthropic was found to infringe because of the torrenting, not because of the training.

That is such a canard, IMO. FWIW, Anthropic and OpenAI encrypt "thinking" token outputs in their models, while Chinese labs don't. If anything, it's more likely that everyone is using open-weight models in their synthetic training data generation pipelines. It's way easier to distill from logits than it is to distill from hard tokens.

https://x.com/EricSimons/status/2099252922098061714

We weep for Dario, that he had to suffer such a devastating attack against his Terms of Service.

What is "illegal" about it?

  • breaking Anthropic TOS and misleading users

    • Breaking TOS isn't illegal per se. It just allows for denial of services, and may define terms by which the provider can reclaim costs.

  • Are you joking...? Sorry if so! Just in case: It's illegal in both the PRC and the USA.

    In the PRC, they[1] leaked tons of national secrets on the PRC's latest AI campaigns, the inner workings of their "opinion monitoring" (read: performative panopticon) and "stability" (read: violent oppression) departments, Chengdu's whole CCTV network, direct-energy weapons plans, espionage activities in Syria to hunt down Uyghur refugees, and god knows what else that Anthropic didn't divulge to us common folk.

    In the US, it's very clearly an attempt to rip off a competitor. I'm not sure how else you could possibly see it. Even if you're a distillation fan in general (which A. why and B. plz don't), they did this through a network of Japanese and Signaporean shell accounts, presumably at least some of which were abusing Anthropic's subscription service in a ToS double-whammy, as it would be exorbitantly expensive otherwise. They also had to hack around Anthropic's API to get CoT traces, which seems impossible to explain away as anything innocent.

    I've been beating the "China isn't necessarily an enemy, it's gonna take us all to handle AI" drum for literally years, but this attack was just... gross. Gross in scale and gross in arrogance. Not a good sign for the dawning alignment crisis, to say the least :(

    TL;DR: Use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS. So... buyer beware, I guess.

    [1]: For clarity, Z.ai was not alone in this, nor were they most egregious attack -- Moonshot.ai (kimi) took that coveted prize. DeepSeek was involved, too.

    • What does any of this have to do with the legality of distilling Claude?

      > use these services if you want, but know that you're supporting aggressive escalations and companies that very clearly don't give a flying fuck about violating the law, much less your ToS

      From my European point of view the same risk/concerns apply when using US providers

    • I'm always wondering when "distillation" comes up how feasible it is, or if it's just BS.

      The Antrophic article mentions "16 million" conversations, GLM models are in the 700-300 billion parameter ranges and while the frontier sizes aren't know but Gemini suggests Astra and Mythos are at around 10 trillion. That'd amount to extracting 40k parameters per conversation without a lot of errors if it was just a distillation (from an unknown source/algorithm as opposed to distilling your own model).

      Now, I can imagine these conversations being used as a verification step that they're not missing stuff in their training, and that their models are capable of most of the same things, but that's mostly confirming that they've stolen the same data from the public as Antrophic/OpenAI has stolen already.

      Or am I missing something here that makes real "distillation" feasible?

    • > alignment crisis

      Alignment is meaningless; as you've noticed, humans aren't all that "morally aligned".

      If the tool needs safety measures it should be kept in a safe enclosure like we do with CNC machines, furnaces, and so on.

    • Like Anthropic and OpenAI are? After all, didn't they distill all the information in the world into their model(s)?

      I mean, if they get to distill other's IP, why can't others distill their IP?

If you understand what they have achieved here, then the notion that they are bottle-necked on training data is absurd.

I wonder how you imagine that China built their own space station? Reliant on using American made duct tape, perhaps?

Do you realize how reasoning models are being trained nowadays? You design/build simulation environments to run agents in, with the environment providing the RLVR "verification" scoring. So why won't Ziphu use GLM to build their own RL training environments? Do you think they are not doing this?

Eh, even if this was true, then they're merely stealing from thieves. Anthropic did break a ToS or two to get training data themselves.