← Back to context

Comment by delichon

11 hours ago

If you want to believe that the success of Kimi is about distillation attacks, ignore this.

I'd kindly suggest that we could also stop calling them "distillation attacks".

  • Agreed. I think when it comes to light that Claude has been known to say “I’m DeepSeek” that everyone has had their hand in that cookie jar. Moreover, paying for API calls hardly seems like an attack; ToS violation to be certain but not in the same category of law as criminal activity like hacking.

    • wait, is there evidence of this? I've not observed it. It sounds like the kind of thing that I want to be true because it would be hilarious but that makes me suspicious.

      2 replies →

  • Why on earth wouldn't you? It's clearly a forcible, aggressive, non-consensual attempt to take something. That's an attack in any other terms. It's totally fair if you approve of the attack, and want the attack to succeed. But your preference doesn't stop it from being what it is.

    • > It's clearly a forcible, aggressive, non-consensual attempt to take something

      Distillers are not taking anything, they are just making their model learn from better ones - isn't that the whole AI training doesn't violate IP argument?

      They just aren't using the tool in compliance with the terms of service. Anthropic could ban them or take them to court maybe. Not an attack still.

    • Fine, then let's refer to the initial data collection/training as 'compilation attacks' going forward.

      ...or maybe we stop defaulting to adversarial paradigms for every conceivable situation.

    • "They are using and paying for an API I am selling...But they save the data and use it for something I don't like! I'm being attacked."

False dichotomy right?

Are Chinese labs impressively innovating? Clearly.

However this doesn’t rule out possible gains due to distillation.

I don’t know the degree of the latter but both things could certainly be true.

  • Didn't Anthropic train on our collective data just to sell it back to us for $100/month? On top of that, Apple is suing them over alleged IP and trade secret theft by ex-Apple employees. Hard to feel too sympathetic, and I’m not an Anthropic hater in particular…

  • Fable was available for a few weeks before Kimi K3 came out. If it was a distillation attack, then that's a truly groundbreaking technological feat to distill a model like Fable in 2 weeks

  • If they can distill fable into a full model post training run in ~15 days without the real thinking traces, yet we know Claude chats degraded with the thinking traces removed (chat resume bug from earlier in the year they reported stripping thinking to shed load as being the cause of degradation), how big can this degree be?

  • All LLMs are based on distillation broadly defined. Western models began distilling texts. If Chinese models are distilling Western models, they are taking information that Western models don't own anyway - but that doesn't mean the Chinese models aren't also taking information from text as well (which they probably also don't own). And none of this means Western and Chinese companies aren't innovating by creating very elegant methods of distillation.

  • “Distillation” is just indirectly pirating the largely pirated training data used to train the original model.

    “You stole my warez!”

The distillation complaints to me sound like when a casino complains about card counting

I stil don't understand them. I want the US to "win the AI race" but I have trouble understanding how most of all inventions today aren't "distillations" of past knowledge. Is Anthropic claiming the data they stole as trade secrets?

  • i want china to win so that i get access to ai and not restricted and censored.

    the chinese models are less censored, you'd better believe it.

    try asking claude about its 'guardrails' (restrictions), very high chance anthropic will censor it.

  • Anthropic is claiming that training an LLM to mimic another LLM is materially different and worse than slurping up stuff written by humans (even if that material is stolen).

    Basically, they want IP protection for Claude. This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

    • Their claim is even stronger than that, they have complaints about their models being used as a validation step for other model output, which is standard practice in the industry.

    • >This is a nakedly hypocritical stance, but completely understandable from a company-needs-to-make-money standpoint.

      No, it's perfectly reasonable once you get down to reality.

      China is not going to care about IP. That's just a fact. So either nobody cares about IP (at the very last in this context) and any AI company can just do whatever with data, or Chinese companies have to be held up to scrutiny.

      We don't have the privilege to be able to hold western companies to higher ethical, legal, and environmental standards and not risk competitiveness.

      That there is a whole lot of people right now who insist on doing above and still somehow praise China at every turn is something historians or news outlets will have to make sense of in some 5 years time.

      5 replies →

You can’t build a frontier model with one single thing. This is an incremental improvement but it doesn’t explain the entire success of the model. The training set is immensely important, regardless of how you feel about distillation.

It can easily be both. Also, they didn't use this innovation in K3 - K3 pre-training would have started months ago and the paper only mentions a 48B model. The people working at this level may not even be heavily involved in shipping a new iteration of K3, or at least theory contributions to it were done many months or even a year ago and after that it is all engineering.

Well said.

The distillation theory does not even make sense as Fable was only around for days (effectively) before Kimi was released.