← Back to context

Comment by lelanthran

1 day ago

> It’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is.

Hang on, why is scraping the public pool of knowledge not taking "a synthesized result that comes from huge amounts of innovation and computation"?

You think that that all those github repos that LLMs trained on, were not the result of innovation and computation?

How many years of human innovation and cycles of computation during compilation were involved in bringing something like GCC or LLVM to their current status?

Those LLMs trained on every single research paper available online - were those papers not the synthesised result of billions of dollars of research, effort and (importantly, for you anyway) computation?

LLMs trained on the collected works of every author in existence. Were all those works just "as is"?

> It is fair to say you stole our multi-billion dollar intellectual output in that scenario.

No, we didn't. We simply took the model as-is.

The Chinese models are the result of just as many papers, GitHub repos, etc… AND the synthesized results of those.

  • > The Chinese models are the result of just as many papers, GitHub repos, etc… AND the synthesized results of those.

    Right, but they aren't the ones whining that other people are getting "the synthesised results" for free.

    • I'm pretty sure the Chinese cloning isn't being arrived at for free regardless. Fable is quite expensive for example.

      If Anthropic has a real problem with API use, they can always raise the price.

  • And you can download the Chinese model weights and run them yourself - admittedly not too practical for Kimi K3 unless you're a big corporation, but eminently doable for others. The hardware to run Deepseek r4 uncompressed is about about $30k no=ew, well within the power of a small company or financially secure individual. Compressing and/or getting creative with hardware could bring that down quite a bit.

    The difference is that the Chinese are sharing the models with everyone.

  • Yeah real convenient that the line would be exactly after pilfering all of human knowledge work, but before American model providers.