← Back to context

Comment by _aavaa_

8 days ago

> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation

Sounds great to me; live by the sword, die by the sword.

Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs.

But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?

  • This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.

    • I'm not doubting it's possible to pass such a law, I'm doubting that's it's a practical or worthwhile goal.

      The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.

      If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.

      4 replies →

  • > Seems only fair

    "You're trying to kidnap what I've rightfully stolen!" -- Vizzini

  • It's pretty common to have such laws. OpenAI can put whatever they want in their ToS, but they cannot go back and sue someone for violating those terms if the government has ruled that clause to be unenforceable.

Forbidding distillation is like forbidding using a compiler to make another(perhaps better, more efficient) compiler.

  • Lots of software licenses have “non-compete” clauses that forbid you from using it to develop a competing product. Wouldn’t surprise me if there was a compiler or two out there with that restriction, most likely niche languages.

    • It's been common in electronic design automation tools to have license terms like that (forbidding use to create a competing product). However, competing companies have often found workarounds, either by finding loopholes or just breaking rules and hoping not to get caught.

    • If a person were to receive data from someone subjected to such restriction, is the receiver bounded by the same restriction?

    • balmer told us that gpl is cancer, but true cancer is us model of licensing

    • How the hell is non-compete legal in market economy? Competition is one of its core strengths. Why would anyone let anyone opt out of this, even a little bit?

      2 replies →

  • let's call it for what it really is, only companies "entitled to legally stolen data, don't steal from us now" are crying about distilation

The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US).

It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.

It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.

Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.

  • > Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning

    "Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."

    I don't know why you're trying to downplay it.

    European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.

    • >I don't know why you're trying to downplay it.

      Ignoring that I have literally zero trust in anything Anthropic has to say on this -- they have been doing the hysterical routine and trying to get every bit of government granted monopoly they can[1] -- those numbers still simply aren't that impressive.

      >European models are so far behind because...

      What a non-sequitur. Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States. European efforts on this are poorly funded, poorly capitalized, and marginal efforts.

      China is very much not Europe. China is looking to leave the US to the dustbin of history, and their efforts are a little more concerted.

      [1] Surely Americans are aware that Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models, right?

      3 replies →

    • > European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.

      You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.

      1 reply →

  • Which Chinese model was it that identified itself as Claude 15% of the time?

    • Claude Opus identifies itself as Qwen if you ask the question in Chinese. So who's really distilling who?

    • Models don't have some self identity, beyond what is explicitly handed to them via a system prompt. There have been many, many cases of models identifying as different models by different makers as a basic identity hallucination. They train on enormous volumes of data including lots of people talking about certain makers and models (ChatGPT was actually a super common one given that it became the kleenex of the LLM world). Hence why vendors have to specifically tell it to override that, and if they don't you get lots of funny cases of identity confusion.

      This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.

Yup, fair's fair. Anything else stinks of 'rules for thee but not for me' (a maxim the frontier labs seem worryingly happy to apply, on several counts).

Don’t know much about how distillation works so please enlighten me here.

> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models

If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?

  • known-good prompt-response pairs are more useful than random semi-coherent texts presumably

  • You need to do both.

    A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.

  • Because the model can output data in a manner optimized for training a new model, including outputs that were post-trained like RLHF and RLVR.

Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them.

ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.

So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).

  • The government absolutely can pass laws that ban particular contract previsions. They do that all the time. In your analogy for example while they can require you to wear red, they can't require you to be white.

  • Governments can do anything they want by passing a new legislation. In your example, they could easily pass a law that states that any ToS cannot reject service to a customer based on the color of their attire. In the USA, it's obviously already illegal for a business to reject service to a customer based on some protected classes like race.

  • That's just not true. You can absolutely have terms of service that are illegal, and the government can enforce them.

Why would reading copyrighted material ever be an issue anyway? Wouldn't copyright law only apply to what you create and publish using the model? Training on every comic book should already be perfectly legal, as long as you accessed them legally, right? But publishing your own Batman comic using that training is copyright infringement.

What I'm saying is, doesn't the law already cover 1?

  • Fair use requires more than you accessing the material legally.

    In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.

    If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

    Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.

    • > If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.

      Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?

      1 reply →

[flagged]

  • This is a silly perspective, inaccurate, and out of bounds framing.

    Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear.

    But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in the classroom, it's not their place to discriminate against anyone/thing (even those like itself (other robots)) that also show up in the classroom.

    But you can't teach at a university using only knowledge learned from the library. you need a degree. You are free to teach at the park, where anyone can hear you. public in -> public out.

    • If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.

      8 replies →

  • No, but the students that learn and distill what the professor teaches are not obligated to use that information only how the professor wants them to.

  • 'why is reselling stolen stuff bad'

    if the professor took all human knowledge, much of which was explicitly not free, and used it to make a for-profit knowledge machine that extrudes unreliable summaries of that knowledge, then yes, being obligated to teach for free would be a fitting punishment.

  • More like, is a professor who learned from books prohibited from writing his own books on the subject?

    • He is prohibited from regurgitating source material, of course! But if he generalized from the books he read and really learned the subject--and even made new connections between ideas--then he is free to write his own book.

      1 reply →

  • 1) No one is asking Anthropic to give tokens for free, but at market rates.

    2) Any professor who tried to ban students from posting lecture notes online would be immediately mocked.

Making an LLM from raw data is value-add.

Distillation is just value extract.

It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.

I think we start by recognizing that ... and then try to figure it out from there.

'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.

  • > Making an LLM from raw data is value-add. > Distillation is just value extract.

    There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.

    • I think that's kind of fair, but it still fits within the context of 'some things are value add' and 'more or less than others'.

      We ought to identify that and integrate that into our thinking.

  • What makes the Internet raw data in a different way? wasn't it mostly worked on by people first?

    • There is value add in AI irrespective of how the data got to what it is.

      Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.

      'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.

      6 replies →