Comment by bluegatty

7 days ago

Making an LLM from raw data is value-add.

Distillation is just value extract.

It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.

I think we start by recognizing that ... and then try to figure it out from there.

'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.

> Making an LLM from raw data is value-add. > Distillation is just value extract.

There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.

  • I think that's kind of fair, but it still fits within the context of 'some things are value add' and 'more or less than others'.

    We ought to identify that and integrate that into our thinking.

What makes the Internet raw data in a different way? wasn't it mostly worked on by people first?

  • There is value add in AI irrespective of how the data got to what it is.

    Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.

    'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.

    • Okay, so if the chinese models are used everyday, do they become a value add? Like what's the line you're drawing here. Amount of value it creates?

      5 replies →