← Back to context

Comment by elar_verole

15 hours ago

I think this can't work because an LLM needs too much data, and before the internet there probably just wasn't enough to get close to what we have now

Even simpler: Can GPT-2 anticipate and build Gwen/Deepseek? I think the answer is almost trivially "no", so I wonder what changed?

  • Lots of things changed, GPT-2 is small (1.5e9) and is also a base model, so it is only doing next-token/autocomplete rather than prompt-response like even the first ChatGPT-3.5 was doing.

    • Just for the sake of clarity: all LLMs up to today are still only doing next-token/autocomplete. The training process got additional stages to shape the model weights, but standalone LLMs are still deployed essentially identically.

      6 replies →

  • That's not really a fair comparison, no? Modern LLMs are much more capable than GPT-2. We'ld need a modern LLM trained on exclusively old data, and that might be impossible

Why couldn't an LLM, if it was smart enough, generate and consume its own data?

I know the answer: because it leads to model collapse. But why is that? Wouldn't a smart model not collapse? It's seeming like they keep getting smarter because we keep pouring more of our own knowledge into them, not because they are actually getting smarter. And yes, sometimes a dumb but persistent bruteforcer can make new discoveries.

  • > if it was smart enough

    and i think this is exactly the crux;

    the really big models need really big datasets

    and current gen LLMs get a lot of training data beyond "all books + all of the internet"

    the objection is then that producing this additional data would already confound it with pre "virtual cutoff date" knowledge (since the training data probably implies mathematical and SWE concepts that were developed post "virtual cutoff date")

  • It's because LLMs are entropy generators. That's not a bad thing for what people are doing.

    But to prevent model collapse you need a way to pump down the entropy. Much like in thermo, it's an expensive and slow process.

  • If it is smart enough to generate data it can consume to train itself better, it is already smart enough to not need to do that.

    • If a human is smart enough to do the Michelson-Morley experiment, they are smart enough to not need to do that.

Maybe we can synthesize large amounts of limited information. I thought that new training data is mostly synthetic anyway.