Comment by somenameforme

21 hours ago

Where some see intelligence, others see token prediction. It's a very good question how token prediction could achieve this, but I think there's a simple explanation. No human can hold more than a negligible percent of all knowledge in his mind at once. LLMs have no such limits and so can reliably connect 'obvious' dots that we miss simply for lack of storage capability.

Well isn't that just semantics? Surely connecting dots in a novel and meaningful way is intelligence regardless of how it's achieved. The thing is that humans didn't get to where we are by connecting obvious dots. Go back to before humans had invented language and when bleeding edge tech was literally that - 'poke him with the pointy end.' Train an LLM on that corpus of knowledge. Even given infinite processing power and infinite time - it's not going to discover the secrets of the atom, put a man on the Moon, or do much of anything besides remix what we'd already done at the time.

I expect there's still much LLMs can achieve simply because of this initial problem. But I expect that they will ultimately start to plateau once these dots have been mostly matched and we reach a point where 'creation' again becomes the missing link. Though even there LLMs will play a major role as tools. For instance Einstein had to spend a significant amount of time in 'retrieval' rather than 'creation' research to develop the field equations for general relativity. If he had access to LLMs trained on all knowledge of the day, he could likely have achieved his goal much more quickly.

I mean, even if you buy the idea that all LLMs are really doing under the hood is insanely good interpolation, the results produced by that interpolation are still novel and still get incorporated into the knowledge corpus of the next training run. I guess the implicit question there becomes whether that expansion allows the knowledge corpus to continuously grow or whether it eventually settles into a steady state.