Comment by amelius
20 hours ago
Nobody knows how it works, really. It just turned out that if you try to predict the next word then you get intelligent behavior, depending on amount of training data, and the size and topology of the network. But again, nobody knows why, and what the limits are.
Agreed. We went this direction for our golems, djinns, and other mechanistic minds because we believe it sort of reflects the primitives of our own neurons (which we also don't fully grok).
Linus Torvalds:
~"Predicting the next token is not an insult. It's pretty much what we all do."
That's of course absurd. The human brain doesn't represent information as discrete tokens, nor does it form sentences autoregressively.
> The human brain doesn't represent information as discrete tokens, nor does it form sentences autoregressively.
I don't know about you, but I tend to speak one word at a time...
2 replies →
Our brain decides every moment what action to take next, that best fits with what we did and felt so far. Language is kind hard to see because it's virtual, but imagine all other things you do. And then language works the same as everything else, same circuits, just no direct connection to muscles.
I heard someone who studies this sort of thing say basically what biological neurons are trying to do is predict as well. Predicting what exactly? I’m not sure. The next time they should fire or something. I can’t find the YouTube video now.
In the last few decades, there has been an increased interest in the role of prediction in language comprehension. The idea that people predict (i.e., context-based pre-activation of upcoming linguistic input) was deemed controversial at first. However, present-day theories of language comprehension have embraced linguistic prediction as the main reason why language processing tends to be so effortless, accurate, and efficient.
https://www.earth.com/news/our-brains-are-constantly-working...
https://www.psycholinguistics.com/gerry_altmann/research/pap...
https://www.tandfonline.com/doi/pdf/10.1080/23273798.2020.18...
https://onlinelibrary.wiley.com/doi/10.1111/j.1551-6709.2009...
Interesting.
And I also found the video I was referring to https://www.youtube.com/watch?v=FHQfmJEpRmU
Predict activations in adjacent neurons, roughly; the relevant keyword is "Hebbian learning" (and "predictive coding" at a higher level).
Also called surprisal minimization in Karl Fristons free energy principle.
Predicting reality, under the "controlled hallucination" framing, corrected by sensory error signals. The brain has no access to ground truth, only input data that helps correct the hallucination.
Based on lots of human interactions, I think there are a lot of human beings out there who mentally aren’t much more than “next word predictors” who happen to be made of meat+neurons instead of silicon+code.
That’s my and probably most people’s understanding.
I have a feeling we know more than that about how it works.
We don't know how our own intelligence works, much less anything else's
We built LLMs.
We didn't build our brain.
Typically when you build something you have a decent idea how it works.
2 replies →