Comment by adamddev1
3 hours ago
Yes, exactly. And this is why I don't think that the LLMs can make significant process beyond what humans have done and published.
"But the math proofs," people will say. A lot of those seem to be spam-solving things with a huge swath of existing lemmas, and a some of these are being debunked and retracted.
Just today I was quizzing ChatGPT about a basic grammar question for a language that has huge training data but for which the grammar was not well documented. It kept giving me confidently wrong answers until I drilled and drilled it and then finally it found/gave back an explanation that perfectly fit a pattern given in one particular grammar, citing that as a source. It doesn't appear to have been able to figure out the inner structure on it's own. It appears only able to pattern match and put things together from what humans have already discovered and written.
Right, there's a difference between statistical interpolation and semantic induction. The whole point is that LLMs can't reason from first principles to drive missing rules. It keeps confidently feeding you approximations until it collides with some source that already mapped it.
Interesting take, given that the whole reason LLMs are interesting is that they're the first system we have that can work in semantic space. Statistical interpolation, that we've solved long ago.