Comment by somenameforme
10 hours ago
Again I think the example of math is good. Many isolated tribes still don't even have numbers. They simply refer to things in broad quantifiers like - none, one, few, some, many. And that's perfectly fine for their needs! Many of the problems that you need math to solve - or that lead naturally to math, like currency, only exist once you've already discovered mathematics.
So try putting yourself in this ancient mindset before mathematics. How did somebody invent it, come up with the concept of numbering everything, further develop the various 'tricks' for manipulating these numbers, and so on? In terms of raw 'complexity' it's far less impressive than the latest LLM models solving some obscure mathematics problem that almost nobody understands.
But in terms 'intelligence', I find it vastly more impressive - because it's again this sort of difficult to describe concept of going from nothing to something. There is no logical baseline that naturally and cleanly leads to math. Almost like a child would say when asked how they learned something, 'Oh I just thought it up.' Except in this case, somebody genuinely did!
If we trained a LLM on such texts that only use "none, one, few, some, many" in their language, wouldn't it likely learn representations of individual quantities and arithmetics anyway?
Provided the training data was extensive enough and training rewarded solving problems that require mathematics.
Interesting question. I don't think so. In spite of what they're achieving right now, LLMs remain token prediction algorithms. We're speaking of going from a world where the concept of 'multiply' simply didn't exist to it being invented and formulated.
I also don't think the people behind the LLM companies think this is the case either. If it were then it'd make so much more sense to drop the current regime and instead move to the most basic systems trained on nothing but the most fundamental first principles and have them try to derive everything from there. It'd ostensibly lead to far more reliable systems with little to nothing in the way of bias. It'd also likely be vastly cheaper than the current practice of trying to train on essentially all consumable knowledge.