Comment by ggreer
1 day ago
Unless you think that AI will quickly hit a wall (which seems odd considering that only a few years ago the best models had trouble doing basic math or counting the number of Rs in "strawberry"), I don't see how that's reassuring. The models will only get cheaper and more capable over time. It seems quite likely that at some point (probably before I hit retirement age) they'll be able to fully replace me at my job.
Is there any specific cognitive task that you are willing to bet that AIs won't be able to accomplish in the next 5 years? Because if not, I'm not sure we're disagreeing about predictions.
Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.
It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.
What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.
Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.
> The letter counting issue was due to how LLMs split text input into tokens
No - this is provably not the issue.
Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).
The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.
The letter counting issue is due to tokenization. And most models still get this wrong often enough, even with reasoning. Probably less so on strawberry given how prevalent it is, and less so than without reasoning, but this not a historical issue. It’s becoming less of one though.
My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.
9 replies →
They still have issues with problems like this actually, and I use all the frontier models from all the major labs, so it's not solved.
1 reply →
> out of date by several years
This is delirious exaggeration. The problem has not even been widely recognized for several years. Fable reported "two rs in raspberry" to me as recently as August. There is some randomness, it's hard to predict which words will trip up the machine, and I haven't been able to do it at all since August. But it was absolutely happening until very recently, and probably still is.
But doesn't that just amount to labs intervening to teach the models to use a particular strategy to mask this one very obvious marker of the difference between their intelligence and biological intelligence? (And similar surface issues like using tool calls / reasoning for arithmetic, even though humans writing on the internet don't typically break show their work for multiplying two numbers)
The deeper architectural difference is still there, which manifests whenever you try to get the models to apply known techniques to modalities and problems outside their training data.
2 replies →
Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)
> Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.
And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...
2 replies →
I was trialing MiMo-V2.6-Pro recently due to its high benchmark scores, and it argued that substring matching the names of audio codecs in a search field was a mistake because "a user searching for 'aac' would get unwanted results for 'alac'." Which isn't exactly counting letters per se, but there are still weird issues with understanding words as strings rather than as tokens.
LLMs already hit a wall. Now it 80% of marketing hype and 20% of retooling and benchmaxing.