← Back to context

Comment by ggreer

1 day ago

Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.

> The letter counting issue was due to how LLMs split text input into tokens

No - this is provably not the issue.

Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).

The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.

The letter counting issue is due to tokenization. And most models still get this wrong often enough, even with reasoning. Probably less so on strawberry given how prevalent it is, and less so than without reasoning, but this not a historical issue. It’s becoming less of one though.

My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.

  • Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?

    But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.

    • I asked both gemini and chatgpt "do frontier models still have trouble counting letters?" and the first word of both responses was yes.

      The fact that no one is taking you up on that bet I don't find to be particularly persuasive. I suspect there will be plenty of cognitive tasks LLMs struggle with in 5 years, maybe even 20. But I wouldn't hazard to guess which, I don't think anyone is capable of that level of foresight.

      5 replies →

    • Solve the Collatz conjecture in the next five years? If humans publish significant advances during that time, and A.I. copies it, then yes. Otherwise, I'd definitely bet money it won't happen. I'll give you 10,000 brownie points if I'm wrong.

They still have issues with problems like this actually, and I use all the frontier models from all the major labs, so it's not solved.

  • I'd love to see some examples of frontier models getting letter counting wrong. Can you share some?

> out of date by several years

This is delirious exaggeration. The problem has not even been widely recognized for several years. Fable reported "two rs in raspberry" to me as recently as August. There is some randomness, it's hard to predict which words will trip up the machine, and I haven't been able to do it at all since August. But it was absolutely happening until very recently, and probably still is.

But doesn't that just amount to labs intervening to teach the models to use a particular strategy to mask this one very obvious marker of the difference between their intelligence and biological intelligence? (And similar surface issues like using tool calls / reasoning for arithmetic, even though humans writing on the internet don't typically break show their work for multiplying two numbers)

The deeper architectural difference is still there, which manifests whenever you try to get the models to apply known techniques to modalities and problems outside their training data.

  • Chain-of-thought reasoning was added for general purposes, not to fix letter counting specifically. It just happens to solve that problem in addition to many others.

    You're in the discussion section of a post about OpenAI releasing hundreds of novel mathematical proofs, and you're claiming that AIs can't apply known techniques to modalities & problems outside their training data? I'm not sure what else would convince you.

    • LLMs are very useful, I use them every day as a software engineer to solve problems and search for information represented within the data available to them. But they are a specific type of intelligence, with many advantages and disadvantages vs human intelligence and it's not clear that just scaling or tweaking them without a theoretical, architectural change will make them more generally intelligent than humans (despite all US AI companies promising exactly that).

      They are fundamentally based in language, and achieving deeper models of the world through language alone is deeply inefficient compared to the way humans model the world for years without any language at all. They do not learn at inference time. They don't have semantic understanding of the difference between their own output and other sources. etc etc.

      That depth is the key for me. Of course they are capable of producing novel sentences that aren't in their training data, but the depth of that novelty is basically within the bounds of language itself. They are capable of more serious depth and more abstract reasoning than that, but I have experienced limits, which it then tries to surpass with tools to convert things it can't understand back into language (unit tests, LEAN) upon which it is trained.

      Because I'm not an AI booster, my account is limited to 5 comments a day. So this is the last reply I'll be able to make today, if you want to continue the conversation we'll have to wait for tomorrow.

Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)

> Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.

And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...

  • The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.

    • There is no such thing as reasoning in models. Any "reasoning" is invented afterwards.