← Back to context

Comment by psvv

1 day ago

Don't frontier models still have trouble counting letters? Or am I out of date? Either way, it doesn't seem to be the same amazing rate of progress we're seeing in other areas.

It's easy to look at a fire burning through a forest and extrapolate that rate of progress across the whole world. But fire doesn't burn everything equally fast.

What other cognitive tasks will be a struggle to make progress on? I suspect there will be some, though which ones they are is anyone's guess.

Your information is out of date by several years. The letter counting issue was due to how LLMs split text input into tokens (usually using BPE). Since 2024, frontier LLMs have used chain-of-thought reasoning to spell out the letters and count them.

  • > The letter counting issue was due to how LLMs split text input into tokens

    No - this is provably not the issue.

    Take any model that fails to correctly count the letters in a word, and ask it instead to spell the word (even a made up word), and it will be successful - they have no problem predicting the letter sequence from the token sequence (and would be shocking if they did - this is what they are built for: seq -> seq prediction).

    The reason LLMs can fail at the letter counting task (depending on model training, prompting) is because of the counting part, not because of any difficulty correctly mapping the input token sequence to the letter sequence.

  • My apologies, I got my info from an LLM. I guess they still have a ways to go in understanding current events.

    • Which LLM specifically? If it’s cloud based, you should be able to share the chat, right?

      But seriously, I am still waiting for someone to wager that AI won’t be able to do a specific cognitive task in the next 5 years. This fact should be evidence enough that we have no idea how far AI capabilities will continue to advance.

      7 replies →

  • The letter counting issue is due to tokenization. And most models still get this wrong often enough, even with reasoning. Probably less so on strawberry given how prevalent it is, and less so than without reasoning, but this not a historical issue. It’s becoming less of one though.

  • They still have issues with problems like this actually, and I use all the frontier models from all the major labs, so it's not solved.

    • I'd love to see some examples of frontier models getting letter counting wrong. Can you share some?

  • > out of date by several years

    This is delirious exaggeration. The problem has not even been widely recognized for several years. Fable reported "two rs in raspberry" to me as recently as August. There is some randomness, it's hard to predict which words will trip up the machine, and I haven't been able to do it at all since August. But it was absolutely happening until very recently, and probably still is.

  • But doesn't that just amount to labs intervening to teach the models to use a particular strategy to mask this one very obvious marker of the difference between their intelligence and biological intelligence? (And similar surface issues like using tool calls / reasoning for arithmetic, even though humans writing on the internet don't typically break show their work for multiplying two numbers)

    The deeper architectural difference is still there, which manifests whenever you try to get the models to apply known techniques to modalities and problems outside their training data.

    • Chain-of-thought reasoning was added for general purposes, not to fix letter counting specifically. It just happens to solve that problem in addition to many others.

      You're in the discussion section of a post about OpenAI releasing hundreds of novel mathematical proofs, and you're claiming that AIs can't apply known techniques to modalities & problems outside their training data? I'm not sure what else would convince you.

      1 reply →

  • Also think it depends on language, literally asked 2min ago from chatGPT (no login so maybe it's a shittier model?)

    > Hur många 'r' I abborre, använd inte web search? Det finns 3 r i abborre.

    And I explicitly had to say not to search the web, because that's what it did by default, to count letters in a word...

    • The free models for ChatGPT, especially without login, do very little reasoning. You should at least log in to set any level of reasoning above Instant, which uses virtually none.

      1 reply →

I was trialing MiMo-V2.6-Pro recently due to its high benchmark scores, and it argued that substring matching the names of audio codecs in a search field was a mistake because "a user searching for 'aac' would get unwanted results for 'alac'." Which isn't exactly counting letters per se, but there are still weird issues with understanding words as strings rather than as tokens.