Comment by fasterik

1 hour ago

"LLMs can't do math" is a pretty hot take in September 2026.

They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.

  • reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)

  • Reasoning models can do math on their own without external tools.

    • Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables.

  • It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)

      TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
      next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
      pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
    

    https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...

That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component.

I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.

  • No it isn't. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago.

They can do math but not arithmetic, which I assume is what the commenter meant