← Back to context

Comment by VCFundedGenYer

1 hour ago

LLMs still can't do math nor count letters in words. Nothing has changed there.

This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model.

But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.

Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: https://www.google.com/search?&q=karney+formula+geodetic+

reference: https://github.com/pbrod/karney

You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.

Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.

"LLMs can't do math" is a pretty hot take in September 2026.

  • They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.

    • reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)

    • It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)

        TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
        next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
        pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
      

      https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...

  • That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component.

    I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.

    • No it isn't. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago.

      1 reply →

  • They can do math but not arithmetic, which I assume is what the commenter meant