← Back to context

Comment by walrus01

2 hours ago

This is true but a sufficiently smart LLM (run in a harness like opencode, no special MCP, no customization done whatsoever) will quickly turn out a basic 1 to 2 page sized python script to do the math. They can't do the math with any guarantee of accuracy with their own internal reasoning since it's a language model.

But, for example, if you ask deepseek v4 flash 0731 to produce a python script to calculate the distance or azimuth directions between two points on an oblate spheroid using the vincenty and haversine geodetic formulas, it'll turn out the factually accurate vincenty and haversine formulas which has a perfect 100% correlation with what is hard coded into human-written GIS software. These things are clearly in its training data set from whatever whole-internet-crawl/scrape built the training set.

Heck, just for fun I asked a reasonably smart LLM to re-implement the Karney formula (which is considerably more complex than Vincenty), just in case I ever had a need to calculate the distance between two points down to the nanometer, and it did it: https://www.google.com/search?&q=karney+formula+geodetic+

reference: https://github.com/pbrod/karney

You still have to be skeptical of its results and capable of understanding if it's gone off on a hallucinatory path, but saying LLMs can't do math isn't really a hundred percent accurate anymore. More precisely it's that they can't do the math internally but they're quite capable of producing the tool that does the math. And often producing a basic one-off tool that does the math takes less than a few seconds, then it runs it, and will spit back the results.

Deepseek v4 flash 0731 (a somewhat randomly chosen example) isn't even particularly sophisticated, large, or capable compared to a GLM5.3 size model or Kimi K3 size thing.

You know what also works to get the Karney formula into a program? You can download Charles Karney's free software (MIT license) implementation in several [1] programming languages and then just make a library call – the API is straightforward. If you have comments or questions you can read his several clearly written papers describing the problem, its history, and his algorithm, or you can directly email him: he's a very nice guy, and pretty responsive.

[1] https://geographiclib.sourceforge.io/doc/library.html#langua...

  • Right, it was really more as a test of how much was contained in the training data set. For my purposes Vincenty is quite accurate enough. This isn't for millimeter level precision land surveying or measurements, but for distance in meters between microwave or millimeter wave band radio sites, point to point links. Even a distance difference of 4 meters plus or minus on a 12 km, 11 GHz band link is going to have no appreciable difference on link budget/reliability calculations, it can be that crude. But not so crude that I just want to throw Haversine at it when Vincenty exists and is not computationally expensive.

    As this was for a test of "what happens if..." I also watched to see if it did any web searches or external data retrieval to build the test script, and it didn't.

    I intentionally didn't give the LLM a direct copy of the software or a link to it, to see what it would do. In my case it was a randomly chosen example I could come up with in 10 seconds of imagination to see "hey what if I ask it to do this...". It also implemented a perfectly usable parabolic millimeter wave antenna gain efficiency calculator based on variable surface smoothness parameters, which is a lot more basic math.

    • As an aside: I'm quite convinced that an extremely precise version can be implemented that is significantly faster than Karney's, roughly comparable in speed to simpler naïve approximations. But for most purposes where the precision matters Karney's implementation is not any kind of bottleneck, so it's not clear it's worth spending significant effort on trying to do better.

      Maybe that's something one of the big LLM companies might want to throw their machines at optimizing if they need to do a lot of geographical calculations.

      1 reply →

This just exposes that they don't even do the thing you said.

Not only is it still true that they can't do math directly, but not even indirectly.

They didn't write a python script to do the math, they found bits of code that are associated with "math" and the supplied arguments.

Someone else already wrote that code and someone else categorized it so that it could be associated with the kinds of problems it applies to.

That isn't an example of idiot at one thing while good at another thing, or solving the same problem just a different way or indirectly. It's being the same idiot at all times. If an actual non idiot thinker didn't write code in the problem domain, and some non idiot thinker didn't tag it as being relevant to that domain, then it wouldn't happen.

It's nothing more than an sql query.

  • I don't know for you, but it would take me more than 30s to find and translate the open source code implementing the formulae/algo into small usable program. The more hesoteric the optimisation in the original code, the more time I need.

    So maybe it is more of a smart completion engine than a SQL answer.

  • > they found bits of code that are associated with "math" and the supplied arguments

    How is this different from a human using an algorithm they have memorized, or reading it from a reference site written by a human and then writing the same formula into a custom one off piece of python code?

    I could have gone and spent a couple of days teaching myself the math behind Karney and reading its reference implementation (very possibly just copy/pasting big chunks of it to save time) and writing a wrapper around it. It would have produced the same result.