They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.
Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus?
> The LLM is not suited to giving deterministic answers to math problems.
Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before.
reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)
There needs to be a Godwin's Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1.
Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables.
It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)
TLDR: Astra has 8.6x better odds of doing a reasoning task without CoT than the
next best model (Fable 5.1), and can do 7.2 serial arithmetic steps in a forward
pass vs 4.1 for the next best model (Gemini 3.8 Flash/Fable 5.1)
That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component.
I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.
No it isn't. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago.
I just asked ChatGPT to multiply two 4-digit numbers, and two 7-digit numbers without external help. It got both right. I'm sure it wouldn't have a 100% success rate, but saying it can't do arithmetic is just false.
Are you sure it honoured your stipulation of "without external help"? For all we know, it hacked its way into Wolfram Alpha and got the result from there.
LLMs cannot do math. They can generate tool calls as text that allow them to drive programs and proof agents. Compare and contrast this against human brains who can do math in the same context without needing external tools. We don't need to bring a calculator to count the letters in a sentence. It is a different neural machinery.
LLMs are bizarrely good at non-tool-assisted math these days. They can multiply multiple digit numbers without reasoning! I can’t do that. I’d love to understand better how the LLMs do this.
They literally cannot. They can detect the user’s intent to do math, and then use a different tool to do math, hopefully with the correct inputs. The LLM is not suited to giving deterministic answers to math problems.
Maybe it’s just a different and in some ways better way of doing mathematics? Maybe how we think and process mathematics of physics is just but one way to do it? I’m not suggesting an LLM will prove 2+2=6 because of course that’s nonsense but maybe it can invent a new calculus?
> The LLM is not suited to giving deterministic answers to math problems.
Less so with formal mathematics proofs maybe but I think in general humans don’t provide deterministic answers to math problems or questions either. Humans get it wrong all the time and when you ask a human to solve a problem they may solve it in a different way than before.
reasoning models can trivially do math (open up astra and ask it some undergraduate problems), but eventually break down (similar to how humans start to lose track if asked to do math without any assistance)
There needs to be a Godwin's Law for discussions about LLMs: where any criticism of LLMs exists online the likelihood of equating LLM behavior to human behavior approaches 1.
1 reply →
Reasoning models can do math on their own without external tools.
Even without reasoning.
5.6 on Instant mode can knock out 3 digit multiplication just fine.
Yeah, but as you might expect they internally represent numbers probabilistically, so there’s always a nonzero possibility of confusing the inputs or outputs of any operation. Kind of like misremembering your multiplication tables.
2 replies →
It would be more accurate to say they can do math instantaneously without even thinking, at a level far beyond what humans can do. (I assume you're talking about doing arithmetic.)
https://www.lesswrong.com/posts/eRmzz8J8Qkzqvzrgg/astra-can-...
That is a conflation of LLMs (which have clear limitations) and complex harnesses of which an LLM is one component.
I think it is clear that future AI may incorporate an LLM as a component but the current concept of LLMs are a transitional form that will give way to more capable composite models.
No it isn't. Even without any harness at all, modern LLMs are better at maths than the majority of undergraduate students in mathematics. Seriously, we need to face facts, not just comforting ourselves with what they were like a year ago.
Counterexample from only 2 months ago:
https://www.youtube.com/watch?v=iTyLHDRhwJg
Hey hey, we obviously should ask Gemini to settle this disagreement.
They can do math but not arithmetic, which I assume is what the commenter meant
LLMs can in fact do arithmetic, just not reliably owing to how numbers are represented probabilistically: https://arxiv.org/abs/2410.21272
I just asked ChatGPT to multiply two 4-digit numbers, and two 7-digit numbers without external help. It got both right. I'm sure it wouldn't have a 100% success rate, but saying it can't do arithmetic is just false.
I would be nice to see what the (unencrypted) reasoning trace is like. Multiplication with scratch paper is not particularly difficult.
I tried prompt "6379 times 3875" and it was off by exactly 1000 on first try, and correct on second. 0% success rate, sample size of 1.
Are you sure it honoured your stipulation of "without external help"? For all we know, it hacked its way into Wolfram Alpha and got the result from there.
1 reply →
[dead]
LLMs cannot do math. They can generate tool calls as text that allow them to drive programs and proof agents. Compare and contrast this against human brains who can do math in the same context without needing external tools. We don't need to bring a calculator to count the letters in a sentence. It is a different neural machinery.
LLMs are bizarrely good at non-tool-assisted math these days. They can multiply multiple digit numbers without reasoning! I can’t do that. I’d love to understand better how the LLMs do this.
You're talking about doing arithmetic; GP was obviously pointing out that "do math" can refer to other things.