Comment by antman
8 hours ago
This argument implicitly makes a few assumptions which will probably not hold in the very near future.
One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.
The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse
Putting hallucination aside, LLM "theory of mind" has gotten worse over time. I feel it peaked in Opus 3 and Sonnet 3.5, GPT 4 and then GPT 4.5 for OpenAI. Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. OpenAI models suffer from this much more than Anthropic models but Claudes have backslid over time too.
This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data.
Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too.
Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost.
Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time.
The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.
> Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally.
This is a feature, and a huge step forward.
If you expect AI to do serious work, you can’t have it guessing what you “really meant”. Every sufficiently advanced task depends on very subtle details in the problem statement, and the correct default behavior for advanced AIs is to solve the task exactly as stated, unless a system prompt or other constraint tells it to do otherwise.
> One is that AI will continue hallucinating in a manner that is not easy to verify
It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.
Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.
> can reduce the risks inherent to LLMs, to a really remarkable degree
To an arbitrary degree.
Just like all of science. Reduce the error to the desired margin.
And why not?
For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.
Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.
Why would this not also be the case for mathematics?
I think OP is saying that hallucination or not is just semantics. There is nothing qualitatively different about hallucinated vs non-hallucinated output.
3 replies →
> It is an old saw at this point
An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.
> assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify
Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.
At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.
I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.
Hallucinations are no longer much of a practical problem in software engineering.
Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.
We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.
Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.
Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
It’s still a common occurrence.
It happens in more subtle ways, but it still happens often enough for me to notice. For example I have had hallucinated checksums show up in lock files as recently as yesterday using a SOTA model.
This is not surprising, since the whole basis of LLM training is to produce output that humans will accept _as a proxy for actual training goals_. In a sense, the training process of an LLM “wants” to produce output that is statistically plausible much more than it “wants” to produce correct output. It’s always going to be a struggle to drive that system towards other goals (and we see this bourne out in practice by the amount of effort that is required to be spent on RL).
I think there will be some threshold of correctness (something like 99.999% of the time) that if the model surpasses it, I can stop needing to check it, but I think we’re still at 99% or something which sounds good, but when you are producing a ton of output you hit that 1% frequently.
> Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.
I 100% agree with this. In fact explaining things to humans is something LLMs are particularly well suited for.
The confidence with which you, anonymous user, keep commenting that "hallucination is not much of a practical problem in software engineering anymore" based solely on your own anecdotal evidence is really remarkable, in not a good way.
4 replies →
"Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.
> but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds
You’re making the following assumptions:
1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs)
2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.
> 2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.
Are they independently wealthy? Or do they have a deal with their local supermarket that they can take food for free?
[delayed]
Even if you somehow have a 100% correct AI, it's not useful unless we can understand and internalise (and communicate) its results.
Who is this "we" you speak of? The professional mathematician community? Were pre-AI results useful outside of this community of people who could understand them?
Yes. Most probably do not understand the notation involved in, and the statement of, the Lindeberg-Levy Central Limit Theorem. But every scientist uses this theorem in one way or another. These ideas have a way of trickling down because to people who work thanklessly to do so.
Something about this sentence makes me think about that Rob Auton bit, that before there were mobile phones, nobody had any reason to tell someone else that they were on a bus.
1 reply →
> One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that.
There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).
No evidence other than the fact that this has been happening steadily in all areas for many years?
You might have a point if the goal was to have LLMs that spit out a correct proof without chain of thought or tool use. LLMs + agent harnesses are more than capable of self verification and course correction.
Is a human deterministically robust? Or is a human also incapable of doing what you claim LLMs will never be able to do?
The direction and pace of capability improvement has already been demonstrated by all models. The latest breakthroughs make that pretty evident, but there have been production systems that are based on probability since the beginning of computing.
What has been demonstrated is a process that outputs lean proofs based on those probabilities. This happened after decades markov chain producing garbled texts and very shortly after gpt2 producing stories about unicorns.
> The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).
An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.
I think people on here tend to somewhat fixate on the determinism issue. Even with a deterministic LLM - stabilising the floating point arithmetic, and choosing from the distribution by a fixed method, or just save the random seeds - there is still a kind of a chaotic unpredictability that can exist between its inputs and outputs. However maybe that is a price that needs to be paid to get creativity.
Deterministic yes, robust deterministic no. The most likely conjunction is not always the best nor representative of what the model is considering unless its certainty is high.
1 reply →
AI cannot explain chess moves it comes up with in an elegant way. What makes you think it will be able to do so for math?
Interesting claim. That's true for old models that simply have no way to explain, LLMs however can. [1]
[1] https://dev.to/natcher/researchers-develop-method-to-train-l...
It's kind of astonishing that after all we have seen in the last years people still find the position that AI will not be able to do an obviously valuable thing likely and it requiring an explanation (instead of the other way around).
No, this is different, and this is coming from someone who has been studying deep learning for the last decade. We are talking about the difference between RLHF and RLVR strategies. The former benefits clarity and explanation, while the latter concerns only correctness. AI was moving in a particularly damaging direction by pushing on the first path, so it was natural to move to the second. But the second will come at the cost of clarity of explanation. It will likely get better at its explanations, but not fast enough to render its most advanced accomplishments readily understandable to the user. The chess example is a pretty good one (that is an RLVR approach).
2 replies →
Some of us still remember 2016, when we had a couple of cars sorta half-driving themselves, and Tesla, Uber and others promised we were only one year or two away from three million people in the US working as drivers being out of a job. And here we are, a decade later. AI is pretty amazing, but companies have a tendency to severely and comically overestimate and oversell it's capabilities, and underestimate the challenges.
6 replies →
[dead]
Expression of tech-faith is not intellectually honest argument.
Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.
Markov chains garbled text to gpt took decades, gpt stories about unicorns to gpt production systems took a few years, gpt production system to gpt astra producing deterministic lean proofs of longstanding mathematical problems happened even faster. Scepticism to the point of requiring proof appears like an academic pursuit while production systems have already been built and are in the process of being enhanced
Did you even RTFA? His argument absolutely doesn't make any assumptions about hallucinations, implicit or not. It's you who assumes Tao must have surely been complaining about hallucinations or some such. You've not addressed any of his arguments and moreover ask questions his post answers.
Here's a longer article which goes into a bit more of the details: https://terrytao.wordpress.com/2026/10/05/the-future-of-math...