← Back to context

Comment by shubhamjain

9 hours ago

A very balanced perspective, and the concerns he raises are reasonable. He acknowledges that AI is going to transform mathematics, but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.

There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics. But what else is to be expected? It's become a maniacal race with too much money. Too much effort is being invested in proving that the exponential curve is still holding.

I wonder if top labs will soon abandon math progress like they did go and chess.

In example of go where I'm more familiar Google deep mind poured large resources to get a super human performance first, establish superiority and abandon it. The community then built their own tools starting from reproducing their papers.

I think similar thing might happen to math. Nobody outside of math cares too much about Hamiltonian cycles in some bizarre graphs or proving lower bounds on complexity of some problem.

Once those results stop being worthy of mainstream media attention, they will abandon math and the progress will be done by mathemicians guiding the models and the community will likely establish some new rules about what makes a valuable contribution. Merely solving not yet solved problem might not be it anymore.

  • The amount of money they're lately ploughing into proving math theorems is inconsistent with how societies and markets have priced pure mathematics. The entire US federal budget for math research is something like $100M annually. A single college football coach can already earn 10 percent of that.

    Pretty much the only enterprise that historically pays some mathematicians handsomely is quant finance, but those people are actually compensated not for proving theorems but rather for statistical modeling and programming skills. And even that industry is so technologically driven these days that pure research mathematicians no longer hold a clear edge over strong programmers with undergrad level probability and statistics at their fingertips.

    • Math is one of the most verifiable domains, esp thanks to LEAN, which also build coding skills.

      The $$$ they're pouring isn't just for marketing. Think of these papers/results more as "useful side effects" from large-scale RL rollouts and post-training. Every token being generated contributes to post-training in some way.

      There isn't a hard boundary between "training" or "inference", modern post-training is arguably inference-bound :)

      3 replies →

    • As long as they continue making headlines they will continue spending. This is just marketing at this point.

    •   Don't hire a straight-A student, unless it's to take exams; or a professor, unless it's to write papers.
       -- Nassim Taleb
      

      How interesting that Anthropic and OpenAI are full of professors and straight-A students!

      4 replies →

  • I don't think it will be the case, maths have real utility. I found something interesting at the intersection of combinatorics and information geometry. To be quite frank I don't understand what I'm doing. And yet, when I ask ChatGPT to use the framework we're developing to write an algorithm, it turns out it has quasi-parity with the state of the art. I have to measure absolute perfs to decide which one is better – theirs, not mine. Ok. Time to keep improving on what I have. And this implies dropping the code and going back to the blackboard doing more super abstract math that are way out of my league.

    • Good luck!

      We are at the point where the way in which humans do math and science changes significantly, and I have no good idea at all in what state is it going to settle down. But you are one of (many, I suppose) people exploring the new wilderness, so I wish you best.

      1 reply →

    • It has utility so some people will pursue it, but it has no immediate business value so I don't believe ai labs will keep spending millions on it.

      Unless they decide that trying p!=np is worth any money.

      5 replies →

  • I think this is an interesting and good theory. They've probably eked out the large majority of the PR benefit at this point, so whether they continue in this vein will tell us a lot about their motivations for this work.

    To take this to the next step, what happened after deep mind pretty much solved Go is that they started looking for the next set of things that hadn't been done yet. It does strike me as very likely that this will follow that same path.

  • Interesting path forwards, and probably partly true, but there are some important distinctions:

    Go was a specialized application. All the math results come as a side effect of reading the whole internet, and it will keep reading the whole internet. It will keep practicing thinking questions. Actually, math might be one of the best ways to keep them contemplating and measure their contemplation abilities, so math will always stay in the loop.

    Also, math might not be useful just for humanity, but also for AI, so the system might actively benefit from new math results itself. (Not sure if any of the recent proofs qualify, but future work might.)

  • > I wonder if top labs will soon abandon math progress like they did go and chess.

    I definitely think that this is marketing, just "with good side effects". My doubt is when they will be able to move to "marketing with better side effects", that is, research with more concrete outcomes (health, materials etc.).

    Problem is, that type of research is much harder. Some doubt that progress in such areas will be quick (https://www.noahpinion.blog/p/wheres-the-intelligence-explos...).

    • They aren't trying to 'solve' chess, go, or mathematical proofs as an end in themselves, but mainly in order to learn more about how to build better systems overall. The goal of AlphaZero was ultimately as a stepping stone towards AGI, and it's the same with LLMs.

      2 replies →

    • Would that be a bad thing?

      While top AI labs no longer focus on chess, the community build way better chess engines.

      Stockfish is probably stronger, than everything the top labs build.

      Wouldn't we expect the same thing for math? That slowly the broader math community would engineer a harness/program... That will surpass the current labs, and be a community ran project

      2 replies →

  • This misses the raw advantage of a good proof. It makes conceptualization simpler. In some ways math is like a hash list of of theorems. This list makes it simpler to prove other calculations, and will always be useful, to both humans and AI models. I can see two new directions 1 - the creation of specialist theorem models; that can answer questions efficiently about one topic and 2 - we probably need to incentivize and codify ownership of theorems; charging a proportion of the compute saved by using them. Ultimately enabling mathematicians to be paid our true market value!

    • Oh boy, please not 2. What if this was a thing already and, since neither Newton nor Liebnitz had kids, we all had to pay some investors who bought the rights to calculus every time we took a derivative.

      3 replies →

    • Considering how essential math and science is for the prosperity of mankind (not even speaking about the cultural value) the question of how to reward people working and contributing in these fields effectively and appropriately is of extreme importance. (And I think the current decline in our societies is to no small degree caused also by our utter failure to address that issue.)

      It is also fascinating, because I don't think there is any solution within our existing system, at least not any I know of. Theorem ownership is not a good solution (and neither are patents in general). Probably the most achievable (or rather the least unachievable) solution is a kind of communist utopia, where people can dedicate their time to a pursuit of any endeavor they see fit, as resources for a decent life are abundant and excessive power capture impossible. (The other option, somewhat dystopian, and which would not require humanity to change too much in its current mode of conduct, would be a totalitarian or caste-like capture of society by the scientific community.)

      Incidentally, if AI proves as powerful as some expect it to become, it could bring about another solution of that issue by making all human science and mathematics obsolete, pushing its true market value to zero.

      (With apologies for rambling.)

    • > It makes conceptualization simpler

      I wonder if it makes conceptualization simpler for models too, given that they're trained already on human-speak. And I'm also curious as to whether humans currently have an innate advantage into simplifying and contextualizing proofs, or will the machines get good at that as well?

  • I had the same idea recently. You've solved all the famous conjectures (all formulated by humans because humans found them interesting), what next? I doubt "AI formulated a math conjecture that nobody else cares about and immediately solved it" will produce that much hype. The actually interesting thing is indeed how mathematicians themselves will use these AI models going forward and how that will shape mathematics of the future.

  • AlphaGo and AlphaZero weren't generalized models. Math capability will presumably keep improving along with the other general capabilities, even if there wasn't a special RL focus for math itself.

  • Yeah chess is a good example. DeepMind came for publicity with AlphaZero. Arranged a match with Stockfish with rigged rules to make AlphaZero look better than it really was (it was amazing but the match wasn't fair) and then just published some games and went home.

    I was bitter about that back in the day as I hoped for more answers, more matches, more "truth" about chess being shown. Soon after that community project Leela Chess Zero was started and not only surpassed original AlphaZero but added few hundred ELO points over it. Then the combination of NN and classical engines happened with NNUE and current Stockfish is again a few hundred ELO points stronger.

    Today we pretty much know the truth in chess for all practical purposes. Human analysts/preparation experts focus on finding interesting path and opponent profiling (what is the most unpleasant for the opponent to face). They don't look for truth anymore. The game is doing great, it's more popular than it ever was.

  • I disagree. Firstly, people in AI likely care about math on a personal level. Secondly math is useful. Playing go or chess is basically a party trick. Being useful gives it staying power.

    But, I do think you are right that there will be some level of moving on. The spotlight is currently on maths and that won't last. It will move to some other area where there is more impact to be had. So while they might shift gears and put less focus on math, it will always be there as part of the portfolio.

  • I think there’s a venue where they start focusing on introducing hypotheses where the model currently can’t solve it, or maybe this is already happening?

    Being able to present useful novel ideas would likely generate a lot of press, for a while. I don’t know how this would look since I’m useless at math, but Im sure there are plenty of unknown problems with massive implications, that once formulated can be solved.

This argument implicitly makes a few assumptions which will probably not hold in the very near future.

One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that. What humans in the loop are doing now is verify the process, propose shortcuts and add legitimacy, through the verification process, if that ends up being succesful its highly likely a lot less mathematicians will be required in the future.

The conclusion that this is not productive focuses on the mathematicians, but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds. Unless it ends up being the greatest hallucination ever ofcourse

  • Putting hallucination aside, LLM "theory of mind" has gotten worse over time. I feel it peaked in Opus 3 and Sonnet 3.5, GPT 4 and then GPT 4.5 for OpenAI. Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally. OpenAI models suffer from this much more than Anthropic models but Claudes have backslid over time too.

    This means when writing documentation, tutorials or commit messages, their output is often a garbled jumble. Assuming shared context, using invented terminology without explaining, leaking conversational states due to improper epistemic boundaries and failing to model the reader. This all usually leads to their freely generated explanations being terrible. Getting good explanations requires chaining questions that force them to line things up properly, which is not easy the less you know. These failures as something LLMs naturally struggle with make sense, given the nature of attention and RL with weak signals from human data.

    Math is not merely a collection of proofs, it's a way of understanding. A proof presented in a manner that cannot be incorporated remains useless. It does not make it's way to physics like Riemannian geometry and matrix math did. This is no less true when done by humans too.

    Your hallucination conclusion, checking if a proof is one, is exactly the counterproductive cost.

    Most of us cannot verify that the claims in the OpenAI lore dump are in fact all correct. It will take tons of work from experts to do this. It took subject expert mathematicians to identify the discrepancy and disconnect in the Navier Stokes proofs, for example. LLMs will struggle to make use of their own proofs or turn them into knowledge that accumulates over time.

    The act of proving is often more valuable than the proof itself. Human constraints and limitations force us to invent tools and abstractions that a 100,000 x 1M context swarm can bypass. The tradeoff from that AI swarm advantage is work that doesn't usually lend itself to being built upon. It's like doing all the side quests and reading all the books of an RPG versus min maxing a straight path with a guide. We might try to identify new abstractions, but the fact that we don't get access to CoT and that much of it will be illegible means mining LLM traces for what human mathematicians produce naturally will be a tedious chore.

    • > Since then, even with Opus 5.5, phrasing has needed careful crafting, in order that it not be taken too literally.

      This is a feature, and a huge step forward.

      If you expect AI to do serious work, you can’t have it guessing what you “really meant”. Every sufficiently advanced task depends on very subtle details in the problem statement, and the correct default behavior for advanced AIs is to solve the task exactly as stated, unless a system prompt or other constraint tells it to do otherwise.

  • > One is that AI will continue hallucinating in a manner that is not easy to verify

    It is an old saw at this point, but what an LLM does still cannot be divided into hallucination and non-hallucination. This is literally an anthropomorphism trap.

    Layers and layers of application-specific verification can reduce the risks inherent to LLMs, to a really remarkable degree, but nothing about what these tools are suggests that this problem will go away; it will just bubble up again somewhere else.

    • > can reduce the risks inherent to LLMs, to a really remarkable degree

      To an arbitrary degree.

      Just like all of science. Reduce the error to the desired margin.

    • And why not?

      For all that I saw over the last few hundred hours with AI on software engineering, hallucinations are no longer a problem at all.

      Not once have I seen a task fail due to what would have been a "hallucination". If they still occur, they can apparently be detected and corrected automatically, or are subtle enough to escape notice with presumably no significant impact on the results.

      Why would this not also be the case for mathematics?

      4 replies →

    • > It is an old saw at this point

      An old saw unless something that's widely accepted, but sadly it seems that many people don't recognize this, even many people working in the field.

  • > assumptions which will probably not hold in the very near future [...] One is that AI will continue hallucinating in a manner that is not easy to verify

    Hold up, that's an even bigger assumption in the opposite direction, and I don't see anything to support it.

    At least in terms LLMs getting all the "AI" hype these days, there is no structural/mathematical reason to believe they won't continue to have the same problem they've always had of generating plausible text over rational text, and I don't think anybody even has a clear idea how it could eventually be accomplished.

    I've seen "then the magic singularity occurs and somehow it solves the problem for itself", but I would classify that more as mysticism than engineering.

    • Hallucinations are no longer much of a practical problem in software engineering.

      Two years ago, hallucinating that the code worked or that a task was accomplished was a common occurrence.

      We have seen that now agent swarms across thousands of agents can coordinate to achieve a result.

      Clearly hallucinations are no longer the problem they once were, since now we can get working results for long horizon tasks that require massive compute.

      Consequently it would seem unwise to assume that current limitations will remain as they are and prevent LLMs from coming up with solutions that they can explain to humans.

      6 replies →

    • "Plausible" text was preferred over rational text when we trained LLMs using RLHF. It's rapidly shifting the other way now with RLVR, which enforces correctness by default.

  • > but it is very productive in terms of hundreds of proofs being produced that had previously consumed uncountable hours of the brightest minds

    You’re making the following assumptions:

    1. the exercise of struggling to find proofs was not productive, but this is precisely how new techniques in math were produced. Brute forcing solutions doesn’t lend itself to the creation of much new mathematics (except maybe the exercise of developing verifiable proofs)

    2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

    • > 2. the point of doing mathematics is to be “productive” in the first place. This is silly. Many people get into mathematics because of the beauty of understanding, for example.

      Are they independently wealthy? Or do they have a deal with their local supermarket that they can take food for free?

      1 reply →

  • Even if you somehow have a 100% correct AI, it's not useful unless we can understand and internalise (and communicate) its results.

    • Who is this "we" you speak of? The professional mathematician community? Were pre-AI results useful outside of this community of people who could understand them?

      3 replies →

  • > One is that AI will continue hallucinating in a manner that is not easy to verify, second is that AI will not be enhanced to produced more simplified amd robust outputs, and third that a human will be required to do that.

    There is literally not a single shred of evidence to indicate either of your supposed eventualities. The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

    • No evidence other than the fact that this has been happening steadily in all areas for many years?

      You might have a point if the goal was to have LLMs that spit out a correct proof without chain of thought or tool use. LLMs + agent harnesses are more than capable of self verification and course correction.

    • Is a human deterministically robust? Or is a human also incapable of doing what you claim LLMs will never be able to do?

    • The direction and pace of capability improvement has already been demonstrated by all models. The latest breakthroughs make that pretty evident, but there have been production systems that are based on probability since the beginning of computing.

      What has been demonstrated is a process that outputs lean proofs based on those probabilities. This happened after decades markov chain producing garbled texts and very shortly after gpt2 producing stories about unicorns.

    • > The core technology of an LLM is sampling from a distribution so there is literally no way to make it deterministically robust (only probabilistically).

      An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.

      3 replies →

  • AI cannot explain chess moves it comes up with in an elegant way. What makes you think it will be able to do so for math?

  • Expression of tech-faith is not intellectually honest argument.

    Where does this "will probably not hold in the very near future" come from? People correctly warn about extrapolating current things onto the future, but then just throw some vague "probabilities" without providing any argument why their "probably" is somehow more grounded than others.

    • Markov chains garbled text to gpt took decades, gpt stories about unicorns to gpt production systems took a few years, gpt production system to gpt astra producing deterministic lean proofs of longstanding mathematical problems happened even faster. Scepticism to the point of requiring proof appears like an academic pursuit while production systems have already been built and are in the process of being enhanced

  • Did you even RTFA? His argument absolutely doesn't make any assumptions about hallucinations, implicit or not. It's you who assumes Tao must have surely been complaining about hallucinations or some such. You've not addressed any of his arguments and moreover ask questions his post answers.

    Here's a longer article which goes into a bit more of the details: https://terrytao.wordpress.com/2026/10/05/the-future-of-math...

My steak is too juicy, my lobster is too buttery, my industry shaking mathematical proofs are coming too quickly

In what world is OpenAI not “genuinely contributing to mathematics”?

I’m getting whiplash from the speed at which people are suddenly accusing them, and AI in general, of not doing enough.

  • In TFA I read (scroll up from the link anchor) this was explained: OpenAI are not giving talks - because they can't answer any questions about the model's work. There is little follow-up activity - the actual elaboration of human understanding of the new ideas is stifled since the problem is solved.

    But you are right, this is not OpenAI's "fault". The problem is - as others have said recently - that many people in mathematics want recognition for solving open questions more than they want the answers to the open questions. Everything about the economics and social environment of Mathematics will have to change.

    I think that this is exactly the same split we see in software: there are those who mainly enjoy the craft aspect of building software, and are uninterested in the product or business they are supporting. Others are primarily interested in the production of useful software or building a platform or company.

    I've always been in both camps myself. When it became obvious that AI was going to destroy the craft aspect - at least two years before it actually could do so - I became very discouraged, even depressed. But once it was actually good at building software, I became very excited about all the stuff I could now build. Sadly, I think a lot of people in our field have never had something they really wanted to build.

  • I'm in the same boat, I guess I had this naive idea that unsolved math problems would mean something if they were solved. But it seems like a lot of them at least were more thought experiments than anything else.

  • Perhaps mathematics was never about proving things? I know it sounds like moving the goal post, and it certainly was what motivated mathematicians on a day-to-day basis, but bear with me for a moment. I think that beyond being a creative activity that humans enjoy, math was about building new tools and systems of thought. Axiomizing things we take for granted, logic, linear algebra and calculus (on which modern LLMs rely so heavily) are such examples. People chose to participate in this field because they found it enjoyable and satisfying in some way. As a side effect, society reaped the benefits every couple of centuries. Will harvesting open problems with AI ever give us these things, or will we just be left laundry pile of Lean formalizations?

  • Mathematics is less about the end product and having a healthy community of people to actually understand the proofs and eventually apply them.

    That community won't exist as many people simply won't even enter the field because it's reprehensible and contemptible, not to mention boring, just to read machine-generated proofs and verify them.

    Collaborating on, or at least working on unsolved problems is what motivates most people.

    AI is like a cheat code in a video game. You get to the end faster but fewer people want to play if the cheat code is always on. You can't turn it off either because the very challenge is to do something unique.

  • What does it mean to “genuinely” contribute to mathematics? Because the definition you use is load-bearing ;) and it might be different from that of others.

  • Because they published a ton of slop papers with terrible English, impossible for humans (even experts) to understand, full of non-standard terms, invented jargon etc.

    So effectively, stuff got proved, but people don't really understand how, so it's mostly fucking useless and done for OpenAI's marketing team, while also pissing off the maths world at large.

    • So if god himself comes down from the heavens and hands you the answers to major unsolved problems, but doesn’t walk you through them and you’ll have to still put in effort to understand them, then that’s “not contributing”?

      2 replies →

> but simply dumping proofs on the math community

So you'd prefer if they kept their work secret? Or you don't want them working on these problems at all? Or they should be required to do the work the way you want them to?

I'm not clear what you see as a better option than dumping.

The biggest problem is LLM tends to produce over engineered, very complicated proofs that are an eyesore even for relatively simple problems. Give it a beautiful Olympiad geometry problem and LLM will tear it apart into ugly algebraic calculations, turns all lines and circles into equations and calculate their intersection points that spans multiple pages because it is a guaranteed way to solve it. Correct, but hardly any use to the user.

  • I’m not deep in to math but the op tweets make sense to me. In that it’s not just the final proof that mattered, but the mind and understanding of the person who arrived at the answer. An LLM dumping the answer can’t elaborate on it, can’t tell the story of how they got there, etc. But it also deprives someone else of that achievement and learning.

    • You can see it in the way we structure college courses: engineering curricula often cover in one semester what mathematicians study over one or two years.

      This is because have fundamentally different goals: being able to use results in calculation versus having a deeper understanding of the subject matter.

  • You know it's valid. You're not working on incorrect assumptions. Surely there's value in that?

    • I don’t know if it is valid. It is unverifiable. I still found some basic algebraic mistakes in top models as late as 3-4 months ago, not sure about it now. But that’s not what I want anyway, so I often put “Do not brute force” in my prompts.

      5 replies →

    • How something is proven is often more important than what is being proven. There are underlying systems and patterns that, when understood properly, improve our model of mathematical (or physical) reality.

      With convoluted and inelegant proofs, AI may fail to uncover those systems and patterns. As a most concrete example, it may fail to recognize some problems as isomorphic to other problems. Brute force solutions are a depth-first search.

      To improve human mathematical understanding, AI is probably best used as a “copilot” (lol) rather than a black box oracle, like these AI companies appear to be doing.

      1 reply →

> simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field.

I tend to agree with this, but what is the alternative? Should OpenAI and Anthropic employ hundreds of mathematicians to do this work? Should they just not solve math problems within their reach?

  • I tend to disagree with OP, for the same reason.

    It's unclear what more could be expected than releasing the presumably already verified results and write-ups for each problem. Should they run a mathematics school too?

    Then the comment goes on to argue AI labs were not interested in actually advancing mathematics, and that investments into AI were manically excessive.

    IMO none of this follows and demand is there to justify the investments.

    The comment then goes further to argue that AI labs were putting too much effort into pretending there was exponential progress rather than actually making progress.

    The factual basis for this claim seems to be that OpenAI released math results and write-ups, and it's not even clear what more they could do on that topic.

    That's a very negative opinion.

  • Look at how actual researchers are currently using the tools: they'll generally use the LLMs to create slop papers, sometimes supported by auto-formalizations, just like OpenAI does. Then they will go through the lengthy process of digesting the results, turning the often incomprehensible and poorly organised outputs into something that humans can understand and build upon, they will then give seminars on the results, further helping with dissemination. Doing so still requires expertise, and probably will for a good while.

    So yes, that is exactly what they should do. Alternatively, if they are too lazy or incompetent to put in the effort themselves, do what AGMAI proposed and fund a third party to help out.

> the grunt work of verifying, refining, and expanding on them

What work do you think mathematicians do normally?

Like they sit whole day and have ideas? And where are the ideas?

The way I see it, _some_ mathematicians enjoy solving puzzles, and now AI is better at solving puzzles.

This does not affect people building new theories.

Also, it's quite prestigious to write a _book_ on some topic. And guess what writing a book entails? Refining and expanding. What you call grunt work.

  • > This does not affect people building new theories.

    Actually, the problem is, somehow the skill of building new theories in math is directly tied to slaving hard over a problem. It's the very experience of slaving away that actually somehow causes ideas to form. Pretty much all mathematicians understand this. Yes, senior mathematicians now can form some new theories, but what about junior ones who will have very little experience in working hard on a problem by hand?

    Of course, they could work on the problem by hand anyway, but they won't because no one will pay them when a machine can do it.

  • My grunt work is different than doing grunt work for a company to fix their problems so that they can make more money.

    • OpenAI does not make money from math papers LOL they just release them for the sake of community. (Because sitting on those results would be considered worse.)

The other gap in AI is it doesn't explain or lay out how it got to the final proof which is often more fruitful for new techniques etc than the final proof by itself.

Right, easy comparison to make the the open source community for software.

And it's not like this is something where we're loaned some top math genius for a limited amount of time and we have to make the most of it. Rather, this is a new high water mark. The accessibility of the results is no longer scarce. The scarcity has shifted, and that's where the focus of the math ecosystem should shift as well. And it doesn't help for a frontier community to saturate and take over messaging pipelines that were typically managed by the math ecosystem. It's not about "stay in your lane" but rather "we need coherence and be careful not to break the system."

Just two cents from someone who could screw up basic cashier math on any given day.

As with any field, convincing people to care about your ideas and your approach is half the battle

Many of the best startup ideas by the best product and engineering minds failed to gain attention and funding. Same with much of the best music - relegated to hard drives with derivative ideas only resurfaced decades later

I would expect much of the recent math dump will be leveraged by other LLM-driven research teams rather than read in depth by a human

I'm sorry I disagree entirely.

The more information the better.

The entire purpose of published work is to remove noise (and perhaps incentivize work through attributing credit).

This information is now out there. You can choose to ignore it if you wish. You may just find yourself a century behind in research.

And on that point most of this research has been looked at by their mathematics panel and comes with lean certificates, it's not exactly noise.

This to me is more the old guard not willing to let go or change their ways.

  • I think in a ideal world your right.

    I think a problem is that math seems like a deeply toxic, ego driven domain.

    I think he argued that e.g because the navier stokes millennium problem ist considered solved now, you won't get any recognition for being the first human to solve.(How would you even proof you solved it yourself and not just regurgitated the ai proof?)

    And since recognition is the main objective, noone would spend time on dissecting the proof, and perhaps finding some unique approach to solving the problem, that could be transferred to other open issues.

    And therefore the problem is now "poisoned". Since it's assumed to be solved noone will research it, and the potential revelations won't be found

    • To be fair I'd hope during peer review this would be apparent.

      During your write up, I'd imagine you would check it's not already out there too. And once complete it's cheap and easy to run it through an LLM and ask is this covered by anything else out there. If its novel and not published it doesn't matter what others say.

      Research is already messy as it stands. Something new can already be dismissed by incumbents as "not novel enough" especially in niche fields where they're likely to be the ones conducting peer review.

  • I think it's quite clear the mathematics panel didn't read most of the papers with the scrutiny it would take to publish it, if only because the lean certificates and the informal proofs are not 100% the same.

    And if it's hard to understand (which seems to be the most common reaction) it's not exactly devoid of noise either

    There was an opportunity for people to work with the AI to produce a proof, now it almost feels they're working against it.

    • I never said it was journal worthy. Just its not all noise. As I say academics are welcome to ignore it, it isn't published in any journals after all.

> Too much effort is being invested in proving that the exponential curve is still holding.

Given sustained exponential growth is mathematically impossible to maintain with finite resources, it's funny to me they're using advanced mathematics to try and achieve this.

Once LLMs pass the threshold to being to invent new general-purpose methods and frameworks, the frontier is irreversibly lost to AI and it becomes simply a hobby that mathematicians pursue. They work through, digest, and maybe write up the proofs for understanding. But the real meat will be growing the LLMs. Who knew software eats the world was so true?

> but simply dumping proofs on the math community and expecting others to do the grunt work of verifying, refining, and expanding on them is hardly a productive way to advance the field

This is a transitive period. In a few years, verification and exchange between model instances will happen faster than humans can follow. Human input will be an ethical question, and not a productivity one, because it will be the bottleneck in any science.

  • How to make sure it doesn't evolve into some sort of Library of Babel of science?

    • That's the billion dollar question. Nvidia currently wants to deploy chips with the sole purpose of monitoring agentic workloads and that doesn't seem farfetched but they obviously have a financial motive to sell more products. It's like a cat and mouse game, like cybersecurity in general.

  • Etic is a factor of productivity, and the larger the contextual window is the weightier it becomes. That’s even integrated within the paperclip parabola.

    If the focus in placed on maximizing some easily measurable output on a narrow perspective, situation is unlikely going to match a sweet spot of holistic equilibrium which is maximizing harmony and happiness through humanity as a whole.

> simply dumping proofs on the math community and expecting others to > do the grunt work of verifying, refining, and expanding

Hm, kinda reminds me of my college days. "Proof trivial, left as home work." was a sentence my Profs loved to say.

"community building"

For what?

If math is just about having a community of other mathematicians to hang out with, it still isn't a career. Nobody is paying money you need in order to to eat, just to hang out in a community.

Just like a software engineer, "Well AI can write all my projects now, but I have my local Rust Users Group to hang out with". Nobody is paying me to hang out and hand code Rust.

Generate and dump on others to verify is how the generative-AI people operate. Be it in maths or just your regular job.

The amount of Confluence pages of "research" that is just a dump of LLM output someone passed to me to review is staggering

I hate this approach, it's unbelievably selfish

Its kinda like the arms race in the cold war. There came out some truly marvelous technologies but the actualy goals were frankly terrifying.

That seems short sighted though. A few years ago models couldn't do this at all, I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities or will remain so.

OAI obviously have a fiscal incentive here, but to presume a year from now we won't see improvements and more succinct work on the results coming from models?

  • The problem is that one a person writes a 60 page proof in theory that person has spent an inordinate amount of time on the proof and can answer questions, describe some insight, etc etc.

    If a random person is given a 60 page proof to digest and not the author, those hidden insights that _aren't_ in the paper might be completely inaccessible. Maybe the AI will "just" be able to provide the insights. Maybe. But pedagogy is tricky work, and despite these AIs being able to do all this fancy math we can't get them to write good cover letters yet, so....

    Ultimately we might be left with just a bunch of intellectually unsatisfying proofs. This means way less drive to simplify the proofs or rework them.

    End result: we generate a layer of "less efficient" mathematics, that won't get built upon. We will not actually have any shoulders upon which to stand.

    • But why should process of discovering mathematical insights be any less attainable to AI models?

      The concern is being raised without evidence, because the evidence points to the gap simply being frontier models have just started to be able to get a raw proof out. Why, given existing progress, should we expect them to be unable to distill insights from those proofs?

      Certainly this even more likely doesn't matter at all for applications: if I can send a radio signal further because my AIs design it a certain way, that's an unambiguous result. Which is really the next step here: turn a proof into a "mechanical" application.

  • > I'm not sure there's any evidence to suggest exploring and refining results is outside their capabilities

    OP didn’t suggest that.

    The bar has been raised. Everyone has to meet it now. An inelegant solution squatted onto the internet doesn’t count as discovery per se, even if it’s impressive.

    • A correct solution verified in Lean will count in perpetuum. It is fine if you want more, but an achievement is an achievement, even if it is by AI.

      1 reply →

  • But what should they do? They got all these proofs, should they just have sat on them?

    • > should they just have sat on them?

      It’s fine that OpenAI posted their findings. It’s not fair to claim these problems have been solved. Not until someone can understand and verify the proof and then communicate the core, novel methodological element to someone else.

      6 replies →

    • >They got all these proofs...

      Your phrasing is illuminating that perhaps they aren't engaged in the creation, understanding, or integration of these proofs by humanity; they just have them. For them, this is a slidedeck they can pass to investors, creditors, the marketing department. Something they can add to the employee onboarding pamphlet.

      What should they do? Hyperbolic maybe, but perhaps engage with humanity.

      1 reply →

    • No need. In the end the mathematicians that don't like this can just not look at the proofs or use them. They have that choice. Just like they didn't "ask" for them, they don't have to even acknowledge they exist.

    • BlackBox: The proof is in the pudding

      This isn't only bad for Math — it's bad for English too.

      'Proof' is going to become the 2026 Most Misapplied Word of the Year.

> There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems) than in genuinely contributing to mathematics

I feel the same way about academia, the papers, the citations, the ego, the narcissism and the taxpayer codependency that got cut off and turns out wasn’t necessary at all thanks to a private sector entity running laps around them

I don’t feel that academics need to pursue the discipline and distributed brain-wracking that has sometimes resulted in the solved math problems, just because more times they find other nooks and crannies to explore along the way. I think the blueprint is enough. Standing on the shoulders of giants is good enough.

and if the concern is that they can’t figure out what to do with a proof, next year’s AI will

> There seems to be more interest in hitting some arbitrary benchmark (we proved X unsolved problems)

> genuinely contributing to mathematics

What's the difference between the two? Proofs are no longer the goalpost?

  • Proofs are valuable but I believe the mathematical community values understanding more. Proofs were previously a great way to develop understanding. Now, less so.

  • Not an expert, but explaining the proof and it being independently verifiable as important as just putting a paper of it. Reminds me of 1000 page proof of Goldbach’s conjecture that a Mathematician reached sometime back. He was told plainly that no one is going to invest time in verifying the proof because there’s a good chance there’s an error somewhere in between.

  • Proofs of open problems are valuable because we are assuming that proving the problem requires some new method or infrastructure in math to prove it. Basically proving open problems isn't actually useful if it doesn't develop new tooling for mathematics, which can help us create new open problems, solve other ones etc.

  • I recommend RTFA, it explains exactly that: why proofs should not be considered the goalpost, and why dumping all these AI-generated proofs might be an overall negative for mathematics as a whole.

Yes, it is indeed a balanced view.

Those few thousand mathematicians are now getting a taste of their own medicine. After all, it was people with extraordinary mathematical talent who developed machine learning and large language models, leaving hundreds of millions of people who earn their living through speaking, writing, or teaching worried about their future job prospects.

Still, I believe almost everyone will be fine. Perhaps AI will also prove good at coming up with new conjectures, and some mathematicians may shift towards applied mathematics or other sciences.

  • Pretty tenuous connection to blame mathematicians for everything bad created in the world, that happened to use math.

This Math 1.0 followed by Math 2.0 framing is wrong. Mathematics did not start a few decades ago when problem solving became the norm. And it will not end now when problem solving turns out to be "easy".

Mathematics will revert to it's main practice, which is to study.

There are lots of weird panic reactions by some prominent problem solvers. See for example the ridiculous cease and desist like statement of AHM shared at Tao blog.

Put this Math 1-2.0 with that AHM statement together and you'll realize that this is a power struggle and that you see only one side of it.

Mathematicians have very diverse opinions about this. I, for example, am for as much as possible automatic harvesting of all these "low hanging" fruits. Should be disclosed as soon as possible, free of any bottleneck, and citable. The mathematics community may do whatever its various members desire to do with these results. Let them decide individually what to do with them. This AI tool is here to stay.

  • This looks indeed like a public relations campaign, it would be helpful to the conversation to contribute your opinion, in a polite way.