Comment by xanderlewis

1 day ago

As Kevin Buzzard recently said:

> In a 2020 piece in the Notices of the AMS, I asked the following question: “If one human had an understanding of all of modern pure mathematics simultaneously, how much further would they immediately be able to see?” Six years later we are beginning to understand the answer to this question.

The other crucial part to this is the ability to actually encode and test the theorem (via Lean). Otherwise, we would be swarmed with a billion lines of theorems that no one will be able to ever understand and verify anyway.

  • A majority of these proofs have not been formally verified yet, I think people are overstating how important lean is to the success of LLMs in mathematics.

    • A paper and a lean proof are always going to be better than just a paper. I think mathematicians generally will not read AI math papers that haven't already been verified, especially since we're about to see a ton more AI math papers. Lean will remain important

      2 replies →

https://news.ycombinator.com/item?id=49919676 - Oct 2026 (156 comments)

  • >I believe that the optimal thing to do ... is to let the machines loose, see what happens, and then begin the journey to where they have stopped. Things are currently moving fast. They cannot move fast forever. But if we get on board now then they will take us to extraordinary new places. And after we have arrived, the new adventure will begin.

    I'm relieved that this time around, they have provided partial reasoning traces and prompts for a small number of problems. Do they now also share data with the other model providers..

    • It's beautiful, but the animals are not thinking this about us.

      They structurally cannot understand what we are doing at the place where we hit our ceiling. Only with our highest technology (well beyond their understanding) do we have the tools to go back for them, and try to bring them along and interface better with us (re: recent work in animal communication)

      15 replies →

    • This makes it sound like OpenAI and other closed source ai companies are an inevitability.

      There is nothing here today that is unpredictable or impossible to control.

      It is everyone's choice to let the greed continue, to let unelected sociopaths capture and feed society to the model.

      It is not acceptable to put others at risk. It can stop and it can be done the right way instead.

      That is, inform the industry that those causing these risks will be prosecuted regardless of their messiah complex.

      The US government must not under any circumstances allow the ai industry to form a cartel.

      We can make some effort to encourage open source models and thus stop the companies from causing hysteria by hiding the model, shrouding it it mysticism and prophesying the end times. China is doing a great service to everyone by making llms available to the public.

Is it possible that we are now dealing with a human that has a complete understanding of whole mathematics while being unable have unique novel thoughts outside of convex hull of training data and their transitive expansions?

You assume that LLMs are just summations of knowledge, implying that they do not create new knowledge. I doubt that this is the case. I mean, it comes down to the definition of knowledge, but as soon as you run LLMs, they can produce knowledge that has not existed before, and from my perspective, this is more like what we call thinking than it is just a reproduction of existing knowledge.

  • Research seems to on balance point towards RLHF&RLVR merely increasing subjective sampling efficiency within the pretraining data.

It doesn't seem clear whatsoever that this is true? Is there evidence that LLMs are very skilled at generalizing across domains of mathematics where the training distribution sees little overlap?

As far as I can tell, this is a victory for verifiable loops using LEAN, reinforcement learning, and oodles of compute. I haven't seen evidence yet that this is proof of broad generalization beyond the training distribution.

  • I think such progress by agents is not a sign of broad generalization but of broad coverage. We have exposure to a subset of deeper scientific subfields and thus can only generate certain attacks to solve a particular problem. Since it is not clear which combination will lead to a solution beforehand it is nontrivial to look at a problem and fill our knowledge gaps. LLMs on the other hand have broad coverage and can generate hypothesis on a wide combination of subfields. With Lean an agentic loop can test these to sift the weak ones. In a way the problems solvable with this setup is also solvable by a human who happens to know the right subfields. These problems are likely to require an esoteric combination so nobody could solve them before. I really am not sure whether all generalization is like this or we can leap and create novelties beyond what an llm can generate. That I guess is the tough question that we need to answer to understand the boundaries of intelligence.

  • Full quote is "Six years later we are beginning to understand the answer to this question. Machines have ingested the mathematics on the internet and are able to manipulate this data in a coherent way. The Erdős unit distance disproof came about because a machine happened to be an expert both in discrete geometry and class field theory; one rarely finds humans who are simultaneously experts in both"

Similarly, there are probably many ideas that have not seen the light of the day because they require deep correlation between seemingly unrelated fields. It is not every day that we get a Isaac Newton or Leonardo da Vinci.