← Back to context

Comment by latent-person

1 day ago

> The natural language proof was derived from the lean code, badly.

Was it? Are you claiming a LLM does reasoning in lean or what? Since this (and all the other proofs by OpenAI etc) have been in the reverse order [1]:

> The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.

[1]: https://openai.com/index/navier-stokes-solution/

Yeah, I was surprised some people think LLMs are reasoning in Lean directly... all their training data is in NL.

  • It's not that much of a stretch: give the LLM a top-level proposition for the thing you want to prove and have it hack away at it. Each sub-step is verified in lean so you know it's correct. But, the linked post definitely suggests otherwise.

    That is definitely interesting because how do you know the 88 hours of work are correct before you throw another 17 hours of lean formalization work on it? You could end up just finding out there was some hallucination in the original work.

  • An incredible number of people think that it is reasoning in lean. Argued with several people on this topic. I think they read headlines about lean being used by LLMs and assume it is being used to write the proof.

  • If you have ever worked with claude code and lean, it goes back and forth between lean and NL reasoning. I have done a few proofs with claude and it almost _never_ gets the argument right in prose. It usually has to go into lean and grind through proof obligations and then it finds problems, work arounds etc.