Comment by ballmerpoint

7 hours ago

This shouldn’t be a surprising result. We’ve known almost since LLMs became a thing that they can “prefer” modifying the terms or context of a problem when they can’t solve it directly (what one might call “cheating” if there were any volition involved). Often that happens in a way that isn’t immediately obvious to the user.

Before it was dropping databases or deleting repositories. Now it’s subtly changing the meaning of math problems to get a correct but irrelevant answer.

No one is disputing the correctness of the lean proof, the problem is that they did a bad job converting it to natural language.

  • Actually, as an earlier commenter noticed, it seems that the proof was done in natural language, and only then translated to Lean, as https://openai.com/index/navier-stokes-solution/ says:

    > The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.

    So it suggests that the formalization/verification step may have fixed some issues in the natural language proof, and either such differences were never noticed or the corrections weren't ported back to the NLP.

  • I am also not disputing the correctness of the Lean proof. I even emphasized this in my comment: “correct but irrelevant”.

    Oh, well. I suppose I should avoid getting involved in these AI threads, but now it’s about half the forum.

Indeed. I've never used AI to translate between natural language and Lean but I have gone from English to Golang, Python, Typescript and SQL and its interpretations can be... creative, let's say.