Comment by fancyfredbot

2 hours ago

When You See One Cockroach, There's Probably More.

This can and should erode our trust in every single proof OpenAI published. The model is clearly faliable despite the lean proof, and clearly the output was't actually checked properly before release. Once these proofs are peer reviewed and published in a journal we might be able to trust them again but until then they are just slop, sadly.

I think OpenAI actually did the right thing by sharing everything with the whole community right now but I also hope that some significant credit will now go to the reviewers who confirm these 'proofs" actually work.

Fantastic. Let’s also apply this reasoning to math papers written by humans too?

Humans produce flawed Lean proofs. Indeed LLMs were successful at finding and fixing many issues in the “core” standard library if I recall correctly. Humans regularly produce flawed papers and have minor issues require fixing. And when it happens it often isn’t as prompt and clear as this.

  • I think you are implying that my suggestion is unreasonable and exceeds the standards applied to science produced by "normal" human processes.

    In fact we already follow exactly this process for human papers and have done for a very long time. Publications without peer review are treated with great suspicion. This is how we end up with journals of varying levels of prestige and rigour.

    The process isn't flawless and there are huge problems with retractions, as well as weird financial incentives and rent extraction but there is definitely an increased level of trust in a paper published in Nature.

  • Super-intelligence that’s going to end mathematics as we know it surely must be held to a higher standard than puny humans?

    More seriously the problem is the complete utter lack of care OpenAI has shown in their desperation to demoralize mathematicians with their new LLM. In their words, it took 3 hours of ChatGPT pro per result, why not spend a hundred hours per result formalizing it, checking if the formalization matches the natural language proof, and whether the argument could be made more simpler and readable. Any human paper has hundreds of hours of work put into it, but OpenAI who is absolutely adamant in demoralizing the mathematical community and demonstrating their superior “intelligence” will only spend 3 hours, write unreadable, inscrutable proofs, not formalize all of them, and then dump it on the mathematical community for some reason.