Comment by solidasparagus
9 hours ago
Every option is going to lead to someone shitting on OpenAI for what seems to be a pretty huge accomplishment. There have been opinions written by some mathematicians that OpenAI should just share the work that they have now so that people who are working on any solved problems can know. Which seems reasonable to me.
I'm up for all the exciting results. But, can't they check once properly before publishing?
I think the answer is literally no, they don't have the ability.
It’s a funny one. I’m not sure what a Lean-less LLM proof even is. LLMs are amazing at bullshitting and skipping key steps and details. I’d imagine a LLM non Lean proof to be generally hard to evaluate - harder than that of a human mathematician perhaps. And the scale effect is against OAI here - the firehose just keeps squeezing out proofs.
Technically,Godel showed you can make proofs say anything. all LEAN does is proof consistency. It does not validate the starting blocks.
I don't think that's true. If what you're doing is building a fuzzer for mathematical proofs then just say that? The fact that it's doing some of the initial work on hard problems is cool. So is the fact that the promise of that new approach is having a social effect of crowd sourcing talented people to follow up on that work. No need to try to misrepresent it as more than that.
This technology has surpassed the the world's best mathematicians at the (explicitly and openly stated) goal by which they measure and reward mathematical progress. Such an advancement that it seems to have thrown the entire field into existential self-doubt. And then someone sits at their keyboard and says it is just a fuzzer.
It’s not unreasonable to want OpenAI to be thorough and rigorously check their work before sharing it though. Sounds like that didn’t happen here.
If say 10% of papers have mistakes, I think they are thorough enough for this kind of volume and complexity.
With the models and scientists they have, I didn't expect this level of mistakes.
10% is ridiculously high compared with published papers in math journals where the number of critical mistakes is near zero.
1 reply →