Comment by viraptor

5 hours ago

So someone ran a different LLM to find an issue they'd find anyway during formalisation? That's not the same as relying on thriving community.

Is that guy just some rando "someone" though?

  • The people who released the papers weren't randos either. OpenAI has a ton of mathematicians on staff, including Jacob Tsimerman, a fields medal winner.

    Scientific progress used to be people debating and correcting other people. Now it's going to be people with AI assistance debating and correcting other people with AI assistance.

    • > Now it's going to be people with AI assistance debating and correcting other people with AI assistance.

      but do these 'people' need to belong to a thriving community or not to be able to do those things?

The bigger question is why there was internal pressure to rush such a historic launch without having someone in the company, anyone, check the proofs first.

This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD nerds weren't confident bosses pushed ahead anyway.

  • What makes you think that nobody checked the proofs first? It's not like someone checking it once without spotting any mistakes means that nobody else will find any mistakes either.

    • Tweeter checked it with Astra. It seems like OAI could have pointed their own instance at it before launch. Because the source of the tip is likely someone at OAI, my guess is that they actually did check. But after the launch.

      3 replies →

    • > What makes you think that nobody checked the [ai output] first?

      Because this is what they say all the time. It's like a badge they have to wear and tell everyone they are wearing, even though we see it.

      You can see the same thing with ANT. Had they looked at Mythos output, they would have realized there were only 76 items, not 79 like the bot claimed. Or the ones that were just a "it crashed" and nothing else (not a cve imo).

      https://www.youtube.com/watch?v=NnV_cWeoo5Q (Linux Kernel team sharing their side of the Mythos "hacking" story)

    • Because people are finding errors using other LLMs. This implies that if they spent a miniscule fraction of the enormous pile of money they spend making this pile of slop they'd find the errors. They didn't want to find errors. They want to build hype for an IPO.

  • My assumption is that they checked the proofs vigurously, but now a way broader community is taking a look with professionals from the relevant subfields, and different agent setups / models.

    • > My assumption is that they checked the proofs vigorously,

      Perhaps with AI.

      A manual human check of each one would take a few month at least. In peer review, there are horror stories in math about more than 1 year before the journal accept the paper. So 3 reviewers x 700 pdf = 2000 mathematicians, that is 10%-20% of the community according to an unreliable count printed by Gemini after scrapping r/math.

      Also, in most cases the only people that can understand the proof in a so short time (let's say a few months!) is the small group of people working in similar problems, i.e. the same group of 20-100 guys/gals that you meet in every conference.

I would say

1. This is expected if you only use a single model family like Claude, eg. we use a different model family for code review than authoring, OAI could have done this too for their math dump

2. Ai needs a good human driver beyond the trivial or mundane, they are expert enhancing machines, not expert creating machines. This is where the community comes in. Reading Tao's ChatGPT session reveals this: https://news.ycombinator.com/item?id=49010345

3. OAI is not trying to be a member of the/any community, this is not the first story to shows this, nor do I expect it to be the last. Perhaps this is them being effective altruists today? /s