← Back to context

Comment by sanxiyn

3 hours ago

Yes: https://x.com/ElliotGlazer/status/2108026240582246600

So someone ran a different LLM to find an issue they'd find anyway during formalisation? That's not the same as relying on thriving community.

  • Is that guy just some rando "someone" though?

    • The people who released the papers weren't randos either. OpenAI has a ton of mathematicians on staff, including Jacob Tsimerman, a fields medal winner.

      Scientific progress used to be people debating and correcting other people. Now it's going to be people with AI assistance debating and correcting other people with AI assistance.

      1 reply →

  • The bigger question is why there was internal pressure to rush such a historic launch without having someone in the company, anyone, check the proofs first.

    This concerns the Hodge conjecture (millennium prize related) paper. Seems to me like PhD nerds weren't confident bosses pushed ahead anyway.

    • What makes you think that nobody checked the proofs first? It's not like someone checking it once without spotting any mistakes means that nobody else will find any mistakes either.

      6 replies →

    • My assumption is that they checked the proofs vigurously, but now a way broader community is taking a look with professionals from the relevant subfields, and different agent setups / models.

      2 replies →

  • I would say

    1. This is expected if you only use a single model family like Claude, eg. we use a different model family for code review than authoring, OAI could have done this too for their math dump

    2. Ai needs a good human driver beyond the trivial or mundane, they are expert enhancing machines, not expert creating machines. This is where the community comes in. Reading Tao's ChatGPT session reveals this: https://news.ycombinator.com/item?id=49010345

    3. OAI is not trying to be a member of the/any community, this is not the first story to shows this, nor do I expect it to be the last. Perhaps this is them being effective altruists today? /s

That’s not proof though is it? If the original LLM output is fallible surely the LLM review of that output is also very much fallible?

  • Proof of what? There is a sign error in one of the proofs, OpenAI acknowledged it and withdrew three papers (two relied on the result).

    I agree LLM review is also fallible (as is human review) but the interesting part to me is that finding this sign error before publication should have been table stakes for OpenAI, it’s their own model that found the sign error.

    I’m curious what was in the original prompt and what was in the prompt that led to finding the sign error, I think it matters a lot for understanding the dynamics here

  • Output of LLM can be infallible* even if LLMs themselves make mistakes. Ditto with humans.

    As much as anything can be infallible.