Comment by stavros

2 months ago

Of course it believes the proof is sound, it wrote it. If you want to check an LLM's output, you should use a different LLM.

Your comment is not substantiated at all.

  • No, the comment is right. The prompt had GPT-5.6 reviewing the proof, and the result, unsurprisingly, survives review by GPT-5.6.

    • Given a new context, why couldn't the same model have a decent shot at reviewing some results? It's not like they identify whether this output is from them and then go "yeah correct", that's not how they work.

      1 reply →

  • If you'd ever tried to get an LLM to review its own code, you'd know.

    • if you get the same session that wrote the code to review it the poor results are entirely deserved.

      and if you get a different instance to review the code then you would know that it works rather well.

Use a human maybe.

Only people can really verify clankers.

Can't trust anything LLM since it will confidently lie too.

It can't take responsibility for verification so it can't verify.