Comment by stavros
2 months ago
Of course it believes the proof is sound, it wrote it. If you want to check an LLM's output, you should use a different LLM.
2 months ago
Of course it believes the proof is sound, it wrote it. If you want to check an LLM's output, you should use a different LLM.
Your comment is not substantiated at all.
No, the comment is right. The prompt had GPT-5.6 reviewing the proof, and the result, unsurprisingly, survives review by GPT-5.6.
Given a new context, why couldn't the same model have a decent shot at reviewing some results? It's not like they identify whether this output is from them and then go "yeah correct", that's not how they work.
1 reply →
If you'd ever tried to get an LLM to review its own code, you'd know.
if you get the same session that wrote the code to review it the poor results are entirely deserved.
and if you get a different instance to review the code then you would know that it works rather well.
Use a human maybe.
Only people can really verify clankers.
Can't trust anything LLM since it will confidently lie too.
It can't take responsibility for verification so it can't verify.