Comment by numeri

17 hours ago

That's such a shit parallel example that it borders on dishonest.

There are hundreds of incredibly strong scientific priors that would have to be disproven for the moon to contribute to the solution.

If a model was trained on this data, even if it was trained using methods that lead you to believe it unlikely to have learned details about the proof (e.g., maybe it was only used to train some kind of reward model, which played a minor role in the overall training and would thus be very unlikely to transfer details of a proof), you wouldn't have to disprove large swathes of known science to be wrong.