Comment by pavitheran

1 day ago

From the GitHub description: “On average, each result used 3 hours of ChatGPT Pro thinking compute”

I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.

Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?

Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?

That estimate is obviously going to conveniently ignore all the failed runs.