I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.
Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
I'd really like some clarity on what that means. E.g. DeepMind has 'cheated' with this in the past, claiming that AlphaZero only took 4 hours to reach super-human chess levels while conveniently leaving out the fact that it was 4 hours x 5000+ TPUs. Sure it's impressive that it only took 4 hours wall-clock but it's very misleading as to cost.
Can we get a number in Blackwell GPU-hours, kWh, or some other compute-scaled metric?
Depends what your metrics are. If you suddenly found a way to have 9 women make a baby in one month that is huge.
I'm not denying that, but I'd still like to know what that cost.
They did say that. "3 hours of ChatGPT Pro thinking compute"
Yes, what does that mean?
1 reply →
>OpenAI has 'cheated' with this in the past, claiming that AlphaZero [...]
That doesn't sound right
Oops, edited.
Oh cool, we will all now have a math genius on our computer.
It was using their internal math model, so not yet for us
I used future tense. It was implied this will be available.
Well, on their computers. But you can rent them for a price.
An open source model will reproduce it 6 months later
1 reply →
Does this imply that it was a one shot prompt with ChatGPT Pro style models (i.e. best-of-N), rather than the agent swarm approach that was used for Navier-Stokes?
It doesn't imply that, it's just measuring the amount of compute.
But if it was an agent swarm wouldn't we expect something like 300k hours of ChatGPT pro compute equivalent instead of 3?
2 replies →
That estimate is obviously going to conveniently ignore all the failed runs.