Comment by arctic-true
1 day ago
Not a doomer but I try not to be a denier, either. These are hugely impressive results. I do not see the technology plateauing at the current level (though I am dubious about an infinite exponential growth).
Before I cope, I’ll note that there are plenty of “doom” scenarios that do not require any improvement in capabilities from what we had before this latest unreleased model. We’re at the point where a determined bad actor with enough compute could compromise critical infrastructure in a way that results in casualties, where this actor would not have been capable of such without LLMs. This may not sound like Skynet, but I don’t see why it makes a difference if I’m one of the casualties.
With that in mind, here is the cope: first, mathematics is an inherently verifiable domain. An LLM can use tools to determine with absolute certainty whether it is correct, and an independent third-party could review and confirm. All of this can be done without any interaction with the physical world or with other minds.
Second, OpenAI is able to marshal compute at a scale that an individual mathematician can only dream of. It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
Third, none of these problems are solved in a vacuum - the reason OpenAI chose these problems is that they are widely discussed and many people are working on them. It’s possible that someone else was close, and OpenAI only contributed the finishing touches. (This wouldn’t need to be plagiarism, to be clear - people publish their work!)
> It’s possible that these problems were lower-hanging fruit (in relative terms), such that they could be resolved simply by throwing a ton of compute at the problem guided by an intelligence that is not itself remarkable in comparison to a human.
See:
> The average result used the equivalent compute of roughly three hours of ChatGPT Pro thinking. (TFA)
Sure and if I make a half court shot after an hour of trying, the result only took 1 second.
Exactly this. If you take the entire start to finish 'agent hours' (measured comparably to man hours) they took to find all discoveries, including the go-nowhere trails that were discarded, and then divide by 90 (or whatever the exact number of results found was) it's almost certainly going to be many orders of magnitude more than 3.
They provided a "snippet" of a prompt here [1] which is not only a beast, but also seems reasonably likely to have been LLM generated. So they're using LLMs to parse a vast body of mathematical work, probably including what people themselves are 'privately' working on with GPT, and then prompting other LLMs to work on such.
[1] - https://github.com/openai/math/blob/main/reasoning_traces/re...
That tells me very little. What was the cost to OpenAI in dollars? What differentiates the high-cost problems from the low-cost problems? And that’s before you consider that OpenAI has strong incentives to downplay its costs while emphasizing its results. A one-liner in a write-up doesn’t change the fact that they have access to massive resources.
In a couple short years we've moved from "AIs can't do anything useful" to "they're lying about the actual cost of the innovative breakthroughs!".
I know that the former and the latter may be discrete subsets of the anti-AI crowd, but come on.
It’s very typical in human math that explaining the final result after years of searching looks very simple too.
Fourth, these hundreds of solved problems are the result of OpenAI attempting tens of thousands of problems and failing. When you hear claims that the average result took about 3 hours of model time, I simply do not believe it. If you account for all the time spend properly, it's probably orders of magnitude more.
I think the announcement says they report the amount of problems attempted somewhere.
Edit: "Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above."