Comment by red75prime
4 hours ago
I mean that it's hard to determine retroactively how much time it would have taken humanity to solve an open problem that was solved by AI. This is a measure that can make ASI look mundane because we don't know how long it would have taken mathematicians to solve a subset of the Erdős problems.
Yeah, I wouldn't purport to use this as a measure of intelligence as it applies to humans. I'd leave that to the philosophers, but my intuition is that models are still a long way off of true human-like intelligence (though benefit from certain unfair advantages).
I'm most interested in these problems as a relative measure of performance for successive model generations. If the prior generation couldn't solve a problem but the current one can, that's useful information, especially when we take into account what the proofs look like.
It's not a perfect benchmark, but I prefer it to many others that I see floating around.
but i feel like its purpose is to measure superintelligence in math. So to be a good measure of it, it can't saturate easily/has to be somewhat mundane at even insanely good levels. (though i do expect that once ais are across all areas/approaches superhuman at math at least 30% will be solved--then probably long-term (like after 2 years) less than 30% will remain unsolved.