Comment by at1as
5 hours ago
Yeah, I wouldn't purport to use this as a measure of intelligence as it applies to humans. I'd leave that to the philosophers, but my intuition is that models are still a long way off of true human-like intelligence (though benefit from certain unfair advantages).
I'm most interested in these problems as a relative measure of performance for successive model generations. If the prior generation couldn't solve a problem but the current one can, that's useful information, especially when we take into account what the proofs look like.
It's not a perfect benchmark, but I prefer it to many others that I see floating around.
No comments yet
Contribute on Hacker News ↗