Comment by stared
3 hours ago
It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010.
I am curious - how does it fare for other benchmarks, or everyday programming?
3 hours ago
It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010.
I am curious - how does it fare for other benchmarks, or everyday programming?
PrimeIntelect is not on official ARC-AGI-3 leaderboard: https://arcprize.org/leaderboard
Good to know!
Is it that it wasn't accepted yet, or are there issues with how it was run?
It’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers.
There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.