Comment by stared

3 hours ago

It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010.

I am curious - how does it fare for other benchmarks, or everyday programming?

PrimeIntelect is not on official ARC-AGI-3 leaderboard: https://arcprize.org/leaderboard

  • Good to know!

    Is it that it wasn't accepted yet, or are there issues with how it was run?

    • It’s a self-improving harness, and ARC-AGI-3 is explicitly a few-shot benchmark. It’s likely that it gave itself more than the maximum number of tries to learn the games, or even hardcoded the answers.

      There’s a lot of improvement to be had from the benchmark harnesses, but sometimes, like with ARC-AGI-3, the limitations are intentional.