Comment by threepts

1 day ago

Why don't they ask their premier model to generate a bench for them?

Jokes aside, a benchmark I look forward to is ARC-AGI-3. I tried out their human simulation, and it feels very reasoning heavy.

Leaderboard: https://arcprize.org/leaderboard

(Most premier models don't even pass 5 percent.)

12 comments

threepts

falcor84 1 day ago

They focus on minimizing the number of moves and don't allow any harness whatsoever, putting the bar extremely high. The current top verified contender (Claude Opus 4.6) is at only 0.45%. But with how new it is, I expect a lot of improvement in the next generation of models.

threepts 21 hours ago
Optimal for judging actual reasoning ability rather than an LLM's ability to regurgitate knowledge from a necropost on HN/Reddit/Twitter from 2018.
- knollimar 21 hours ago
  
  a small harness that stores text files and manages context could be useful, otherwise you lose all ability to measure that skill (and that's important because it represents real world use cases on large code bases)
- jjmarr 16 hours ago
  
  I'm making an LLM agent that can play DS games. The biggest blocker is clicking on the right spot to move things around in space rather than reasoning abilities.
  Arc AGI seems to test that as well. Every game is a rectangular grid to make it as easy as possible yet the AIs still fail.
  I'm fairly certain the way forward isn't through agents directly interfacing with UIs but through agents using scripts and other tools to interact with the interface. That's why harnesses are so critical to performance on tasks like this.
  I would like a version of Arc AGI that tests the agent's ability to dynamically create these harnesses.
  
  3 replies →

sowbug 20 hours ago

Why don't they ask their premier model to generate a bench for them?

It's not a crazy idea. Have the older model interview the newer one and then ask both (or maybe a third referee model) which one they think is smarter. Repeat 100x with different seeds. The percentage of times both sides agree the newer model won is the score.

alansaber 1 day ago

Very (reasoning) heavy benchmarks do seem like the way to go, being the hardest to game.

xtracto 20 hours ago

Can AI write a problem so difficult that even AI cannot solve?

Hehe

ngruhn 18 hours ago

How about prime factorization

therealdrag0 21 hours ago

[dead]