Comment by docheinestages
5 days ago
Isn't this cheating? Or rather, are frontier agents only looking at one question at a time? If I understand correctly, you're looking at all the examples of the exam questions. If the exam was adjusted so that you can only look at one question at a time, you won't get 44% anymore.
Why would that be cheating? That's what humans do when they learn, they look for the signals and patterns that reduce the possible set of answers so they can converge on the solution and narrow the search space.
I mean, I'm interested to know if the frontier models also get to see all questions at once. Then it's more fair game than if they just see one question at a time.
They look at each problem individually.
As the ARC AGI guidelines [1] state, 'a core design principle of ARC-AGI is that the test taker must not know what the test will be.'
[1] https://arcprize.org/policy#dataset-security