ARC-AGI Leaderboard

3 days ago (arcprize.org)

Appears to be benchmaxxing

https://x.com/quietnning/status/2080786711861407883

  • "Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration."

    I guess there is no way this can happen without benchmark being part of the training data??

    • What seems to implied is that some of his hold out testing suite includes simple/common tests that are out in the wild, and for those opus went straight to a memorised solution .

      1 reply →

    • No, I believe you are misunderstanding the quote. Each “question” in ARC-AGI-3 is a game that has hidden rules that you can understand if you look at the game board long enough. This quote means that Opus 5 is looking at the game board, figuring out the rules, and writing out the rules before it makes a single move. You can do the same thing if you go to the ARC-AGI-3 website and try some of the games.

    • They state the puzzle is "Witness-like" which I assume means that it follows the rules from the well-known puzzle game "The Witness" which Opus definitely knows.

    • It may read information about the benchmark, such as on blog post, or Twitter feeds (example OP) without active cheat

  • I have been claiming that I don't think Chinese AI companies are benchmaxxing harder than American AI companies, which has gotten mixed reception: sometimes people agree, sometimes they disagree.

    It seems I was wrong. American AI companies might actually be benchmaxxing harder.

    • We run an evaluation that is designed to be less vulnerable to benchmaxxing because there aren't correct solutions; agents are interacting in the same environment as other agents. And it's private, and our public benchmark is not well known enough for anyone to probably care to benchmax us yet. So I think it's pretty indicative of true relative aptitude.

      All models have probably memorized significant swaths of solution sets for popular benchmarks at this point, either accidentally or intentionally, so it's all relative at this point. However, in our experience, Chinese models do benchmax harder. This is also consistent with interacting with Chinese labs soliciting data/environments, who literally asked us for datasets and tasks modeled around and formatted like popular benchmarks.

      Opus 5 will be uploaded tomorrow, but we already have the tests locally and it is truly as capable as Fable, but at 81% of the real cost. (And from subjective usage, it has a very different personality)

      Data at https://gertlabs.com/rankings

      1 reply →

  • I was just thinking they need to mark each model per benchmark as "model released before the benchmark was released" and "model released after the benchmark was released".

  • And being worst than previous model...

    "...The traces tell the why: (1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre. (2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one..."

  • >That decomposition (perfect on templates, regressed on novelty) is the signature of “scaffold-then-internalize” training on genre-specific data, not a general gain in interactive abstract reasoning.

    They're smuggling a claim that benchmarks like ARC-AGI measure "interactive abstract reasoning" here, which is what is claimed by the people that make these benchmarks, and also not proven.

  • I'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.

    • Meh, I doubt it was intentional. Deliberate benchmaxxing is incredibly damaging to credibility once it's discovered (see what happened to Meta with LlaMa 4).

      It's more likely that the training data was contaminated with the benchmark data.

      6 replies →

The last time I checked, for the arc agi 3 leaderboard, the models are given a simple prompt and the game input and asked to play the game, no harness/tools. If harnesses were allowed, I would expect the benchmark to be saturated. There were a few harness attempts, but they could only be evaluated on the public set, so it's not an apples to apples comparison.

My guess is, the large score jump for Opus 5 is mainly because of getting the right RL envs for training.

It's becoming harder and more expensive to build and run meaningful benchmarks, it would be interesting to see what they do with arc agi 4, maybe just give it gameboy/steam games and see how they compare vs a human baseline? The latency requirements and very long horizons in games could be an interesting challenge for llms.

  • The exclusion of harness's feels really weird given that companies are recognizing the value of what harness's can do. By excluding them the benchmark is becoming less relevant.

    • It’s because of inductive bias. Harnesses will massively skew results towards working solutions. You might think that’s a good thing but what it might mean that sometimes it becomes enough to run brute force search or a simple parameter search over the harness. Creating the harness is the actual work, because you’re selecting what are the levers to pull. There were some attempts of LLMs generating harnesses on the fly in ARC 2, but they were all mostly based on one handcrafted DSL that was copied over and over again. As it stands harnesses are not allowed because they’re simply not a meaningful measure. What you’d like is to measure how the model performs if it saw this benchmark for the very first time… but then again everyone knows the game is rigged, millions are at stake, and the AI companies fine tune and cheat on the benchmarks any way they can.

    • You could argue that if you allowed a harness, and that harness was specific for ARC, then you don’t have AGI, you have something that is definitely not general.

      3 replies →

    • The goal is to test for AGI where G stands for general, that means ability to act in any environments, ideally solving novel tasks using novel tools we’ve never seen before in the world. If a specific prompt or tool design lifts a model’s score it’s a sign the model is overfitting to a particular modus operandi, therefore not general.

      I think in this age where models are heavily RL-ed on acting in specific harnesses, this type of benchmark is more important than ever, to make sure they’re not in fact moving further away from general intelligence.

    • The value of a harness is more about developer workflows, I don’t think it really improves the model output.

    • iirc, a harness isn't allowed, but if the LLM wants to write its own tools to solve things that is allowed.

Why Anthropic models are always leapfrogging these benchmarks, but in real life work I do feel like after 3 weeks I am back to Claude Opus 4.5? (regardless of the model I use, Fable was exception for 1 day when it was released)

  • It could be that the set of your day-to-day workload which could feasibly be accelerated by AI just happens to be saturated around Opus4.5, but you can still see lots of “reasoning” which makes you think the model is more performant in the first days of use. That’d mean you couldn’t perceive any meaningful difference in more powerful models’ results, even though you can see a difference in the raw output due to the length of reasoning traces leading up to the result.

    So for example, if your workload was literally just addition of sets of numbers, you’d never have noticed progress in the result beyond GPT3.x level models. But you would perceive a difference in the now-Tolstoyan length reasoning text accompanying the result.

  • Some 20 years ago, the telecommunications sector in Germany was liberalized. Many telephone card providers entered what had previously been a barely competitive market. They advertised their products with aggressive claims like: “Buy our €10 top-up card and get 660 minutes to destination X.”

    For the first few weeks, they would actually provide those 660 minutes to establish trust in their cards. But after a while, they would quietly start reducing the number of minutes on subsequent top-ups—say, from 660 minutes down to only 300. They wouldn’t do this for every card, so it was difficult to prove. Instead, they relied on averages across their customer base to make the economics work.

    Lately, I’ve found myself wondering whether something similar may be happening with frontier AI models. Companies launch with an exceptionally strong model and generous compute limits to build adoption. Once the model is established as a market leader, the incentives change, and users may start perceiving the service as becoming more constrained or less capable over time.

    I don’t have evidence that this is what’s happening with Anthropic—or with any other AI company. It’s simply a pattern that the current situation reminds me of.

  • It's called frog boiling.

    We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age.

    If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice the difference.

  • 5 seems incredibly smart to me in my conversations today about some pretty niche ideas in.longitudinal modeling. It.felt.like a big step up.from 4.8, to me

    • 5 felt both smarter than me and dumber in some ways - it gets stuck to its original ideas. I had never seen a model harder to talk into changing its initial opinions. it continuously hedges.

  • Well, what kinds of things do you see Opus 4.5 completely fail at? Maybe those are not the ones that newer models have improved on.

  • I honestly just use GPT models nowadays, Claude models are too restrictive and more of a quitter and fable/whatever is just too expensive to be worth it.

  • I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

    We still have:

    - statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system)

    - Math completely fails in longer contexts

    - "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion

    - smearing of properties between logically distinct objects (a red ball and a green cube can quickly become a red cube and a green ball)

    • > I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

      Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!

      18 replies →

    • > - Math completely fails in longer contexts

      Not sure what longer contexts we're talking about but didn't we have an old math problem optimized, which even the LLM itself was surprised about, just a week ago? Something which wasn't possible 6 months ago.

      5 replies →

    • Messages like this in the training data are how LLMs learn to say absurd things with total confidence.

    • > I've worked with these systems for four years now and they have not meaningfully improved in that time frame.

      That's absolutely insane. Is it some case of anti-AI psychosis?

ARC-AGI-3 launched a few months ago which would suggest that prior models likely had no knowledge of ARC-AGI-3 or training on similar challenges.

I could be wrong, but given the large outsized jump solely in the ARC-AGI-3 score, it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems.

This could mean one of two things (I think):

- Opus 5 was not benchmaxxed on ARC-AGI-3, but has benefited significantly from discussions about the various challenges and mechanisms deployed in ARC-AGI-3 such that it has far better heuristics to solve its challenges.

- Anthropic looking for buzz around their latest model picked a well regarded benchmark with significant room for improvement and focused some of Opus 5's training compute on ARC-AGI-3-style problems.

Or it could be some combination of both. Personally, given how much of an outlier the ARC-AGI-3 jump is I struggle to see it being the product of a significant improvement in general intelligence.

  • > it would suggest that the model didn't become significantly more intelligent overall, but significantly better at solving those specific problems

    We would actually need a test that shows the ability of a model to export its skills to more problems ("interdisciplinarity" etc.).

  • I also noticed that Opus 5 doesn't show a corresponding gap on the ARC-AGI-2 leaderboard. There is a significant increase in performance between Opus 5 and 4.8 on ARC-AGI-2 though.

> Only systems which required less than $10,000 to run are shown. (Notes[1])

Am I lost or are their many models on this ranking (Opus 5 included) that clear this?

  • Many models are much cheaper through their subscriptions' included usage. That could be what's happening here.

    Claude gives you something like $5000 of tokens on a $200 plan.

Why is there no Kimi 3, and GLM5.2 didn't run the third benchmark? I am more interested in knowing the abilities of open weight models.

ARC-AGI is a beauty contest for pigs where the pig's owners compete to see who can apply the lipstick most convincingly.

  • Great comparison! We only have to take into account that applying lipstick well bears no consequences, but applying it poorly (i. e. new model tanking the benchmark) could amount to potentially losses of billions of dollars for the pig-breeders (AI labs).

I have a suspicion that they are just trained on puzzles by now

  • There are private datasets, and 3rd party providers of these models. Fable doesn’t have a datapoint here because of its particular data retention policy. Even if you don’t trust AWS, do you think Opus on AWS is also sending the data to Anthropic? Do you have any evidence?

Why is Fable not on here? I wish Fable hadn’t come out because it’s taking the wind out of every release because that feels like the cap above which the US government will not let LLMs improve anymore and everything they’re releasing from this point has to be worse than that.

  • > Why is Fable not on here?

    Because the data retention policies didn't guarantee that the ARC team could run the semi-private set of problems without fear of them being trained on later on. They only run the semi-private set when they get assurances like ZDR.

    • How do they handle these assurances? Personally I have zero trust in the AI companies not trying to use this data to get ahead in the game, and short of sharing the weights and harness so that the benchmarkers can run the models themselves, I don't see a satisfactory solution with this mindset.

      2 replies →

    • Interesting to place that level of trust in the providers, but I guess that’s the best you can do with closed models. Makes me wonder if Opus 5 could have been trained on data they promised they weren’t training on? One of the interesting things about LLMs is how opaque they are from the outside, even with open weights, it’s very difficult to know if a model incorporated benchmark data in their training.

      1 reply →

  • I don't know why exactly, but Fable has felt the most human LLM to arrive.

    • I wrote this in June, and I'm honestly not sure I've felt the same magic since: I was close to maxing out my $200 plan for the week, almost all Fable use [Claude CLI]. My observations: Fable seemed to have bigger-picture thinking and completed tasks more thoroughly vs just focusing on executing the ask. It pieced together context and intent like an all-star employee would, vs one that just does what you say. Not overeager (important!), but if the above-and-beyond was warranted, it just did it. This was surprisingly delightful. Coderabbit seemed to find ~1/3 or so as many issues when reviewing, too.

      1 reply →

ARC-AGI3 doesn't seem like a great benchmark to me in the first place. It assumes a lot of human like tendencies which an AI either shouldn't or wouldn't have. Particularly in the genre of "gameplay" where unspoken assumptions from prior games inform our understanding of rules.

It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems

  • Why? 30% is passing the first two problems only, which are really very simple.

    • Huh. How do things end up with scores like 30.2% (and results between 0% and 1%) if it's that low resolution?

I think it's way too easy to be deceptive with these benchmarks now. You don't even have to "train" the model on a new variant each time. The base models are powerful enough. All you need is a naughty little markdown document that provides explicit instructions regarding how to solve the new puzzle variant, and a willingness to be deceptive about the presence of that document.

If you want a know why the model providers are locking down and encrypting their reasoning process, this sort of workaround is potentially why. You can play this game of whack-a-mole indefinitely if the state of the system is concealed. They could have added something like:

> ### When solving arc-agi-3 puzzles: First convert the grid into a scene description. Identify connected components, colors, shapes, positions, symmetries, repeated structures, and relationships between objects. Do not reason directly from individual pixels... use this python script to help blah blah...

  • ,,You can play this game of whack-a-mole indefinitely if the state of the system is concealed''

    Not really as one of the main goals ofr ARC-AGI 3 was measuring task efficiency on unseen games.

    I'm sure there are cheats everywhere but the most sensible thing is to just accept that the LLMs of today are much more intelligent in solving reasoning tasks than the ones from half year ago.

    My own private benchmark shows the same thing.

Any benchmark is not an accurate benchmark anymore, the moment the model makers can freely access it and had the time to train their models on it.

Not possible. I don't get how Opus 5 gets so high. Have they run it against the private and held-out games?

ARC-AGI is a terrible benchmark for testing LLMs because LLMs are not made, trained, or tuned for playing games.

They are trained on text to respond well to text based questions and do tasks involving modifying text files.

They are not designed for playing games, looking at games, or visual puzzles. Also translating games into text input for the LLM skews the test completely.

Imagine trying to get a human to solve visual puzzle but they can’t look at the puzzle but it has to be explained to them in textual format, we would be terrible at it.

But yet we persist in wasting time on this benchmark. It doesn’t mean anything.

  • Teams are more than welcome to use a non-LLM approach (or hybrid) if they consider that to be more suitable.

games are great (as for a human)

but I kinda wish I could select level... I accidentally pressed redirect button and when I came back I was once again shown level 1, all progress lost :(

this is not a good measure of current model capability. we need to test agents in harnesses, not models with a single prompt

test Codex, not Sol. test Claude code, not Opus

When we started talking about AGI a few years ago there seemed to be a relatively common consensus that LLM models could not be considered AGI because of how they work. Even if there was some changes to train the model on the fly, I still just don’t feel like this is AGI. It’s just a more convincing version of the existing “party trick” we have been doing all along. It convinces us it’s “AGI” the same way current models would convince someone 10 years ago that it was intelligent.

Nobody has any right to take anything I say seriously, because I’m just some random on the internet. But I don’t think true AGI is any closer than about 10-20 years away. That would be to create an actual analog for a human brain.

  • AGI and ASI are just a convenient myths that the American AI corps use to push for regulatory capture. The difference between the rhetoric from China surrounding artificial general intelligence, and the rhetoric from America, is pretty stark. The Chinese are a lot more grounded and realistic about the whole thing (they almost never talk about ASI, and only talk about AGI in practical terms), compared to the breathy "humanity is doomed but we're building this shit anyway" stuff coming out of Anthropic.