Comment by martianvoid

3 days ago

It's actually crazy to see the difference between opus 5 and the next best model on ARC AGI 3 when you actually look at the ARC AGI problems

How believable is this benchmark? EG maybe opus was training on this? (You can try to identify the IP of wherever previous ARC questions came from)

  • That’s why you have a private dataset.

    • Which you have sent to Anthropic/OpenAI/Google's servers when you run the benchmarks for the previous models.

    • Doesn't matter, people built harnesses that solves arc agi 3, so all you need is to train your model to work like that harness by default. That makes a model specialized at solving arc agi 3 without making it smarter in general.

      It is very hard to make a benchmark you can't do that for, but it is very easy to make your own personal test that others can't do that for since now it isn't a benchmark they can target.

      9 replies →

Why? 30% is passing the first two problems only, which are really very simple.

  • Huh. How do things end up with scores like 30.2% (and results between 0% and 1%) if it's that low resolution?