← Back to context

Comment by nerevarthelame

3 hours ago

I feel like I'm completely missing something with ARC-AGI. The tasks are so limited in scale and very black and white, which do not at all map to real-life challenges.

I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.

Anyone taking a single look at the ARC-AGI "challenges" can see things a 5 year old could reasonably solve.

They're just jerking eachother off and sending eachother the elevator back: "independent" ML engineer (worked at <large ML company> and currently runs <ML company looking to be bought out) writes a shitty benchmark (writes a single example and spams an LLM to make more variants) and releases it out as the BRAND NEW FRONTIER IN THINKING.

Every single benchmark has been catastrophically flawed and made by clowns.