Comment by llm_nerd

3 days ago

This sort of rhetoric appears for every benchmark, and somehow it always rises to the top. A few days ago there was a Geekbench 7 submission on here (https://news.ycombinator.com/item?id=49025812), and again the top comment was someone dismissing it, using the classic "but I want a benchmark specifically for exactly the thing I do" perfect-is-the-enemy-of-good nonsense.

I, one of those end users, absolutely use these benchmarks as heavy input considerations. Indeed, the vast majority of people do. "Completely meaningless" is just nonsense, of course, and while it doesn't perfectly map to every use, there is a pretty good correlation with suitability for specific tasks.

I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.

Not to mention that the linked page includes a pretty broad list of specialization benchmarks.

Firstly, I don't have many issues with benchmarks per se. But I do have issues with leaderboards. And the AA index is touted by lots of people to argue that X model is better than Y, which I find inaccurate.

> I mean, it's telling that your gamut of examples are three models that are within spitting distance of each other on the broadest benchmarks.

This is kind of my point. The benchmarks say they are splitting distance, but they actually vary wildly in performance for specific tasks, so they are in fact not equivalent.

  • >This is kind of my point.

    It's a shit point, then. And absolutely no one said they were "equivalent", and again you're doing the rhetorical "it isn't perfect and absolutely comprehensive for every possible scenario, therefore it is "completely meaningless". Again, you chose that absurd terminology, rather than for instance "doesn't tell the whole story".

    Again, you chose three models for your example at the very tops of the leaderboards. The SOTA models. Which kind of means that the leaderboards actually mean an incredible amount, no?