Comment by krm01
6 hours ago
We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
6 hours ago
We can throw benchmarks in the bin by now. Each one I've seen is heavily biased and skewed. It holds very little reliable data points (unfortunately)
If you look at papers on benchmarks, they're usually created to expose gaps in how models are trained. It should be no surprise that models get better on them over time, because you can't get better at what you don't measure.
Cherry picking the benchmarks you present is where the falsehoods lie.
My conclusion is the opposite. If benchmarks were meaningless, surely Meta would be able to find some benchmark that shows they are better than Sol and Fable. The fact that they can't do that tells me that benchmarks still do mean something.
Muse 1.1 performed relatively well according to benchmarks, putting it within spitting distance of the premier models. However, based on the results I got from it and the review videos I watched, it wasn’t even close.
Opus 5 is incredible at making games. Almost like a generation better than other models from my experience. You won't see that if you just look at the popular benchmarks..
You have to test each model on your actual use case to see how well it really performs.
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
Or they did try to game the benchmarks and just didn’t do it well enough.
Benchmarks are one data point, not the only one, but the easiest one to compare.
1 reply →