Comment by nrub
6 hours ago
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
6 hours ago
Or they spent time optimizing their model to real world problems they're facing and didn't waste time trying to game a benchmark.
Or they did try to game the benchmarks and just didn’t do it well enough.
Benchmarks are one data point, not the only one, but the easiest one to compare.
Right, but the point is that you can't conclude that a model is necessarily bad because it's not hitting the same scores on benchmarks. I just don't agree with lacker's conclusion, because their logic doesn't seem to consider that. Scoring lower on a benchmark doesn't strictly mean they have a bad model, but it may be the case. Like you said it's one data point, but being the easiest, and obviously most gamed, means you should probably weigh them less heavily.