Comment by ouz-a
5 hours ago
I stopped trusting benchmarks ever since LLMs started speaking alien like English, if I can't understand what they are saying how can I trust them.
5 hours ago
I stopped trusting benchmarks ever since LLMs started speaking alien like English, if I can't understand what they are saying how can I trust them.
That's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?
This article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing
Isn’t that the same? It’s a sort of recursive version of overfitting specific benchmarks
same here, it reads exactly the same whether the number is real or completely made up, so the confidence stops meaning anything.