← Back to context

Comment by ouz-a

5 hours ago

I stopped trusting benchmarks ever since LLMs started speaking alien like English, if I can't understand what they are saying how can I trust them.

That's the load-bearing smoking gun—should I write a better benchmarks to catch the seams?

This article is about using LLMs to overfit for a specific benchmark (or make a custom software for niche use cases) though. Not about LLMs benchmaxxxing

  • Isn’t that the same? It’s a sort of recursive version of overfitting specific benchmarks

same here, it reads exactly the same whether the number is real or completely made up, so the confidence stops meaning anything.