← Back to context

Comment by namuol

5 hours ago

Can someone please explain how these models aren’t just fine tuned for benchmarks? I’m not plugged in to this space much but it seems like such an obvious problem…

They definitely are - but also people are using them pretty extensively for work. So ultimately you can't really fake "is it good". But there's no real measurements of that when a model is released, so we are stuck with benchmarks.