Comment by fourside
1 day ago
But if the test metrics are fundamentally flawed they might not be useful even for relative comparisons. Like if I told you that Model A scores 10x as many blorks points as model B, I don’t know how you translate that into insights about performance on real world scenarios.
No comments yet
Contribute on Hacker News ↗