Comment by philipwhiuk
1 month ago
Would it be any more applicable?
We see the same problem in database benchmarking.
The good thing about the deliberately-non-real world pelican case is that it gives a general impression of how much the model is improving because it's not likely that it's being specifically targeted at it, rather than a 'real world case' which might have been specifically optimised for.
No comments yet
Contribute on Hacker News ↗