Comment by fwlr
3 hours ago
The relentless progress towards saturation of benchmarks is, I suspect, at least partly a similar story. Whatever holdout questions are used to evaluate GPT-x will be in the GPT-x conversation logs, and therefore in the training set for GPT-x+1.
I can't think of better invention than one that can saturate a benchmark of every interesting problem.