← Back to context

Comment by fwlr

3 hours ago

The relentless progress towards saturation of benchmarks is, I suspect, at least partly a similar story. Whatever holdout questions are used to evaluate GPT-x will be in the GPT-x conversation logs, and therefore in the training set for GPT-x+1.

I can't think of better invention than one that can saturate a benchmark of every interesting problem.