Comment by imtringued
1 hour ago
You mean any repeatable benchmark will be saturated.
The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.
No comments yet
Contribute on Hacker News ↗