Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.
Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).
You mean any repeatable benchmark will be saturated.
The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.
Disagree.
Examples:
- predict a coinflip: easy to verify, hard to learn
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.
Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.
these just need more compute:
- earn $100: easy to verify, hard to learn
- increase paid subscriptions in an A/B test: easy to verify, hard to learn
but we both know these examples go against the spirit of my point
Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).
2 replies →
You mean any repeatable benchmark will be saturated.
The problem is that there is a huge perverse incentive. The intelligence is in the training layer not in the model parameters, but the intelligence is really good at remembering things, so if you let it take the test, it can RL it.
Then I propose the tomjen-1 benchmark: prove the N vs NP problem formally undecidable.
Yes, but not necessarily under tight budget constraints.
theres no budget constraints for AGI
Yup. Only subjective taste remains.
Nope that will be commodified in short order.