← Back to context

Comment by mikert89

12 hours ago

Anything you can verify to be right or wrong can be done by a model. All benchmarks will be saturated

Disagree.

Examples:

- predict a coinflip: easy to verify, hard to learn

- earn $100: easy to verify, hard to learn

- increase paid subscriptions in an A/B test: easy to verify, hard to learn

I won't get into it, but there are many properties beyond verifiability that are needed to saturate a benchmark.

  • Doesn't "saturated" mean that essentially there won't be any more progress in the benchmarch? Also of note is that two of your points only mean something on an occidental capitalist system.

  • these just need more compute:

    - earn $100: easy to verify, hard to learn

    - increase paid subscriptions in an A/B test: easy to verify, hard to learn

    but we both know these examples go against the spirit of my point

    • Perhaps, but I think a bigger problem than lack of compute is the cost of rewards. Games like Chess and Go were solved long before self-driving, partly because it's incredibly cheap to acquire the reward of a bad board game decision, relatively to how expensive it is to acquire the cost of a bad driving decision. With driving, acquiring the reward can cost you $20/hr for human supervisors to generate disengagements, or $100k if you crash, or $30B if you crash the car into a person in a way that causes your company to collapse (e.g., Cruise).

      2 replies →

Then I propose the tomjen-1 benchmark: prove the N vs NP problem formally undecidable.