← Back to context

Comment by anon373839

6 days ago

All models are benchmaxxed, period. ”Jagged frontier” is the euphemism du jour, I believe?

Anthropic/OpenAI were touting PhD-level intelligence three years ago. And they’re still shipping models that aren’t smart enough to realize things such as the need to drive the car to the car wash (because they hadn’t yet hill-climbed that particular brain-teaser).

> aren’t smart enough to realize things such as the need to drive the car to the car wash

When this first went viral, I immediately tried it on Opus (whatever version was latest at that time), and it got it first try. Tried a few more times in fresh sessions and it got it every time.

Sonnet did screw it up, though.

Jagged frontier is not the same as being benchmaxxed. Benchmaxxed is à la Goodhart's Law "when a measure becomes a target, it ceases to be a good measure." Jagged frontier is about how models that seem superhumanly intelligent at one category of tasks (e.g. coding web applications" can seem toddler level or worse at another category (spatial reasoning) because the training corpus doesn't generalize to there.

Anthropic/OpenAI were touting PhD-level intelligence three years ago

No, they weren't. GPT-5 was where OpenAI started talking about PhD-level, and that was less than a year ago.

  • (March 2024) Antropic claimed Claude 3 Opus had "graduate-level expert reasoning" with GPQA results of around 60% showing a roughly phd level performance.

    (Sept 2024) OpenAI claimed o1 was phd-level in their launch post.

    You're kinda both wrong. :)

    • They claimed the model was PhD-level, but they never mentioned the university the model graduated from... :)