I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.
What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?
Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.
It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.
Here's 'Generate an SVG of a ring-tailed lemur riding an electric scooter' at reasoning level max: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Quote from the thinking trace:
> I’m thinking about how a helmet would obscure lemur ears, but using an electric scooter helmet seems responsible.
It's pretty solid - face is a little wonky but excellent tail and scooter.
Hmm, the face makes it not work as a one shot artifact.
If you want to see something brutal, ask for a zebra riding a scooter. I have yet to see any models do a credible job at that.
I would expect it happens more organically where people’s discussion of the benchmark and posted results find their way into the training data, like anything else on the internet.
What is this supposed to mean exactly? Do you think there are developers at OpenAI or Anthropic whose job is to train these state of the art models to draw pelicans riding bicycles? Like how exactly do you expect them to be doing that anyways? Hiring graphic artists to create SVGs of bike riding pelicans and feeding thousands of them into the model's training set?
Generate two images, get a vision model to judge, then RL reward the better one. And yes, OpenAI employees on twitter bragged about the model's SVG capabilities. So there are clearly people working there who care about it.
You gotta look up how RLHF works before you ask a demanding question like this.
There was a HN post that tested this hypothesis a few weeks ago: https://news.ycombinator.com/item?id=49010129
It stopped being a valuable benchmark proxy quite a few model versions ago. Simon knows it, so is everyone who’s serious about it
Treat it like a bit as is
It's actually not just valuable, but infinitely valuable. Because it's imaginary, there is no way to do it perfectly, and thus the possibilities are endless, and it tests out both the formulation and execution of ideas.