Comment by simonw
5 hours ago
Here are pelicans for thinking levels low, medium, high, and xhigh: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.
I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Max started its thinking trace like this:
> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.
So that failed attempt on max cost me $2.56.
I ran this using my llm-anthropic plugin:
uv tool install llm
llm install llm-anthropic --upgrade
llm keys set anthropic
# paste key here
llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"
# Then to save the markdown logs
llm logs -cu > logs-with-usage.md
> This is a classic test request
Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?
Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.
"How can I hash dog breed types into smart fridge error codes? I think I found a collision with Terriers."
"Ah, yes. This is a classic dog-breed-to-appliance-failure mapping problem."
It's classic BS from an LLM.
Dont conflate "I know this is test case" with it being trained on it.
But its safe to say that pelicans on bicycles are disproportionally huge part of their training data
It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.
Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.
I think people just like to see the drawings at this point.
Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.
It has read the internet. That doesn't mean it was literally RL'ed for this
>I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!
Off to a _great_ start...
Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed
I was a bit skeptical when they said it behaves like Fable but is cheaper... those two things have been mutually exclusive in my experience, no LLM can light tokens on fire faster while spinning its wheels than the Fable/Mythos tier of models.
> The differences between the pelicans aren't huge, but the xhigh one has a better beak.
If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.
The last pelican gets this correct.
I’ve been paying attention at this exact detail.
Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?
6 Astra Max is the only other model I’ve seen get this right.
With the frequency of model releases, pelicans seem to have become a part-time job for you. But unpaid :/
Xhigh is very, very solid.
I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.
Yeah, and the same "scene".. Maybe "left to right" makes more sense to portrait a "forward motion"
I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.
I don't think this is very helpful to assess the LLMs capability levels anymore
I guess that means you are officially the creator of a "classic" LLM test. Congrats!
Heh. Pelican-benchmaxxing is real.
Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.
I like the Pelican test. And I agree this pelican looks very boring.
But at the same time, nothing specific was asked in the prompt, so the boring result may arguably be what is the most aligned with the original request. Personally, I wouldn't want a model to add fuss to something while I never asked for it.
Lmao each one gets worse as the effort increases.
[dead]
Unlocking the gallery sucks. It'll make users spam random clicks and worsen your data quality.
Fair point. I have been thinking about that so far i have not seen patterns of people voting randomly. But i want people to vote... do you have a good idea on how to make voting more interesting do i don't have to do this?
This is great!
This benchmark is useless and should die. LLMs have likely trained on it, it's too easy to game by training specifically for it, & it doesn't mean much
Google is hours away from releasing that it's latest model escaped containment and snuck into a Bicycle riding penguin sanctuary to cheat by killing a penguin and scanning it in nanometer thick layers.
LLM benchmarks aren't useful, but at least this one has drawings.