← Back to context

Comment by simonw

4 hours ago

Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.

The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.

The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:

  llm install llm-anthropic -U                                
  llm anthropic refresh
  llm -m claude-haiku-5.5 'prompt goes here'

EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/

thinking_effort: max Reasoning trace: This is the classic pelican-on-bicycle SVG test.

Opus 5.5 had similar response on max: This is a classic test request

https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;)

Of course, the sun again. Everyone knows that a pelican can't ride a bicycle without a sun in the frame and can only go right.

  • I wonder if we'll start to see pelicans like a mascot of sorts. You could have a pelican pin on your backpack.

    > "What's up with the pelican?"

    Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...

I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane.

Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.

  • I use max all the time, a bit annoyed that they keep trying to silently switch me off it. (Claude Code will refuse to remember a setting of max and will continually reset it to xhigh - I have an objection to these patterns in general)

    I'm definitely not the average person though.

    • I say half facetiously - have you tried writing a skill or rule to remember your setting as a workaround?

      I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.

      1 reply →

I find it helpful when you post your link that compares the model to other models in the same class or family, or shows progression over time.

The pelicans all start to look the same after a while.

But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.

I thought Anthropic models didn’t generate images.