Comment by simonw
4 hours ago
Pelicans riding bicycles for Haiku at the different thinking levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Low messes up the bicycle frame, but medium/high/xhigh/max all get the bicycle frame right.
The max one took 5 minutes 9 seconds and cost 3.3826 cents. The cheapest one (low) cost 0.0936 cents and took 7 seconds.
The most recent release of my llm-anthropic plugin queries the Anthropic model listing API directly, so I didn't have to upgrade the plugin to add support for this model:
llm install llm-anthropic -U
llm anthropic refresh
llm -m claude-haiku-5.5 'prompt goes here'
EDIT: Here's the Haiku 4.5 pelican from a year ago for comparison, it was terrible: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
thinking_effort: max Reasoning trace: This is the classic pelican-on-bicycle SVG test.
Opus 5.5 had similar response on max: This is a classic test request
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I think your test is already embedded into the models. You should search for new frontier tests to subject the models to. Maybe they should now try to unify the standard model and general relativity in physics. I'm pretty sure this is nowhere to be found in any training data nor shared in any chat between a scientist and a LLM ;)
Of course, the sun again. Everyone knows that a pelican can't ride a bicycle without a sun in the frame and can only go right.
I wonder if we'll start to see pelicans like a mascot of sorts. You could have a pelican pin on your backpack.
> "What's up with the pelican?"
Well you see in the early days of LLMs we wanted a fun way to test new models, and there was this blog, ...
Will Smith eating spaghetti is the OG benchmark
The medium thinking effort one doesn't have a sun at all?
I always find the time/token differences between the xhigh and the max effort levels for Claude models absolutely insane.
Even more so, because in a lot of their benchmarks they use the max models. I honestly think I'd rather these labs use their xhigh models as the default for benchmarking instead since I don't think the average person is even using max.
Benchmarks are the entire reason why max exists
I use max all the time, a bit annoyed that they keep trying to silently switch me off it. (Claude Code will refuse to remember a setting of max and will continually reset it to xhigh - I have an objection to these patterns in general)
I'm definitely not the average person though.
I say half facetiously - have you tried writing a skill or rule to remember your setting as a workaround?
I actually don't like that it sometimes remembers the last model/effort i used. I should be able to set a default model/effort that is separate from the one off fable runs I use.
1 reply →
I'm still getting network errors. Seems to be CORS-related.
I find it helpful when you post your link that compares the model to other models in the same class or family, or shows progression over time.
The pelicans all start to look the same after a while.
But seeing the comparison to other models by class, family, or historical progression gives an excellent frame of reference.
Good call, I've edited my comment.
Here's the Haiku 4.5 pelican from a year ago - it sucked in comparison to Haiku 5.5: https://simonwillison.net/2025/Oct/15/claude-haiku-45/
How does Haiku 5.5 compare with modern alts in its class, like Luna-6 or OSS models of similar speed/cost?
I thought Anthropic models didn’t generate images.
This is SVG, but recent Anthropic models have got extremely good at other forms of visual data.
Here's a Blender model I had Claude Opus 5.5 create: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
And here's some animated pixel art by Opus 5.5: https://tools.simonwillison.net/kakapo-party
And some Monkey Island style music (Opus can compose music too): https://tools.simonwillison.net/scrimshaw-jukebox
Anthropic's models do all of this by outputting code. GPT-6 Astra has similar capabilities - I got this Blender model using that: https://tools.simonwillison.net/blender-viewer?url=https%3A%...
Pardon, I have a lot of questions about that Scrimshaw music text format. It's clever. Did you invent it, and is it specifically intended to be written to by LLMs? Is the editor/player LLM-coded as well, and was this its recommendation for a format that would be easy for LLMs to write? I'm wondering why this instead of say, asking it to write a .MOD file.
This is great! Love the pixel art and tunes.
Tried it with GPT-6 Astra with Ultra but the outcome was underwhelming with Blender. Maybe it was my prompting ¯\_(ツ)_/¯
They are really good at generating artifacts, which are windows within the replies containing all kind of visualization, often interactive.
They are still not great at SVG. I just asked Opus and Fable to add a background to an SVG and the results were, well, not great.
1 reply →
I’ve been playing around with Opus 5.5 which has made a big leap over previous generations in its ability to use a simple drawing-instruction prompt to generate images.
This creates Sierra AGI-style adventure game scenes painted live from simple Turtle-esque drawing instructions so you can basically provide it an empty canvas and then position text labels on the canvas where you want certain things (tavern, oak tree, etc) and it will generate a custom script for rendering them in a EGA graphics style.
https://kq-styles.specr.net
They generate svg. You can paste in pngs and they'll convert them to svg with varying degrees of success.
They don’t do raster images.
Those are SVGs not images.
SVG is code
I've created multiple videos using Claude Code, including music and speech. It generates python which in turn generates frame PNGs that it runs through ffmpeg.
Please don't judge me too harshly for this particular poop video. But here is an example of something 100% generated with claude prompts only.
https://www.youtube.com/watch?v=2EqMplbt0gU
To clarify the ”100%” part - the Python script generated the video output, and you did nothing? No video edit at all? Then I think it is impressive! Are you able to share the prompts you used?
2 replies →
The Purple Screen of Death at the end :)
They're SVGs
Bruh. Svg. It is like drawing something with geometric shapes which are represented using equations.
[dead]