Comment by simonw
12 hours ago
Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.)
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
12 hours ago
Pelicans for 3.6 Flash and 3.5 Flash-Lite (Cyber isn't available to me through the API yet.)
https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
I am growing tired of these pelicans posts every time a new model is published. Feels to me like low effort personal brand promotion. Just sharing my 2 cents.
You and a few other people, but enough people still appreciate the bit that I'm going to keep doing it.
They're easy enough to skip - click the little "-" icon and you'll collapse the entire sub-thread.
+1. First thing I look for in a model announcement thread. I actually came across this one an hour ago and was sad there were no pelicans yet.
It's a decent heuristic because the better models generate better pelicans. That's all. Nobody sane is going to make a bet on a model based on a pelican. But it's cool, it's tradition by now, and it's a semblance of a good first impression for new models.
Imo it's a very stupid test, and should not be used for model performance.
But you bet my ass I check everytime to have a look at see how that pelican looks, it's just a fun check and also interesting to see the cost/results.
I love and appreciate you doing and sharing them, with stats and details.
thank you.
I love the pelican escapades.
[flagged]
10 replies →
But how else am I supposed to know when we've reached AGI, until I see an absolutely flawless pelican?
All of the pelicans so far have had really weird flaws / quirks so I am always a little interested to see how well these models perform at this task, since I've seen all the past pelicans and have some anchoring.
Seeing a truly flawless pelican would tell me that the model has true visual reasoning capabilities as well as good taste.
I think you're underweighting the Pelican test.
Not only does it give you a super easy-to-grok understanding of the model quality just by looking at the image, but when you compare tokens and costs (both input and output), you really get a good, simple COST x QUALITY evaluation across models.
Simon explains it well: https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l...
Simon, you should put up a summary table page that you update after every release.
I feel the same way. It was fun at first but has gotten tiresome. Does anyone actually use these models to generate SVGs?
Yes? And even for simple interactions, they use SVG by default.
Yeah I don’t get it. It tells me which model can draw an svg of a pelican riding a bicycle. It does a great job at that and the presentation is good.
But why is this an indication of literally anything else?
1 reply →
Yes
At this point, it's kind of a hackernews thing. Simon posts them as a single comment in the relevant thread. It's okay for this place to have a little bit of a sense of community, and you can just ignore the comment.
Its a nice benchmark. Like hearing the ice cream truck on a summer day.
It's both.
I agree to rednb that at this point it feels like rather obvious brand building, but also, I agree with you that some value is in it.
It does not feel all that authentic though, and it's good to react allergically to lack of authenticity. Bad for a lot of business models, but good for humanity.
2 replies →
More like living next to an ice cream truck car park
Every parent groans haha
At this point it does not show anything as models are fine tuned on all kinds of benchmarks.
1 reply →
Yeah. It's something I can do myself in a couple seconds if I want, also on more varied SVG scenes. If this is going to be a benchmark people turn to I'd like to see more effort put into it than just a one-sentence prompt.
Here is a different opinion. I’m always looking forward to see the pelican whenever a new model is released. It’s plain simple to understand, memorable, subtle enough in terms of details, and my favorite part is that you have been doing them consistently long enough for it to be useful for comparing almost everything with anything.
Disagree. They're a nice tradition, but besides that, they're a useful way of eyeballing improvements. I realise labs are likely to be training for Pelicans - but if they're all training for them, the differences in the results are as indicative as they were before labs trained for them.
The 3.6 Flash pelican is just about the best I've seen.
My 2 cents: you don’t have to look at the pelican if you don’t want to.
All the models do this well. It's a test that tell us nothing at this point.
Hard disagree. I love a little bit of whimsy (which I feel the world is lacking more and more everyday) from Simon everytime a new model is announced.
Perhaps freshen it up and extend the test by feeding the model the rendered output so it can iterate once. Assuming a multi-modal model.
I'm happy for Simon to post what he wants, when he wants. He's earned it.
I love them. Keep 'em coming.
Vibe code an extension that autocollapses any post mentioning pelicans and by simonw?
I've been close to writing one that will automatically upvote ALL downvoted posts. I'd call it something like Anti-echochamber.HN
Do something instead of complain
Your comment reads very pedantic with a hint of jealousy. The pelican and xbox controllers are great ways to see how well it can follow direction dealing with svg a difficult format for LLMs to use and testing their spatial vision awareness.
It's just how he is. Prior to LLMs he was cramming a datasette link into every thread. Downvote and move on.
I find Simon's work informative and entertaining; the last thing he can be accused of is low effort. The Pelicans are just a bit of fun icing on top.
Sending a one sentence prompt to an LLM and posting it to hackernews constantly isn’t low effort? Today I learnt something new
I'll split the difference. When it's a blog post there's usually an interesting observation or two, but if it's totally automated? Maybe just do the ones with a post.
I like seeing the pelicans, it's a tradition.
Yeah and it's surely in the training data by now. Long past time to stop.
You say that, and yet 3.5 Flash-Lite produced an SVG without a pelican.
[flagged]
A new model arrives. The pelican, with uncanny commercial instinct, is never far behind.
Sponsored blogs and paid newsletters are after all, notoriously poor at subsisting on silence :)
Linking directly to the rendered markdown as opposed to a post on my blog is a poor way to promote my blog.
A piece of the frame is missing between pedals and back wheel. The frame of the bike passes through the bird. It also puts a cap on the bird's head, and a fish in it's mouth.
The fish and the cap where always added when I asked an llm to improve it's first attempt.
This continues the trend in LLM progress of better=more stuff
Edit: I wonder if this is a function of the reasoning training, where more tokens/ stuff is rewarded.
I generated a very stylish Pelican using the webapp. Hard to put a judgement on it relative to yours https://share.gemini.google/XSfmve2mEGDV
3.1 Flash Lite has a better pelican that 3.6?
You should be banned for your constant spamming of this. Its ridiculous. Every single AI post! Constant personal promotion.
Flash-lite did the John Cena Pelican
[dead]