Comment by stingraycharles

1 hour ago

I'd turn that question around: why bother with the pelican tests at all? Simon explained why he started them in the first place:

"I chose that because a) I like pelicans and b) I'm pretty sure there aren't any pelican on a bicycle SVG files floating around (yet) that might have already been sucked into the training data."

That's no longer true. Every new model gets a blog post with its pelican SVG in it, and that ends up on the web like everything else. Haiku 5.5 and Opus 5.5 now say "this is the classic pelican benchmark" in their reasoning traces. People keep pointing this out and it keeps getting dismissed.

I, for one, don't see the point anymore. The reason for running the test is gone. I just don't get why we still treat the results as meaningful.

https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle

I continue to do the test because I still learn something new from it every time.

This time, just seeing the difference between Haiku 4.5 (a year ago) and Haiku 5.5 (today - and 1/10th the cost) was worth it alone.

Same for Mistral the other day - the leap from Mistral Large 3 (their previous best model) to Mistral Large 4 was similar to the Haiku 4.5 to 5.5 jump.

I wrote some more thoughts about what value we can still get from the pelican test back in July - https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l... but I've actually become MORE confident in its ongoing value since then. Using it to compare reasoning levels is proving particularly useful at the moment.