Comment by simonw
2 hours ago
I continue to do the test because I still learn something new from it every time.
This time, just seeing the difference between Haiku 4.5 (a year ago) and Haiku 5.5 (today - and 1/10th the cost) was worth it alone.
Same for Mistral the other day - the leap from Mistral Large 3 (their previous best model) to Mistral Large 4 was similar to the Haiku 4.5 to 5.5 jump.
I wrote some more thoughts about what value we can still get from the pelican test back in July - https://simonwillison.net/2026/Jul/16/kimi-k3/#what-can-we-l... but I've actually become MORE confident in its ongoing value since then. Using it to compare reasoning levels is proving particularly useful at the moment.
> because I still learn something new from it every time
Can you share in what ways these learnings affect your decisions or behaviors?
I get a good initial intuition about how much they are going to cost, how much the reasoning efforts affect their output (and their duration and cost), and how much they have changed since the previous release in the same model family.