Comment by simonw

8 hours ago

I finally got access. Here are the pelicans! https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

The "max" one at the bottom took 4 minutes 2 seconds and cost 63.206 cents.

For comparison, here those new Astra pelicans are in a grid with the GPT-5.6 pelicans: https://static.simonwillison.net/static/2026/gpt-6-and-5.6-p...

Interesting that the composition is stable across different runs. There is much more variety in other models.

  • Especially interesting that once Astra gets to high, xhigh, max it uses the same aesthetic.

I think the Gemini 3.8 Flash ones the other day were superior pelicans. Particularly the murderous one who was going to squish his tiny cousin.

Interesting that terra xhigh effort is cycling right left, in opposition to the prevailing pelican left to right direction.

The most interesting part to me is the bike. I don't think any of them would actually work, but the mistakes feel somewhat human.

Technically on par with gemini flash 3.8, but I give Astra more points for style, and for breaking from the pack by not adding the headgear, and the fish in a basket.

Even the "low" pelican is better than most other models.

  • It's subjective but I think only gpt-5.6-sol XHIGH is comparable to gpt-6-astra LOW in quality. And it's 24.11 cents vs 9.55 for astra low.

    • Something I've noticed especially with fable at work is it seems to be smarter, so it takes fewer tokens and burns less of my usage than smaller models would for similar tasks.

63 cents is quite cheap isn't it? If you compare vs Fable5.1 Max @ a whooping $3.30

The "max" pelican looks very serious, almost as if it's determined to win the race!

  • Yes! And do you get the impression, like I do, that model effort and pelican effort seem correlated?

    • Yes that's how the vectors work.

      This overall issue has resulted in hilarious missteps in the past, including grok going off the rails and claiming to be mechahitler.