Comment by simonw

4 hours ago

The fact that they've heard of the test doesn't seem to help them draw a good picture of a pelican riding a bicycle.

That said... here's "Generate an SVG of an armadillo in fishnet tights jaywalking on Mars" on xhigh for comparison: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... (and here's the same thing from other models: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...)

Why even respond to these comments anymore?

Every new model you make this post, every time there’s someone who posits it might be trained on, and every time the answer is “maybe but probably not” it’s not a worthwhile conversation to have at this point, either the models can do some arbitrary thing or they can’t.

  • I'd turn that question around: why bother with the pelican tests at all? Simon explained why he started them in the first place:

    "I chose that because a) I like pelicans and b) I'm pretty sure there aren't any pelican on a bicycle SVG files floating around (yet) that might have already been sucked into the training data."

    That's no longer true. Every new model gets a blog post with its pelican SVG in it, and that ends up on the web like everything else. Haiku 5.5 and Opus 5.5 now say "this is the classic pelican benchmark" in their reasoning traces. People keep pointing this out and it keeps getting dismissed.

    I, for one, don't see the point anymore. The reason for running the test is gone. I just don't get why we still treat the results as meaningful.

    https://simonwillison.net/2024/Oct/25/pelicans-on-a-bicycle

Whoah!!! gemini/gemini-3.8-flash is what I've been hoping to eventually see with the pelicans!