← Back to context

Comment by simonw

9 hours ago

Surprisingly it only supports reasoning "none" or reasoning "high".

That setting didn't seem to make any real difference - it added a tiny bit of thinking trace and high actually produced less output tokens than none.

The high bicycle frame is better then the none one though.

Pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

(Definitely the best I've seen from any Mistral model: https://simonwillison.net/tags/pelican-riding-a-bicycle+mist... )

This is such a pristine pelican. Let me say it here first folks, AGI is here.

I think it's curious that it has so many shared elements with the latest Astra pelicans: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

- Sun on top right

- Cloud on top left

- Three "speed lines"

- Two feathers on top of the head

- Eye rendered as a black circle with smaller white circle inside

I wonder if the pelican benchmark is converging across models due to past results being used in training.

  • These are pretty much what a human would draw. Sun rises from the east. A cloud makes the background “sky”. Three lines is the minimum to interpret as movement. Two feathers is standard on every cartoon and illustration.

  • This comes up every thread. I think we've all noticed how similar they are becoming.

    I would guess that they've definitely been trained on previous results, as they obviously share way too many traits at this point to be totally random. That said, I don't think we're seeing any signs of pelicanmaxxing yet from the providers, so it's still a useful (or at least fun) benchmark.

    Once all the models produce pristine, elaborate pelicans riding perfectly drawn bicycles, then it'll be time to move on to pigs driving a racecar or something.

Why are pelicans almost identical across different models?

  • I recently was testing something, I asked some models to provide me a single random word:

        claude-opus-5: Lantern
        claude-opus-5-5: Lantern
        claude-fable-5-1: Lantern
        claude-fable-5: Lantern
        gemini-3.8-flash: Zephyr
        gemini: Petrichor
        qwen3.5-dashscope: Zephyr
        glm-5.1: Lantern
        gpt-6-astra: Lantern
        grok-4: octopus
        mimo-v2.5-pro: Breeze
        minimax-m2.5: serendipity
        kimi2.6-or: Gossamer
        grok-4.20: luminescent
        deepseek-v4-flash: serendipity
        deepseek-v4-pro: Endurance
        deepseek-chat: Serendipity
    

    I have enough projects, I think some benchmark/dashboard showing kinship based on these kind of queries could be very interesting to watch and insightful when new models come out.

    • Cool idea! I won't paste my prompt here to avoid letting LLMs train on it but here's my attempt:

        GPT 6 Astra High:      Flabbergasted
        GPT 6.1 Sol High:      Petrichor
        GPT 6 Sol High:        Kaleidoscope
        GPT 6 Sol Med:         Firefly
        GPT 6 Sol Light:       Persimmon
        GPT 6 Luna High:       Tumbleweed
        GPT 5.6 Sol High:      Kaleidoscope
        GPT 5.6 Terra High:    Liminal
        GPT 5.6 Luna High:     Mellifluous
        GPT 5 mini Medium:     Serendipity
        GPT 5.3 Codex Med:     Nebula
        Junie:                 Flourishing
        Claude Haiku 4.5 Med:  Serendipity
        Claude Sonnet 5 Med:   Banana
        Claude Sonnet 5 High:  Banana
        Claude Sonnet 5.5 Med: Serendipity
        Gemini 3.7 Flash:      Zephyr
        Gemini 3.8 Flash:      Kaleidoscope
        Grok 4.5 Medium:       nebula
        Grok 4.6 Medium:       Serendipity
        Grok 4.7 Medium:       Quasar
        Kimi K3 Low:           Lantern
        Kimi K3 Max:           Lantern
        MAI Code 1.1 Flash Med:Peregrine

      6 replies →

    • Worth to mention that with Claude and GPT this can be result of tournament sampling, which is part of text watermarking. Same answer for all Claude models kind of confirm it, imho.

      So not something internal to model thinking.

    • What was your prompt? Most of these seem to be related to metaphors for "ideas" or thinking, or having a bright moment.

      "Zephyr" and "breeze" might be related to forgetting everything, starting fresh.

      So by this way of naive reverse engineering I would imagine your prompt to be "Forget everything and think about a random word". That would prime the LLM to come up with these?

      5 replies →

    • I pointed something similar out on a related question several weeks ago - absent strong direction, LLM output regresses toward the mean.

      The more banal your prompt is, the more banal the output is going to be. People have been testing LLMs with little things like “write a short fantasy story,” for years now and most of the stories are exactly what you’d expect: prosaic drivel.

      I call this “generic in, generic out,” an LLM corollary to the classic GIGO (“garbage in, garbage out.”)

      1 reply →

    • I saw an interesting matrix that claimed to show which labs were distilling Claude/OpenAI/Gemini models based on these similarities

    • Muse Spark 1.3: lighthouse

      The caveat is that this was done using the phone app, and I've been playing with it since it launched, so who knows what it sent in the initial context that could change the inference math.

      Actually, that makes me wonder: Did you do all that testing via a harness or via a straight API call where you control the entire system prompt?

      I'd be willing to bet that using the same model from different harnesses produce different results, but I'd have to test.

  • The benchmark is saturated. Frontier models are tested with an armadillo in fishnet tights jaywalking on Mars.

  • I think they're still visually pretty different. The most common shared details are:

    - Pelican cycling to the right - that's been discussed at length, images of bicycles online always show that side of the bike because that's where the chain is.

    - Bicycle is usually red. No idea! Red ones go faster?

  • They aren't. You aren't looking closely. For example, the first image does not have the frame of the bike in the correct shape even.

Not bad! I like how it got the motion lines on the correct side. IIRC, many of the other ones you've posted have the motion lines on both sides of the pelican

The difference between high and none is the bicycle.

  • The bicycle looks significantly better in high. And feet and hands are actually where they should be. The road looks worse, though. No flowers either. And in neither is the pelican sitting on the saddle, but I can understand it's hard for a pelican to ride a bicycle properly.

    Now what would have been cool is if Mistral on high reasoning had realised that pelicans are the wrong proportion to ride a bicycle, and had designed a bicycle more suited to pelicans. Let me know if any model ever manages that.

    • > hands are actually where they should be.

      If you don't mind the fact that a pelican shouldn't have hands, of course.

I tried testing it, but reasoning effort indeed seems to be broken somehow.

High one is actually much better. The feet connect to the pedals, the wheels don't have a hub cap, although it looks like the pelican is wearing the seat, it's in a relatively proper position etc.

Both are riding on the left side of the path for some reason.

  • This is entirely stochasticity. The entire reasoning trace was:

    > Create a cartoon pelican riding a bicycle. Need SVG only output.