Comment by lwansbrough
2 days ago
Perhaps this isn’t a new observation but the problem with LLMs is very clear with these. It’s a nice microcosm. The LLM will draw a fish companion (unprompted!) with a nice gradient but won’t get the pelican’s feet right.
It’s obviously a problem of fundamental understanding and demonstrates that reasoning is more “directionless rigour”.
Have you seen what happens when you ask 500 humans to draw a bicycle? No pelican, nobody riding it, just the bike.
https://foerstel.com/memory-bikes/
I'd like to see a human create a better pelican SVG without being able to look at the results. LLMs are language models, not (natively) vision models*. That's why it's a good benchmark: it engages LLM's logical reasoning in a way that we can check visually. The fact that they make errors which can be spotted visually, doesn't prove that LLMs are so far behind humans on language/logic tasks.
Besides which, "fundamental understanding" isn't a binary. Humans can also have a deep understanding of a subject and still make errors. (I do, anyways.) I do not think "fundamental understanding" is itself well enough understood to say definitively that LLMs do or do not have it.
*Some models have multimodal capabilities but these are separate weights, I believe it's only engaged in processing images. From the reasoning trace, the model never rendered the result -- only instructed the user to do so. I welcome correction if I'm wrong, I only have a surface-level understanding here.
AFAIK the only well-known LLM with no "conventional" multimodal encoders required for audio and images is Gemma 4 12B.
Hmm, doesn't seem to translate to an especially coherent pelican: https://xcancel.com/TeksEdge/status/2063108620842356970
Or a matter of having limited time but spending it thinking about the wrong thing. Which, not to anthropomorphize it, but it's not like humans don't do that, or at least I do.