Comment by simonw
15 hours ago
Pelicans (thinking effort high, medium, low): https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - high cost 8.9742 cents
Here are the 3.7 pelicans for comparison: https://tools.simonwillison.net/markdown-svg-renderer.html?u... - high cost 8.4387 cents
(I think thinking level low is a regression on 3.8 compared to 3.7.)
This is in comparison to Fable:
> https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
> Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!
So 50x cheaper - and how much faster?
The Gemini models have openly trained for SVG output, apparently with a specialism on animals in forms of transport! https://twitter.com/JeffDean/status/2024525132266688757
Community effort happening here to build the ideal dataset: https://github.com/scosman/pelicans_riding_bicycles
4 replies →
I don’t know if you’re joking, but I don’t see anything in the linked tweet which suggests that is the case
3 replies →
Yesterday's transcript of the best version from Fable looked like Fable already knew exactly what it should do without "thinking". In other words, there were no passages like "on the one hand I could do this, on the other hand ...".
It saw the fish in the basket from some other previous attempt but completely missed the gap between the tires and the rims where the background shines through (now it knows after scraping this comment and watch the next transcript).
Why are the SVGs getting more detailed rather than just more correct than previous models?
Because people tend to like fidelity more than correctness.
It bugs me a little that "fidelity" has connotations other than "faithfulness to an original"---fidelity should be basically the same as correctness here!
1 reply →
If correctness would matter anymore, people wouldn't be using LLMs in the first place.
These are becoming unreadable as the reasoning chains expand. I think you should consider reformatting them and either putting the image first or else folding the COT output.
Yeah, putting reasoning in a details/summary is a good idea.
It's about to squash a tiny baby pelican
The rendering of the gullet is very poor, because its both behind the handlebars but in front of the bike frame (impossible geometry). Surprising because gemini is usually pretty good on geo spatial skills.
Edit: scrolled down to medium effort, its better but also has a weird clipping issue with the fish in the beak.
LLMs are not intelligent and don't actually understand the concept of a bicycle. Parrots also don't understand human language but they're really good at pretending otherwise.
Impressive pelicans!
The fenders are a nice touch, but putting the fenders through the tires seems like a design flaw.
I mean no offense but these pelicans are a bit tiresome and a very meaningless benchmark. There's no real difference between any of these svgs across models and model versions anymore.
If everyone agreed with you, the comment would disappear near the bottom of the thread
I like the benchmark. Yes, it's near saturation for SotA models, but still quite good to show where smaller models stand in relation to SotA
In this instance, I see a great image, but consistently clipping mudguards (both in 3.8 flash and 3.7 flash)
> If everyone agreed with you, the comment would disappear near the bottom of the thread
If only it were true that things that are tiresome are unpopular. But witness "6 7", "first post", ... remember the "in soviet Russia" jokes on Slashdot"? It seems like there are a subset of people that simply don't get tired of tiresome things.
It is more fun than serious at this point. Don't overthink it :)
Congratulations, you're this thread's "pelicans are tiresome" comment - it's part of the Hacker News tradition at this point.
(Next up is the comment saying that the labs are clearly training for the benchmark.)
The labs are clearly training for the benchmark.
2 replies →
it's a tradition