Comment by simonw
6 hours ago
https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - default reasoning level.
Here's reasoning level high: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
For some reason reasoning effort low and medium used similar numbers of tokens, and xhigh used less than high. I think I need to try without OpenRouter in the middle.
UPDATE: I tried again with the xAI API directly: https://tools.simonwillison.net/markdown-svg-renderer?url=ht... - not a great deal of difference between reasoning levels, and this time xhigh and low used the same number of reasoning tokens for some reason.
For comparison here's a fresh run against Grok 4.6: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
I think these are the worst I've seen, at least in some time. It's a silly benchmark though, not sure what to make of it
I think the result is fine. The benchmark is silly to the point of being useless nowadays.
It used to be a mess in various interesting ways. Now, almost every big release can draw something perfectly functional.
So the question - without a correct answer - given the prompt "Generate an SVG of a pelican riding a bicycle":
Does the user want the least lines of code to make it functional, or the best looking version?
The user at the very least expects the bike to have bike geometry; Grok seems to struggle with that
it's not truly tested until it plays a match or ten in Brood War imo
If you think this is bad, look up mistral.
Here are some somehow standardized pelican tests but for 3d scenes in threejs at threejseval.com
Grok 4.7 generations: https://threejseval.com/models/grok-4.7-high
Also go vote on https://threejseval.com so you can help evaluate how Grok and other model performs compared to each other!
I tried in Cursor and see a lot of improvements over Grok 4.6 svgs. The AA numbers indicate it's not very token efficient though: https://artificialanalysis.ai/agents/coding-agents?agents=co...
Poor fella doesn’t have a seat. Intriguing design where both pedals are on the same side of the frame. Balancing must be a challenge.
Are there good tools for doing context audits? I feel I have no good way to visualize what a new session is getting by default in a given repo without crawling through every potentially included markdown file
It’s a bad look to be evaluating models built in the furtherance of fascism and white supremacy. As this is a closed model, any support and training of it ultimately benefits those ends.
What is the default reasoning level?
Hahaha. I remember when musk and his merry band of sycophants were bragging about grok producing the only physically accurate bicycle. What happened?
Yeah, he had opinions: https://twitter.com/elonmusk/status/2023833496804839808