← Back to context

Comment by embedding-shape

6 hours ago

At the bottom it says "Score 98.58", what measure is used for this score? It's kind of horrible, the perspective is all off (legs of the table makes that very obvious), the mouse/hamster has two mouths, a stub for a right paw, looks like left hand holds a melon on a stick or something, and there are pluses in the background for some reason. Not sure it'd call it "close to perfect" which the score seems to want to indicate.

Do you prefer the fable one?

It's more "correct" but looks a lot worse in my opinion:

https://aibenchy.com/compare/openai-gpt-6-astra-high/google-...

Yeah, that's confusing, the score is for the entire benchmark, not for SVG generation only.

Good point about the mouths, I just noticed, lol

Imo, it's still better than most models, I personally like the stylized perspective.

You can view here all generations for all models: https://aibenchy.com/showcase/

I've replaced "Score" there with model ranking, to reduce confusion, thanks for the feedback!