← Back to context

Comment by vova_hn2

5 hours ago

Website looks very cool, Fable's octopus-organist looks very cute, but I feel like this benchmark (generate an SVG by a short and slightly ridiculous description) in general has been completely Goodharted [0].

I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.

[0] https://en.wikipedia.org/wiki/Goodhart%27s_law

“I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.”

This has been the plan since the start of all this, they regurgitate code in every better forms but they still aren’t inventing new things yet.

I don't think a bunch of similar tasks can really saturate the "create a SVG of X", because the model should have a quite good spatial understanding of the world and how everything interacts.

For example Gemini 3.8 Flash seems very impressive at first glance but the results are not actually very coherent, this shows that its "world model" is not particularly great (compare to SOTA models).