← Back to context

Comment by Epitaque

5 hours ago

That benchmark might have some issues. You prompted the models to generate an image of striking a ring against a crucible. Then you (presumably, manually?) scored the images that depicted an anvil higher than the ones striking something resembling a crucible.

That’s a good catch. Yes, all scoring is done through manual review since relying on a VL model for these kinds of meta-metrics is a sort of loose equivalent of gödel's second incompleteness theorem.

I’ll have to think about this one. When I crafted the prompt, I wasn’t really thinking about the differences between a crucible and an anvil. It was more the visual of an archangel smelting halos for newly arrived heavenly beings.

  • I'm not sure why one would even strike metal against a crucible! It's a container for liquid metal. One of the outputs shows it being smashed by the manoeuvre, which is probably the most realistic outcome of all of them.

    Sorry, I'm not trying to nitpick. I'm just joining in because I'm interested in how the models dealt with the request.

    • Well this is HN - original home of the "ummm actually..." - so I appreciate when people pick all the nits. :)

      Even though I prompted for a crucible in the prompt, I think the fact that the prompt also contained terms like “blacksmith” and “hammer,” caused it to lean towards anvils over crucibles in some of the pictures (which as you brought up makes more sense anyway).