Comment by jjcm

5 hours ago

Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not.

Haiku 5.5: https://html.non.io/lcars-haiku-5.5/

Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5

Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...

One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.

Pac-Man Bench:

Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.

TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5

All results: https://jonclegg.github.io/pacman-bakeoff/

  • Interesting that you have gpt-6-luna at $0.01 vs. claude-haiku-5-5 at $0.16 for this task. I see the score disparity though and I played them briefly. My takeaway from this is that the choice between Luna and Haiku 5.5 may remain nuanced. Luna may be a lot cheaper still and good enough for some jobs. Is that your read of the results?

    • Actually, I misspoke. At least as far as Pac-Man Bench, Luna does about as good of a job. The ghost logic's not quite as good, but it also makes a map that doesn't have nonsensical sections in it. So maybe call it a wash.

      1 reply →

  • Something is not right there. DSv4.1 flash shows $1.89 for tens of thousands of tokens? What am I missing?

> likely best used for small subagent tasks / tightly scoped work.

Hasn't this always been the case with Haiku?

To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it.

  • Sure, but presumably Haiku was distilled from the same training data. Part of this is seeing how much the capabilities degrade as their model size goes down.

Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.

  • It's not really about whether the design looks good. It's about if the model can take the design given to it and replicate it in code. Opus 5.5 matches the designs almost to the pixel. Haiku built something else entirely.

    • I guess I'm giving GP feedback about their product diffui.ai, not really about Opus' performance.

  • Totally fair, but I'd encourage you not to look at the design so much as the task. This was a design that's part of a benchmark test suite specifically for image->html conversion. The dense visual noise / complexity / flowing svg shapes are things that most LLMs have trouble with.

    It's meant to be a good test, not a good design.