← Back to context

Comment by d1l

4 hours ago

At work we use haiku 4.5 for a handful of latency sensitive tasks that are fairly simple. It performs well. Just started testing 5.5 as I’ve been anticipating a nice improvement since it was teased. Results so far are trash. Prompt leakage even. And it’s slower. I guess it’s cheap but I think they got the balance wrong on this.

Exact same thing here.

Both evals and Human pairwise tests for our use case are giving Haiku 4.5 first place in pretty much all tests.

No we'll try understand if we need to change our prompts to match performance ...

edit: maybe this will help: https://platform.claude.com/docs/en/build-with-claude/prompt...

  • Thx I’m digging into it but am somewhat comforted not to be alone in this. Turning up the level helps some but then we lose the speed. I don’t know that we’ll switch to 5.5 and may shop a different provider.

    Anecdotally we ran sonnet 4.6 for our more complex stuff and sonnet 5 was a LOT worse. 5.5 seems to have fixed it and we cut over our customer workloads. It’s strange, really.

Curious as to why. Haiku 4.5 has been far away from pareto frontier for a long time. Maybe you need to update your prompt for the newer model in your workflow.