Comment by minimaxir

5 hours ago

Pricing is...a bit weird.

    Input
    $0.10 / MTok for prompts up to 100,000 tokens
    $0.50 / MTok for prompts over 100,000 tokens

    Output 
    $0.50 / MTok for prompts up to 100,000 tokens
    $2.50 / MTok for prompts over 100,000 tokens

100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.

In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])

Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.

For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.

These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.

In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.

  • I said this in another comment, but Artificial Analysis has the cost per task of Haiku on max roughly equal to that of Sol on medium, and the latter is significantly more intelligent. (And I'd wager that Sol probably finishes tasks more quickly, even with Haiku inference being faster.) So Haiku really only makes sense on lower reasoning levels, and only if you care about intelligence and speed more than you do about cost effectiveness (where Luna currently dominates). And that's without even bringing Chinese models into the mix.

  • noticeably smarter remains to be seen in practice. For now, Haiku is a bit more expensive than Luna on < 100k token, but I just don't have any agentic work below 100k, so this is going to be 5x more expensive than shown on these charts. It's hardly competitive ...

  • > it’s really impressive how much intelligence per dollar has grown in just a few short months.

    Open weights models giving a distant salute from afar

  • The benchmarks are very long form logic, knowledge, and coding tasks though. I'm very interested in Haiku 5.5's performance on ObviousBench where Luna 6 is currently SotA.

There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.

You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc

  • > ...this tokenizer, the same input text produces approximately 30% more tokens on Claude Haiku 5.5 than on Claude Haiku 4.5.

    So, it is might be even worse.

    • No, it's just Haiku 4.5 is so old that it predates the new Claude tokenizer change in Claude 4.7+

If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful.

I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.

  • >If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue.

    "less context is better and if you can't get stuff done with less yur bad" is the worst argument ever.

    it might be pure luxury to your eyes, but it's great to not require the use of a special custom harness that transcribes everything into emoji and compresses everything into barcode images.

    it's great to have a million token context to throw a large project into. If I need 100k just about any current gen consumer GPU in the world has very good models that I can self host for 100k context, limiting myself to 100k on someone elses machine seems to be missing a lot of the point unless the model itself is extraordinary.

It's actually existing flat per-token pricing that is weird.

Neither encode nor decode are linear in compute, so providers need to price for average expected length.

This is just getting closer to the true cost of generating tokens.

  • My theory here is that providers cover the non-constant costs of output tokens as context length caries using the cache input fees.

  • Flat per-token pricing is likely just logistically easier, particularly if these closed models are also picking up the kv cache efficiency improvements seen in recent open weight models.

  • Flat pricing is weird too but jumping up 5x at one cutoff is surprising in the other direction IMO

Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.

For this application 100K token input is plenty.

Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.

  • I think the 2.5 times cost but actually pays off in terms of intelligence compared to jev and the general capability of using it beyond classification

    • The classification performance remains to be seen, but presumably we'll soon start to see classification benchmarks.

      For other tasks like summaries (another suggested usage) it's good to see Haiku and Luna now competing against each other on cost.

      I'd love to know how the business automation market breaks down by volume of call type though - hard to imagine that decision making (e.g. branching, triage) isn't a very large part of it, greater than these other suggested Haiku use cases.

Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low".

Luna does as well, but just at a higher limit.

From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.

  • I think they are also trying to make sure Deepseek and other chinese models don't eat their lunch. They need something price competitive.

If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models.

You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget.

I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).

I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.

I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.

>> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents

Your vibes don't appear to be supported by facts. From the announcement:

>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.

encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens.

There are plenty of workflows like translations where you'd easily be under the cap.

Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?

  • That's not what any benchmarks that look at cost per task or similar says in terms of cost. The Chinese models, generally speaking, might be cheaper per token but need a lot more tokens to get there.

    • Except for the new MiMo V2.6 models, which appear to give some of the best value right now, at least on paper. (I haven't tried them so I can't speak from experience.)

  • Some people/organizations are ideologically opposed to using Chinese models. Not me, I use GLM-5.3-Flash for almost everything (the subscription-subsidized pricing on a legacy Z.ai plan makes it the best value model by a wide margin), along with some MiMo and DeepSeek. Still, I use Luna for certain tasks where speed is more valuable than performance; I can see this new Haiku displacing Luna for those. If you mean Haiku 4.5 though I agree, that model was a waste of time and money.

    • Luna is not really the fastest. You need to use it in high/max to get the good output for what it is good for: summarizing. And that is already close to two minutes per task...

    • I’m on the Legacy v2 plan and same: nothing comes close to 5.3 Flash’s value on it. It’s crazy, no wonder they discontinued them!

  • Isn't the point of this release that it's comparable?

    AAI Index // Input // Output

    Haiku 5.5: 43 // $0.10 // $0.50

    Mimo 2.6 Pro: 46 // $0.43 // $0.87

    Mimo 2.6 Flash: 38 // $0.10 // $0.28

    Seems competitive to me? Plus then I don't have to manage multiple providers

  • Where do you get this 10% number? Checking providers I know/respect, and GLM 5.3 flash is $0.15/m. Haiku is $0.10/m.

  • Presumably everyone who doesn't bother integrating a third party API key into their harness, which would probably be most of the Claude Code users.

  • On subscription pricing a $20 Anthropic subscription gives >$500 equivalent tokens, which is not so different, and you get smarter models. API pricing has decent margins.

    And Opus 5.5 is really good.

  • Well, unless you're using OpenCode Go, it's per-token costs (even if already super low), while Haiku falls under the Claude sub. It's just more straight forward and you aren't feeling a "loss" with the sub.