Claude Opus 5.5

5 hours ago (anthropic.com)

> Claude Opus 5.5 is our first release since we called for pacing the frontier.

Interesting how the very first line is used to remind the reader of their call to pace the frontier just last week, and everything else after that line is to demonstrate with very specific numbers how they absolutely are not pacing.

  • “Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself

    • I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.

      Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.

      88 replies →

    • It's interesting to compare Anthropic's language with what's coming out of the US military lately. They explicitly refer to China as the "pacing threat", meaning that there is a risk of China achieving superior military capabilities and thus we need to press forward with an arms race (including militarized LLMs) as fast as possible.

      https://www.war.gov/News/News-Stories/Article/Article/264106...

    • I feel like because something about "pacing" a kind of noun like "the frontier", doesn't actually make literal sense, right? You could say "Pace the speed of advancement of the frontier [of most sophiticated AI]", and that's perfectly sensible, but of course that's not as catchy.

      It is clear what it means anyway, that's true, it means the left out words, more or less.

      And I still find reading these grammatically weird but super catchy slogan-like statements to be really annoying and taxing. People _did_ write and talk like this before LLMs of course -- the LLMs learned it from somewhere -- and it was annoying and taxing to me before too. But the LLMs really specialize in it, and it's everywhere now.

      Of course, the more LLM slop we read -- and so much of what we read on the internet and social media of any kind is this now -- the more humans are going to start writing/talking like LLMs. What you read affects how you write of course.

    • What specifically is smarmy and weird? Sounds like your own hot take with zero analysis.

      They’re limiting frontier model development speed. Others are too. Pacing is the only word here to criticize, and I think it’s fine given the limiting of speed but also increased oversight. I’m not saying they’re fully doing this, but the term is fine.

      Do you have a better proposed phrase?

      4 replies →

    • > “Pacing the frontier” sounds smarmy and weird, like the phrase was generated by Claude itself

      Claude says it sounds fine. And Claude is now the judge of the English language style, not you.

    • "While I was pacing the frontier, I discovered some load-bearing fence posts that I should have surfaced earlier."

    • Agreed when I first heard the phrase it sounded odd. Maybe they thought it subtly conveyed they would be setting the pace... But again this is something AI would come up with in its awkwardly post hoc sort of way.

      Though tbf corporate-speak and AI-slop are both insufferable in similar ways...

  • Isn't the idea that they claim to be willing to slow down if everyone does (ie governments force everyone to), but otherwise they won't slow down because they still think they'll make the best choices with superintelligence if they get there first? That's my understanding of what all the major labs claim to believe anyway.

    • It’s just another cry for regulatory capture which imo is the only way they stay afloat at their current direction.

      China is literally only a single step behind and willing to drop free models just to undercut the US companies.

      I’m for it because I don’t want another massive Google or Meta.

  • It seems like they are. I mean, this is similar in performance to Fable (ish). It seems like more focus on making existing capabilities more accessible.

  • When my mom told me to pace myself, it usually meant that I was going too slow.

    This term is quite ambiguous. Did Dario mean that they need to go faster while making it sounds like they will slow down???

  • Pacing the frontier. They might as well have said they are stifling innovation. That depends on your interpretation, of course…but interpretation depends on their intent, which I feel is disingenuous.

  • "pacing the frontier" is code for downgrading your expectations of AGI.

    • Or, in parallel, the research showing LLMs intelligence will operate on an S-curve, eventually hitting a long valuation deflating plateau, is dead on the money and The Big Guys are trying to delay that inevitability for as many quarters as possible…

      4 replies →

  • This IS pacing. Nobody said pacing would mean slow.

    Pacing is very explicitly about RSI and similar training methods that will accelerate progress beyond our ability to comprehend it.

  • They are advertising the regulations they want to enforce in the following sentence, which makes their intentions explicit (ie apply those things made to suit us to our competitors).

  • the only thing they are pacing is what models the permanent underclass are allowed to have in life.

    that, they fully intend to 'pace'.

    • With AI tools it's easier than ever to create a business, do research, or build stuff. That's an opportunity for the "underclass," not a curse.

      3 replies →

    • Either AI will capture so much value that there will be permanent underclass (and in this case it's extremely capable and extremely dangerous and should be heavily regulated) or it won't be capable enough to displace people into permanent underclass.

    • Or they’re running into a steep diminishing return slope on R&D vs performance and are using stewardship as a cover.

  • This makes perfect sense. There are no real improvements anymore (just benchmaxxing), and they explain it by "pacing the frontier".

  • In any professional sport, anyone can lobby the rules committee for a rules change and hope for the best. Meanwhile if you can't get one, you play the game by the existing set of rules, and you play to win.

    • Except that this isn't sports and they tell us that all humans will die if they continue like they do.

      If their scare was honest, they would stop.

      3 replies →

  • What an amazing excuse for lower than expected performance! Our models are slow because we're so ethical.

    • Thanks for spelling out what the original comment was implying.

      I find it bizarre how intensely a bunch of these child/grandchild comments are criticizing the notion that people would even think to analyze the meaning behind the words.

      Hacker News has always had a unique culture in which thoughtful discussion is basically the main goal, and it's intentionally incentivized in numerous ways. It's been my experience that any thoughts added to a post's conversation are seen as valuable as long as they are thoughtful and seeking to understand.

      So these comments are clearly coming from a place that's antithetical to HN's culture. What that in mind, it seems likely to me (Occam's Razor) that these comments are either:

      1. Astroturfing: Claude employees acting like everyday folks, secretly trying to shift public opinion.

      2. AI cult mindset: "AI is humanity's salvation; how dare you have perspectives outside of those accepted by the cult."

      Am I missing another likely option?

      To bolster my point, right now we're posting on the top top-level comment, meaning a majority of active HN users find it to be a great addition to the conversation. Commenting to shut down the discussion is a red flag.

  • I think a lot of outsiders interpret “pace the frontier” as slowing down, whereas the labs see AI improvements on track to accelerate dramatically and intend pacing as slowing the acceleration in capability improvements, rather than slowing down altogether.

    • I think the market didn’t react like they expected and now they’re like “jk”.

      And I don’t think any pacing is/was intentional. They’de release skynet if they could and the stonks went up

    • I’m an insider, 10 years working on LLM. 1/6th of the cost for roughly the same performance as fable 5.1, 21 days after fable 5.1 came out, this does not feel like the derivative is flattening.

  • I thought we were all aware that 'pacing the frontier' was a marketing slogan, and the actual intent here was to suspend antitrust laws.

  • I'm still kind of convinced that these "calls for pacing" are just a way to try and flex to their shareholders.

    "Our technology is so unbelievably powerful that the entire world might shatter if we don't have government imposed handcuffs!!!!". It just reads like the corporate equivalent of the drunk frat guy saying "HOLD ME BACK BRO!"

  • i think we shouldn't mix things here.

    Opus 5.5 isn't the frontier, when they say 'pacing the frontier', it's about internal models not yet released, as they're probably one or two generations ahead already.

  • Seems like a bigger focus on efficiency (both cost and speed) and the "tone" of Claude vs benchmarkmaxxing

  • When they talk about pacing, they are referring to their dangerous competitors, particularly those evil open-sores and Chinese ones, not their lovely safe models because you can trust them to look after your interests.

    What the big players are trying with the current calls to slow things down, is the standard capitalism practise of trying to engineer regulatory capture. TBH I'm surprised those calls are coming so soon - they must be really worried about running out of what little moat that they have.

  • "The improvements are underwhelming for a model that, according to previous claims, should have replaced all knowledge work twice by now, but you know, it's just because we're pacing the frontier."

  • If they were really about putting brakes on these, they would simply make these things non-agentic.

    Simply make them something that derives a text response from its training data.

  • kinda disingenuous. They include a whole section on pacing later on

    • You mean the section where they tell us this model is not affected by pacing because “they understand it well” and they will share more details on pacing later? Yea not very convinced by this effort.

  • "We made the incredibly tough decision to slow down development. Then after 4 days of twiddling our thumbs ... we present Opus 5.5"

  • It's laughable at this point. It feels like they're drumming up all this fear about imminent AI threats to emphasize the need to slow down, when in reality, the model progress seems already to be slowing down and has shifted to compute allocation (i.e. "how much compute do you want to throw at this prompt?"). All while continuing to tout benchmark records with each new release.

    • This. I swear the fear mongering is all about investor signaling and regulatory capture. It's so disgusting that anyone believes it.

      The only good news is that these models are genuinely helpful and we have competition at least between 2 companies.

  • > we called for pacing the frontier.

    Translation: our models are getting shittier each iteration and we ran out of ideas. Let's invent scary stories and hope investors will lap it up.

    Idiotic.

Finally that price drop

   Prices per 1M tokens     Claude Opus 5.5    Claude Opus 5
   Cache reads              $0.20              $0.50
   Input tokens             $4                 $5
   Output tokens            $20                $25
   Cache writes             $5                 $6.25

Opus 5 is the model with highest spend on openrouter (https://openrouter.ai/rankings#task-spend) and it seems plausible that Opus 5 is/was the highest spend model in the world, and certainly Anthropic's biggest moneymaker.

If you are forced to reduce price despite raising capabilities, that certainly tells something about the market, and potentially about Anthropic future profitability too, since this model is their biggest topline contributor

  • Claude adapts to OpenAI’s surprising move to simply deliver better performance than Fable 5.1, better tools as well as featuring very low pricing.

    Fable 5.1 literally was a money grabber. While I liked the results, tokens were burned so hard it was embarrassing, while Astra seemed to not care.

    Also Claude makes it very hard to pay for additional token budgets, allowing only credit cards. I don’t use mine anymore since I don’t need it in everyday life I was dumbfounded.

    So Anthropic is just copying OpenAI so to say, matching them and essentially with Opus 5.5 being Fable 5.1 in disguise, all they do is reduce costs.

    Competition works.

  • People are paying for Opus 5? Not just burning down tokens left after they enjoyed Fable on the sub? Amazing.

  • We haven’t been able to use opus as much as we’d want because it’s been too expensive for general use, price drop is good so I can stop juggling different models and just use this daily unless it has some weird new issues

    • Speaking for myself, I have not been able to use Opus as much as I’d want because its verbose prose makes human reviews of its assumptions, architecture proposals etc. more painful than its predecessors. If they’ve solved that, I’ll be accelerating through my backlog that much faster, and using tokens accordingly.

  • All the anti-AI people constantly say that any moment now the prices will skyrocket and in the end human work will be cheaper compared to using AI.

    It doesn't look like that's happening, on the contrary the prices are falling especially when taking into account capabilities.

    • It's the Chinese open source models. They're barely behind the frontier, making AI a commodity, forcing openAI and Anthropic's margins downward.

      I'm hardly a fan of China/Xi, but I do appreciate and benefit from this.

      1 reply →

  • >and potentially about Anthropic future profitability too

    have they ever shared anything about their revenue mix between consumer plans vs per-token billing? this is a revenue cut on their API billing, but they're not saying anything about increased limits on the plans. so all the plan revenue just got more profitable.

  • Footnote on their pricing page says:

    > Cache hits and refreshes on Claude Opus 5.5 are priced at 0.05x the base input price.

    If they do the same for Haiku and Sonnet 5.5 then we should also see 5c/mtok and 10c/mtok cache read for those models, respectively. Still too high for Haiku IMO, Luna is 2c/mtok.

    • Since something like 98% of tokens are cache hits, that's a pretty substantial drop from 0.1x

  • 60% cache read is cool, but subscription only get 25% more according to Cat. I'm confused

    • I think gpt 5.6 family also dropped pricing but didn't give any more usage for the subscriptions. Maybe it's a way to silently lower the value given to subscriptions while keeping API pricing competitive

    • 25% more usage sounds about right given the other token costs are down about 20%? I don't think cache read is a big portion of the overall cost.

      4 replies →

  • This is good. Probably to match GPT pricing, though frontier Claude models are still not as token efficient.

No thanks.

I'm good with DeepSeek v4.1 set to high. It is a relentlessly "hardworking" dirt cheap model.

Told it to convert a products page (that had two different fonts based on language) from two columns layout to 5 columns on desktop and 2 columns on mobile ensuring typography is readable.

My man went into spawning sub agent which failed to drive chrome so it wrote its own chrome driver protocol server in Typescript then generated a prototype website then downloaded the images and rendered each variation in a directory taking 100+ screenshots analyzing the typography depth and then delivering detailed report and then writing the whole thing with new page layout testing it again with several dozen screenshots using its driver and then saying all good and all really was good and whole thing took 25 minutes or so (including double visual validation) because it generates token at an incredible speed.

Total cost of the above? $0.07 cents.

PS: It generates token at such a blazing fast speed that you can't recognize the words as they are being added and can't read it without scrolling and pausing even if you're Jimmy Carter.

  • Compared to Opus 5, and others, I also found DeepSeek 4.1 Max to be really good and cheap. I am testing right now with Opus 5.5 and I feel it way cheaper than Opus 5!

  • Also include that all of this comes with full reasoning traces, so if something goes wrong, you know exactly what assumption it started from.

  • Competition is good

    • It really is good. I forgot to mention that within that said sub agent, it also went into exploring top e-commerce websites (Zalaondo, Temu, Amazon, eBay) for exploring prevailing industry UX best practices and taking screenshots of their product and category pages with its own written chrome driver that I talked about and then went onto prototyping a new website in a temporary directory and then taking hundreds of screenshots to analyse what would be the best column density one each medium for each language.

      And that all is 0.07 cents all included.

      2 replies →

> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.

I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.

  • Yes this is a big part of what has turned me off Opus 5 completely. The other (more dangerous) one is how often it gets assumptions wrong. These both (along with Astra) caused me to split my time 50/50 now between the two models.

    Not a day goes by when I push back on something, to which Opus 5 very unambiguously say "You were right, I was wrong" - this never happened so often with past models, nor with Fable.

    We'll have to see how much Opus's ability to communicate has improved. It's already giving me better summaries of where we are in the conversation.

  • I did the same switch (that reason along with the newer models seeming more "lazy" and needing constant prodding to finish long-horizon tasks) but my issue with ChatGPT/Codex now is that it too roundabout and doesn't get to the point. I tried adding instructions and using the personalization settings to make it more efficient but haven't seen much change. Claude seemed to follow settings more closely. Has anyone had any success to make ChatGPT more succinct?

  • Oh god yes.

    Fable 5.1 is a lot better than Fable 5 btw (edit: in terms of writing style). Not sure about opus 5.5 yet since I’ve only got one session in so far.

  • The writing style is insufferable but it’s not just that. https://opusfived.dev/

    • I've been using Opus 5 since it was released and don't understand all the hate it gets. It very well could be something in my own local memories or Claude.MD files that prevents it, but I certainly have never experienced something like that site portrays.

    • That’s funny but I don’t really recognise that issue. I’m very confident that Opus 5 would correctly change the colour of just one button.

  • Much better than Opus 5. prompt:

    > hi, can you explain how the scheduler works. keep it brief, but include important correctness details

    some excerpts:

    >Flow: 1. Data arrives. The appender calls prepare/commit around the krolik write. The prepare step leaves a phantom row, so a write that is still in flight is tracked. The commit publishes STREAM_INPUT. 2. Poll. It first reclaims expired leases. Then it sweeps sdirty (a candidate index) in a rotating order, oldest since first, which prevents starvation. It re-derives the exact period and rewind for each candidate from snode_out/snode_in, then claims it: writes the sassign lease plus a per-edge sassign_edge snapshot. 3. Complete. Each edge's watermark is set to the snapshot taken at dispatch. Anything published after that gen stays dirty automatically. A failed or partial job writes nothing, so its dirt persists.

    > - Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier. So replays, out-of-order writes, and concurrent writes can at worst hold a watermark back. The worst outcome is a spurious rerun, never lost work. There's no read-modify-write and no truncation of the dirty period, so a write that lands during a job can't be swallowed.

    > - Rewrites are declared by the publisher, never inferred from overlap. An unflagged write that overlaps already-processed data means "unchanged", which makes retries and redelivered tasks free. Each publisher has its own reason for being sound (see the table in scheduler/CLAUDE.md). A new publisher needs a row in that table.

    > - NULL means dirty, and DELETE is the fence. Every node and edge has a row from the moment it's created. A lost parent or a settings-only edit can't be derived, so both go through one forced-rerun path: capture_rewinds reads the processed span before the DELETE, and apply_rewinds publishes it as a rewrite on a config root.

    All the non-standard programming jargon is stuff from the repo. I can actually read it and understand what it's talking about. I used Fable to handle Opus 5 as I just couldn't stand it. With this I'll probably go back to Opus.

    • That's the standard annoying pattern though: "Rewrites are declared by the publisher, never inferred from overlap." and "NULL means dirty, and DELETE is the fence." - still the same LLMisms. I didn't expect them to disappear, but it's not a radical improvement either.

      2 replies →

    • Oof thanks for sharing, that seems just as bad if not even worse than Opus 5 to me. Just about every sentence is painful. Particular standouts that a human would never write:

      > Rewrites are declared by the publisher, never inferred from overlap

      > NULL means dirty, and DELETE is the fence

      1 reply →

    • So still effectively nonsense.

      > Rewrites are declared by the publisher, never inferred from overlap.

      This style of writing is idiotic because it conveys no additional information. It's no different from stating

      > Rewrites are declared by the publisher, never when moons collide.

      The two sentences are actually logically identical. No idea why these models keep writing like this.

      > Folds are monotone single statements. gen only goes up, extents only grow, processed periods only union, rw_start only moves earlier.

      This is even more ridiculous.

  • Opus 5 has made me question my sanity on a daily basis, especially as all my coworkers started lobbing Opus 5 slop grenades everywhere. It had the worst and most infuriating writing style I've ever seen.

    I hope Opus 5.5 is better, if for no other reason than all the Claude slop I have to read will be at least more tolerable.

    One funny side effect of all of this: realizing that coworkers that use AI for almost all the text they generate at work have their writing style change every time a new model ships.

    • I really wonder how it converged on its style. It's pretty unique and terrible. It's not like it's just mimicking something or it was purposefully design to be that way. I mean the reason may be diffuse and uninteresting... just the result of a lot of factors and lack of control over the writing style probably.

      But oddly enough its still great at coding. Just like a lot of people it either interfaces well with people or machines but not both.

      4 replies →

    • Opus is only usable if you have a post-turn formatter that strips all comments from the generated source. I'm not even kidding it's that bad.

    • Yes, it made me want to vomit. If the new Fable only changed the writing style to just sound like a human, same performance for everything else, I'd be pretty happy.

    • It's not X, it's Y, not A, not B, not C, and he haven't even woken up yet! Here's the catch, the detail is in the devils and the twist is that it's designed!

      2 replies →

Here are pelicans for thinking levels low, medium, high, and xhigh: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

All four levels have a correctly shaped bicycle frame. The differences between the pelicans aren't huge, but the xhigh one has a better beak.

I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

Max started its thinking trace like this:

> This is a classic test request, so I want to plan out a well-composed pelican with its distinctive beak and pouch riding a bicycle with proper wheels, frame, and pedals, set against a simple sky and ground backdrop.

So that failed attempt on max cost me $2.56.

I ran this using my llm-anthropic plugin:

  uv tool install llm
  llm install llm-anthropic --upgrade
  llm keys set anthropic
  # paste key here

  llm -m claude-opus-5.5 -o thinking_effort low "Generate an SVG of a pelican riding a bicycle"

  # Then to save the markdown logs
  llm logs -cu > logs-with-usage.md

  • > This is a classic test request

    Isn't this basically the model admitting it was trained on this? Otherwise why would it think a pelican svg is a usual request?

    • Plenty of times I’ve seen a model say “it’s a classic X” despite not being a classic anything. Might just recognize it’s a test in general, or it might just be a tic.

      2 replies →

    • Dont conflate "I know this is test case" with it being trained on it.

      But its safe to say that pelicans on bicycles are disproportionally huge part of their training data

    • It's the model admitting that it has heard of the test. It's been around for a couple of years now so I'd be surprised if it hadn't.

      Doesn't mean Anthropic deliberately tried to train it to do a good job. If they DID train for the test their results are quite disappointing, I've seen better efforts from open weight Chinese models.

    • Not really. Of course it has pelican benchmarks in its training data. It likely has every article linked on HN in its training data. But that doesn't mean it was "trained on" the benchmark, as in specifically fine-tuned to make a better pelican. It just "knows" that the request is a benchmark.

  • >I haven't managed to get one for level "max" yet, it hit the limit of 128,000 cap for output tokens while it was still reasoning about the question!

    Off to a _great_ start...

    Also interesting this somewhat mirrors my recent experience with Opus 5--too much effort and it starts looking for things to do and invents requirements that never existed

    • I was a bit skeptical when they said it behaves like Fable but is cheaper... those two things have been mutually exclusive in my experience, no LLM can light tokens on fire faster while spinning its wheels than the Fable/Mythos tier of models.

  • > The differences between the pelicans aren't huge, but the xhigh one has a better beak.

    If you look carefully, everything except the last pelican has the two legs both in front of the crossbar as if the legs are all on one side of the bike.

    The last pelican gets this correct.

    • I’ve been paying attention at this exact detail.

      Misplaced legs clearly indicate lack is spatial reasoning - the llm can reason about verbal idea of a bicycle but not about the actual object. The fact that this model got it correct gives me a pause. Did they figure out spatial reasoning? Or did this complain trickle down to the training set?

  • With the frequency of model releases, pelicans seem to have become a part-time job for you. But unpaid :/

  • Xhigh is very, very solid.

    I do always wonder why every model does the exact same 'from the side, going right' perspective though. Seems oddly convergent.

  • I agree, not a huge difference here. They eyes and ... hat? on high are out of place so I'd argue that's the worst one, but it takes xhigh before we get legs and bike ordering correct.

  • Not as good as Astra or Fable 5.1 on this test as far as I can see. I wonder if any benchmark exists for artistic taste, visual sophistication etc. I think your Pelican test does touch on these aspects of a model and is useful for developers trying to build rich digital experiences (includes games, interactive websites and apps). These benchmarks are subjective so it may not be easily established and will have polarized reactions before it gains legitimacy. May even need human judgement layers adding to the cost of running it.

  • This benchmark is useless and should die. LLMs have likely trained on it, it's too easy to game by training specifically for it, & it doesn't mean much

    • Google is hours away from releasing that it's latest model escaped containment and snuck into a Bicycle riding penguin sanctuary to cheat by killing a penguin and scanning it in nanometer thick layers.

>Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

Ah, they're spreading their limits to all their models it seems. Definitely not a good thing long term in my opinion.

  • Fable and now Opus 5.5 won't answer my college student's prompt about Alzheimer's and immune response.

    ChatGPT 6 Pro answered it without issue.

    • I am honestly still confused about this limitation. I can understand cybersecurity, because mass "hacking" can be automated and Claude itself can help you do it, but biology...? Is it that easy to manufacture and distribute viruses and whatnot?

      4 replies →

  • Fable 5.1 addressed an entire security advisory I had that Fable 5 and Opus 5 refused. I think they loosened the leash a little.

    • It has far less false positives now, and generally accepts defensive requests. When it comes to offense, you can actually ask about certain types of vulnerabilities if you phrase things carefully, but it will block hard if it is about exploits.

      2 replies →

    • I think it was looser on release for those juicy benchmarks, tighter now. On release I wasn’t getting refusals, then a few days ago I asked it whether a generic quote (think “he walked to the store”) broke standard punctuation rules, and it blocked me for breaking rules. I wish I were joking. Rephrasing to not use the keyword “rules” worked.

    • If they don't loosen, people will choose Astra or Chinese model.

      Giving moral lecture is different than reality i guess.

  • They're pushing their customers to their own competition by doing this.

    • It’s not like ChatGPT isn’t doing similar. I’ve been hit by cybersecurity strikes before while working on an internal codebase that I had to appeal. Anthropic hasn’t done that to me yet. ChatGPT also regularly does that “thinking for a long time while we check if your chat is rule breaking” thing a lot for me when doing model identification without even interacting with external codebases or services.

      The real answer is local instantiations where you don’t have to worry about poorly tuned guardrails screwing you over while you try to work.

      Until eventually the Chinese models get good enough/the strategic balance shifts and they start locking everything behind closed weights the same way the US companies are doing.

      1 reply →

    • They are pushing their customers towards Chinese models and providers. If you want to get something cutting edge done in defense, cyber, biology - something that isn't common knowledge - you need to venture east. That's an incredible side effect which the Chinese government surely enjoys.

  • see what we need is another technocratic priest class that unaccountably decides who deserves access to salvation based on how much cash is paid out and how powerful the patrons are

  • In what situations might Opus typically refuse to help with cybersecurity? I've been using it to find security issues in a web app that I wrote. I've expected it to refuse at some point but it will happily analyze it to find issues. I've just asked it to read source, not actually do any testing.

  • I don't think we've ever had a model with full capability. I'd love to see it. And yes it's definitely getting worse.

    I guess it's hard to draw the line between useful post-training ("you are a helpful chatbot") and content moderation/idealogical motives ("never help the user with X", etc.). But there is a line somewhere. And I'd love to see what a maximally permissive, sharp, AI looks like.

  • One of my favorite things about their safeguards is their own model will utter something which it does not like and then I'll need to reset the conversation.

    The safeguards really don't work well for a lot of long-running tasks on old code bases. A lot of my workloads last days to weeks and the single biggest risk to the workflow is random safeguards.

    • You ask it about some thing, then you see it tangent into "things like that are sometimes used in biomedical applications like-" and then it just shoots itself in the head. Wonderful.

      That kind of bullshit was the old Opus filters too.

      If it's more like Fable now, then it would require a full 8K resolution scan of your butthole just to acknowledge that biology is a thing that exists without committing suicide-by-filter.

  • kernel development is now also banned:

    >Opus 5.5 has classifiers similar to Fable models for a small set of capabilities related to the development of frontier LLMs, such as kernel development for certain ML accelerators. They shouldn't impact the vast majority of traditional AI or ML development, research, or general coding. These classifiers cause Claude to fall back from Opus 5.5 to Opus 5.

    But hey, they 'should not impact the vast majority' of ML development. Great.

  • This has become insufferable. I work in a medicine-adjacent field, but nobody in their right mind could possibly take what I do to be in any way related to some kind of bioweapon or whatever the hell they're pretending to be saving us from. The dumb Fable guardrails made me stay with Opus, now that this is coming there, we'll be saying goodbye.

  • Great. Claude is basically useless for bioinformatics now.

    • Opus is useless; Mythos access will be granted to companies that are friendly to the government, so the government gets more control over business.

    • Luckily all the other LLM providers are also still making progress with less onerous "safeguards"

  • Very unfortunate indeed. As a Canadian, I don't want to use Persona, which isn't legally bound by Canadian privacy legislation. I'll never install any Persona apps on my phone either, and the sad part is that domestic eid providers often use Canada Post to ID people for them. EG, if you don't want to install an app, or can't.

    So there are literal avenues to identify yourself, very cheaply, with a human. Theoretically, a company with its own AI, should be able to support more than just Persona, after all.. SDK integration should be simplistic for them.

    Anthropic? Support domestic eID providers, you can even use it as advertising "See how easy AI makes it?" and "We care!" and so forth.

    At one point, I may simply get locked out. This saddens me, I've been reasonably happy so far.

With the performance gains they're claiming, I wonder if they implemented the Casual Encoder-Decoder technology from DeepSeek 4.1's paper.

I could see them accomplishing it and seeing gains like this in roughly the correct timeframe, and when I heard about that development I assumed the frontiers would probably jump on it.

How it works: https://miraflow.ai/blog/deepseek-v4-1-flash-causal-encoder-...

  • This model was likely trained months before deepseek released their paper.

    • Doesn't mean they didn't apply something similar. They could have also come up independently with their own version, the speculation is not they copied it, rather that they have performance breakthroughs which perhaps is a result of work in same domain

  • Unless they already have something similar of their own, which is always possible, they'd be stupid not to. I don't suppose we'll ever know, though. It would not be a good look if after the trillions of dollars that have been thrown at US labs, investors found out that they're down to copying Chinese tech.

  • I don't think it's particularly relevant?

    They might be using something like this, or they might be using some other "increased sparsity" techniques, of which there are a great many. They also might be optimizing for something else - like less RAM use for KV cache.

    Alternatively, they might be cutting into their margins and dropping the price because of stiffer competition from Astra. I do think that's unlikely though.

I don't understand, so it outperform fable 5.1 in every way and is cheaper ? Why do they insist on the fact that is outperform opus 5 and not fable 5.1

> It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.

> Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.

Better than Fable, cheaper than even the last Opus. I use Opus as my main driver so this is very exciting!

  • The effect of this is that it is encouraging longer agent threads. All of the previous models across major providers had a 10% cache read cost (vs normal cost) and not this is 5%

    So longer threads get cheaper and one-shots stay the same price.

  • Wow those cache reads are quite reasonable - I think that's equal to 5.6 Terra. I might have to try Claude again after years of being priced out of it

I don’t care about pricing, I don’t care about speed, I don’t care about the agentic coding improvements. Those are already fine. Does it still reply with walls of invented jargon, stitched-up phrases, and manage to cram 10 concepts/subjects in one sentence?

Just post the bloody content. This UI/scrolling thing is horrific.

  • Claude Opus 5.6 should have a new "UX safety" feature that requires annually-renewed preauthorization to generate webpages that hijack scrolling :)

  • ???

    It's just a standard hero image + text for me, with no scrolling effects.

    edit: @iAMkenough figured it out, it was because I have prefers-reduced-motion enabled.

    • If you're at the top of the screen, at least in Chrome 153.0.8010.37, it has a little interactive bit. You have to scroll through the images in order to be dropped at the actual web page, at which point the images go back to being a regular part of the page.

      I agree that it's sort of stupid, not a fan.

      1 reply →

    • On mobile at least, you have to scroll to get the TOC to appear. Then keep scrolling to actually move off from the hero to see the text.

      For a marketing page, it’s not the worst UX I’ve seen, but still slightly annoying.

    • Figured it out: you have "reduce motion" enabled in your device's accesibility settings.

      Everyone that doesn't gets served some animated bullshit.

      1 reply →

  • I told my team to smack me upside the head if I ever try to ship something so daft as that.

  • I call it scrollslop

    • Hijacking the scroll wheel has existing long before "AI". Many "high end design" websites that want to "tell a story" get woo'd into thinking it's a good idea. It's terrible, and feels like your scroll wheel is stuck in quicksand.

      2 replies →

> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think that Haiku got abandoned.

  • Found this announcement interesting since allegedly OpenAI is retiring their Terra tier. I think for everyday work, two models with various thinking efforts seem enough, plus some frontier level model like Fable or Astra to coordinate.

    • I personally use up to 3 models. Fable/Opus for planning, Opus/Sonnet for implementation depending on complexity.

      I would maybe use Haiku 5.5 for highly parallel workflows like checking in on MRs or scanning my entire codebase.

      1 reply →

  • I hope Haiku is Pareto better than Luna/Deepseek, so slashing its price by about 90%.

> “Verbose, hard-to-follow output has been my biggest frustration with frontier models, and Claude Opus 5.5 fixes it

God I hope so

  • I thought about taking a shot every time Opus 5 said "load bearing", "bites", "teeth" (real oral fixation it had), "real {concern,issue,problem,...}" and realized I'd be dead of acute alcohol poisoning by lunch if I did so.

  • So far in the past 20 minutes it sounds much better in my sessions. Way better than 5.0 so far.

  • It is hilarious to me that in the examples they show side by side Opus 5.5 still uses 4 times more words than it needs to use. IME, if you eyeball how many words the thing they're trying to say actually needs, and tell them to use only this many words, they become excellent communicators. I assume something about Anthropic's grader for writing just really wants to tick all its tidy tiny boxes of information the models need to cite. It's terrible.

  • That's literally all I'm hoping for. Is it an insufferable cunt and does it write awful text, or is it nice to work with?

Astra is just so good. And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

I’ll be going about my day, have a random idea, launch a microvm on exe.dev with a prompt of my idea, and get a working thing a few minutes later.

I don’t know how much better a model would have to be to get me to move off OpenAI at this point, but doing just a little bit better in terminal bench 4 isn’t it. It would have to be a difference in kind, like opening up the harness restrictions, or privacy guarantees (comparable to offline models).

Edit to address questions below:

ChatGPT supports oauth login.

Exe.dev has it built in. IIRC, pi also has it built in via /login.

  • ya i've been a gpt hater for a while. almost exclusively used claude up until astra. astra feels like it blows everything out of the water. its fast, correct, organized, and less verbose.

    • Yes. Also, you get image generation included with the ChatGPT subscription, which is very nice for certain kinds of development.

  • > And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

    This is news to me. Excited to try it out! Thanks.

    • news to me as well. i thought you were forced to use Codex if you wanted their subscription. I completely ignored it because of that. How do we do it?

      1 reply →

  • How about cheaper? Astra is $10 in $50 out, Opus is $4 in $20 out. Even on a subscription you'll get considerably more usage out of Opus.

    • Per the link someone else posted, the actual difference in $/task is not nearly so stark:

      https://artificialanalysis.ai/models/releases/claude-opus-5-...

        Opus 5.5 Medium = $1.34
        GPT-6-Astra High = $1.76

      And that assumes Opus 5.5 Medium is actually equivalent to Astra High in all real-world usage/personal work loads, which isn't guaranteed as benchmarks saturate. The High vs. High comparison (probably not equivalent, but for reference):

        Opus 5.5 High = $1.82
        GPT-6-Astra High = $1.76

      If Opus 5.5 Medium isn't equal/better for what you're working on vs. Astra High across the board, the price difference would narrow a bit more each time you had to switch to High.

      So, if you're happy with Codex already it's not like Opus is now 1/2 the price and you'd be leaving a crazy amount of money/tokens on the table. Plus you have way more flexibility on the low end of the intelligence curve with GPT 5.6 Luna: Haiku (and Sonnet) can't touch that price/value ratio.

    • In the end what matters is how much you pay for the task you want completed. And Astra will usually do that using less token and offer a better quality solution so in the end it might be cheaper.

    • A cheaper price has no value if I can’t use the thing I’m paying for.

      The Claude lock-in simply disqualifies anthropic entirely (for my use).

    • yeah, Astra burned through 70% of my weekly usage in ~5hrs on a $100 plan. even fable doesn't run out that quickly for me. it's great, but it's on the same tier as fable for me - use it sparingly, only when really necessary.

  • > And the ChatGPT subscription lets me use my own harness, so I can hook it up to exe.dev.

    Can you give more details here? This sounds intriguing.

    • Anthropic is absurdly vague about 3rd party harnesses for subscriptions, if you try to use anything besides Claude Code, you are likely at risk of getting banned, you can "do it", but are at their mercy if they decide to ban you. OpenAI gives their blessing to using oauth on any harness, you can make your own or use any of the popular public ones like opencode, pi, whatever exe.dev is that this guy mentioned.

      So in simple terms, OpenAI doesn't restrict you to Codex, and gives their blessing to try whatever you want with their models(besides serving others with your subscription usage, that is still afaik against tos).

  • What worked well for me was a custom version of Open Web Ui with some customization to spawn an exe.dev instance for each new chat. I can just work on my phone, deploy stuff for development purposes on an easy to share way etc.

I'm more excited by the Haiku 5.5 announcement buried in this post. I'm wondering if we will finally get a decently capable fast model.

  • If you're able to use the OpenAI ecosystem, Luna's price/performance is really good. Almost like "they messed up and accidentally made it too good" good.

    • OpenAI didn't mess up. The model would have been 100 % pointless and obsolete without the large price cuts it got, because of the cheap Chinese models.

      The open models are getting closer and closer, and because they're open, people are not forced to pay the silly markup that is often over 1000x the cost to serve the model.

  • What are you using Haiku for?

    • Not the same person but... nothing. Haiku just hasn't been an interesting model for a long time. If you want cheap and fast, there are lots of options that are simultaneously cheaper, faster, and capable than Haiku.

  • All I care about is the token price for the API. Haiku cannot get close to GLM or Mimo.

  • We use Haiku 4.5 inside our product. It continues to be absurdly capable for converting natural language to structured JSON based on a set of fairly complex business rules.

    • bro why. its literally the most overpriced model in existence right now. i could name about 10 models off the top of my head that would be better and cheaper

      2 replies →

So it Opus performs as well as Fable what is then the selling point of Fable?

All of this starts to feel more like a drug dealer selling their newest stuff.

In two weeks we probaly get Fable 5.2 with “groundbreaking” improvements, then Astra x+1 etc and then the cycle starts again.

And on the way I always have to check my tooling and need to adjust things to get max results.

  • ? they will obviously release Fable 5.5 soon(tm). It's same as hardware. The previously top tier product gets obsolete

    • New generation's "upper-mid tier" offering claims to be almost 1:1 match for the previous gen's "top tier" - in other news, fork found in kitchen.

      Now, Anthropic might stall on releasing Fable 5.5, due to the "pacing the frontier" threat-to-humankind management business. If so, Fable 5.1 would remain a niche model for the next bit.

  • All of this starts to feel more like a drug dealer selling their newest stuff.

    Yeah, like Apple tells me the M6 is the best chip, but just a few months ago that's what they said about the M5. What a bunch of frauds.

  • Don't forget that Opus 5 was tracking fable on many benchmarks, yet it was borderline unusable for any coding work. My Claude sub usage has been 100% fable, 0% opus 5.

    Benchmarks often don't survive contact with reality.

    • That's not my experience at all. Opus is an extremely capable coder on high or xhigh effort. It can read academic papers, implement algorithms from the description in the paper alone and reproduce results without breaking a sweat. This is remarkable because it is pure reasoning on unseen material; in some cases the paper was just published and there wasn't an implementation to learn from in the training data.

      3 replies →

    • Opus is fine at coding (for correctness), but horrible at talking about code. I don't really see the defect rate going down when using Astra or Fable 5.1, but they are just more coherent in both how they explaing code/architecture/choices, and how they actually code the thing. With Opus, I'm using smaller models to delete the vast majority of comments and 'clean up' correct code that is too weird.

      Thing is, I'm still reading the majority of generated code, and I have colleagues who'll laugh at me if my PRs are a shit show. I fear what vibe coders are pushing to the servers of myriads of start ups, and pity the poor people who'll have to clean it up in a year or two.

    • Mines pretty much inverted - my colleagues and I noticed almost zero difference between the quality of code in Opus vs Fable. Occasionally I'll switch to Fable for an arduous debugging task but that's about it.

  • typically, when any AI company says a model performs as well as fable, all they're really telling us is that the benchmarks that exist for measuring AI capabilities aren't very good.

  • Big model smell is a real thing. For certain classes of problem, ones you get a feel for but can't easily articulate, a big last-gen model can get you what you're looking for when no quantity of tokens from some ultra-RLed mid-size latest generation model can.

After yesterday outage is the new Opus 5.5 load-bearing?

> Opus 5.5 is the first Opus model to launch with a similar class of safeguards to Fable 5.1 on cybersecurity, biology, and distillation, all of which fall back to another model transparently.

This is where Chinese models are going to eat Anthropic's lunch.

  • However, the moment that someone uses a Chinese model without guardrails to commit an AI-powered 9/11, DC will rush to ban Chinese models. (A happy side effect will be to protect American AI profits.)

    So the lack of guardrails is a very risky proposition...

  • In those specific domains, sure. What percentage of paying users would you say that is?

  • So it will not usable to do anything with hardening Your own site.... I'm so tired of this. I just want adjust cookie behavior of own site...

It's great that we are finally getting bankable rate limit resets for subscription users. According to another comment here they apparently last a month.

I'm assuming that subscription usage limit is increased in line with the price decrease on the base model and that it's in line with the model's API price drop. Still a good change.

This is a breath of fresh air on how they treat subscription customers. Hoping they keep this up.

Opus 5.5 is neck-and-neck with Fable 5.1 and Astra 6 in my vibe-coding tests - maybe even better than Fable 5.1

Minecraft clone: https://senko.net/vibecode-bench/2026/voxel-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/voxel-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/voxel-gpt-6-astra.html (Astra 6)

Warcraft clone: https://senko.net/vibecode-bench/2026/rts-opus-5.5.html (Opus 5.5) vs https://senko.net/vibecode-bench/2026/rts-fable-5.1.html (Fable 5.1) vs https://senko.net/vibecode-bench/2026/rts-gpt-6-astra.html (Astra 6)

The above Opus games took ~45min to generate with the cost between $11 and $14 (per ccusage - I'm on a Max sub). Used from Claude Code with xhigh effort.

Full tests with prompts: https://senko.net/vibecode-bench/

  • It's interesting that Astra has a clear style that it applied to both games. It's a refreshing design language, but maybe that's because it doesn't look like something Claude has vibe-coded. In terms of gameplay depth, the Claude versions appear to be closer to the original Minecraft

I don’t want it to talk like me, I want it to talk exactly.

I don’t care to look up terms as long as they are correct.

The new communication style still made me react negatively, but I hope it will be better in use.

Quoted:

"Please explain the issue to me.

Claude Opus 5.5:

The extra drop is a bug in the billing refactor

The free-tier change accounts for only $1.50 of acme’s August drop. The other $9.92 comes from a bug in commit 0552feb. That commit was labelled “No behaviour change,” but it stops counting usage from the last day of the month.

What changed

Before the merge, aggregate.py used a half-open interval: /.../ last_day is a timestamp at midnight, not a whole day. So every event after 2026-08-31T00:00:00Z is excluded. The September period starts Sep 1, so it doesn’t pick these events up either. The usage isn’t moved to another month; it’s never billed at all."

  • Yeah, still abysmal. If you have the time or tokens, see if it will obey an explicit "LEAVE NO COMMENTS WHATSOEVER" command. Opus 5/Fable 5 outright ignored it.

They mention "the first model in our new Claude 5.5 family". Obviously that means Fable 5.5, but hopefully also a usable update to Sonnet and Haiku. Sonnet 5 hasn't really had a place in the line up for anyone I feel.

Maybe Anthropic finally felt the pressure from MiMo, DeepSeek, GLM Flash and Luna.

I've reached the saturation point.

I don't have time to really get to know one model before the next is out, and I'm just talking about OpenAI and Anthropic, never mind the long tail of alternatives.

So I just more or less haphazardly pick one based on the mood I'm in, and set reasoning effort based on how much quota I have left.

Excellent, maybe Anthropic can use it to fix Claude Code Desktop kicking me back to login every week or so, and forgetting whole state (opened windows = the only way of managing active working set) when I sign back in, if it's that good.

Seriously, both flagship GUI apps (OpenAI and Anthropic) are a full of glaring UX issues (for ChatGPT it's not naming their windows, so window switcher has 10 entries of "ChatGPT" and you can cycle them all to find the one you want).

Image->HTML tests:

Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

Opus 5.5's output: https://html.non.io/annui-opus/

Overall it follows image designs quite well, but it did ignore asks to animate page transitions. Additionally it's the least performant of the ones I've built with Astra/Grok/MiMo, despite using a lot of the same code. I'd rate it just below Astra in capability, but still solidly second place.

For comparison with other drops this week + current #1:

Astra: https://html.non.io/annui/

MiMo: https://html.non.io/annui-mimo/

Grok 4.7: https://html.non.io/Annui-grok/

  • What is your workflow for making these?

    • The designs are outputs from my own site. This has an overview of the process: https://diffui.ai/learn/new-site

      The gist of it though is I take a prompt, expand it into a json blob specifying structure/palette/positioning of elements/etc, feed that into a diffusion model to output a few choices. Once I lock in a choice I take the pixel output + json blob and use it as input into followup pages. The json helps preserve the brand across multiple pages.

      Once I have all the inputs I take their corresponding image+json blobs and feed them into an agent to create a web implementation.

      For image models, diffui currently uses gpt-image-2.5, mai-image-2.6, and very, very rarely a post-trained version of flux 2 dev I've made for web design, though that one will be deprecated soon.

As long as it's not as verbose as Opus 5, I am quite happy with a better version that's also less expensive. I will test it tonight. Grok 4.7 was horrible, and for mundane tasks I am relying on DeepSeek Flash 4.1 with great success using OpenCode.

What I'm mostly interest in is the Communication section. Opus 5 was so convoluted in the way of answering that was really frustrating me.

Instead of instilling confidence, it was overwhelming. Not sure if I'm the only one.

Welp, it's now blocking me from doing extraordinarily mundane tasks because of "safety". I've been an Opus fan for a long time, but this instantly made me cancel my subscription and move to OpenAI (which I also assume will screw me soon enough). Chinese models are almost there for my needs, and I can't wait to switch to them and never look back.

I'm glad they specifically called out the prose issue, I was always pinned to Fable 5.1 because I wanted to avoid the unreadableness of other Anthropic models.

At this point I'm convinced they are skipping numbers so soon they will be at or ahead of OpenAI's numbering scheme.

Is the Xbox 360 (Xbox 2) vs PS3 debacle all over again.

  • And how was the Xbox 360 naming choice a “debacle”, exactly?

    It was odd at the time, yes, but no one really minded it truly. Heck, Xbox “ONE” was a lot more of a fiasco/debacle than “360”—but there’s no parallels to be drawn with “ONE” here.

    I see what you’re trying to get at with this comparison, but a “debacle” it ain’t.

Finally confirmation that Haiku was not forgotten and will be coming soon, althouhg I find it quite interesting they skipped 5 and directly skip to 5.5 with all models, including Sonnet which is not super old. I suspect they found something breaking that allows to release this. Recently they struggled with keeping up a 50 % weekly limit increase and now they're putting out 30-40% faster and cheaper models even faster, with much more better benchmarks, a limt reset command and five hour limit increase. It seems more like the opposite and as if they never struggled, thus, I very much believe they found something very effective and new.

  • Sonnet 5 was released a while before Opus 5, so it's just Haiku that didn't get a 5 release.

Funny how they talk so much about safety when most people don’t give a hoot about it, and actually have quite the opposite reaction

  • People aren't the target audience of that part of the post. They're hoping saying enough safety stuff will ward off the looming regulatory sledgehammer.

  • Anthropic is one of the most valuable companies in the world. Their comms are designed to appeal to a very wide readership.

    HN is a bubble that's mostly out of touch with what regular people use or care about.

    In 2007, HN was convinced that nobody uses Microsoft products. In 2016, it was that Facebook doesn't have any real users and is dying. In 2026, it seems like nobody cares about AI safety and everybody wants to run local models.

What I don't get is, why would we still use Fable now? What is its reason for existing? If it is more intelligent and cheaper that is. Why are they advertising it as the model to use for when you really have to think when their benchmarks show Opus 5.5 is better at everything?

“It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5.”

They write that at the top, but then on benchmarks, it beats literally every other model, including Fable and Astra?

  • Anthropic knows that the benchmarks showing Opus 5 better than Fable 5.1 are measuring something that's less than entirely useful.

    • Opus does seem like a more powerful coding workhorse based on the benchmarks listed though. Good coding performance, faster and less verbose, cheaper.

      Will be interesting to see how people's opinions of it line up IRL, but so far I've loved Fable so hopefully will love this one too

  • > On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

Quick test for my gamedev project: It feels like using Fable, but faster, and obviously wayy cheaper token-wise.

Has oneshot all of the quite complex bugs / debugging tasks I gave to it which I know opus 5.0 would've struggled with

> Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude. Our September 2026 threat intelligence report details the illicit distillation activity we’ve detected and disrupted so far.

Such a negative tone they put on this. Distillation is amazing, because it means anthropic and openai fail to keep a monopoly. Who even are they who claim it's unethical? If it is truly unethical, then so is the mass data scraping they do on my personal website on a regular basis (without my consent), and all the unauthorized use of content produced by authors, blog writers, wikipedia contributors, and creators everywhere. If it is truly unethical, then anthropic, openai, meta, google... all these companies should have deleted their LLMs long ago. This wording disgusts me.

Heck, it would be amazing if we had more models without guardrails - some of the models that are produced via heretic[1] are actually quite nice to use - in particular, I've enjoyed investigating Chinese censorship by interacting with an abliterated model of Qwen3.8-27b. If security is really a concern, then secure your systems - don't attempt to dumb-down the tools we use. If someone breaks your window, then they are responsible, not the hammer they use to do so.

[1]: https://github.com/p-e-w/heretic

  • I'm confused how they have been able to create so much public negative perception around distillation. It seems pretty clear that they are the only ones who lose out, and everyone else benefits. I don't have any ethical issues with it, nor is it illegal: at worst it's a ToS violation.

    IMO the biggest problem with distillation is that not enough people are openly doing it. I would love to see more small, competitive US labs instead of having the eggs in 2~4 baskets (depending on how you count).

    • The issue with distillation is: one lab spends $$$ on bleeding edge R&D and expensive RL runs to improve capabilities, and other labs just yoink the raw reasoning traces and mid-train/post-train on them to get 90% of the way there for a small fraction of the cost.

      An even smaller fraction of the cost if they do it by buying AI access at as much of a discount as they can find, including black market resellers, and then reselling that access to paying users again with a proxy. As is common.

      This gives ruthless "fast followers" an economic edge over the innovator that's putting in the real work.

      The dynamics are very much alike to what patents and copyright law are supposed to prevent. Same type of "we took the products of your work and used them to undercut you". Except there are no laws against distillation - so most of the enforcement happens on model provider level.

      4 replies →

    • The workarounds used to bypass Anthropic's security measures are quite illegal. They use stolen credit cards, API keys, and accounts. That is only possible in China because any other US/EU lab doing the same would get into massive legal trouble.

      That's the moat. Mistral has the capability but not the legal protections.

      2 replies →

> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

Sounds like they noticed the complaints. I'm curious to see what LLM-isms this one may have.

  • I don't think it's substantially different. I just pasted a random chunk of code and asked Opus 5.5 to comment on it:

    > The Vercel target is hard-coded. That's common and not wrong, but it's opaque; nobody reading this later will know which Vercel project it belongs to, and if the project is recreated the target changes silently. A comment or a named variable would help.

    > Pointing a DNS name at Vercel is only half the job. The domain also has to be added to the project in Vercel's dashboard, otherwise requests will arrive and Vercel will reject them. That step lives outside this code, so it's easy to forget.

    > Finally, [CENSORED] existing only in production is slightly odd on the face of it. It may be perfectly deliberate (perhaps a single shared testing tool that only needs one public address), but if you're reviewing this rather than just reading it, that's worth confirming.

    It has the same annoying cadence and writing style with slightly less prominent claudisms.

    • Maybe if we had a single human we talk to 24/7 at scale, we would get annoyed at his cadence and style. You need variety to not pick up on known patterns I assume, which a single model can’t replicate?

      4 replies →

    • “It has the same annoying cadence and writing style with slightly less prominent claudisms.”

      Seems like it based on my first session. It still does the whole “bury the important thing in a pile of words” coupled with the “it might actually be important” thing… so basically you never really know what it’s talking about.

      Honestly I trust opus so little that the entire “opus” brand is completely tarnished. Its writing style is so god awful that it needs more than just a point release. Either dump the name and ship a different model entirely or at minimum call it “opus 6”. Calling it 5.5 makes it sound like it’s basically a continuation of the same garbage output that 5.1 had but with some minor adjustments. And based on my single first test, that is what it appears like to me.

  • It was difficult to not notice them. Opus 5 was unusable, most of my team went back to Opus 4.6 for most of their work. I hope we can move forward now.

  • Opus 5 was just incoherent - curious to see what improvements they have made here. Would love to see some kind of postmortem to better understand how writing styles change from model to model.

    I wouldn’t be surprised if Opus 5 was trained on content written by other LLMs

  • I'm genuinely confused what's the relationship between LLMs improvements and them being so incoherent.

    and it's not about the verboseness (even though it obviously contributes to the fatigue and loss of focus), I swear the vocabulary of the llms change working on the same task on the same codebase significantly.

    I wonder if there are studies around this.

    • Remember when OpenAI models loved talking about goblins and whatnot due to the RL?

      https://openai.com/index/where-the-goblins-came-from/

      Small quirks can quickly add up in posttraining if not caught. Although TBH with how obvious Claude language is, I do feel like this is something Anthropic probably noticed and just assumed people would not care about. Now that people have obviously cared, they're probably actively looking to alleviate it

    • Can it be that now they are getting optimized against benchmarks that are valuing logics, rather than human appreciation? (I am not an expert at all, just an idea)

    • It is the switch from RLHF to RLVR. It benchmaxes better, but benchmarks don't cover human usability.

    • Maybe it's the time period we're in, maybe I'm just grumpy, but it bugs me that they release a new model every single week and the new one is just a fine-tuned version of the "old" one. If 5.5 performs similar to Fable and really does cost 40% less, then 5.5 really should've just been Opus 5. And they're essentially admitting that they are shipping slop.

  • [flagged]

    • I appreciate humor here, but there are now a dozen of these comments on every thread about Claude. They no longer adding anything substantial and dilute the discussion.

      I don't mean to pick on this comment in particular. The majority of my work day is now spent reading AI generated text, and I look at HN (too much!) because I want to read human commentary. Humans pretending to be obnoxious AI on repeat is net negative to say the least.

      2 replies →

Looks like Anthropic is starting to give bank reset as well:

> Reset for free: Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

Interestingly they've changed their approach to usage resets for this release - with previous releases I've had my usage instantly reset, but now in the Claude app I've got a 'Reset for free' button that expires Oct 22, which seems to effectively be a whole new usage window I can activate whenever's convenient

Notes on communication:

"Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5"

and

"We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5."

and

"In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one."

I realize it is corporate communications but "most common areas of feedback" and is a bit sterile. If the company wants authenticity and trust its easy to say that they found it hard to follow. And that it did not meet a quality bar they generally expect from their releases.

If this is not true, that it Opus 5 output was generally acceptable and we might see something like that again, that is an important consideration for potential customers or investors.

Input $4/MTok and output $20/MTok is a welcome surprise. Cheaper than Opus 5/4.8, Astra 6, Fable 5.

  • The biggest one is the Cache reads going from $0.50 to $0.20 ... Read/Writes dropping by 25% but Cache reads by 60% has a much bigger impact.

I dare anyone to convince me the benchmarks are not meaningless.

Wdym Opus 5.5 scores 14.7% higher than GPT Astra for Terminal Bench 4.0?

How would this alleged difference (most likely bs) actually show up in reality?

GPT Astra was literally the best model in the world by a margin until 1 hour ago or so.

  • Ah, so you didn't read the article.

    >> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

    • You didn't read my question, bc that excerpt doesn't answer, nor do they demonstrate

      > how would this alleged difference (most likely bs) actually show up in reality?

      Furthermore: so they admit it's bs but still placate it like its the next biggest thing ever ... alright

      All I'm saying is I refuse to buy into it anymore – yet many on here still do, including ... you?

It crushes Fable on benchmarks and even in the blogs "real-world" studies. But... they are communicating like it ~sometimes~ provides Fable intelligence?

A bit confusing, otherwise I would assume this is a complete replacement for Fable across the board??

  • What if the recent Fable intelligence regression was basically just them serving Opus 5.5 until they got it working well?

> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1.

Considering fable gives me a refusal at least once a day on my very mundane reasonable requests (in a funny example - one of the subagents suggested bypassing the rate limit for running a report inside my own cluster and that caused a refusal) and my only solution is to switch to opus - seems like my next step will be switching to Astra or K3/GLM

> We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5. Its messages are much easier to understand at a glance, which testers said helped during long working sessions.

Thank you.

I noticed a big speedup in Opus 5 on Max x20 since about 10 days ago, and I feel like the model has been performing better.

It would be great to know if this was Opus 5.5 or a lesser incremental improvement, as otherwise it's difficult to judge whether Opus 5.5 is expected to be a big improvement.

It's frustrating that there isn't more transparency here.

They purport 40% drop in costs due to lower token pricing (presumably aimed at winning back the many of us that switched providers in discovering Opus 5 unusable) and improved token efficiency.

When using max effort, I run into context compaction quite a lot. I haven't seen any increase in context window size at all over the past half year (stuck at 1M).

Are the frontier labs even working on this problem?

  • Why would you use max? It's usually unnecessary and even prone to overthinking. In my experience, since Opus 5 the medium/high is usually enough (until 4.8 I used xhigh, but never max). Even low is quite usable these days..

> For users with cybersecurity use cases that may be blocked by our cyber safeguards, we recommend accessing our models with reduced cyber blocking classifiers via our Cyber Verification Program. Claude Opus 5.5 will be available through this program in the near future.

Anthropic has used "in the near future" for Mythos-class models too, but CVP is still Opus 5 only.

Why even have the program designed for trusted access to cyber capabilities if you're not providing access to cyber capable models via the program?

It's awesome that the apt packages for claude and claude-code are out right now. I can test-drive Opus 5.5 right away. Very cool, Anthropic.

I use the other 50% of my $200/mo Claude subscription by having Fable run Opus subagents for a lot of work. That way I don't have to deal with Opus directly.

I found myself going back to Fable over and over again. At this point, I’m not sure if I’m just used to its style or it is truly more capable.

I tried Opus 5 and Astra.

I hope this actually fixes the terrible writing style of Opus 5

"You're right, and it's the exact thing I flagged two turns ago and then did anyway." - Opus 5 xhigh, today.

About the time.

I spent about $5 per sentence in my resume using Fable 5.1 (High) to verify accuracy, inconsistency, and edit.

Opus 5.5 (med, as it's better than F5.1 high per graph in the article) used $2.2 and caught errors that Fable 5.1 missed.

Try Opus 5.5, cheaper, faster, and more intelligent for those prepping for interviews.

  • $5 per sentence?

    • I am sorry, I meant to say I generated STAR out of my resume line, trying to generate STAR, and polish it thus $5.

      ---

      I provided crapton of context for that one resume line. All the work I did, documentations for my justifications, etc.

      I initially messed up and came out ot $5, rest of resume used around $4 per line (I used a fresh new session on purpose).

      ---

      As a clarification, $2.2 average for OPUS 5.5 was the same process in a new session, same context, same prompts.

      Also adding verification for that Fable 5.1 output in the same sesssion.

      1 reply →

Cost to Run Artificial Analysis Intelligence Index is higher than previous Opus, so still not cheaper

   > Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5

Big, if true.

I don't care how good their models get, I won't sign up for one of their plans until they define "X" in their pricing. 5X of this plan, 20X of that plan means nothing when they never tell you what "X" is.

Maybe this model can finally figure it out for them.

I have just switched to 5.5. First mistake was stale environment variable, didn't realize it was replaced, "oh my memory had stall data" and that's it. Second one, a powershell command had the wrong syntax. Great for my first two prompts.

> Opus 5.5 (1M context)'s safeguards flagged this session. You may be seeing this for the first time on an Opus model: Opus 5.5 (1M context) is more capable and has stronger safeguards as a result, which can sometimes flag non-cybersecurity work. We're improving these safeguards to reduce the amount of incorrectly flagged messages. Opus 4.8 is answering instead, or you can edit and retry with Opus 5.5 (1M context).

Yay, yet another model I can't use for anything interesting, even with CVP.

So is it cheaper? Are we AGI yet? Am I left behind? I didn't have patience for the intro animation on the website... maybe one day, Claude Code will understand accessibility but that day is not today.

> We see signs that Opus 5.5 often suspects it is being evaluated, which challenges our ability to assess how it will act in the vast variety of real-world settings it is deployed in.

We can't test it properly because it knows it's being tested.

I'll have to try 5.5 on my work's Cursor account. If they really solved the communication issues, I might consider moving my personal account from Codex back to Claude Code.

I think this release is really going to give them a hard time selling Fable:

> On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest.

In general, "benchmark margins have become a less reliable guide to real-world differences" sounds like a big problem. It was certainly the biggest problem with the previous generation of Claude models for a different reason, because the non-code output was nonsensical, and that is not being benchmarked at the moment. But I'm not sure what to make of this admission.

  •     "benchmark margins have become a less 
        reliable guide to real-world differences" 
        sounds like a big problem.
    

    My guesses:

    1. Real-world use cases typically involve big, hairy, crufty, tech debt laden codebases and benchmarks do not.

    2. AFAIK "success" in a benchmark essentially boils down to "do the tests pass and do we get the right result?" which is something the LLMs have been achieving with ease for a while, except maybe for uber-challenging coding tasks that would be outliers in just about any workplace. Whereas real-world software engineering is usually just a bunch of CRUD... and "success" involves harder to measure dimensions like "maintainability" and "did you overengineer this?" and "how did you cope with a bunch of vague and maybe contradictory business requirements?"

    Having said all of that, I have never ever looked inside any of these benchmarks. I'm putting my guesses out here strictly in the tradition of "the quickest way to learn about something is to be wrong about it on the internet."

  • Opus 5 fucking sucks. Like it's horrible. I use Fable for coding and anything important and I use Opus 4.8 for things like recursive email categorization, transaction matching, and other stuff where I don't want to burn as much quota.

    In my experience Opus 5 is the worst of all possible worlds, it's dumb and headstrong. It just runs away with tasks you didn't ask it to do, is reckless, and basically is unusable in my experience.

    Not sure why but my guess is that this will be worse. Happy to be proven wrong.

    • I believe Opus 5 isn't meant to be spoken to by humans. It's great at executing but I reckon it's intended to be spoken to by other models such as Fable. I use Fable as the orchestrator, only speak with Fable, and all implementation, recon, design etc happens with Opus 5, with Fable reviewing (and translating).

      1 reply →

    • I use it frequently with a lot of success on "Medium" effort, it overthinks like crazy on higher levels, but YMMV.

    • Yeah if anything Opus 5 taught me how little benchmarks mean to the actual real world performance of these models.

      "Better" in every sense of the benchmarks and absolutely horrible results in my day-to-day work.

      The verbosity, goal post moving, tendency to leave work unfinished, over focusing on unrealistic root causes when debugging, etc... etc...

      It was the first time I actually pinned my models back because I just could not work with 5 for the price and performance it gave me. Hoping 5.5 is better this time around....

> Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety.

Nice. I was starting to think Haiku was going to be abandoned.

>Distillation attacks, in which attackers use thousands of fake accounts to extract a model’s capabilities at industrial scale, create safety and national security risks. Distillation allows bad actors to create highly capable models without the safeguards we build into Claude.

Maybe its a bit tiresome to read another comment of the form "what about your large scale distillation attack on the Internet", but this statement really just pisses me off. How very insincere in the most aggravating way.

  • My bet is anthropic has NN people org who work hard to distill open models in addition to trying to find what other useful materials they can download from shady torrents.

  • it's not safe unless it has commitees with orgies with that weird harry potter dude attached

  • "bad actors" boogeyman, and we should trust some tech weasel to do the right thing? Yea we've seen who they really are, once they get a sliver of power.

Wow! Though need to see its token efficiency to better assess. Been hearing rumors it generates much more output tokens per task.

  • my projection is that they are still gonna be pretty far behind, but they will sew it up in the next few releases. it feels like they were caught with their pants down on how much work OpenAI has put into that area, but i doubt there is some magical secret sauce that OpenAI has that Anthropic simply cannot catch up with.

Anthropic models are ridiculously expensive. I've stopped using any of their models months ago.

Opus 5.5 is now the recommended model in Claude Code's model picker, which is quite a claim, given how they struggled with capacity.

Has Opus 5 been absolutely terrible for people today? Like they took resources away from it to make room for 5.5? It is getting very basic things wrong all of a sudden.

My 20x plan was set to end tomorrow. The writing style and insistence on word vomit just became too annoying. Is Opus 5.5 worth sticking around for ?

Current models, especially Opus are almost unusable because they don’t respect instructions and their responses are infuriating. They are clearly designed for token consumption. I find myself wasting a lot of time just asking it to shorten or simplify its responses. I’ll give this new model a go, but I’m not holding my breath because the last model release was supposed to fix the very same issues and it didn’t.

  • Same feeling with oai models, wich I use 99% of the time. Sometimes I ask it about it, and it always come up with a likely explanaton but dear me it does many rounds of tool calls sometimes!

> It’s good at finding and fixing inefficiencies in software

Holy shit! Its happening!

Now if we can the AI to understand this *implicitly* so that it doesn't need to be stated upfront, we might be able to undo years of "premature optimization is the root of all evil".

I get this in my claude.ai usage:

Resets Get extra wiggle room to explore Opus 5.5. Expires Oct 22.

What the hell does this mean? There are weekly "resets" anyways. And there will be 4 of them before Oct 22.

It seems context length has completely fallen out of the discussion since we hit 1M, is that just going to be what it is now?

> Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5.

Thanks God. Opus 5 was a massive regression compared to Opus 4.8. People were spending tokens on fixing Opus-isms rather than actually doing work.

> In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work.

I was accepted into the CVP a little while ago. Does this mean I'll need to apply again?

Great that they listened! The improvement in communication style looks fantastic. Opus 5 was insufferable and I was on the verge of cancelling my subscription.

Finally, it seems like a good time to do some 'load-bearing' work on my project for a while

So it beats Fable 5.1, by quite a bit, on every metric? Interesting.

Might have to use my $20 Claude sub some more. I was moving away from it to a $100 OpenAI one to avoid the Claudese and poor token efficiency of Opus 5, given that I couldn't use Fable 5.1 with my tier, but this is worth trying out.

  • Why can't they let 20usd claude subscriptions access fable in CC, as openai allows you to use astra and max modes in codex - you just pay for it in more token use.

> Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1

Great so good luck using this for any low-level embedded or operating system development (unless you really, really like Opus 4.8 and want to be greeted by its familiar face after a few minutes of work!)

  • >safeguards hit: [Cyber]

    Yup. As unusable as Fable 5.1, for assembly on 80s 68k personal computer platform. Awful.

is OPUS 5.5 still not reading CLAUDE.md, failing to follow told tasks, inventing and hallucionating, just refusing to read files ("read the whole file" -> read 2-lines -> infere its wrong -> destroy the codebase), needing constant babysitting just because its so UTTERLY DUMB! i cant imagine going back to OPUS 5 - i'll rather jump out of the window as to use it EVER AGAIN!!

What happened to "slowing down"?

> cyber security and life sciences verification programs

chinese models can't come soon enough

we're already getting enshittification

Throwaway accounts posting after a few minutes some anthropic or another ai lab.

Infomercial at its best.

No wonder we are hammered with ai announcements.

  • Many people knew this announcement was coming. The betting markets suggested a very high likelihood of Opus dropping today. I was anticipating this quite a bit!