← Back to context

Comment by AIblemblio

10 hours ago

For sure people who don't grasp the difference between models, might be stuck in 'good enough' models.

But Opus 5.5/GPT is such a game changer in comparison to sooo many others, its still a moat for now.

You've got to think the comments about "good enough" are people who have not yet tried Opus 5.5. I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.

As for Mistral - I got really excited when they said Large 4 was focusing on being #1 in cybersecurity, because that's somewhere that they genuinely could edge out Anthropic & OpenAI. Have it actually solve problems, instead of Anthropic flagging "you tried to find a null pointer exception bug in your own code, we're now reporting you to the US government". But on the Mistral benchmarks I'm seeing, this looks very disappointing, but at least they haven't entirely given up. I genuinely thought Mistral had given up on new general models. They need to learn the bitter lesson all over again.

  • I feel like the "good enough" argument isn't about how big the gap between models is but about how good they are at solving the tasks at hand.

    The capabilities of all models increasing so much all the time means there are simply less and less tasks you need a frontier model for.

    Even if Opus 5.5 is 500x better than Deepseek, if deepseek can solve all my problems, why do I need to pay for more?

    • Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed.

      Obviously any model will do if you use it as a better autocomplete.

      I believe that there is a large gap in expectations between different workflows.

      Until the AI like reads my mind and produces perfectly production ready apps with minimal intervention from my side, there is still going to be room for improvement.

      2 replies →

    • If Opus 5.5 is 500x better than Deepseek, but Deepseek can solve all your problems, maybe you need to work on better problems. If you don't, and you're in business, your competitors will work on the better problems. If you're an employee, your employer might prefer to pay Anthropic instead of you. If you're doing projects you're interested in, you can tackle more ambitious projects with a more capable model.

      This morning I elicited a microkernel operating system from Opus 5.5. Well, mostly. It doesn't implement task switching yet; we'll see if it runs into a wall at some point. But it boots in QEMU, and it's running a user process in ring 3 and serving web pages.

      6 replies →

  • I use Opus 5.5 daily for my job. I am aware (and in awe of) it's capabilities.

    Look at the context in which I used that term 'good enough'.

    What i was saying is that there are tasks for which a dumber model can be good enough, and for organizations with sovereignty/ privacy concerns, those concerns can be strong enough to incentivize the use of a dumber model.

  • > I haven't been this struck by a step change since 4.5/4.6. It's a bigger jump even than when Fable first arrived.

    I had the exact same experience. And unlike Fable, it doesn't gobble up your entire usage limit in a few hours.

    I always wonder what the "good enough" people are actually using it for.

    • “Good enough” as in we don’t see any point in 1-shotting everything we want to build in lightning speed. If Opus 5.5 can 1 shot it then that product is essentially commodified, no point in any one building it except as an internal tool.

      If Opus can’t 1-shot it, then it must rely on our human intelligence which can be complemented well enough with a dumb model as a frontier model.

      1 reply →

    • The thing is, for how long? Are they going to keep giving you "so much intelligence" for a "small" subscription dollar amount? When they really, for real, need to start making money to cover their costs, what do you think it's going to happen? Suddenly you will start having tasks that a "good enough" model is going to be fine.

      1 reply →

I agree that Opus and GPT are almsot surely better, but so many real users are nervous enough about giving Anthropic and OpenAI access to all of their internal information that they may be willing to stomach worse models if it gives them more security.

The real question is if this model is good enough that it can still accelerate work, and not be a hindrance to real work like older Mistral models often were.

If they can do that, they'll have customers.

  • The Navier-Stokes fiasco made me push for local/controlled models very hard. "Can't rule out" that they stole data (backed up by their backdoor offers of sharing credit).

    If these companies will steal from deep pockets like Disney or Sony (some of the most infamously litigious copyright trolls to ever exist), they won't think twice of stealing every bit of code you upload to them.

    If your code passes through an AI company's servers, you can assume you just gave it to them. In turn, when your competitor tries to copy that new feature you just added, the AI is now trained in exactly how to copy you and eliminate your competitive edge. Unlike your employees, the AI isn't bound by the same rules and even if it were and violated them, your company probably doesn't have enough money to prove it in court (and that's if we somehow reverse some of the stupid "AI is the most transformative use of copyright I've ever seen" judges who have drunk the coolaid).

    Most companies could build the compute to run GLM or Kimi models for way less than the potential loss due to IP theft from using third-party systems.

    • I would also add to this, there are ways to use customer data to improve your model outside of just using it as “training-data”.

      A simple loophole, use the code to create an RLVR environment where the resultant code is the end goal / max reward. Technically the customer data is never trained upon, but effectively you’re using it. Even better, use the code as a seed to generate synthetic data similar to it and use that synthetic data as rewards in an RLVR model.

      Unless you can host the ChatGPT model on your own servers, which I know some enterprises are doing, I don’t think there’s any hope of protecting your data / competitive advantage from these frontier companies. Better to be paranoid, than be commodified by these companies.

    • tbh to me if the AI company writes all of your code & your eng don't even review it anymore then… the AI company _controls your company_. maybe that's ok if you make widgets but less ok if you do anything in dev tooling, security, or [insert market they may suddenly decide to compete in].

No it is not. Only maybe for the noobs or vibe coders.

People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.

  • I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.

    Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.

    Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."

    Yeah, not that smart.

    • You're not looney at all. Frontier models do dumb things all the time, especially on mature codebases. Just yesterday Opus 5.5 butchered the OOP model in a codebase I work on - it duplicated a load of classes that should have just been subclasses. A junior checked it in very satisfied that it was perfect. The LLM review passed it, the tests were fine, and it implemented the feature successfully. It's just the code design had poor taste and poor long-term maintainability.

      I keep seeing this kind of thing over and over, and honestly it's not got _that_ much better since the big breakthroughs about a year ago.

      For sure I happily vibecode stuff without worrying about it when it's a greenfield project, and if the LLM has written it entirely from scratch then usually it's well structured and sane. But making changes in messy, mostly human-written mature codebases is still a minefield.

  • Try a bigger code base or more complex stuff and you will easily see that the solution, speed and amount of problems Opus5.5 solves vs older models is relevant.

    • I am frequently running agents on a multi-microservice application workspace where I really need the 1M context windows, because they are filled to the brim when implementing features that require changes on several services and APIs.

      This works fine with Opus 5.5. But it also works fine with GPT 6.1 Sol, Kimi K3 and MiMo 2.6 Pro.

      It doesn't work equally well with Sonnet 5.5, interestingly.

  • I’ve been saying that. When you have no idea what you’re doing, you *need* the latest greatest model because it’s the only way to reduce errors.

    For people who have some expertise, the models accelerate the grunt work, but you’re the one validating it.

Please, give it another 6 months and they catch up. The American labs are currently trying everything they can to block others instead of advancing their models, trying to build an artificial moat. The American models are not that great, they are good, and they have a lot of agentic workflows in the back, but its basically a hardware limitation at this point. Once the HW makers catch up, and we can move away from the Nvidia monopoly, things will speed up quite a lot IMO.

  • We've been hearing the line about them only being a few months behind for a year now, during which time O/A have grown their revenue like 10x, haven't they?

    • We've also been hearing we're 6 months from AGI for about three years, and here we are.

      "Now, here, you see, it takes all the running you can do, to keep in the same place. If you want to get somewhere else, you must run at least twice as fast as that!"

      3 replies →

    • Those are two different things. The market is expanding, so even if competitors are catching up, you can have your own revenue, in absolute terms, grow.

      2 replies →

    • The thing is O/A have been much louder on pacing the frontier, lately.

      And yes, open weights are still behind, but are catching up.

It's an improvement, but game changer might be a bit of a stretch. If I lost access to Anthropic or OpenAI models tomorrow, I would be annoyed, but would reach for a slightly inferior model. Last year I wouldnt be able to say the same, and rhe challenge is that the moat is drying up fast. Whether its general improvements in model training by other competitors, or straight up distillation of SOTA models, the moat is shrinking and the available capital and spend for American model providers is going to dry up quickly as competing good enough models are adopted by more consumers.

It's especially the case as more non-Americans look to self hosted models and domestic cloud inference providers using open models that the US providers who are still leading the charge need to drastically drop their prices and find a path to profitability in order to maintain their lead and retain the advantage they had as AI turns into a commodity (which is happening faster than I think even the frontier labs initially predicted).

I use Opus 5.5 at work.

I use MiMov2.6Pro, DeepSeekv4.1Flash, GLM5.3, Hy4, Qwen3.8 and KimiK3 at home. Opus5.5 is not a game changer.

  • I do a broad amount of diverse experiments/projects I always wanted to do and throwing Opus5.5 against it just works

    I have to admit, Sonnet got really good too.

    But Opus just uses tools, a broad spectrum of it, etc. it feels like sure if you add some router behind it you could split it up if you need to but if you give me the choice, its opus allll day long.

> X is such a game changer

I hear this literally every other week about whatever the newest FoTM model is.

Unless you can provide concrete examples of things you can do with them that you simply couldn't do with last week's model, it's absolutely meaningless.

Yeah, I agree with this. I think the "the models are good enough" narrative is a myth. I've heard it so many times over the last year, but the model number keeps changing...

There is no ceiling on what you can accomplish with more intelligence, so there will always be a market for the best models, and that market is likely to just keep growing. If Opus 13.5 can one-shot a profitable company or discover a new disease treatment or whatever you can think of that a swarm of relentless super-geniuses could accomplish, companies (and governments) will throw money at it.

I also think there will always be a market for many sub-frontier models that will continue to grow rapidly as well, because "good enough" is definitely a thing for a given task.

  • One could still argue that models are good enough for a given task. I primarily use Opus at work for writing code and I realized that for my usage the intelligence of Opus 4.8 is more than enough. Sure the newer models are better but I can still do my work with having access to newer models

Despite Opus 5.5 got really bad the last days for me. Looks like they nerfed it again. This is extremely unreliable.

  • Or maybe they secretly believe you are trying to distill their models and are deliberately degrading your experience. Who knows with them?