Comment by Systemerror7A69

11 hours ago

I feel like the "good enough" argument isn't about how big the gap between models is but about how good they are at solving the tasks at hand.

The capabilities of all models increasing so much all the time means there are simply less and less tasks you need a frontier model for.

Even if Opus 5.5 is 500x better than Deepseek, if deepseek can solve all my problems, why do I need to pay for more?

Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed.

Obviously any model will do if you use it as a better autocomplete.

I believe that there is a large gap in expectations between different workflows.

Until the AI like reads my mind and produces perfectly production ready apps with minimal intervention from my side, there is still going to be room for improvement.

  • How I feel it Opus 5.5 and Astra 6 are both quite disappointing. They still write shit Rust. And they are very slow.

    Now my provider serves DeepSeek around 300 tokens a second. The code is shit but I can have a few more rounds of corrections.

    DeepSeek was maybe a dollar for the full PR. Opus/Astra over 50 dollars. And double that if you use their fast variants which matches the DeepSeek speeds on certain providers.

    Yes, it is so good I rarely test new models anymore. Or think about cost.

  • > Many on HN still have the opinion that you must understand every line of code in the project, and that all is lost should you merge code that wasn't reviewed.

    Uhm, yes?

    • Last week I elicited a parametric 3-D design model of a sand battery, a simplified thermal circuit model of it for rough exploration, a backward Euler multiphysics solver to evaluate its performance more accurately, and a volumetric visualization for the resulting temperature distributions, including some custom WebGL shaders: http://canonical.org/~kragen/sw/sandbattery/

      It's about 12000 lines of code. I could probably understand every line of it if I spent a couple of months on it, although I'd need to learn some things about WebGL, numerical computation, and solid-state physics. But I elicited it in four days. Probably I'd be better off spending the next couple of months doing something else instead.

      So, I was faced with the question of how I could keep the artifact thus created from being completely valueless. My solution was to export the heat-equation solution produced by the solver as (gzipped!) CSV, so that I can use my own code (that I do understand every line of!) to check that the solutions found by the solver are, in fact, solutions.

      If that check checks out, the backward Euler solver may be of some value even if I don't understand every line of it.

If Opus 5.5 is 500x better than Deepseek, but Deepseek can solve all your problems, maybe you need to work on better problems. If you don't, and you're in business, your competitors will work on the better problems. If you're an employee, your employer might prefer to pay Anthropic instead of you. If you're doing projects you're interested in, you can tackle more ambitious projects with a more capable model.

This morning I elicited a microkernel operating system from Opus 5.5. Well, mostly. It doesn't implement task switching yet; we'll see if it runs into a wall at some point. But it boots in QEMU, and it's running a user process in ring 3 and serving web pages.

  • > maybe you need to work on better problems

    I have enough real problems in life. I don't need to invent new ones just because a new technology is available.

    Many of my problems in life are fully solved far past my satiation point by a 3b model that costs me nothing to run.

    Many others are not.

    But in either case, when I am acting and living wisely, almost all of my problems exist prior to the existence of technological solutions to those problems.

    This is also true for the customers and employers that I care to work with. This has changed in me over time, but I now try my best to avoid inventing new problems. The world has enough big, important problems already.

    > This morning I elicited a microkernel operating system from Opus 5.5.

    This is cool but also a good example. I don't need a personalized microkernel just because it's possible to have one.

    Maybe I need one and I don't know it, but the problem statement definitely isn't "I have inherent desire for a personalized microkernel".

    • You've misunderstood what I was talking about; that's a different meaning of the word "problem". Possibly you did not intend to start a merely semantic argument, but that's what you ended up doing, so I am unfortunately going to have to point at the dictionary.

      You're talking about definition 1 in https://en.wiktionary.org/wiki/problem, "A difficulty that has to be resolved or dealt with," with the examples given being racism, addictions, and lack of access to health care. Those aren't the kind of problems AI can help with.

      The kind of "problem" that AI can help with is definition 2, "A question to be answered, schoolwork exercise." That is the character of engineering "problems", although they are more open-ended than schoolwork exercises, because there are many defensible tradeoffs. "How can I build a bridge here?" or "How can I improve the fuel efficiency of this vehicle?" is a "problem" in the sense of a question to be answered with knowledge and effort, not in the sense of being similar to racism or addiction.

      If the questions you're thinking of are so easy to answer that they can be easily answered by a hypothetical AI model 1/500th as good as Opus 5.5 — well, think up some harder questions.

  • The best problems to work on are not necessarily the hardest ones, nor the ones that need the most intelligence. They're the problems you, or other people, actually have. Are you going to give up on painting your deck because it's too easy and you don't need a 500x genius to do it?

  • Yeah no, if you elicit Opus 5.5 , anyone else can, and you have no moat either.

    But if on the other hand, I mostly use my human intelligence and just need a dumb model to complement my human intelligence at low cost and high speed (say review every commit to catch obvious bugs), I have a much better chance of building an actual moat than you do.

    But outside of coding, it’s even more clear that you don’t need frontier intelligence. My customer service agent is very happy with a 100B param Deepseek flash model, thank you!

    • Yes, obviously a microkernel operating system that you can vibecode in a morning is not a salable product. But it's still a perfectly good operating system, and sometimes people need those, and they're a huge pain in the ass to write, especially to debug. But sometimes off-the-shelf OSes aren't good enough. Something like 50% of the embedded operating system market is still "Other/Custom".

    • > say review every commit to catch obvious bugs

      I’m using subscription models for exactly that, better models catch more subtle bugs, and they catch them faster. It works out far better in terms of work-hours saved.

      Also A/ then OAI slashed token pricing by 2x~5x on their latest models