Comment by wg0
9 hours ago
No it is not. Only maybe for the noobs or vibe coders.
People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.
9 hours ago
No it is not. Only maybe for the noobs or vibe coders.
People who aren't afraid of rolling their sleeves into any code base? The difference is practically zero.
I agree; yes, I can see that they need a bit less hand holding each cycle, but I also see these "frontier" agents do some absolutely dumb shit that I have to correct and then I'm wondering if I'm the looney one here.
Maybe it's because people stopped watching what their agents are doing and stopped looking at the quality of the output. But I still see agents being absolutely mindless like a junior dev.
Recent example: it updated an an API to add newly released models to the backend. There's a list of models that require specific configuration for the reasoning effort and temperature or the API call fails. GPT 6.1 Sol misses this and code fails at runtime because the newer models need to be added to the list for special handling of temp and reasoning. Fixes it for one model and tests it for that model using an E2E test. But doesn't test the other models that were added for the same error condition...I had to explicitly ask it to do so and it finds them and adds them to the list and says "that's on me."
Yeah, not that smart.
You're not looney at all. Frontier models do dumb things all the time, especially on mature codebases. Just yesterday Opus 5.5 butchered the OOP model in a codebase I work on - it duplicated a load of classes that should have just been subclasses. A junior checked it in very satisfied that it was perfect. The LLM review passed it, the tests were fine, and it implemented the feature successfully. It's just the code design had poor taste and poor long-term maintainability.
I keep seeing this kind of thing over and over, and honestly it's not got _that_ much better since the big breakthroughs about a year ago.
For sure I happily vibecode stuff without worrying about it when it's a greenfield project, and if the LLM has written it entirely from scratch then usually it's well structured and sane. But making changes in messy, mostly human-written mature codebases is still a minefield.
Try a bigger code base or more complex stuff and you will easily see that the solution, speed and amount of problems Opus5.5 solves vs older models is relevant.
I am frequently running agents on a multi-microservice application workspace where I really need the 1M context windows, because they are filled to the brim when implementing features that require changes on several services and APIs.
This works fine with Opus 5.5. But it also works fine with GPT 6.1 Sol, Kimi K3 and MiMo 2.6 Pro.
It doesn't work equally well with Sonnet 5.5, interestingly.
I’ve been saying that. When you have no idea what you’re doing, you *need* the latest greatest model because it’s the only way to reduce errors.
For people who have some expertise, the models accelerate the grunt work, but you’re the one validating it.