Comment by swat535
13 hours ago
People tend to not realize how far these models have become.
I remember when they would always hallucinate APIs that wasn't there or make up fields that didn't exist.. those problems are virtually solved now.
So with that in mind, why wouldn't AI be able to write better code?
The code would have to be maintainable by AI itself (operating based on the assumption that the future will be Agentic Engineering)
> I remember when they would always hallucinate APIs that wasn't there or make up fields that didn't exist.. those problems are virtually solved now.
I love how people have said this for over a year now, and no matter how often I try it, it's still just as broken.
Try it with any task that isn't in the training data. Say, a custom protocol where you need to cross-reference multiple implementations and the specs to really get it, and with no answers on stackoverflow or medium.
At first it's hilarious, but after a while it just gets exhausting. For all these real-world tasks you need to put so much effort in that it's much easier to just write the code yourself, even with the latest (current Gemini) and greatest (Opus 4.6/4.7) models.
If the parts for your project don't already exist, AI can't help you either. And if they do exist, why spend money on AI if the GitHub search is free and you can just fork and modify what you need?
I don't use Claude, but isn't the consensus right now that Opus 5 is worse than previous generations? I suppose you could just commit to always using Fable and never use Opus, but it
> might decide your 8yo is actually trying to build a cleverly disguised bomb so her request gets silently downgraded to a dumber model
My contention isn't that models haven't improved or won't keep improving – my position is that the business goals of our American AI firms (Anthropic especially) aren't necessarily aligned with continuing to make those improved models available to the public forever. We need only look at the Mythos/Fable split for evidence of this happening already.
https://www.reddit.com/r/ClaudeAI/comments/1vgpyni/my_opus_5...
https://www.reddit.com/r/ClaudeAI/comments/1vgq0jm/opus_5_af...
https://www.reddit.com/r/claude/comments/1vfvdgz/anthropic_l...
Yeah, the models aren't improving because Reddit told you so, and the businesses aren't incentivized to make better models. Galaxy brain take.
> Yeah, the models aren't improving because Reddit told you so
Is that the most charitable interpretation of my comment you could come up with? I don't think you're engaging in good faith.
> the businesses aren't incentivized to make better models
Anthropic and OpenAI are incentivized to pursue regulatory capture. It doesn't take a galaxy-sized brain like mine to understand this.
2 replies →
Explain how you square this with leaderboards showing models are clearly improving in ELO. These are blind taste tests, not benchmaxxing.
> Explain how you square this with leaderboards showing models are clearly improving in ELO
If I had to guess, the benches aren't measuring what people care about. But you'll have to ask the people in those Reddit threads why their vibes don't match up with your benches, I don't use Claude and don't give a shit.
> So with that in mind, why wouldn't AI be able to write better code?
> Since this is Claude, the models of the future might be more expensive, or more locked down, or might decide your 8yo is actually trying to build a cleverly disguised bomb so her request gets silently downgraded to a dumber model, etc.
I have no idea what you're trying to add.