← Back to context

Comment by brink

5 hours ago

I also have found that AI has not lived up to many of its promises and have dialed back what AI gets control of. My projects were turning into unmaintainable messes. The people who say coding is solved aren't paying attention.

Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.

  • I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.

    Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance

  • I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.

    • It is also easier than ever to build specs and have nice easy to maintain projects. It just doesn't happen magically via few shot prompts :)

  • Have you considered that maybe this is a reflection of your skills rather than that of the LLM?

    • It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.

      I’m saying it’s probably multiple factors and both you and GP are right.

    • Have you considered it isn't?

      Save your "you're holding it wrong" if you're not going to suggest how to hold it.

      Cult speak escape hatches are intellectually lazy.

      7 replies →

  • I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.

People, especially those who don't do software development, often conflate coding and software development.

The models aren't good at architecture and design. But they take direction on architecture and design and design well and can refactor code quite effectively. AI agents can absolutely be used to clean up vibe coded code bases once you figure out if the investment is worth it. The mess can be avoided if you give them sufficient guidance on architecture and design upfront.

That said, doing so purely in text form doesn't feel great right now. I've been thinking about UML lately. The problem with that was the roundtrip after the code was generated and then the implenetation happened. I don't necessarily think UML is the solution, but neither is walls of dense text.

  • >The models aren't good at architecture and design.

    Can you elaborate to back up this claim? WHat exactly is your yardstick for "being good at SW design and architecture"?

    Because I found the current SOTA AI models being great at architecture and design, much better in fact than most average real-world devs. Is your yardstick just the John Carmacks of the world by any chance? Because most devs are not John Carmack. They are also not Linus Torvalds, they are not Stallmann, etc.

    Maybe your LLM experience is still stuck in the 2023 era of ChatGPT?

I've had the opposite experience. Spec driven development really makes things much easier to maintain. If you're just yolo'ing it and throwing prompts around things will fall apart fast.

Were you reading through the code it was generating to make sure the flow was intuitive and comprehensible for each PR? Projects only turn in to maintainable messes if you blindly merge in unmaintainable messy code.

As usual, when people post comments like this they never specify what model and when they used it. There is a huge difference between GPT 3 and 6 for example.

Late 2025 also had a step change when agents could largely code autonomously without handholding like previously, and to be honest it's not worth hearing opinions about AI from before that time, that's how significant the change was.

Oh my god. I am sad to have to acknowledge this even if the models have “theoretically” gotten better. I am not a programmer but I am super opinionated with the design and structure of the code and need to make sure whatever I write/ask to generate and use for my own tasks is understood by me at least once. I wasn’t this way in 2023. I am constantly reading/trying for an agent to help me write simple code without spending too much time. But so far nothing has worked. I wish someone figures it out otherwise, I personally would not be able to realize the AI agent productivity benefits that others are supposedly seeing.

  • If you specify a structure, pattern, or design, it will stick to that design after a code review phase. Just tell it what standards you have, and after two rounds, it will have written what you described.

I treat AI like an intern or junior team member. With enough guidance, they can contribute a lot but you can't let them loose without supervision or they usually will produce a big mess. As of now, you are still responsible for overall architecture. One strong indicator that something is going wrong are big pull requests where the AI has rewritten large sections of the code.

AI is a useful tool but it absolutely needs a lot of human guidance.

Unlike a compiler it won't give up at the first sign of trouble but that just means it left alone it will dig bigger and bigger holes.

Treat prompt engineering as a discipline and refine your technique. When it produces garbage throw out the work and start over until you figure it out.

Vibe coding has always had and always will have one fundamental flaw: you didn't write the code it produced, therefore you don't have a good understanding/mental model of it.

On the other hand asking these clankers "review the feature branch I wrote" and "review my entire codebase for bugs" or "help me debug this" has saved me months of prospective work.

And more recently most major models have been getting _really_ good at RE, for example you can have OAI models (and maybe A/'s if they don't refuse) use idalib MCP and reverse-engineer stuff from start to finish, then follow up with GLM 5.3 for vuln assessment and exploit PoC.

Stuff that used to take weeks or months now just takes a few hours, or less.

Were you reviewing the changes? I've seen lots of projects with this problem. But if AI PRs have to pass the same review bar as any other PR then it shouldn't be an issue in theory (if you can actually maintain the discipline).

  • Isn't it tiring to keep up with considering the speed of the generated code? And if you want to be careful with the review you lose a significant part of the speed advantage.

    • Yes, but it's the only option if you want to retain a maintainable code base. And you can always slow down. You'll probably still be a bit faster than you were before.

    • Assuming we are measuring time and

      `total = dev + review`

      If dev approaches zero, but you review at the same pace as you always have, are you in a better position? Yes.

      Will you potentially have a backlog of code waiting for review? Also yes.

      Would you prefer to be waiting for the dev team for all of the time instead, then still have the same amount of reviewing to do at the end of it? Absolutely not.

    • It's the same as any other PR. Keep the changes small and contained to that feature. AI can do this if you instruct it to, there is no need to vibe code some 40k line monstrosity.

I am not saying you are right or wrong, but I don’t think you can make such a broad conclusion simply because of your own experience.

You say you used AI and your projects turned into unmaintainable messes, so your conclusion is that it means AI is not living up to its promises.

I guess if the argument is “AI makes it so you always get a great result no matter how you use it”, then your argument is sound. Your projects not working out proves that AI doesn’t always work no matter what.

However, that doesn’t mean you can’t use AI to create sustainable and well organized code. Failing to do something doesn’t mean it’s impossible and anyone who thinks they can is not paying attention.

I can’t run a marathon. If I went out and tried to run one, I would get a few miles and collapse, failing completely.

I don’t think it would be reasonable, though, at that point to say “running a marathon is impossible, anyone who says they can do it clearly lying. I tried and didn’t even make it 5 miles!”

I wish people would stop assuming their experience with something is the only possible truth.

  • Stop being overly polite to these people. It is seriously coming down to the point of utter delusion.

    As ai becomes better these people will begin changing their story because it’s utterly obvious what’s happening.

The people saying programming is solved couldn't program in the first place is what I've noticed. So to them it really does feel solved, things "work" and they don't have to learn what they don't know about programming.

  • I've been programming professionally for decades. LLMs are extremely useful. At this point if you haven't figured out how to get value out of them, you're either holding them very wrong or you're being willfully ignorant.

    • I think it depends which LLM tool you're using. If you're using an older, worse model (the kind that you can use for free), the experience is significantly more frustrating. On the other hand, I'd say that the current best models are very useful with a skilled operator.

      1 reply →

  • There are a couple of notable counterexamples here (nobody sane thinks Carmack doesn't know how to program, for example), but by and large I agree with your observation. The people excited about programming with LLMs are, on average, people who weren't good at programming to begin with. Still, given that these counterexamples do exist I try to avoid painting with an overly broad brush for the sake of nuance.