Comment by lucianmarin
6 hours ago
Yes. I started two projects with AI from scratch. Both abandoned, complete mess. Projects without AI are so easy to manage, maintain, add/remove features, etc. I use AI as a search engine on my projects instead of Google. I ask what's wrong with my code and change or improve it myself based on my experience.
I have the opposite experience. I don't have the patience sometimes to clean up my code and stick to coherent conventions and organization even though the will is there. With my AI projects I watch it and if the AI starts drifting I ask it to go through and look for convention/directory structure violations and it happily cleans everything up in about 10 or so minutes.
Are you in Python by chance? Python has a lot of crazy hidden/inexplicit/spooky action at a distance stuff (especially in the frameworks) that can make LLMs gunk up code by defensively programming or just burn context chasing data provenance
I've been a vibe-coding skeptic for years, but because of the math breakthroughs of the past few weeks I decided to experiment with the latest models on some test projects. They're a lot more capable than I thought they would be. I agree that it's easy to create an unrecoverable mess, especially when you're one-shotting a lot of features without detailed instructions. But I find that as long as I'm strict about the API boundaries and force the agent to work in small chunks, it's pretty effective. As one example, I got it to write an SVG renderer in a few hours (not the whole spec, but most of the path features and text rendering), which would have taken me at least a week just for the coding part, plus extra time to learn the algorithms.
It is also easier than ever to build specs and have nice easy to maintain projects. It just doesn't happen magically via few shot prompts :)
You could have also copied an SVG renderer that implements the whole spec from whatever open source project the model copied it from.
It didn't copy any source code from any external projects. I had it write a stratified sampling renderer for ground truth, then had it implement feature by feature by matching the pixels. Unless you mean it "copied" it in the sense of third-party code being part of the training data. I don't think that definition of "copy" makes any sense given how these models represent embeddings. It would also imply that humans are "copying" the things they've learned from.
Have you considered that maybe this is a reflection of your skills rather than that of the LLM?
It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.
I’m saying it’s probably multiple factors and both you and GP are right.
Have you considered it isn't?
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
There's a blog entry https://nousresearch.com/refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.
I guess they know how to hold it?
4 replies →
I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
What was your process?
3 replies →
Meaning OP is a better programmer then a statistical model which produces the most probable results?
Meaning a bad workman blames his tools.
2 replies →
I'm reading y'all comments and it seems we're still finding our footing, and will for some time. I have the luck that I have access to pretty much all frontier and other models alike. I've been extremely negatively biased towards any LLM use in the start. Then the influx from juniors came, then the revolt of doing PRs on such slop came, then some structured methods how to do LLM work came, then vibe projects came, etc, etc. I literally have all the described experiences you've guys mentioned. From good, to bad, to ugly. It's like there's no one particular way about doing this, and no two projects share the same approach - just like ye olde times.
When and what model?