Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
I find it funny that the original complaint was: "AI made a mess of the codebase".
Your response was: "Well you're not doing it right, but these hermes devs know what they're doing".
But the blog post you linked to shows their prompt, which is:
> I want god files broken up. I want simplification across the board. I want unification of helpers and methods that can be reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want interpretability of the codebase and how things connect to each other up.
So it sounds like AI made their code a mess too. They then tried to make the point of how much money they saved cleaning up the code with AI, that AI made a mess of to begin with.
And if you look at the merged PRs on that project, a ton of them are bug fixes... to the code the AI wrote. And that's been my personal experience too: AI creates a huge amount of churn in a codebase. Just vast amounts of PRs fixing code that the AI itself wrote.
You're using quantity metrics to answer a quality question.
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
If you really want to check some quantity metrics to try to reason about code quality, look at whether "fix" PRs are overall LOC neutral or negative (not counting tests). In this project, almost every "fix" is an addition. Worse, almost every fix is more branching.
If almost every PR is some sort of fix, and most of them add branching, and there's thousands of them weekly... That leads to only one place and I want to be nowhere near it.
I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
Do you mind sharing any code from what you have produced? People talk about LLM successes and failures, but what's there to really talk about when the code can speak for itself?
In case it is unclear, I am genuinely curious. I have great success with chatbots, but vibing coding has never gotten me further than a proof-of-concept.
It could be that small variations in prompting lead to large differences in quality of output, especially over longer horizons.
I’m saying it’s probably multiple factors and both you and GP are right.
Have you considered it isn't?
Save your "you're holding it wrong" if you're not going to suggest how to hold it.
Cult speak escape hatches are intellectually lazy.
Sure, happy to provide you with an example of how to hold it (turns out Steve was right) =D
https://github.com/NousResearch/hermes-agent is 99% (just a guess) LLM generated. 1140 closed pull requests this week. 1.5k closed issues. The github insights page for commits doesn't load for me presumably because it can't handle this scale of commits. But I estimate ~1K commits per day on average.
There's a blog entry https://nousresearch.com/refactoring-hermes-with-1393-agents that details some work that was done by LLMs to refactor and improve the code.
I guess they know how to hold it?
I find it funny that the original complaint was: "AI made a mess of the codebase".
Your response was: "Well you're not doing it right, but these hermes devs know what they're doing".
But the blog post you linked to shows their prompt, which is:
> I want god files broken up. I want simplification across the board. I want unification of helpers and methods that can be reused. I want less if-if-if-if-if-if-else routing. I want code legibility up. I want interpretability of the codebase and how things connect to each other up.
So it sounds like AI made their code a mess too. They then tried to make the point of how much money they saved cleaning up the code with AI, that AI made a mess of to begin with.
And if you look at the merged PRs on that project, a ton of them are bug fixes... to the code the AI wrote. And that's been my personal experience too: AI creates a huge amount of churn in a codebase. Just vast amounts of PRs fixing code that the AI itself wrote.
You're using quantity metrics to answer a quality question.
I had a look at the kind of issues that are reported at that project (there's 15k of them, so I can at best assess a couple). It looks like a complete mess: A lot of concurrency and resource mismanagement issues and edge cases that in a better-managed project would have been avoided by construction. They will now will likely be solved by more defensive programming, driving overall complexity ever upwards.
If you really want to check some quantity metrics to try to reason about code quality, look at whether "fix" PRs are overall LOC neutral or negative (not counting tests). In this project, almost every "fix" is an addition. Worse, almost every fix is more branching.
If almost every PR is some sort of fix, and most of them add branching, and there's thousands of them weekly... That leads to only one place and I want to be nowhere near it.
1 reply →
[flagged]
I’ve had the best luck by spending quite a bit of time going over the big picture architecture up front and then diving into the modules to further refine the details, making sure to generate step-by-step chunks of work in Markdown format for implementation. I’ll spend literally a couple of days doing this before starting any coding.
Edit: My latest project is all GPT-6 Astra High. It takes a lot of steering to keep it from adding a bunch of, while useful, features that are not strictly enough to the point. That main issue is it’ll use a lot of extra tokens in the process!
What was your process?
Do you mind sharing any code from what you have produced? People talk about LLM successes and failures, but what's there to really talk about when the code can speak for itself?
In case it is unclear, I am genuinely curious. I have great success with chatbots, but vibing coding has never gotten me further than a proof-of-concept.
2 replies →
Meaning OP is a better programmer then a statistical model which produces the most probable results?
Meaning a bad workman blames his tools.
A bad workman blames his tools.
A good workman shuts up and finds better tools without complaining.
1 reply →