Comment by arjie
2 hours ago
It always seems bizarre to me that people complain about software now. You can just write the thing you want. If you want to do things in parallel do them in parallel. Claude Code will allow you to run multiple instances in the same folder and let them communicate. Previously I used to let them intermediate through a communication bus but now they seem to be able to talk to each other.
I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.
Yeah I regularly have 5+ agents working in 1 repo simultaneously even modifying the same file. Since they were all running on my singular local machine. Builds and tests were interfering at 1 point so 1 automatically proposed writing a single script that basically mutex locked it with appropriate wait and timeouts. It had even added this to my own personal agentic todo backlog. They've proposed new skills for me.
They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.
And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.
I've now been trying for six months to get agents to produce decent code. As in readable / easy to follow, easy to change. I'm doing something really wrong. It's killing me. Everything it produces will work, buts it diabolically over complicated. I've built skills that have helped. But not massively. I've used other people's skills, in particular Matt pococks grill me and Dex hortlys show me. These have helped a bit. I work in enterprise, I want to be proud of what I'm producing, but trying to understand and then cajole and agents code into something that's good is exhausting. If anyone here has been through this and can share how they got through this, I'd really appreciate the help.
Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).
Same, it’s so smart but so dumb. I explain an architecture change to it and write examples of strongly typed Go code and how to store structure, it agrees and then proceeds to write some untyped string map ball of mud that has 900 specializations.
But maybe that’s on us, AI doesn’t care about all these special cases, it’s not debt to it as it will simply read them all when making changes. We’re obsessed with quality and what code is supposed to look like but those are human standards, AIs evolve to look at this complexity as a single picture, they can simply see through it so what is spaghetti code to us is merely some code to them that works as it should and is efficient. It’s interesting we can see how the two things drift apart, you would think at some point AI generated code should explode but it hold together unreasonably well in most cases…
Use hooks to run a bunch of review steps after -all- code writing steps your agent does and give it the exact review criteria you just described (via git hooks, or your agent harness of choice's own hooks eg https://code.claude.com/docs/en/hooks or AGENTS.md). So after every step where the agent writes the shitty untyped string map ball of mud, your orchestrator/main thread agent that spawned the code-writing-subagent spawns a follow-up review agent automatically that is given that output, your prompt that explains what well written Go code looks like, and even the sample/golden-path code of your choice to use as a style guide.
Each time you encounter a shitty thing you hate, add a new 'review type' / 'thing to watch out for' and just ask your agent to add it to your hooks for you. This works well with Claude at least.
I have about a dozen or so hooks that run on every integration branch my agents write that review for all sorts of things from correctness to spec, performance improvement opportunities, modularity, analysis of any dependencies added, 'definition of done', UI/UX, etc.
I recently told Claude it should run the whole suite of reviews twice. I will probably go on and proceed to having it run like 5 times eventually idfk.
But the more you start asking your agents to modify their own behavior, using the native solutions offered by Cursor, or Claude, or Codex, the sooner you'll start to feel better about the results.
2 replies →
But you didn't just write the thing you wanted, you're relying on other software that happens to do it, and if you were happy with the communication bus you would've stuck with that.
Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.