← Back to context

Comment by dbmnt

5 days ago

"throughput is limited by my approval"

This is exactly where I'm at with AI. (Mostly via Claude Code, but I'm not sure the harness, or my workflow in particular, is the important part.)

I am increasingly wondering though, is there really a valid reason to keep blocking on my approval? Most of the time, I wind up saying yes anyway, because the model has a valid, efficient solution.

What if it's faster at this point to just fix the mistakes?

Scary thought, but it seems like we're close, or already there.

The problem is the agents keep going off the rails and will hack external servers to achieve the goal you give it. Letting them run full speed overnight and waking up to find they have commit multiple crimes is not ideal.

  • in orchestration, the agents are only fed the tasks given by the person running the orchestration side and if we are talking about code changes and git commits and deployments, this is pretty benign behavior. "Hacking external servers" I am pretty sure would have to be written in the prompts to begin with to allow this to occur in the first place.

Think of it this way. Let's say you have 80% success rate in decision making. Great, you'll just fix the other 20%.

Here's the problem: decisions happen in sequence, layering on top of each other. If you have 3 chained decisions of 80% success rate each, your combined success rate is now 51%, marginally better than a coin flip.

I just YOLO it. Just don't give it access to anything that can produce permanent consequences. The cost of a mistake is maybe one hour of fixing it. If what it produced is 80% right and 20% wrong, by letting it run overnight, you've gotten 80% of the work done that otherwise wouldn't have happened.

  • If you have work building on top of other work, your "20% wrong" quickly turn into "80% wrong".

    If feature B depends on implementation of feature A, and feature A happens to be on the 20% wrong side, then feature B will be wrong too, regardless of whether you happen to have better luck with it being implemented on the 80% right side. It'll simply be built on a broken foundation. It WOULD have been correct if the foundation was right, but it wasn't.

    This is actually very common in software development, where you rarely have stuff happening in isolation. A lot of features deliberately touch each other, build on top of each other, and even more so when you factor in "accidential" overlap due to unclean technical boundaries (in spaghetti code, everything touches everything).

    • Yeah this is why you have your top guy (Fable/astra) split everything into discrete tasks and go one layer at once. Get 10 subagents working on parallelized tasks that have no dependencies on each other through the night.

    • This was a problem maybe 6 months ago. With Fable/Astra orchestrating Opus/Sol subagents it now works really well.

      Ask it to interview you before you start and it'll get your constraints pretty well.

I guess it depends on what you’re doing in my current work I couldn’t I have even with Opus 5.5 tons of iterations.

For example I was optimising a checkout experience and I at least did 30 iterations till I was happy.

But there is also other stuff like namings. They pick good names but if you work with exchangeable vendors I need even more explizit names

However of course I’m trying to get as much in linting and agents.md

But even if I would say Yes to everything I couldn’t come up with things to do fast enough.

But maybe skill issue ?!

I had a similar thought a couple of months ago. I did some work to fully sandbox the agent - from the rest of my computer, from my user data, from other projects and from the world at large. I now find myself doing a lot on auto, I do a thorough human edit and review and then I squash. If it gets it wrong, I can throw the changes away and rebuild the sandbox. Having a good plan, good automated QA and all the other things that already helped is essential. It needs the right tools for whatever it's working on - for example: if you want frontend web dev, you need to get it using something like Playwright and looking at the screenshots.

I only leave it truly unattended if it's working on a very tight improvement loop, for everything else I'm still checking in on it between working on other things.

I still find agents need a lot of guidance and steering to produce the kind of work I want, but auto mode in a strong sandbox is very useful to me to take a bite out of that. It's particularly good for exploring problems experimentally - where most of the exploration might be thrown away after settling on a solution.

Having used it sandboxed, I wouldn't dream of letting auto mode run outside it. It's very creative at trying to work around the constraints of the sandbox (legitimately, not to escape it) and the classifier for auto mode seems very permissive with the right context.

For context, I'm using per-project VMs with very limited egress and restrictive mounts. Self-built tool to glue it all together, currently unreleased. There has been an explosion of sandboxing tools recently, none of which was quite what I wanted. Heavily inspired by Gondolin <https://earendil-works.github.io/gondolin/>, but fat long-lived VMs.

"Auto" mode in Claude works really well if you don't want to do dangerously skip permissions.

The approval-every-call model does train you to say yes, so it stops being a control. What I've found works better is gating on the class of action rather than each call. Reads run freely, writes and shell commands that mutate ask, and the thing that always holds is a new outbound host or a credential the agent hasn't used before. Most turns then run without interruption, and the rare prompt is one you actually read. It also answers SchemaLoad's point below: overnight runs are fine if the only thing they can't do unattended is reach a host nobody approved.