Comment by kinduff

1 day ago

Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.

I'd say that the first "mistake" we made was having it run in auto-merge mode. It was incredible to see what it could produce, and the speed with which it worked, but while the results seemed good, they were being produced at a rate that we couldn't human-verify. This is not to say that they were bad, but that we had no way to convince ourselves that they were good.

Once we started to review the code in more detail (and put the hospital into "human approval mode" to stop the momentum) we found one major thing that we corrected.

When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We've since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how/when the deferral should be actioned. This has prevented several items from falling through the cracks, but we're constantly refining what is a valid deferral and what should be fixed immediately.

There are several other things that we've corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.

  • > When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped.

    Could you structure the DAG so that after each node that contains the work, you have one dependent node that verifies each distinct requirement was implemented as expected?

    This way if a single requirement is dropped the system alerts you rather than it being silently dropped.

    It makes sense intuitively that if a task has nothing that depends on it the LLM might accidentally attempt to drop it (even purposefully as an optimization).

    • I like this. Today the parent issue closes automatically once every child is checked off, and nothing re-reads the parent’s original ask against what actually shipped. There is a deferred-scope auditor that catches the deferrals an agent admits to. Your "verify node" would catch the ones it doesn’t, which is the more dangerous kind. I think it would fit naturally as one final child per decomposed parent.

      2 replies →

  • Is a code comment and ledger the best way? Should the agent just fill out a form or something and attach it to the sub issue. This is how the hospital would work.

    • It often does that too, but we keep the ledger and the code comment as well, to ensure that if it doesn't get resolved by the sibling, that it's not lost. If the sibling does resolve it, it's removed from the ledger and the code.

      1 reply →