Five months treating bugs like patients and coding agents like a medical team

2 days ago (cockroachlabs.com)

This feels to me like a lovely example of both high-quality systems design and information theory. From my perspective, one of the more abstractly interesting thoughts surfaced by this write-up is the degree to which you've "encoded" a highly complex set of relationships through use of metaphor.

That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.

Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.

Thank you for sharing and the care you put into writing this!

  • Thanks for the thoughtful response. We too were surprised by how deeply we could take the model of the teaching hospital to software engineering. Whenever we think to expand the system in one dimension or another, the teaching hospital model seems to have a nearby analogy.

    I will confess however that some people internally find the model confusing. For example, one user couldn't remember that to get an issue actioned, they needed to put it into the "waiting room". They wanted to be able to action their issues without learning how complex medical systems operate. There seems to be more work on our plates to make this resonate with all users.

    • I'll be very interested to read any followups about the part you described at the end about getting it to do more "teaching," both from the perspective of how we can use this to help newer devs out but also from the perspective of helping ourselves understand the software being built.

      One thing I've been worrying about recently is the notion of comprehension debt and making sure the humans can still understand the system (so as to be able to do incident response or something if the LLMs are down). The slower, more methodical workflow described here should definitely help by allowing checks like proper docs whenever a piece of the architecture changes for instance, but I'm curious what else could be done to help the humans understand as much of the implications of what the AI is doing as possible.

      Edit: having the code author and reviewer be completely separate like you described below is also probably pretty helpful for making sure changes are comprehensible from the outside.

    • Now imagine what it is like to run an actual hospital with real people whose lives (or the lives of their loved ones) are often at stake, with underpaid and overworked staff and with messy biological creatures as the subjects instead of bits and bytes. If anything this whole exercise should also give you a much deeper appreciation of the people that feel themselves called to help others.

    • I had a similarly useful analogy of the legal system emerge, in a similar way. I think there's a great analogy between policy in the legal system and policy in software engineering, and they have some awfully good (and very, very historical) ways to think about e.g. amending, repealing, and adjudicating things based on those policies.

  • i think role-play and utilizing the full power of language, stories, character, and narrative will unlock very sophisticated use-cases and in general a new dimension to agentic systems much in the way you're describing.

    probably what will work best in the future is specific training or fine-tuning against curated datasets of narrative fiction and/or texts in general? not really sure.

  • I feel like this is something we do with humans as well. Things like "scrum", "sprint", "the clean coder", etc.

  • yeah this is really cool. i've been thinking about sort of a mirrored idea, where you end up with wizards, mages, clerics etc and fantasy terminology used. idk if it would be as effective as this is, but maybe it would be more creative somehow?

    i kinda want to try to build this off of github. it's essentially just an event bus / message queue that workers (agents) tap into.

I can think of nothing I'd like to use less than a database or filesystem vibecoded by Gas Town-flavored psychosis. Roleplaying with LLMs is not the secret to producing amazing code.

The next role: Insurance Rep.

Ensures all other agents are operating efficiently and within reasonable levels of token usage given the expected level of effort to 'resolve' the patient. In moderate to severe cases, may lead to patient defenestration or agent revolt.

Or perhaps, the author considered this role but found it typically costs more than it saves.

A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?

  • I posted in another threat about this: I am seeing a lot of people building their own little bespoke factories. They introduce endless quality gates until it slows development down, and then they add more agents to decide when to run certain actions, and on and on. The end result from what I've seen and personally participated in, is that it often winds up providing negative value in the software development lifecycle. It creates a whole lot of heat, but IMO is not helping the teams utilizing them to ship value any faster than they would with a more limited setup.

    • I just finished ripping out one of these “dark factory” setups. Removed about 70k lines of code and 750k words of generated documentation. For what is effectively a 5-screen app.

    • Funny that we've automated meeting hell. Definitely a truism that we reimplement the org chart in software.

    • I've gone through this cycle recently. I think static lint/type checks are super useful, but agentic code review loops can easily go off the rails.

  • This is Microsoft scale, I'd be surprised if it made any difference these labs. It's more likely it's plain managerial mishandling of the infra in chasing new profit heights.

Rafi and I, who authored this post, will be hanging out here for any questions people may have.

  • Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?

    I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.

    • No, we haven't performed any ablation studies yet - it's a good suggestion.

      The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.

      I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.

      2 replies →

  • This, with an extremist take on code quality via linting in ci, is no doubt the future, at least for maintenance and extending the interface kind of work.

    1. How does this work with greenfield lifts where the scope and final vision are not yet figured out?

    2. Sorry if I missed this in the post, but will you open source this system?

  • The blog post links to an issue in the Sinai repo, but it’s private or just doesn’t exist?

  • Which inherent limitations did you recognize in the metaphor before commencing this research?

    • One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have "treatment" start on the cheapest possible model and see how it ended up, only escalating to more capable/expensive models as necessary. We rejected this idea because we suspected that it would end up in more time/tokens on the capable models to fix any problems.

      In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.

I am curious why so many of these systems are based on GitHub issues. Why not use a proper ticket system?

I use Redmine (lightly customized) and have found that to work very well. Especially for performing review / UAT of the work or providing response to the agents questions when it wants me to pick an option.

The underlying data model of a ticket / project management system is so rich and well suited to bodies of work.

I set different kinds of structures for different projects, Complete with setting and ambiance. It’s like agents work better if they are role-playing. It’s extremely disorienting and people with marginal mental stability are going to really have a bad time. What have we wrought?

Since we are all building our version of this, what mistakes did you made until you arrived to a well-balanced solution? I really like that you know how much an issue is, because then you can start optimizing.

  • I'd say that the first "mistake" we made was having it run in auto-merge mode. It was incredible to see what it could produce, and the speed with which it worked, but while the results seemed good, they were being produced at a rate that we couldn't human-verify. This is not to say that they were bad, but that we had no way to convince ourselves that they were good.

    Once we started to review the code in more detail (and put the hospital into "human approval mode" to stop the momentum) we found one major thing that we corrected.

    When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We've since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how/when the deferral should be actioned. This has prevented several items from falling through the cracks, but we're constantly refining what is a valid deferral and what should be fixed immediately.

    There are several other things that we've corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.

    • > When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped.

      Could you structure the DAG so that after each node that contains the work, you have one dependent node that verifies each distinct requirement was implemented as expected?

      This way if a single requirement is dropped the system alerts you rather than it being silently dropped.

      It makes sense intuitively that if a task has nothing that depends on it the LLM might accidentally attempt to drop it (even purposefully as an optimization).

    • Is a code comment and ledger the best way? Should the agent just fill out a form or something and attach it to the sub issue. This is how the hospital would work.

      2 replies →

Great post. Thank you for sharing it on HN.

Modern best practices for a medical team (standard roles, responsibilities, consequences, processes, etc.) have evolved through trial and error, since the advent of civilization, to prevent costly human error.

In hindsight, it's not too surprising that the same best practices can be applied by teams of AI agents to minimize costly AI error.

It's still incredible. We sure live in interesting times!

  • You could also run them like an army, like an artist commune, like Valve (add virtual tables on wheels which they can pull wherever they want), like a lawyer's office, etc.

If the correct analogy for a team of SW agents isn’t a team of SW engineers why is that?

  • For example, we don't have a clear analogue for "triage nurse". What would be it? Oncall engineer for production issues; "product owner" maybe for the regular tickets?

    Thing is, software engineering is still (relatively) young. And more of a craft than true engineering. So the "good practices" are _somewhat established_, but at the same time, not quite. Like, we have glimpses into what works and what doesn't, but not a universally-established workflow to follow. And AI is upending what we already knew; it's really not surprising that people are looking elsewhere for workable metaphors.

    • > For example, we don't have a clear analogue for "triage nurse". What would be it?

      Triage nurse seems well defined in software engineering, perhaps just without a single title. There are certainly engineers that focus on triaging GitHub issues and then fixing immediate showstopper bugs and/or prioritizing the issue for further attention from other engineers.

      That role can be communicated to agents by writing down that description. The risk with saying that the role is “triage nurse” is that you can’t be sure what other roles and allowances the agent will infer it has.

If I wanted to build something like this, at least conceptually with the roles, where’d I start? My first guess would be to give Claude your blog post; but any other pointers to make it work reliably? Do you happen to have the system open source?

Maybe I missed it, but I didn't really see anything about the long term quality or maintainability of the code. All I see is agent agent agent.

  • It's true that we don't have long-term maintainability data just yet. We've just shipped the first product of this model to customers and likely won't have any detailed maintainability data for several months (and for good data, several years). We hope to author more blog posts on this experiment in the future.

    • Perhaps doing a random sample of the steps by hand will be a good way to notice maintainability issues?

One role I didn't see was patient advocate/representative? That might be another approach to non-convergence - "how is this going?" and escalation.

  • We actually have a /sinai-advocate skill, where a human can advocate on behalf of a stuck patient. We use it every once in a while when the labels get screwed up, or a workflow fails for some reason.

I’m surprised at the reasoning behind using the most expensive model which likely led to the “doing too much”. I think it would have been nice to have A/B tested different hospital setups in order to maximize not only efficiency but to see where the strength of each model started to emerge in a given role.

Fable feels like overkill for this also.

Humans: $160k / 9 months = $600/day

AI Software Factory: $4172 / 2 days = $2086/day

This seems unsustainable, unless you're also generating 3x the revenue.

  • I don't get your maths? [1] 160/9 months != 600 for any given number of days a week e.g. 5,6,7 (something between 6 and 7), but where did the 9 come from, is it some kind of adjustment for weekends? Anyway there is also a concept of "fully loaded employee cost".

    In any case the way to think of this is not "/day".

    The reason is simple. If you buy 1000 barrels of oil, you buy 1000 barrels of oil not 17.4 days of oil. There is now a disconnect between work done and time. Infact you would be sane if you said "that result I can get in 2 days for $4172, if you can get me that same result in 1 hour, I'd pay $8344". See where this is going?

    Yes a lot of thought work is now a commodity, and if you want the commodity faster (last minute booking, uber to come quicker etc.) you pay more not less. Value being $/hour is over.

    [1] The ? acts as both a question and a regex.

  • I have qualms and disagreements with the whole "software factory" concept, but your math doesn't help the argument. The cost of human labor is more than salary—or even hourly rate, if you're thinking of paying a contractor.

  • Only one of the two trends is downwards.

    • ... Token cost is still heavily subsidized by the frontier labs while they're private companies and raise billions whenever they decide. The downward cost in tokens is unsustainable long term without revolutions training or inference of models.

You give a very precise measure of redundancy in the skills. Can you give a bit more detail in how you decided you needed to audit them and how you went about it? Was it fully agent driven? Mostly human?

I absolutely love how you are able to pull so much latent behavior from the underlying LLM. I wonder what other analogies can be pulled into agentic coding that come baked into the existing weights.