Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.
No, we haven't performed any ablation studies yet - it's a good suggestion.
The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.
> One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I use a similar process for a local agent swarm approach (locally hosted Qwen-3.8-27b or Qwen-Flash-Next), used so far for personal-grade projects. Through tool and process accretion, the review stages are told to check both the work and reasoning of the implementation stage, via processing the pi.dev session log.
Through parsing the jsonl log, the reviewer sees the subagent prompt, the tool call sequence, and non-thinking narration along the way. That allows the reviewr to audit the implementer's process (e.g. were tests run?) and spot procedure violations or gross hallucinations. The reviewer also independently runs the test suite, so even a hallucinated pass is caught.
It's relatively expensive both in tokens and time spent running ideally duplicate tests, but the independent workflow has nonetheless caught errors that would very likely have been missed by a same session, same context review.
Could you fix a bug in this codebase yourself, without the help of agents?
Maybe trying this would be a good experiment? Say route 10 or 20% of any future issues to the “Chief” to troubleshoot and fix directly, and record those for further work?
This, with an extremist take on code quality via linting in ci, is no doubt the future, at least for maintenance and extending the interface kind of work.
1. How does this work with greenfield lifts where the scope and final vision are not yet figured out?
2. Sorry if I missed this in the post, but will you open source this system?
Sorry about that. The repo is internal to Cockroach Labs for now. It was an issue about adding a progress rollup comment to decomposed parent issues. We’ll get the post updated.
One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have "treatment" start on the cheapest possible model and see how it ended up, only escalating to more capable/expensive models as necessary. We rejected this idea because we suspected that it would end up in more time/tokens on the capable models to fix any problems.
In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
Did you do any ablation studies on what is actually useful vs what happens to just work because these systems can work around whatever people do?
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.
No, we haven't performed any ablation studies yet - it's a good suggestion.
The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.
> One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I use a similar process for a local agent swarm approach (locally hosted Qwen-3.8-27b or Qwen-Flash-Next), used so far for personal-grade projects. Through tool and process accretion, the review stages are told to check both the work and reasoning of the implementation stage, via processing the pi.dev session log.
Through parsing the jsonl log, the reviewer sees the subagent prompt, the tool call sequence, and non-thinking narration along the way. That allows the reviewr to audit the implementer's process (e.g. were tests run?) and spot procedure violations or gross hallucinations. The reviewer also independently runs the test suite, so even a hallucinated pass is caught.
It's relatively expensive both in tokens and time spent running ideally duplicate tests, but the independent workflow has nonetheless caught errors that would very likely have been missed by a same session, same context review.
[flagged]
Yes agreed, using a system very similar without the hospital vibe :)
Could you fix a bug in this codebase yourself, without the help of agents?
Maybe trying this would be a good experiment? Say route 10 or 20% of any future issues to the “Chief” to troubleshoot and fix directly, and record those for further work?
This, with an extremist take on code quality via linting in ci, is no doubt the future, at least for maintenance and extending the interface kind of work.
1. How does this work with greenfield lifts where the scope and final vision are not yet figured out?
2. Sorry if I missed this in the post, but will you open source this system?
How much of this is a snake eating its own tail against vs the control group?
are any of these artifacts public and available for inspection/use?
Not currently, but open sourcing this has been discussed. Stay tuned.
The blog post links to an issue in the Sinai repo, but it’s private or just doesn’t exist?
Sorry about that. The repo is internal to Cockroach Labs for now. It was an issue about adding a progress rollup comment to decomposed parent issues. We’ll get the post updated.
Which inherent limitations did you recognize in the metaphor before commencing this research?
One thing that teaching hospitals have is a longer ladder of experience. For example many hospitals have medical students, interns, residents, chief residents, fellows and attendings. One idea we contemplated was to have "treatment" start on the cheapest possible model and see how it ended up, only escalating to more capable/expensive models as necessary. We rejected this idea because we suspected that it would end up in more time/tokens on the capable models to fix any problems.
In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
[flagged]