← Back to context

Comment by ajstorm

1 day ago

No, we haven't performed any ablation studies yet - it's a good suggestion.

The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.

I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.

> One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.

I use a similar process for a local agent swarm approach (locally hosted Qwen-3.8-27b or Qwen-Flash-Next), used so far for personal-grade projects. Through tool and process accretion, the review stages are told to check both the work and reasoning of the implementation stage, via processing the pi.dev session log.

Through parsing the jsonl log, the reviewer sees the subagent prompt, the tool call sequence, and non-thinking narration along the way. That allows the reviewr to audit the implementer's process (e.g. were tests run?) and spot procedure violations or gross hallucinations. The reviewer also independently runs the test suite, so even a hallucinated pass is caught.

It's relatively expensive both in tokens and time spent running ideally duplicate tests, but the independent workflow has nonetheless caught errors that would very likely have been missed by a same session, same context review.