Comment by Majromax
9 hours ago
> One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I use a similar process for a local agent swarm approach (locally hosted Qwen-3.8-27b or Qwen-Flash-Next), used so far for personal-grade projects. Through tool and process accretion, the review stages are told to check both the work and reasoning of the implementation stage, via processing the pi.dev session log.
Through parsing the jsonl log, the reviewer sees the subagent prompt, the tool call sequence, and non-thinking narration along the way. That allows the reviewr to audit the implementer's process (e.g. were tests run?) and spot procedure violations or gross hallucinations. The reviewer also independently runs the test suite, so even a hallucinated pass is caught.
It's relatively expensive both in tokens and time spent running ideally duplicate tests, but the independent workflow has nonetheless caught errors that would very likely have been missed by a same session, same context review.
No comments yet
Contribute on Hacker News ↗