← Back to context

Comment by rafiss

6 hours ago

I think this is definitely a legit failure mode with this model. We’ve hit some parts of it. For example, a plan reviewer kept pushing a technically valid but negligible security finding through five review rounds and about $70 before a human closed it as won’t-fix. The "Research Department" ran every three hours and was filing around 8 new issues a day, more than we could review. Review nits used to spawn new issues of their own.

What helped us maintain a degree of sanity is that every stage records tokens, time, and cost in a ledger, so these show up as numbers rather than vibes. The fixes were mostly about giving a stage permission to stop. Plan review can now propose won’t-fix when the exposure is negligible. Research runs daily, with a cap on unreviewed proposals. Small fixes now ride along in the PR instead of becoming new issues.

The other thing is that the quality bar should match the domain. Our migration tools move customer data into a database, so our question was “would we be embarrassed to have merged this?” For small apps I’d agree most of this is overkill.

Fellow HNers: this is one of the authors of the post we are commenting on.

Please do not flag him into oblivion again. Sheesh.