Comment by james_marks
1 day ago
A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?
1 day ago
A hint at why GitHub actions have been unreliable. How many teams are running factories like this on GH infra now?
I posted in another threat about this: I am seeing a lot of people building their own little bespoke factories. They introduce endless quality gates until it slows development down, and then they add more agents to decide when to run certain actions, and on and on. The end result from what I've seen and personally participated in, is that it often winds up providing negative value in the software development lifecycle. It creates a whole lot of heat, but IMO is not helping the teams utilizing them to ship value any faster than they would with a more limited setup.
I just finished ripping out one of these “dark factory” setups. Removed about 70k lines of code and 750k words of generated documentation. For what is effectively a 5-screen app.
These software factories are just s/code/specs/g. 5000 lines of (agent generated and poorly reviewed) specs for 2500 lines of code. All because the CTO dreams of getting a few likes from Elon Musk.
Funny that we've automated meeting hell. Definitely a truism that we reimplement the org chart in software.
I've gone through this cycle recently. I think static lint/type checks are super useful, but agentic code review loops can easily go off the rails.
> not helping
It does! It helps getting promotion with tokenmaxxing.
lol, no doubt.
I think this is definitely a legit failure mode with this model. We’ve hit some parts of it. For example, a plan reviewer kept pushing a technically valid but negligible security finding through five review rounds and about $70 before a human closed it as won’t-fix. The "Research Department" ran every three hours and was filing around 8 new issues a day, more than we could review. Review nits used to spawn new issues of their own.
What helped us maintain a degree of sanity is that every stage records tokens, time, and cost in a ledger, so these show up as numbers rather than vibes. The fixes were mostly about giving a stage permission to stop. Plan review can now propose won’t-fix when the exposure is negligible. Research runs daily, with a cap on unreviewed proposals. Small fixes now ride along in the PR instead of becoming new issues.
The other thing is that the quality bar should match the domain. Our migration tools move customer data into a database, so our question was “would we be embarrassed to have merged this?” For small apps I’d agree most of this is overkill.
Fellow HNers: this is one of the authors of the post we are commenting on.
Please do not flag him into oblivion again. Sheesh.
This is Microsoft scale, I'd be surprised if it made any difference these labs. It's more likely it's plain managerial mishandling of the infra in chasing new profit heights.