Comment by alamsterdam
4 hours ago
It's been hours :(
I have sympathy for the on-call team trying to resolve it, most of us have been there done that.
But seems something is systematically going wrong at GH
4 hours ago
It's been hours :(
I have sympathy for the on-call team trying to resolve it, most of us have been there done that.
But seems something is systematically going wrong at GH
> But seems something is systematically going wrong at GH
Yes, we call it: Microslop.
It's honestly insane how terrible their reliability is. Over the past 4 weeks, 8 days with GitHub Actions outages, many of them multi-hour outages: >3 hours on each of July 9th, July 20th and today, and 1.5 hrs on July 23rd.
Outages happen, but this many outages so close together, and so many of them so major/long lasting, something is systematically wrong for sure. It's been seriously hamstringing our ability to ship code at my company.
Mon and Dad (MS) are (generally) pretty solid with uptime.
What is happening at GH?
Rate of change trying to keep up with new challengers? Over-reliance on AI? Engineers trying to debug slop?
Yeah who knows, would be interesting to hear an inside take if any readers here are also GH devs!
They're at 93.91% uptime over the past 90 days, according to https://mrshu.github.io/github-statuses/ , and that doesn't even include today's outage yet.
A glorious one nine of reliability.
3 replies →
> What is happening at GH?
In large part, the move from AWS to Azure. Azure's just bad.
2 replies →
"most of us have been there done that" - so true. That feeling in your gut when you realize something you just did caused an outage is pretty unique.
They're one rewrite in Rust away from fixing everything. (jk)
Time to touch some grass. Better use of my time and energy than twisting the remaining things on my todo list today to make more progress on them than I have managed. I should have taken a long lunch but I rebased the hell out of a PR instead.