Comment by Twirrim
5 days ago
I got known as a "Chaos Monkey" at AWS, and have sort of carried on that title. It's not exactly accurate, but tech stuff just breaks around me, and it's very, very rarely my fault. I'm pretty cautious and resist any urge to "I wonder what happens..." with anything that could possibly have an effect beyond me.
The "Chaos Monkey" effect at AWS was so pronounced you could literally see on our sev2 count graph when it was my on-call week, because I'd get dramatically more sev2s than any other engineer.
The service I worked for at AWS was amazingly stable and reliable. A large majority of the alarms that ever fired, fired because of problems with another service we depended on in some way. This billing thing is a good example. I was on-call when there was a major SQS outage in a region, through a DynamoDB outage, S3 outages, all sorts of stuff. We used to joke that it would probably be a net benefit for AWS if I wasn't on-call, just so my spooky-at-a-distance wouldn't happen.
I eventually lost my "most sev2s in a week" record toward the end of my time there when someone else was on-call and DynamoDB had a major outage. DynamoDB held all of our metadata so everything broke and every alarm we had fired over the space of a few hours.
No comments yet
Contribute on Hacker News ↗