Comment by sethammons

6 years ago

That is a fantastic comment. I was reading through Google SRE stuff. Lots of good info. Lots of very, very bad ideas too. Lots of it is "at scale, you can't make everyone happy, so don't try." You end up with SLOs and SLIs based on percentages of availability, errors, etc. If 99.99% of requests are good, don't worry about the rest. However, at scale, 0.01 can be literally millions of real, breathing people with feeling and priorities and things to do. When you give away your product for free, these are good trade offs. When you sell your product, the same assumptions may not apply.

I cannot afford to ignore 0.01% of incoming requests because we charge per request. At a certain point, yes, money and pros and cons yadda yadda yadda. It just feels that, once again, people are saying, "but Google does it!" -- and that does not mean you should.

100% is unachievable, especially at Google scale. Your car or coffee shop or whatever also fails 0.0x% of the time. Heck the energy net in western countries doesn't reach 99.999% uptime and I would call it very dependable.

I'd say the biggest problem it the non-existent support. It's fine if 0.01% fails, as long as there is human troubleshooting or help available if necessary.

I think you are very confused about how availability is tracked. If a class of users is always getting an error this would be investigated. There is literally no way, at any scale, to guarantee 100% availability.

  • >I think you are very confused about how availability is tracked. If a class of users is always getting an error this would be investigated.

    Oh, please, you are confused. I worked at Google, most teams just don’t have metrics like this.

    While I don’t think it’s acceptable, reality is that it’s expensive to compute such metric, i.e. you need to slice you monitoring stream per each user, has privacy implications, and often it’s unclear what action should be taken.