Comment by jeffrallen
3 hours ago
I work on the same kind of thing, and while we think hard about bootstrap problems, we always find new surprising ones. The problem is you never know until you do it, and creating a faithful test of restarting giant systems is economically impossible. Because if you say to the boss, "look, I need 1 million now to test against a maybe 100 million loss, maybe in 10 years" they don't give you the money (and rightly so).
Even if you did the $1MM test there is very low likelihood that the $100MM event would be fully mitigated 10 years down the line (after who knows how many changes - physical, logical, and even in the org chart).
The only way to approach readiness here is repeated investment - like one team doing the deep dive and another pulling cables and then constantly doing pre- and post-mortems.