Comment by readthenotes1
5 hours ago
A lot of people don't check their backups until they need to restore.
A lot of people are incompetent.
5 hours ago
A lot of people don't check their backups until they need to restore.
A lot of people are incompetent.
Nobody wants backups as a feature. The feature is restore.
There are plenty of people who are happy to check off "Backups" boxes, or talk about their backup strategy, or whatever - while hoping the day never comes when they actually need that tricky "Restore" feature.
A backup without a restore test isn't a backup at all
You don't always have a copy of your hardware to restore onto. And the test's entire purpose is that you're not yet sure whether your restore will truly work. So you can't just run a backup and restore on your true prod system, because you're not sure it won't wreck it. So you need extra money to have a second system onto which you try to restore. If you don't have a lot of money, you will want to actually use your disks for storage, not to put them into a second testing server. Of course I'm not talking about very professional companies with super critical data. Just simpler smaller scale places or consumers.
No, even professional companies with critical data balk at this.
I worked at a company worth a few billion and the leadership balked when they told engineering they wanted a near instantaneous failover system and our department informed them that would require paying for a second environment that could be rolled over to.
It is rare to find leaders who can accept the cost of redundant infrastructure that is there for emergency backup.
What puzzles me is why they can’t accept it when they are perfectly fine with insurance costs and I can’t see much of a difference between the two when looking at a spreadsheet of costs other than possibly tax differences between the type of expenditure.
1 reply →
I've worked on backup/failure systems since the mid-90s, and I've found there's one universal truth: If you don't fully test your backup/failure system, you don't have a backup/failure system.
There's generally two wrong responses: (1) We spent a lot of money on 'blah blah blah', a lot of other companies use it, so yeah, we've got a backup/failure system. And, (2) inadequate testing - either, we tested 1 of 50 services, and it worked, so the whole system can be restored; or, we gracefully tested, and it worked, so it will obviously work during not-graceful incidents.
And the root cause of this is generally that no one gets promoted for implementing an adequate backup/failure system, or it's extremely rare.