← Back to context

Comment by kbenson

4 hours ago

That's useful if the system is still up and you can trust the state information, but when you have a disk failure, or there's enough system state corruption that you don't want to rely on that state tracking, you're then going to need something else.

Also, are you pinning your package versions in your config? For every single case of that, is the original package still available? Were they coming from remote sources which may not keep prior versions always available?

Did you build anything from scratch? Did you supply a source or binary archive to install from? Is that source still available, of the exact same version?

There are numerous things you need to worry about when rebuilding from what the state should be based on a description of that state (which may include arbitrary actions), as opposed to restoring from a recorded state that includes everything needed in the recorded data.

The more you can record the better. Should you also record nix configuration state and use that to restore when it makes sense? Sure! If it's simple and easy and doesn't use an inordinate amount of space, why wouldn't you? Just don't confuse nice to have with sufficient based on the specific requirements and goals your backup system needs to meet.

We have entire system images in an immutable remote store, but that's not what we go to first, it's what we go to when our quicker and easier systems are insufficient. If we were running NixOS, we'd probably keep everything we're doing now, and then also just add the NixOS state as a separate thing that's backed up and saved on every change (and in fact, along these lines, we also run etckeeper on every host and track all changes to /etc with daily commits if something is changes but also an automatic commit at the beginning and end of every ansible run we do that applies our host configuration).

More is better. If I want to see configuration changes, I can usually just look in etckeeper on the box, if I want to restore a tree to the exact state it was at a specific point that a backup was made through rsync, I can do that easily enough (or just navigate the directory tree on the backup server, if I just want to see what it was, or diff between different states at different times). If I need to restore the whole system image from a known good backup, I can use Veeam and the local repository for that. If the local repository is dead, I can use the replicated off-site repository. If an attacker has destroyed both those, I go to the immutable off-site backups, which I'm probably doing to use after rebuilding my backup infra from scratch to restore from. Each of these protects from slightly different scenarios, and offers different trade-offs for security and integrity and ease of use.

Sure, you may lose out of the "faster" part if you have to pull from a remote cache (both inputs and outputs) instead of your local one. And maybe just blindly relying on the community maintained caches isn't right for you, so now you're still in a position maintain a cache and back it up through traditional means. It's not a magic bullet, probably doesn't pay off for a lot of use cases.

But what it buys you is a sort of verifiability that you can't really get any other way. Anyone can build any part of the system from its declared inputs and say "hmm, I got a different hash than you, maybe something is up." If you're restoring from an image which somebody has tampered with, then the image becomes the source of truth, rather than the source code, and it gets a lot more difficult to scrutinize its validity because the pool of available scrutinizers is just the people who care about your particular image--that's likely to be a lot smaller than the set of people who care about whatever sources you're relying on.