Schrodinger Backup
Let's imagine that you carefully design your backup strategy (refer to Backup Strategy, Backup Types), deliver it to production, configure schedule to trigger it regularly, store backups on another region for DR purposes.
Can you feel safe after it?
No 😱.
The problem is that the backup is there, but not really...
Until you have a process for regularly restoring production data, you have no guarantees that it works. It's not possible to test restoration on a real production environment, so this procedure should retore data on another environment and execute at least basic sanity checks.
With this idea in mind I decided to check what's in the industry has there: I asked about testing backup procedure in X DevOps community, checked what public clouds offer and looked for the suggestions over the Internet.
Key findings:
🔸 Most teams have never tested the restoration of production backups. They verify only procedure itself on some test environments.
🔸 GCP and Azure recommend to test production restoration, but you should prepare e2e procedure on your own (or I was not able to find it quickly).
🔸 AWS offers automatic testing procedure for its managed storages with an ability to create custom validation workflows.
🔸 Uber has a great article where they shared their continuous backup\restore approach.
Surprisingly, there are not much practical information about how to implement regular restoration testing.
Most probably there are 3 reasons for that:
- It's expensive
- The process is very env and company specific
- It may be more relevant for big tech companies where data loss is a critical business risk
So don't assume that no errors and existent backup files mean that you have a backup. You don't really know until the real incident .
#engineering #backups
Post #219
230
- 🔥 4
- 👍 1
- 💯 1