← blog · September 18, 2026

Backups are not enough: run restore drills

A green backup job gives false confidence, and an untested backup is not a backup. We cover the silent traps we hit, valid but incomplete archives, stale locks, and tool mismatch in the recovery environment, and why automating a restore drill is the only guarantee that actually holds.

There is a line that gets repeated for years: if you do not have a backup, you do not have your data. True, but incomplete. The harsher truth we learned running our own infrastructure is that an untested backup is not a backup. A backup job that shows green gives you false confidence, and you pay for that confidence at exactly the moment you need it most.

Our backup jobs run hourly, the archives are kept in two separate off-site locations, and older ones are pruned after a while. The dashboard was all green. Then one day we tried to restore an archive and found that half of it was not there. The backup job had finished without an error, because the file was technically valid, just incomplete. A compressed archive can still open from the start up to a point even if it was cut off in the middle. A quick integrity check does not say corrupt, because it reads as far as it can and stops. We had trusted a file that looked valid without checking its real size.

The first lesson from this: seeing a backup finish is not enough, you have to restore it and measure it. Ran is not the same as restorable. The only meaningful test of a backup is restoring it into an empty place and counting what comes out. For a database we check row counts, for a set of files we check the file count and the checksums of a few samples. If it does not match what we expected, that backup does not exist.

The second lesson is about where you restore. In our early attempts we tried to restore right next to production, on the same machine, and immediately ran into collisions. Same container names, same ports, same data volumes. If a restore risks breaking the live service you are trying to protect, the drill itself is a hazard. The fix was to restore into an isolated space: a separate namespace, ports that do not clash, throwaway volumes. When the test finishes, that space is deleted entirely. That way the drill never touches production.

The third lesson was to not leave this to a human. "Let us test a restore once a month" is a good intention, and it gets forgotten in the first busy week. We automated the verification. On a regular schedule a job takes the latest backup, restores it into the isolated space, checks the counts and checksums, and records the result. If something does not add up, an alert goes out. So the answer to "is our backup sound" depends on a running check, not on someone remembering.

There are side traps we hit often. One is a stale lock. Backup tools place a lock so two jobs do not write to the same repository at once. If a job is interrupted, the lock can be left behind, and every job after it fails silently with "repository locked". The dashboard may show no error, because the job never started. You have to watch the age of locks and clear ones older than a threshold. Another is tool mismatch. If at recovery time all we have is a small system, the compression tool there can behave differently from the one that produced the backup. One day a minimal gzip on a small recovery box could not open an archive that the full version opened fine. You have to take the recovery environment as seriously as the backup environment.

Finally, location diversity. Keeping a backup in one place means that when that place is gone, the backup is gone with it. Two separate locations, ideally two separate providers, is far more robust. But even here the main message does not change: having two copies does not mean either one restores. Test both copies.

If we compress all of this into one sentence: a backup is not an action, it is a promise, and do not trust a promise you have not tested. A backup being green does not tell you it will work on that bad day when you finally restore it. The only way to learn whether it works is to rehearse it regularly, without waiting for that bad day. The rehearsal is cheap, the real disaster is expensive. Since we automated the rehearsal, we stopped looking at the color on the dashboard and started looking at the date of the last successful restore. That is the number that actually gives confidence.