← blog · September 5, 2026

Actually Measuring RPO and RTO Targets in Disaster Recovery

A green nightly backup job does not mean your RPO and RTO targets are actually met. Only a real, repeated restore drill produces the true number. Here is how to build one, and where it quietly lies to you.

What RPO and RTO actually are, and why they stay on paper

Every disaster recovery plan centers on two numbers. RPO (Recovery Point Objective) defines how much data loss you can tolerate, and RTO (Recovery Time Objective) defines how long you have to bring a system back up. Setting these numbers is easy; writing "RPO 15 minutes, RTO 4 hours" in a meeting takes an hour. Knowing whether those numbers actually hold is the hard part. Most teams never measure it. Because the backup job turns green every night, they assume the target is being met. That assumption is where the most expensive surprise in a real outage comes from.

Taking a backup is not the same as meeting your RPO

A backup tool completing successfully every night proves none of three separate claims: that the data captured is consistent, that it can actually be restored, and that the restore finishes within the target time. Take a database dump captured while the filesystem is mid write: the file size looks correct, but the content can be corrupt. The backup tool reports zero errors, and the first query after restore fails. The only way to catch this is to occasionally run a real restore. A log line that says "success" is not sufficient evidence.

The second problem shows up in retention and pruning logic. Most backup tools delete old snapshots according to a policy. That deletion step can silently fail because of a stale lock file or an interrupted prior run, while the tool still reports "backup complete", because what it measured was the backup step, not the pruning step. The result: the repository quietly grows, nobody notices, until the disk fills up or a restore is actually needed one day. This is why backup monitoring should not only ask "did the backup job succeed", but also "is repository size on the expected curve" and "what is the age of the oldest valid snapshot".

The variables people skip when measuring RTO

Measuring RTO as "how long does the restore command take" is misleading. Real RTO is the sum of getting access and authorization (who has which key, which console, and will that be remembered during a crisis), transferring the data (bounded by network bandwidth, scales linearly with data size), restoring it, starting services in the correct order, and waiting for dependencies such as DNS, certificates, or external services to become ready.

The most common mistake here is running the restore test against the small dataset from initial setup and recording that time as the RTO. If the data grows tenfold over six months, restore time grows proportionally, but because nobody repeats the drill, the number on paper never changes. This measurement is not a one-time certificate; it is a number that needs to be refreshed alongside data growth.

The second skipped variable is the human factor. Which document to consult during a crisis, which system that document lives on, and whether that system is even reachable during the outage, are usually never tested. If the runbook itself lives somewhere that becomes unreachable in the disaster scenario, such as a wiki hosted on the system that just went down, it does not matter how fast the restore steps themselves run.

Three testing approaches and where each one fits

Tabletop exercise: the team gets together, talks through a "server X is down" scenario, and walks through who does what. It is cheap and surfaces communication gaps, for instance nobody knowing where the backup keys are kept, but it does not measure real time. It is a fine starting point for small teams and should never be used on its own as proof of RTO.

Automated restore drill: at regular intervals, weekly or monthly, the latest backup is actually restored into an isolated environment, put through a health check, and the elapsed time is recorded. This produces a real RPO and RTO number because there is little room for human intervention or assumption. Cost is moderate: it requires an isolated environment and an automation script, but once built, upkeep is low. For most teams this is the right balance. It matters that the hardware and network profile of the restore environment is not wildly different from production; a test over a fast internal network in the same data center hides the transfer time a real disaster, restoring from a different region, would actually take.

Live failover test (chaos engineering, game day): actually shifting some or all of production traffic to the backup system and creating a real outage to test against. This gives the most realistic number because it includes factors no drill reproduces, such as network behavior, DNS TTL, and cache warm-up time. It is also the highest risk and highest cost method, and should not be attempted without mature monitoring and rollback infrastructure already in place.

Practical recommendation: start with an automated restore drill for small and medium systems. It gives you a real number without production risk. As criticality increases, for payment systems or single-point databases, increase frequency and eventually move toward live failover testing.

Steps you can apply

Do not set one RPO/RTO target for the whole organization. Define tiers based on data criticality: user data and financial records might warrant an RPO of 15 minutes, while logs and cache data might tolerate 24 hours. Set a separate test cadence per tier; weekly for high criticality, monthly or quarterly can be enough for low criticality.

Collect restore duration as a metric, not just backup job duration. Record the time from the start of the restore to the moment the service passes its health check, and track the trend over time. As data grows this number will grow too; seeing that ahead of time lets you adjust capacity or backup strategy, for example incremental backups or parallel restores, accordingly.

Verify data integrity after restore. Check not just that the command exited with zero errors, but that the restored data has the expected record count, passes a checksum, or clears an application level smoke test. "The restore command returned 0" and "the application is actually running on correct data" are two very different claims.

Test key management separately for encrypted backups. A backup can succeed while the decryption key is lost during rotation, meaning the real RPO is infinite, discovered only during an actual disaster. Regularly test whether the key still actually opens the backup, and make sure the key itself is backed up independently.

Keep the runbook reachable under the outage scenario, and actually follow that document during the drill rather than relying on the memory of whoever wrote it. A drill run from memory hides exactly the places where the document is incomplete or out of date.

When this level of investment is not worth it

Not every system deserves a weekly restore drill. For low criticality systems that are easy to reproduce, such as a static cache or a log archive, a simple and infrequent check is enough. Keep test frequency and depth proportional to that system's actual outage cost; otherwise the team spends most of its time testing low risk systems while the drill for the truly critical one gets neglected. RPO and RTO targets are not a checkbox, they are a contract for how many minutes of data loss you have actually accounted for in a real outage. Test the contract before you sign it.