Skip to content

What is Disaster Recovery (RPO and RTO)?

DevOps & Reliability, explained by the engineers who build it. Definition, how it works, use cases and common questions.

RPO and RTO definition

Disaster recovery (DR) is the set of policies, tools and procedures for restoring systems and data after a major disruption, such as a cloud region outage, ransomware or accidental deletion. It is defined by two targets: the recovery time objective (RTO), how quickly service must return, and the recovery point objective (RPO), how much data loss is acceptable.

RPO vs RTO

The recovery point objective is measured backward from the incident: an RPO of one hour means you accept losing up to an hour of data, so backups or replication must capture changes at least that often. The recovery time objective is measured forward: an RTO of four hours means the service must be running again within four hours of the disaster being declared.

Both should come from business impact, not technology preferences. A marketing site might accept a day of RTO and RPO, while an order system may need minutes. Lower targets cost more, so many companies set tiers for different systems. Disaster recovery complements high availability, which handles smaller, routine failures automatically.

Disaster recovery strategies

Cloud providers describe a ladder of strategies, each faster and more expensive than the last. Most organizations mix them, matching each system to the cheapest strategy that still meets its agreed RTO and RPO, and documenting who may declare a disaster and trigger the switch.

Whichever strategy you choose, recovery depends on more than data. Secrets, DNS, certificates, container images, third-party allowlists and access for the on-call team all need to exist in the recovery region, and they are the items most often missed in a first drill. The four common strategies are:

  • Backup and restore: regular backups copied to another region; cheapest, with recovery measured in hours
  • Pilot light: core data replicated and minimal infrastructure ready, scaled up during a disaster
  • Warm standby: a scaled-down but running copy of the full environment that can take traffic quickly
  • Multi-site active-active: full capacity in two or more regions serving traffic, with near-zero RTO and RPO

Backups, ransomware and the 3-2-1 rule

Replication alone is not disaster recovery. If someone deletes a table or ransomware encrypts data, replication copies the damage everywhere within seconds. Point-in-time backups that are immutable and stored in a separate account are what let you rewind to before the incident. The classic 3-2-1 rule still applies: three copies of data, on two different types of storage, with one copy offsite.

Protect backups as carefully as production: separate credentials, object lock or immutable vaults such as AWS Backup Vault Lock, and alerts on deletion attempts. Ransomware groups often target backups first, precisely because backups are what make paying a ransom unnecessary.

Testing and documenting a DR plan

An untested plan is a guess. Schedule restore tests and failover drills, measure the actual recovery time against the RTO, and fix whatever slowed the team down: missing runbooks, expired credentials, DNS TTLs that were too long or dependencies nobody remembered. Infrastructure as code makes recovery far more reliable, because the environment can be recreated in another region from version-controlled templates.

Nexzem helps clients set RTO and RPO per system, implements backups and cross-region recovery on AWS, Azure and Google Cloud, and runs recovery drills, so the first full restore does not happen during a real emergency with customers waiting.

RPO and RTO: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

What is the difference between RPO and RTO?

RPO, the recovery point objective, is the maximum acceptable data loss, measured in time before the incident. RTO, the recovery time objective, is the maximum acceptable downtime after it. An RPO of 15 minutes and an RTO of two hours means losing at most 15 minutes of data and being back within two hours.

Is backup the same as disaster recovery?

No. Backups are one ingredient. Disaster recovery also covers where and how systems will run during an outage, who decides to fail over, how DNS and users are redirected, and how long each step takes. A company can have perfect backups and still need days to recover without a tested plan.

How often should a disaster recovery plan be tested?

Critical systems should have restore tests at least quarterly and a fuller failover exercise at least once a year, plus a test after major architecture changes. Many teams automate backup restore checks weekly. Each test should record the achieved recovery time and any problems, then update the runbook.

Keep exploring the devops & reliability glossary

Need RPO and RTO in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.