Disaster Recovery: RPO & RTO
AWS for Interviews: lesson 15 of 18
Pay for exactly as much standby as your RPO and RTO demand.
Lesson 15 of 18 · 7 min
Disaster Recovery: RPO & RTO
Step 1 of 9
Two numbers decide disaster recovery. RPO: how much recent data you can lose. RTO: how long you can be down.
The Idea
RPO is how much recent data you can lose; RTO is how long you can be down. Four strategies, each costlier and faster: backup and restore (RPO hours, RTO up to 24 h), pilot light (data live, servers off: minutes, tens of minutes), warm standby (a scaled-down working copy: seconds, minutes), multi-site active-active (near zero for both).
Real-World Example
A shop needs RPO 15 minutes and RTO one hour. Pilot light fits: an Aurora global database secondary in the recovery Region, launch templates and images ready, Route 53 failover records. On failover it promotes the database, scales out the app tier and shifts traffic.
The Tradeoff
Replication copies corruption too, so every strategy keeps point-in-time backups. Fail over with data-plane actions such as Route 53 health checks rather than control-plane calls, and test it before you need it.
Hands-On
# illustrative — ARNs are placeholders; copies a backup to the recovery Region
aws backup start-copy-job \
--recovery-point-arn arn:aws:ec2:us-east-1::snapshot/snap-0abc \
--source-backup-vault-name prod-vault \
--destination-backup-vault-arn arn:aws:backup:us-west-2:111122223333:backup-vault:dr-vault \
--iam-role-arn arn:aws:iam::111122223333:role/backup-copy
Your turn
Sort the strategies from the longest RTO to the shortest.
- Warm standby
- Multi-site active-active
- Backup and restore
- Pilot light
Mini quiz
1 / 3
Which strategy keeps a scaled-down but fully working copy that can take traffic at once?
Sources
- REL13-BP02 Use defined recovery strategies to meet the recovery objectives — Reliability Pillar, AWS Well-Architected Framework
- Disaster recovery options in the cloud — Disaster Recovery of Workloads on AWS