Skip to content
BytePatterns

DOP-C02 · Domain 3: Resilient Cloud Solutions · 15% of the exam

Task 3.3: Implement automated recovery processes to meet RTO and RPO requirements.

Meeting RTO and RPO by design: backup and restore, pilot light and warm standby, cross-Region AWS Backup copies, failover tests for RDS, Aurora and Route 53, and load balancers that route around failed targets.

Study it

  • Disaster recovery strategies against RTO and RPO

    Lesson coming

  • AWS Backup across Regions and accounts, and failover testing

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 2

A company backs up Amazon EBS volumes and Amazon RDS databases in a production account with an AWS Backup plan. After a ransomware drill, the security team requires a copy of every backup in a separate backup account of the same organization, and no one, including the root user of either account, may delete or shorten the retention of those copies for 35 days. Which actions should the DevOps engineer take? (Choose TWO.)

  1. ACreate Amazon Data Lifecycle Manager policies in the production account that take EBS snapshots and keep them for 35 days
  2. BAttach a vault access policy to the destination vault that denies backup:DeleteRecoveryPoint to every principal in the backup account
  3. CLock the destination vault with AWS Backup Vault Lock in governance mode and a minimum retention of 35 days
  4. DAdd a copy action to the backup plan rule that copies each recovery point to a vault in the backup account
  5. ELock the destination vault with AWS Backup Vault Lock in compliance mode and a minimum retention of 35 days
Show the answer and why
  • ACreate Amazon Data Lifecycle Manager policies in the production account that take EBS snapshots and keep them for 35 days

    Incorrect

    Data Lifecycle Manager automates EBS snapshots in the same account, so an administrator there could still delete them, and it does not cover RDS.

  • BAttach a vault access policy to the destination vault that denies backup:DeleteRecoveryPoint to every principal in the backup account

    Incorrect

    A vault access policy controls only the AWS Backup APIs; EBS and RDS snapshots can also be reached through those services' own APIs. Vault Lock in compliance mode is what stops every user.

  • CLock the destination vault with AWS Backup Vault Lock in governance mode and a minimum retention of 35 days

    Incorrect

    A lock in governance mode can be removed by users with sufficient IAM permissions.

  • DAdd a copy action to the backup plan rule that copies each recovery point to a vault in the backup account

    Correct

    AWS Backup can copy backups to other accounts in the organization as part of a scheduled backup plan, with a destination vault and key in that account.

  • ELock the destination vault with AWS Backup Vault Lock in compliance mode and a minimum retention of 35 days

    Correct

    After the grace time, a compliance-mode lock cannot be changed or deleted, and AWS Backup denies any user, including the root user, who tries to delete a backup or change its lifecycle.

The copy puts the backups out of reach of the production account, and the compliance-mode lock makes them immutable for their retention period. A governance-mode lock can be removed with enough permissions, and an access policy covers only the AWS Backup APIs.

Question 2 · choose 1

A company runs 40 Windows and Linux servers with a mix of commercial applications in its own data center. For disaster recovery to AWS, the business requires an RPO of seconds and an RTO under 30 minutes. Running a full copy of the servers on AWS all the time is too expensive, and the applications cannot be changed. Which solution should the DevOps engineer recommend?

  1. ANightly backups of every server with AWS Backup to a vault in AWS, and restores to EC2 instances when a disaster is declared
  2. BAWS Elastic Disaster Recovery replicating the servers continuously to a low-cost staging area, and launching recovery instances when needed
  3. CA warm standby copy of every server running on smaller EC2 instances, kept in sync by application-level replication that the teams build
  4. DEC2 Image Builder pipelines that turn each server into an AMI every week, with instances launched from the AMIs during a disaster
Show the answer and why
  • ANightly backups of every server with AWS Backup to a vault in AWS, and restores to EC2 instances when a disaster is declared

    Incorrect

    Nightly backups give an RPO of up to a day, far beyond the required seconds.

  • BAWS Elastic Disaster Recovery replicating the servers continuously to a low-cost staging area, and launching recovery instances when needed

    Correct

    Elastic Disaster Recovery uses affordable storage and minimal compute, with an RPO typically in the sub-second range and an RTO measured in minutes, depending mostly on boot time.

  • CA warm standby copy of every server running on smaller EC2 instances, kept in sync by application-level replication that the teams build

    Incorrect

    Running every server all the time is the cost the business wants to avoid, and building replication would mean changing the applications.

  • DEC2 Image Builder pipelines that turn each server into an AMI every week, with instances launched from the AMIs during a disaster

    Incorrect

    Weekly images give an RPO of up to a week, and Image Builder builds images from recipes rather than copying running on-premises servers.

Elastic Disaster Recovery replicates block-level data continuously to a staging area in AWS and keeps compute to a minimum until a drill or recovery, when it converts and boots the servers as EC2 instances.

Question 3 · choose 1

A bank runs an Amazon Aurora PostgreSQL global database with its primary cluster in us-east-1 and a secondary cluster in us-west-2. A regulation requires the bank to move the primary to the other Region twice a year and run there for several months to prove the recovery procedure works. Both Regions are healthy when this happens, and no committed transaction may be lost. How should the DevOps engineer move the primary?

  1. APerform a switchover of the global database to the secondary cluster in us-west-2
  2. BPerform a cross-Region failover of the global database to the secondary cluster in us-west-2
  3. CDetach the us-west-2 cluster from the global database, promote it to a standalone cluster, and point the application at it
  4. DRestore the latest automated snapshot of the primary cluster in us-west-2 and switch the application's endpoint to the restored cluster
Show the answer and why
  • APerform a switchover of the global database to the secondary cluster in us-west-2

    Correct

    A switchover, previously called managed planned failover, is meant for planned operations with healthy clusters. It synchronizes the secondary before promoting it, so the RPO is 0.

  • BPerform a cross-Region failover of the global database to the secondary cluster in us-west-2

    Incorrect

    Failover is for unplanned outages, and its RPO is typically a non-zero value in seconds, depending on replication lag.

  • CDetach the us-west-2 cluster from the global database, promote it to a standalone cluster, and point the application at it

    Incorrect

    Detaching stops replication without first synchronizing, so writes not yet replicated can be lost, and the clusters are no longer one global database.

  • DRestore the latest automated snapshot of the primary cluster in us-west-2 and switch the application's endpoint to the restored cluster

    Incorrect

    A restored snapshot misses every transaction committed after the snapshot, and the restore adds downtime.

Aurora Global Database offers two ways to move the primary Region. Switchover is for planned, healthy situations such as regional rotation and loses no data; failover is for real outages and can lose the writes still in replication lag.

Question 4 · choose 1

An Auto Scaling group behind an Application Load Balancer runs a web application. When the application process hangs, the target group marks the instance unhealthy and stops sending it traffic, but the instance keeps running and passing its EC2 status checks. Capacity stays reduced until someone terminates the instance by hand. The team wants such instances replaced automatically. What should the DevOps engineer do?

  1. AShorten the target group's health check interval and lower its unhealthy threshold so that hung instances are detected sooner
  2. BRaise the group's health check grace period so that new instances have more time to start before they are checked
  3. CTurn on Elastic Load Balancing health checks for the Auto Scaling group
  4. DAdd a CloudWatch alarm on the target group's UnHealthyHostCount metric with an EC2 reboot action for the instance
Show the answer and why
  • AShorten the target group's health check interval and lower its unhealthy threshold so that hung instances are detected sooner

    Incorrect

    Faster detection only takes the instance out of the load balancer sooner. The group still uses EC2 status checks and does not replace it.

  • BRaise the group's health check grace period so that new instances have more time to start before they are checked

    Incorrect

    The grace period keeps a new instance in service for a minimum time before it can be replaced. It does not make the group act on what the load balancer reports.

  • CTurn on Elastic Load Balancing health checks for the Auto Scaling group

    Correct

    The default EC2 health checks do not see an application failure. With ELB health checks turned on, the group treats an instance the load balancer reports as unhealthy as unhealthy and replaces it.

  • DAdd a CloudWatch alarm on the target group's UnHealthyHostCount metric with an EC2 reboot action for the instance

    Incorrect

    EC2 alarm actions can be added only to alarms on per-instance EC2 metrics. UnHealthyHostCount is a target group metric, so there is no one instance to reboot.

Auto Scaling groups use EC2 status checks by default, which only see the instance and its hardware. Turning on ELB health checks lets the group act on the load balancer's view of the application and replace the failed instance.

Question 5 · choose 1

At 10:42 a faulty migration corrupted a table in an RDS for PostgreSQL DB instance. Automated backups are enabled with a seven-day retention period, and a read replica that serves reports replicates without any configured delay. The team needs the database as it was at 10:40, must not lose transactions committed before then, and must keep the damaged instance unchanged for the investigation. What should the DevOps engineer do?

  1. ARestore to 10:40 with point-in-time restore into a new DB instance
  2. BRestore last night's automated snapshot into a new DB instance and keep the damaged instance running
  3. CPromote the read replica to a standalone DB instance and point the application at it
  4. DBacktrack the DB instance to 10:40 with the RDS console
Show the answer and why
  • ARestore to 10:40 with point-in-time restore into a new DB instance

    Correct

    Point-in-time restore creates a new DB instance at any time within the retention period and leaves the source instance unchanged.

  • BRestore last night's automated snapshot into a new DB instance and keep the damaged instance running

    Incorrect

    This also keeps the damaged instance, but every transaction committed after the snapshot would be lost.

  • CPromote the read replica to a standalone DB instance and point the application at it

    Incorrect

    The replica applies the source's changes asynchronously, with no configured delay, so it already holds the corrupted table.

  • DBacktrack the DB instance to 10:40 with the RDS console

    Incorrect

    Backtracking is a feature of Aurora MySQL DB clusters and is not available for RDS for PostgreSQL DB instances.

Automated backups and transaction logs allow restoring to any time within the retention period. The restore always creates a new DB instance, so the application is pointed at it after the restore and the original stays available for analysis.

Question 6 · choose 1

A recovery runbook creates 40 database volumes from the latest EBS snapshots in a known Availability Zone. Tests show that after recovery the databases are slow for hours because blocks are fetched on first access, which breaks the RTO. What should the DevOps engineer do?

  1. AUse io2 volumes with the highest IOPS when creating the volumes from the snapshots
  2. BTurn on fast snapshot restore for the snapshots in that Availability Zone
  3. CCopy the snapshots to another Region so that the volumes are restored closer to the data
  4. DArchive the snapshots to the archive tier so that restores are faster
Show the answer and why
  • AUse io2 volumes with the highest IOPS when creating the volumes from the snapshots

    Incorrect

    Higher provisioned IOPS do not remove the first-access latency of restored blocks.

  • BTurn on fast snapshot restore for the snapshots in that Availability Zone

    Correct

    Volumes created with fast snapshot restore are fully initialized and deliver their provisioned performance at once.

  • CCopy the snapshots to another Region so that the volumes are restored closer to the data

    Incorrect

    Copies elsewhere do not change how blocks are initialized.

  • DArchive the snapshots to the archive tier so that restores are faster

    Incorrect

    Archived snapshots must be restored first, which takes longer.

Fast snapshot restore is enabled per snapshot and Availability Zone. Volumes created from such a snapshot in that zone are fully initialized at creation.

Question 7 · choose 1

An active/passive application runs replicas in two Regions behind Route 53 failover records. During an incident, operators want to move all client traffic to the other Region with a single highly reliable action of their own, without editing DNS records and without depending on endpoint health checks. What should the DevOps engineer set up?

  1. AA Lambda function that updates the DNS records when operators run it from the console
  2. BA lower TTL on the failover records so that DNS changes take effect faster
  3. CWeighted records whose weights operators change from 100/0 to 0/100
  4. DARC routing controls with their health checks on the failover records
Show the answer and why
  • AA Lambda function that updates the DNS records when operators run it from the console

    Incorrect

    Editing DNS records during an incident is what the operators want to avoid.

  • BA lower TTL on the failover records so that DNS changes take effect faster

    Incorrect

    A lower TTL still leaves operators editing records or depending on health checks.

  • CWeighted records whose weights operators change from 100/0 to 0/100

    Incorrect

    This is still a manual change to DNS records.

  • DARC routing controls with their health checks on the failover records

    Correct

    Routing controls are on-off switches that reroute traffic through routing control health checks on DNS failover records.

ARC routing controls are grouped on control panels in a cluster. Turning a routing control on or off changes its health check, which shifts traffic between Regional replicas.

Question 8 · choose 1

A backup plan in AWS Backup protects DynamoDB tables. A new DR requirement says every recovery point must be copied to a second Region. The copy rule in the plan does not copy the DynamoDB recovery points. What should the DevOps engineer do?

  1. ATurn on DynamoDB point-in-time recovery for each table
  2. BTurn on AWS Backup advanced features for DynamoDB in the Region
  3. CConvert every table to a global table so that AWS Backup copies it
  4. DExport the tables to S3 every day and turn on Cross-Region Replication for the bucket
Show the answer and why
  • ATurn on DynamoDB point-in-time recovery for each table

    Incorrect

    PITR protects tables in their own Region; it does not copy recovery points.

  • BTurn on AWS Backup advanced features for DynamoDB in the Region

    Correct

    With advanced features, DynamoDB backups support cross-Region and cross-account copies.

  • CConvert every table to a global table so that AWS Backup copies it

    Incorrect

    Global tables replicate data but do not change how backups are copied.

  • DExport the tables to S3 every day and turn on Cross-Region Replication for the bucket

    Incorrect

    This builds a separate export path instead of copying recovery points.

Advanced DynamoDB backup features in AWS Backup unlock cold storage tiering, cost allocation tags, and cross-Region and cross-account copies.

Question 9 · choose 1

During a data center outage, a company recovered 30 on-premises servers to AWS with Elastic Disaster Recovery and ran production there for a week. The data center is now repaired, and the workloads must return to the original servers with the changes made while running in AWS. What should the DevOps engineer do?

  1. ARebuild the on-premises servers from the backups taken before the outage
  2. BKeep running in AWS and delete the original servers
  3. CUse Elastic Disaster Recovery to fail back to the source servers
  4. DCopy the instances' files back to the servers with a nightly script
Show the answer and why
  • ARebuild the on-premises servers from the backups taken before the outage

    Incorrect

    Old backups miss the week of changes made in AWS.

  • BKeep running in AWS and delete the original servers

    Incorrect

    The workloads must return to the original servers.

  • CUse Elastic Disaster Recovery to fail back to the source servers

    Correct

    After the disaster is mitigated, Elastic Disaster Recovery can fail back to the original source infrastructure.

  • DCopy the instances' files back to the servers with a nightly script

    Incorrect

    Scripts copy files without a consistent, managed failback.

Elastic Disaster Recovery launches recovery instances in AWS during a disaster and supports failback to the original source infrastructure once it is available again.

Practise domain 3 →Practise all domains →