Skip to content
BytePatterns

SAP-C02 · Domain 1: Design Solutions for Organizational Complexity · 26% of the exam

Task 1.3: Design reliable and resilient architectures.

Meeting an RTO and an RPO: backup and restore, pilot light, warm standby and multi-site, Elastic Disaster Recovery, automatic recovery from failure, and backups that can really be restored.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company runs a critical ERP system on 40 VMware virtual machines in its data center. It needs disaster recovery on AWS with an RPO of seconds and an RTO of under one hour. The application must not change, and the company wants to pay as little as possible for the recovery site while no disaster is in progress. Which solution meets these requirements?

  1. AReplicate the servers continuously with AWS Elastic Disaster Recovery and launch recovery instances only during a drill or disaster
  2. BBack up the virtual machines to AWS with AWS Backup every hour and restore them to Amazon EC2 during a disaster
  3. CRun a full-size copy of every server on Amazon EC2 at all times and keep the data in sync with database and file replication software
  4. DExport the virtual machines to Amazon Machine Images every night with VM Import/Export and launch them during a disaster
Show the answer and why
  • AReplicate the servers continuously with AWS Elastic Disaster Recovery and launch recovery instances only during a drill or disaster

    Correct

    Elastic Disaster Recovery replicates source servers continuously into a staging area that uses affordable storage and minimal compute, and it launches recovery instances only when needed. AWS states that it enables RPOs of seconds and RTOs of minutes.

  • BBack up the virtual machines to AWS with AWS Backup every hour and restore them to Amazon EC2 during a disaster

    Incorrect

    Scheduled backups give an RPO equal to the backup interval, an hour in this design, which misses an RPO of seconds.

  • CRun a full-size copy of every server on Amazon EC2 at all times and keep the data in sync with database and file replication software

    Incorrect

    Running the whole recovery site at full size all the time meets the objectives but is the most expensive option while no disaster is in progress.

  • DExport the virtual machines to Amazon Machine Images every night with VM Import/Export and launch them during a disaster

    Incorrect

    Nightly images give an RPO of up to a day, far from seconds, and importing large virtual machines takes time.

Seconds of RPO without a running copy of the site is what continuous block-level replication into a staging area gives. Backups and images are bound to their schedule.

Question 2 · choose 1

An order database runs on Amazon Aurora MySQL in us-west-2. If the Region becomes unavailable, the database must accept writes in us-east-1 within minutes and lose at most a few seconds of data. Until then, a reporting application in us-east-1 must read the data locally. Which solution meets these requirements?

  1. AAdd Aurora Replicas in two more Availability Zones of us-west-2 and point the reporting application at the reader endpoint
  2. BCopy automated snapshots to us-east-1 every hour with AWS Backup and restore the latest one during a Regional outage
  3. CConvert the cluster to an Aurora global database with a secondary cluster in us-east-1 and fail over to it
  4. DTurn on Aurora Backtrack with a 72-hour window and rewind the cluster after the Region recovers
Show the answer and why
  • AAdd Aurora Replicas in two more Availability Zones of us-west-2 and point the reporting application at the reader endpoint

    Incorrect

    Aurora Replicas protect against the loss of an instance or an Availability Zone within one Region. They do not help if us-west-2 is unavailable, and reads would cross Regions.

  • BCopy automated snapshots to us-east-1 every hour with AWS Backup and restore the latest one during a Regional outage

    Incorrect

    Restoring an hourly snapshot can lose up to an hour of data and takes longer than a promotion, so it misses both the RPO and the RTO.

  • CConvert the cluster to an Aurora global database with a secondary cluster in us-east-1 and fail over to it

    Correct

    A global database replicates to secondary Regions with typical lag under a second, secondary clusters serve local reads, and a secondary can be promoted to take writes when the primary Region is impaired.

  • DTurn on Aurora Backtrack with a 72-hour window and rewind the cluster after the Region recovers

    Incorrect

    Backtrack rewinds a cluster in place to an earlier point in time. It creates no copy in another Region, so it cannot keep the database writable during a Regional outage.

Cross-Region writes within minutes with seconds of data loss and local reads meanwhile is the Aurora global database pattern.

Question 3 · choose 1

A reporting application uses Amazon RDS for PostgreSQL. During business hours, read-only dashboard queries keep CPU above 90%, and the writes of the order service slow down. The instance already uses the largest class the company will pay for. The dashboards can read data that is a few seconds old and can use their own connection string. Which change solves the problem?

  1. APut Amazon RDS Proxy in front of the database and point the dashboards at the proxy endpoint
  2. BAdd read replicas and point the dashboards at them, adding more replicas as dashboard use grows
  3. CConvert the instance to a Multi-AZ DB instance deployment and send the dashboard queries to the standby
  4. DChange the instance to the next larger instance class during business hours and back again at night
Show the answer and why
  • APut Amazon RDS Proxy in front of the database and point the dashboards at the proxy endpoint

    Incorrect

    RDS Proxy pools and shares connections. The dashboard queries would still run on the same instance and use the same CPU.

  • BAdd read replicas and point the dashboards at them, adding more replicas as dashboard use grows

    Correct

    Read replicas take read-only queries off the primary instance and use asynchronous replication, which matches dashboards that accept a few seconds of lag. Scaling out with more replicas avoids a larger class.

  • CConvert the instance to a Multi-AZ DB instance deployment and send the dashboard queries to the standby

    Incorrect

    In a Multi-AZ DB instance deployment the standby does not serve read traffic; it exists for failover.

  • DChange the instance to the next larger instance class during business hours and back again at night

    Incorrect

    This is scaling up, which the company has ruled out on cost, and a class change can cause downtime while the instance reboots.

When scaling up has reached its limit and the load is read-only, scale out with read replicas.

Question 4 · choose 2

A company has 60 accounts in AWS Organizations. It must back up Amazon EBS volumes, Amazon RDS databases and Amazon DynamoDB tables in every account on one schedule. Copies must be kept in a separate account so that a compromised workload account cannot delete them, and for one year nobody, including the root user of the backup account, may delete or shorten the retention of those copies. Which actions meet these requirements? (Choose TWO.)

  1. ACreate Amazon Data Lifecycle Manager policies in every account that copy the EBS snapshots to the backup account each day
  2. BLock the backup vaults in the workload accounts with AWS Backup Vault Lock in governance mode
  3. CApply AWS Backup backup policies from AWS Organizations that copy each recovery point to a vault in a central backup account
  4. DTurn on S3 Object Lock in compliance mode for the backup vaults in the central backup account
  5. ELock the vault in the backup account with AWS Backup Vault Lock in compliance mode and a minimum retention of 365 days
Show the answer and why
  • ACreate Amazon Data Lifecycle Manager policies in every account that copy the EBS snapshots to the backup account each day

    Incorrect

    Data Lifecycle Manager automates EBS snapshots and EBS-backed AMIs. It does not back up RDS databases or DynamoDB tables.

  • BLock the backup vaults in the workload accounts with AWS Backup Vault Lock in governance mode

    Incorrect

    A vault in governance mode can be changed or removed by users with sufficient IAM permissions, and these vaults stay in the workload accounts, which the copies must not depend on.

  • CApply AWS Backup backup policies from AWS Organizations that copy each recovery point to a vault in a central backup account

    Correct

    Backup policies apply one backup plan to the accounts of the organization, and a copy rule can send recovery points to a vault in another account of the same organization.

  • DTurn on S3 Object Lock in compliance mode for the backup vaults in the central backup account

    Incorrect

    Object Lock is a feature of S3 buckets. Backup vaults are protected with AWS Backup Vault Lock instead.

  • ELock the vault in the backup account with AWS Backup Vault Lock in compliance mode and a minimum retention of 365 days

    Correct

    AWS Backup denies any user, including the root user, who tries to delete a backup or change its lifecycle in a locked vault, and once the grace time ends a compliance-mode lock cannot be removed while the vault holds recovery points.

Organization backup policies give one schedule everywhere and copies in a separate account; Vault Lock in compliance mode makes those copies immutable even for the root user.

Question 5 · choose 1

An enterprise has 80 applications with agreed resilience targets. Leaders want each application assessed against its targets, with recommendations based on AWS best practices, in one central place instead of spreadsheets. Which service fits?

  1. AAWS Fault Injection Service
  2. BAWS Trusted Advisor
  3. CAWS Resilience Hub
  4. DAmazon CloudWatch dashboards
Show the answer and why
  • AAWS Fault Injection Service

    Incorrect

    FIS runs fault experiments; it does not track targets and assessments for a portfolio.

  • BAWS Trusted Advisor

    Incorrect

    Trusted Advisor checks accounts against best practices but does not assess applications against their own resilience targets.

  • CAWS Resilience Hub

    Correct

    Resilience Hub lets you define resilience goals, assess applications against them and implement recommendations based on the Well-Architected Framework.

  • DAmazon CloudWatch dashboards

    Incorrect

    Dashboards show metrics, not assessments against resilience goals.

Central resilience assessment against goals is AWS Resilience Hub.

Question 6 · choose 1

A multi-account application fails over to a second Region with a long runbook: scale up compute, promote databases, and shift traffic, in a fixed order. Engineers make mistakes under pressure, and leaders want the recovery orchestrated and observable from one place. Which capability fits?

  1. ARegion switch in ARC with a plan for the application
  2. BZonal shift in ARC for the application's load balancers
  3. CRoute 53 failover records with health checks
  4. DAWS Backup restore jobs in the recovery Region
Show the answer and why
  • ARegion switch in ARC with a plan for the application

    Correct

    Region switch orchestrates large-scale, complex recovery tasks across accounts from a centralized, observable solution.

  • BZonal shift in ARC for the application's load balancers

    Incorrect

    Zonal shift moves traffic for a resource away from an impaired Availability Zone to healthy zones in the same Region. It does not orchestrate a failover to another Region.

  • CRoute 53 failover records with health checks

    Incorrect

    Health checks move DNS traffic but do not orchestrate scaling and database steps.

  • DAWS Backup restore jobs in the recovery Region

    Incorrect

    Restores recover data but do not orchestrate a whole Regional switch.

Orchestrated, observable Regional recovery is Region switch in ARC.

Question 7 · choose 1

An Aurora MySQL cluster has a db.r6g.4xlarge writer, two db.r6g.4xlarge readers that the application reaches through the reader endpoint, and one db.r6g.large reader that a BI tool queries through its instance endpoint. In the last writer failure, Aurora promoted the small reader, and the application stayed slow until the team failed over again by hand. The team must make sure a failover always leaves a writer of the current size, without paying for more instance capacity, while the BI tool keeps querying live data on its own small reader. Which change meets these requirements?

  1. AChange the BI reader to db.r6g.4xlarge so that whichever reader Aurora promotes is as large as the current writer
  2. BAdd a third db.r6g.4xlarge reader so that large readers outnumber the small one when Aurora picks a failover target
  3. CGive the two large readers failover priority tier 0 and set the BI reader to tier 15 so that it is promoted last
  4. DRemove the BI reader from the cluster and give the BI tool an Aurora clone of the cluster with one small instance
Show the answer and why
  • AChange the BI reader to db.r6g.4xlarge so that whichever reader Aurora promotes is as large as the current writer

    Incorrect

    A larger BI reader would make any promotion safe, but Aurora charges for each DB instance by its class, so this pays for a much bigger instance that only serves reports.

  • BAdd a third db.r6g.4xlarge reader so that large readers outnumber the small one when Aurora picks a failover target

    Incorrect

    Aurora does not pick a target by how many readers of each size exist; it promotes the reader with the highest priority. The extra reader adds cost and leaves the small reader's priority where it is.

  • CGive the two large readers failover priority tier 0 and set the BI reader to tier 15 so that it is promoted last

    Correct

    Aurora promotes the replica with the highest priority, from 0 for the highest to 15 for the lowest. Ranking the large readers first ensures one of them becomes the writer, and changing a priority neither triggers a failover nor costs anything.

  • DRemove the BI reader from the cluster and give the BI tool an Aurora clone of the cluster with one small instance

    Incorrect

    A clone is a separate, independent volume that starts from the source's data, which suits analytical queries on a copy. Writes made to the source after cloning do not reach it, so the BI tool would lose live data.

Upsizing the BI reader or adding another large reader spends money, and a clone gives the BI tool a copy that stops tracking production. Aurora picks the failover target by replica priority, and only among replicas in the same tier does it prefer the largest. Putting the two large readers in tier 0 and the BI reader in tier 15 fixes the promotion order at no extra cost, and the BI tool keeps its own reader.

Question 8 · choose 1

A claims platform runs on Amazon ECS on AWS Fargate behind an Application Load Balancer in us-east-1. Its database is already an Aurora PostgreSQL global database with a secondary cluster in us-west-2. The disaster recovery targets for a Regional outage are an RPO of 1 minute and an RTO of 2 hours. Finance will not pay for application compute or load balancers in us-west-2 except during drills and real failovers. All infrastructure is defined in CloudFormation templates. Which disaster recovery strategy meets these requirements at the lowest cost?

  1. AMulti-site active/active: serve traffic from both Regions, and turn on write forwarding so the secondary cluster passes writes to the primary
  2. BBackup and restore: replace the secondary cluster with hourly AWS Backup copies in us-west-2, and rebuild the stack from templates after failover
  3. CWarm standby: run a scaled-down copy of the ECS services and the load balancer in us-west-2 at all times, and scale it out after failover
  4. DPilot light: keep the secondary cluster, and deploy the ECS services and load balancer from the same templates only when failing over to us-west-2
Show the answer and why
  • AMulti-site active/active: serve traffic from both Regions, and turn on write forwarding so the secondary cluster passes writes to the primary

    Incorrect

    Active/active runs the full workload in both Regions all the time, which costs the most and breaks the rule against running compute in us-west-2 outside drills. Write forwarding is useful when secondary Regions must accept writes, which this design does not need.

  • BBackup and restore: replace the secondary cluster with hourly AWS Backup copies in us-west-2, and rebuild the stack from templates after failover

    Incorrect

    AWS Backup can copy recovery points to another Region on a schedule, but with hourly copies up to an hour of committed data can be lost, which breaks the 1-minute RPO. Continuous replication is what keeps the RPO that low.

  • CWarm standby: run a scaled-down copy of the ECS services and the load balancer in us-west-2 at all times, and scale it out after failover

    Incorrect

    Warm standby keeps a scaled-down but fully functional copy of production running in the recovery Region, so it pays for exactly the compute and load balancers that finance refuses to fund outside drills.

  • DPilot light: keep the secondary cluster, and deploy the ECS services and load balancer from the same templates only when failing over to us-west-2

    Correct

    In a pilot light, the data stores stay always on and replicated while application servers are not deployed until a drill or failover. Aurora replicates to the secondary cluster with latency typically under a second, which covers the 1-minute RPO, and redeploying the stateless ECS tier from IaC fits within 2 hours.

Read the cost rule as the deciding constraint: nothing that serves traffic may run in us-west-2 until a drill or disaster. Warm standby and active/active both break it. Backup and restore would satisfy the cost rule but loses up to an hour of data with hourly copies. Pilot light keeps only the replicated data always on and switches the rest on from the same templates, which meets the 1-minute RPO and the 2-hour RTO at the lowest cost.

Question 9 · choose 1

About 2,000 employees on Windows desktops map a drive to a share on a Single-AZ 2 Amazon FSx for Windows File Server file system. During a recent Availability Zone issue, the share was down for more than an hour. The company now requires that the share fail over automatically to another zone in the same Region, that users never have to remap drives after a failover, and that no acknowledged write is lost when a zone fails. Which solution meets these requirements?

  1. ATurn on shadow copies with an hourly schedule so that users can restore earlier versions of their files themselves
  2. BGive the file system a DNS alias, and after a zone failure restore its latest backup in another zone and move the alias there
  3. CCopy the share with Robocopy every 15 minutes to a second Single-AZ file system in another zone, and switch users there if a zone fails
  4. DMove the share to a Multi-AZ FSx for Windows File Server file system with file servers in two Availability Zones
Show the answer and why
  • ATurn on shadow copies with an hourly schedule so that users can restore earlier versions of their files themselves

    Incorrect

    Shadow copies let users restore previous versions of files and folders, and recover deleted files, from Windows File Explorer. They are stored alongside the file system's own data, so they do nothing for availability during a zone issue.

  • BGive the file system a DNS alias, and after a zone failure restore its latest backup in another zone and move the alias there

    Incorrect

    A DNS alias can be associated when a backup is restored to a new file system, so users would keep their drive mappings. But the restore is a manual step after the outage, and writes made since the last backup are lost.

  • CCopy the share with Robocopy every 15 minutes to a second Single-AZ file system in another zone, and switch users there if a zone fails

    Incorrect

    Robocopy can copy the files and their metadata to another FSx for Windows File Server file system, so a second copy would survive a zone loss. The switch is still manual, and up to 15 minutes of acknowledged writes would be lost.

  • DMove the share to a Multi-AZ FSx for Windows File Server file system with file servers in two Availability Zones

    Correct

    A Multi-AZ file system replicates data synchronously between two zones and fails over automatically to the standby file server, typically in under 30 seconds. Its DNS name stays the same, so Windows clients resume without manual action.

Three constraints decide it: automatic failover, no drive remapping and no lost acknowledged writes. A restored backup behind a DNS alias and a Robocopy copy in another zone both keep the data, but they need a manual switch and lose recent writes. Shadow copies help users undo changes, not survive a zone outage. A Multi-AZ file system keeps a synchronous standby in a second zone behind the same DNS name.

Practise domain 1 →Practise all domains →