Skip to content
BytePatterns

SAA-C03 · Domain 2: Design Resilient Architectures · 26% of the exam

Task 2.2: Design highly available and/or fault-tolerant architectures.

Surviving the loss of an instance, a zone or a Region: Multi-AZ designs, DNS failover, the disaster recovery strategies and their RPO and RTO, durable copies of data, and removing every single point of failure.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company runs its order system in one AWS Region. For a Regional disaster it needs an RPO of minutes and an RTO of tens of minutes at the lowest ongoing cost. Data must be replicated continuously, but the application servers do not need to run in the recovery Region until a disaster. Which disaster recovery strategy fits best?

  1. ABackup and restore, with backups copied to the recovery Region
  2. BWarm standby, with a scaled-down stack that is always running
  3. CMulti-site active/active, with full capacity in both Regions
  4. DPilot light, with data replicated live and servers turned off
Show the answer and why
  • ABackup and restore, with backups copied to the recovery Region

    Incorrect

    Backup and restore is the cheapest, but RPO and RTO are measured in hours because everything is rebuilt from backups.

  • BWarm standby, with a scaled-down stack that is always running

    Incorrect

    A warm standby keeps a working copy of the whole stack running, which costs more than the requirement needs.

  • CMulti-site active/active, with full capacity in both Regions

    Incorrect

    Active/active gives the lowest RPO and RTO at the highest cost, far beyond what is asked.

  • DPilot light, with data replicated live and servers turned off

    Correct

    Pilot light replicates data continuously and keeps the core infrastructure ready while application servers stay off until failover, minimizing ongoing cost.

Continuous data replication with idle compute is the definition of pilot light. Warm standby and active/active run compute all the time.

Question 2 · choose 1

A web application runs in us-east-1 behind an Application Load Balancer. A static maintenance website is hosted in an S3 bucket in us-west-2. Amazon Route 53 must send users to the maintenance site automatically, and only when the application is unhealthy. Which routing configuration meets this requirement?

  1. ASimple routing with both endpoints as values of one record so that clients pick one
  2. BFailover routing: the app as the health-checked primary, the S3 site as secondary
  3. CGeolocation routing with the application for North America and the S3 site as default
  4. DWeighted routing that sends 50 percent of the queries to each endpoint all the time
Show the answer and why
  • ASimple routing with both endpoints as values of one record so that clients pick one

    Incorrect

    Simple routing returns the values in random order and does not check health, so users reach the maintenance page at random.

  • BFailover routing: the app as the health-checked primary, the S3 site as secondary

    Correct

    Failover routing answers with the primary while its health check passes and with the secondary only when it fails.

  • CGeolocation routing with the application for North America and the S3 site as default

    Incorrect

    Geolocation routing chooses by the user's location, not by the application's health.

  • DWeighted routing that sends 50 percent of the queries to each endpoint all the time

    Incorrect

    Weighted routing splits traffic by weight, so half the users see the maintenance page even when the application is healthy.

"Only when unhealthy" is active-passive failover: a failover record pair with a health check on the primary.

Question 3 · choose 2

A serverless application runs thousands of concurrent Lambda invocations against an Amazon RDS for PostgreSQL DB instance in a single Availability Zone. The database runs out of connections at peaks, and the company also needs automatic failover without data loss if the Availability Zone fails. Which changes meet these requirements? (Choose TWO.)

  1. AConvert the DB instance to a Multi-AZ DB instance deployment
  2. BPut Amazon RDS Proxy between the functions and the database
  3. CAdd a read replica of the DB instance in a second Availability Zone
  4. DRaise the reserved concurrency of the functions to the account maximum
  5. ERestore the latest automated backup in another zone if the zone fails
Show the answer and why
  • AConvert the DB instance to a Multi-AZ DB instance deployment

    Correct

    Multi-AZ keeps a synchronous standby in another zone and fails over to it automatically.

  • BPut Amazon RDS Proxy between the functions and the database

    Correct

    RDS Proxy pools and shares database connections, which suits many short-lived Lambda connections, and it shortens failover for applications.

  • CAdd a read replica of the DB instance in a second Availability Zone

    Incorrect

    Replication to a read replica is asynchronous and promotion is a manual step, so recent writes can be lost and failover is not automatic.

  • DRaise the reserved concurrency of the functions to the account maximum

    Incorrect

    More concurrent functions means more database connections, which makes the connection problem worse.

  • ERestore the latest automated backup in another zone if the zone fails

    Incorrect

    A restore creates a new DB instance and takes time, and changes since the last backed-up transaction log are lost.

Multi-AZ handles the zone failure with synchronous replication, and RDS Proxy handles the connection storm from Lambda.

Question 4 · choose 1

A vendor's licensing server runs on a single EC2 instance with EBS volumes. Its license is bound to the instance's private IP address, and the software cannot be changed or clustered. The company wants the instance to come back automatically, with the same private IP address and data, if the underlying host fails. What should a solutions architect do?

  1. ACreate an Auto Scaling group with a minimum, maximum and desired capacity of one, from a nightly AMI
  2. BRun a second instance in another Availability Zone and move an Elastic IP address to it by script
  3. CRely on automatic instance recovery, for example with a CloudWatch alarm on the system status check
  4. DLaunch the instance into a cluster placement group so it restarts on healthy hardware in the group
Show the answer and why
  • ACreate an Auto Scaling group with a minimum, maximum and desired capacity of one, from a nightly AMI

    Incorrect

    Auto Scaling replaces an unhealthy instance with a new one that has a new instance ID and private IP address, and data since the last AMI is lost.

  • BRun a second instance in another Availability Zone and move an Elastic IP address to it by script

    Incorrect

    The second instance has a different private IP address, which breaks the license, and its EBS data is not the same.

  • CRely on automatic instance recovery, for example with a CloudWatch alarm on the system status check

    Correct

    A recovered instance keeps its instance ID, private and Elastic IP addresses, and attached EBS volumes, and moves to healthy hardware.

  • DLaunch the instance into a cluster placement group so it restarts on healthy hardware in the group

    Incorrect

    Placement groups only influence where instances are placed. They do not recover a failed instance.

When an application cannot change, automatic instance recovery keeps the same instance, addresses and volumes on new hardware.

Question 5 · choose 1

A media company runs Amazon Aurora MySQL in eu-west-1. It needs to recover from the loss of the whole Region with an RPO of about one second and an RTO of minutes, and users in North America need fast local reads. Which solution meets these requirements?

  1. AConvert the cluster to an Aurora global database with a secondary cluster in North America
  2. BAdd Aurora Replicas in other Availability Zones of eu-west-1 and fail over to one of them
  3. CCopy the automated snapshots to a North American Region every hour and restore after a disaster
  4. DMove the data to Amazon DynamoDB global tables and point users at the nearest Region
Show the answer and why
  • AConvert the cluster to an Aurora global database with a secondary cluster in North America

    Correct

    A global database replicates to secondary Regions with latency typically under a second, serves local reads there, and can fail over across Regions.

  • BAdd Aurora Replicas in other Availability Zones of eu-west-1 and fail over to one of them

    Incorrect

    Aurora Replicas protect against an instance or zone failure inside the Region. The loss of eu-west-1 takes them with it.

  • CCopy the automated snapshots to a North American Region every hour and restore after a disaster

    Incorrect

    Hourly copies mean an RPO of up to an hour, and restoring a cluster from a snapshot takes far longer than minutes.

  • DMove the data to Amazon DynamoDB global tables and point users at the nearest Region

    Incorrect

    Global tables are multi-Region replication for DynamoDB. Using them means rewriting a MySQL application for a different database model.

Second-level RPO across Regions with local reads is what an Aurora global database is for. In-Region replicas do not survive a Regional outage.

Question 6 · choose 1

A company must replicate every new object in an S3 bucket to a bucket in another Region. A compliance rule requires replication within 15 minutes backed by a service-level agreement, and alerts about any object that replicates late. Which solution meets these requirements?

  1. AS3 Cross-Region Replication without S3 Replication Time Control
  2. BS3 Versioning in both buckets and an S3 Lifecycle copy rule
  3. CS3 Cross-Region Replication with S3 Replication Time Control
  4. DAn hourly S3 Batch Operations job that copies new objects
Show the answer and why
  • AS3 Cross-Region Replication without S3 Replication Time Control

    Incorrect

    Replication without RTC has no 15-minute commitment and does not send events for objects that miss that threshold.

  • BS3 Versioning in both buckets and an S3 Lifecycle copy rule

    Incorrect

    Versioning is a prerequisite for replication, but lifecycle rules only transition or expire objects. They do not copy them to another bucket.

  • CS3 Cross-Region Replication with S3 Replication Time Control

    Correct

    RTC replicates most new objects in seconds and 99.9 percent of them within 15 minutes, backed by an SLA, and adds metrics and event notifications for objects that replicate late.

  • DAn hourly S3 Batch Operations job that copies new objects

    Incorrect

    An hourly job cannot meet a 15-minute target, and RTC does not apply to batch copying.

A replication time commitment with late-object alerts is the definition of S3 Replication Time Control.

Question 7 · choose 1

A company uses a warm standby in a second AWS Region. In a failover test, the Auto Scaling group in the standby Region could not scale from 4 to 60 instances, although the same size works in the primary Region. What should a solutions architect do to prevent this during a real failover?

  1. ARequest service quota increases in the standby Region ahead of time to match production
  2. BDo nothing, because service quotas apply to the whole account across every Region
  3. CBuy a Compute Savings Plan large enough to cover the 60 instances in the standby Region
  4. DTurn on AWS Trusted Advisor checks so that the group can exceed its limits in a failover
Show the answer and why
  • ARequest service quota increases in the standby Region ahead of time to match production

    Correct

    Many quotas apply per Region. AWS's disaster recovery guidance says to set the recovery Region's quotas high enough to scale up to production capacity.

  • BDo nothing, because service quotas apply to the whole account across every Region

    Incorrect

    Some quotas are account-wide, but many are Region-based, so the primary Region's values do not carry over.

  • CBuy a Compute Savings Plan large enough to cover the 60 instances in the standby Region

    Incorrect

    A Savings Plan is a pricing commitment that lowers the bill. It neither reserves capacity nor raises a quota.

  • DTurn on AWS Trusted Advisor checks so that the group can exceed its limits in a failover

    Incorrect

    Trusted Advisor inspects the account and makes recommendations, including on service limits. It does not raise any limit.

A recovery Region must be ready for production scale, and that includes its service quotas.

Question 8 · choose 2

A web application runs on one EC2 instance in one Availability Zone, with a self-managed MySQL database on the same instance. Users reach it through an Elastic IP address. The company wants the application to survive the failure of an instance or of an Availability Zone. Which changes should a solutions architect make? (Choose TWO.)

  1. ARun the web tier in an Auto Scaling group across two zones behind a load balancer
  2. BMove the instance to a larger instance type with more memory and faster networking
  3. CAdd a second EBS volume and mirror the database files across both with RAID 1
  4. DCreate a nightly AMI of the instance so it can be launched again by hand
  5. EMove the database to Amazon RDS for MySQL with a Multi-AZ deployment
Show the answer and why
  • ARun the web tier in an Auto Scaling group across two zones behind a load balancer

    Correct

    The group replaces unhealthy instances and can launch in another zone if one fails, while the load balancer sends traffic only to healthy targets.

  • BMove the instance to a larger instance type with more memory and faster networking

    Incorrect

    A bigger instance is still one instance in one zone, so it is still a single point of failure.

  • CAdd a second EBS volume and mirror the database files across both with RAID 1

    Incorrect

    Both volumes stay attached to the same instance in the same zone, so an instance or zone failure still stops the application.

  • DCreate a nightly AMI of the instance so it can be launched again by hand

    Incorrect

    A manual relaunch from an AMI is slow and loses all data written since the image was taken.

  • EMove the database to Amazon RDS for MySQL with a Multi-AZ deployment

    Correct

    A Multi-AZ deployment keeps a synchronous standby in another zone and fails over automatically.

Remove each single point of failure: spread the web tier across zones behind a load balancer, and give the database a standby in another zone.

Question 9 · choose 1

An internal application runs on EC2 instances with EBS volumes and an Amazon RDS for MySQL database in one Region. For a Regional disaster the business accepts losing up to 24 hours of data and being down for up to 8 hours. It wants the lowest ongoing cost, with backups managed in one place. Which solution meets these requirements?

  1. AAWS Backup plans that take daily backups and copy every recovery point to a vault in a second Region
  2. BA warm standby in a second Region, with a smaller copy of the application running and a cross-Region read replica
  3. CRDS Multi-AZ for the database and an Auto Scaling group that spans two Availability Zones
  4. DDaily EBS and RDS snapshots kept in the same Region with a 35-day retention period
Show the answer and why
  • AAWS Backup plans that take daily backups and copy every recovery point to a vault in a second Region

    Correct

    Daily backups meet the 24-hour RPO, restoring in the second Region fits an 8-hour RTO, and AWS Backup manages EC2, EBS and RDS backups and their cross-Region copies in one place. Nothing runs in the second Region until a disaster.

  • BA warm standby in a second Region, with a smaller copy of the application running and a cross-Region read replica

    Incorrect

    A warm standby meets the targets easily, but it keeps servers and a replica running all the time, which costs far more than these loose targets need.

  • CRDS Multi-AZ for the database and an Auto Scaling group that spans two Availability Zones

    Incorrect

    Both protect against the loss of an Availability Zone. Neither keeps a copy of anything outside the Region.

  • DDaily EBS and RDS snapshots kept in the same Region with a 35-day retention period

    Incorrect

    Snapshots that exist only in the failed Region cannot be restored anywhere else. Regional disaster recovery needs copies in another Region.

An RPO of a day and an RTO of hours are the backup-and-restore tier, the cheapest disaster recovery strategy. The copies must live in another Region to survive a Regional event.

Question 10 · choose 1

Web servers run in an Auto Scaling group behind an Application Load Balancer. Sometimes the application on an instance hangs. The load balancer marks the target unhealthy and stops sending it traffic, but the instance keeps running and the group never replaces it, so capacity stays reduced. What should a solutions architect do?

  1. ATurn on detailed monitoring for the instances so that their status is checked every minute
  2. BAdd a scheduled action that replaces every instance in the group once a night
  3. CTurn on Elastic Load Balancing health checks for the Auto Scaling group
  4. DRaise the health check grace period of the group so that hanging instances have more time to recover
Show the answer and why
  • ATurn on detailed monitoring for the instances so that their status is checked every minute

    Incorrect

    Detailed monitoring sends metrics more often. The group still judges instances by EC2 status checks, which pass while the application is hung.

  • BAdd a scheduled action that replaces every instance in the group once a night

    Incorrect

    Scheduled actions change capacity at set times. They do not detect a failed instance, so capacity would stay reduced until the next run.

  • CTurn on Elastic Load Balancing health checks for the Auto Scaling group

    Correct

    With ELB health checks on, the group also uses the load balancer's health status and replaces an instance that fails it.

  • DRaise the health check grace period of the group so that hanging instances have more time to recover

    Incorrect

    The grace period only delays health checks for new instances. It does not make the group look at the load balancer's view of the instance.

By default an Auto Scaling group checks only EC2 status. Adding ELB health checks lets it replace instances whose application has failed.

Question 11 · choose 1

A ticketing application stores bookings in a DynamoDB table in us-east-1. The company is adding a European Region so that users on each side of the Atlantic write to a nearby Region, and either Region must keep accepting writes if the other fails. Which solution meets these requirements with the LEAST operational overhead?

  1. ACopy on-demand backups of the table to Europe every hour and restore them there during an outage
  2. BConvert the table to a global table with a replica in a European Region
  3. CStream changes with DynamoDB Streams to a Lambda function that writes each item to a table in Europe
  4. DMove the data to an Aurora global database with a secondary cluster in Europe
Show the answer and why
  • ACopy on-demand backups of the table to Europe every hour and restore them there during an outage

    Incorrect

    Backups give a copy, not a second writable table: Europe could not take writes until a restore finished, and up to an hour of bookings would be missing.

  • BConvert the table to a global table with a replica in a European Region

    Correct

    Global tables replicate automatically between replicas in different Regions, each replica accepts reads and writes, and the application can keep using one Region if the other is unavailable.

  • CStream changes with DynamoDB Streams to a Lambda function that writes each item to a table in Europe

    Incorrect

    This is replication you build and run yourself, one direction per function, with your own handling of conflicts and failures.

  • DMove the data to an Aurora global database with a secondary cluster in Europe

    Incorrect

    An Aurora global database has one primary Region that takes writes, and moving from DynamoDB to Aurora is a large migration.

Writes in two Regions with automatic replication and no custom code is the global tables feature of DynamoDB.

Question 12 · choose 1

A request to an ordering API passes through Amazon API Gateway, three Lambda functions, an SQS queue and a DynamoDB table. Some requests take several seconds, and the team cannot tell which component adds the delay. Which AWS service should a solutions architect use to find the slow component for individual requests?

  1. AAWS CloudTrail, to list the API calls that each component makes
  2. BVPC Flow Logs, to capture the network traffic between the components
  3. CAmazon CloudWatch metrics, to compare the average duration of each Lambda function in the request path
  4. DAWS X-Ray, with tracing turned on for API Gateway and the Lambda functions
Show the answer and why
  • AAWS CloudTrail, to list the API calls that each component makes

    Incorrect

    CloudTrail records who did what through AWS APIs, such as creating a function. It does not break a request's latency down by component.

  • BVPC Flow Logs, to capture the network traffic between the components

    Incorrect

    Flow logs record IP traffic metadata for network interfaces in a VPC. They show neither how long each service spent on a request nor calls to managed services outside the VPC.

  • CAmazon CloudWatch metrics, to compare the average duration of each Lambda function in the request path

    Incorrect

    Metrics are aggregates per function. They can hint at a slow function but cannot follow one request across API Gateway, the queue and the table.

  • DAWS X-Ray, with tracing turned on for API Gateway and the Lambda functions

    Correct

    X-Ray traces individual requests as they pass through the services and shows how much time each segment takes, which points at the slow component.

Per-request latency across several services is a distributed tracing question, which is what X-Ray answers.

Question 13 · choose 2

An Amazon RDS for PostgreSQL DB instance runs in a single Availability Zone. The company needs automatic failover with no data loss if the zone fails, and the ability to recover in another AWS Region with an RPO of a few minutes if the whole Region fails. Which steps meet these requirements? (Choose TWO.)

  1. AConvert the DB instance to a Multi-AZ deployment
  2. BCopy the automated snapshots to the other Region once a day
  3. CAdd a read replica in another Availability Zone of the Region
  4. DCreate a cross-Region read replica in the other Region
  5. EPut Amazon RDS Proxy in front of the DB instance
Show the answer and why
  • AConvert the DB instance to a Multi-AZ deployment

    Correct

    Multi-AZ keeps a synchronously replicated standby in another zone and fails over to it automatically, so committed data is not lost.

  • BCopy the automated snapshots to the other Region once a day

    Incorrect

    A daily copy gives an RPO of up to a day, far from a few minutes.

  • CAdd a read replica in another Availability Zone of the Region

    Incorrect

    A read replica is updated asynchronously and is not promoted automatically, and a replica in the same Region does not help when the Region fails.

  • DCreate a cross-Region read replica in the other Region

    Correct

    A cross-Region replica receives changes continuously, usually seconds to minutes behind, and can be promoted to a standalone database during a Regional outage.

  • EPut Amazon RDS Proxy in front of the DB instance

    Incorrect

    RDS Proxy pools connections and can shorten failover for clients, but it keeps no copy of the data in another zone or Region.

Multi-AZ answers the zone requirement with synchronous replication. A cross-Region read replica answers the Region requirement with an RPO of minutes.

Practise domain 2 →Practise all domains →