Skip to content
BytePatterns

SAP-C02 · Domain 2: Design for New Solutions · 29% of the exam

Task 2.4: Design a strategy to meet reliability requirements

Designing for failure from the start: Multi-AZ and multi-Region patterns, loosely coupled components with queues and workflows, managed services that fail over by themselves, DNS routing policies, and service quotas that cap growth.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A new order intake API will receive bursts of up to 2,000 orders per second during flash sales. The downstream fulfillment system can process at most 50 orders per second and cannot be scaled. No order may be lost, and customers must get an immediate confirmation that their order was received. Which design meets these requirements?

  1. AUse API Gateway throttling at 50 requests per second so that the fulfillment system never receives more than it can process
  2. BSend the orders to an Amazon SQS queue consumed by a Lambda function with a maximum concurrency on its event source mapping
  3. CPublish each order to an Amazon SNS topic that the fulfillment system subscribes to over HTTPS
  4. DCall the fulfillment system synchronously from a Lambda function with reserved concurrency of 2,000
Show the answer and why
  • AUse API Gateway throttling at 50 requests per second so that the fulfillment system never receives more than it can process

    Incorrect

    Throttled requests are rejected with an error, so during a burst most orders would be refused instead of being kept for later.

  • BSend the orders to an Amazon SQS queue consumed by a Lambda function with a maximum concurrency on its event source mapping

    Correct

    The queue stores every order durably and absorbs the burst, and the maximum concurrency setting of the SQS event source caps how many function instances call the fulfillment system at once.

  • CPublish each order to an Amazon SNS topic that the fulfillment system subscribes to over HTTPS

    Incorrect

    SNS pushes each message to the subscriber as it arrives and retries for a limited time, so the burst would reach the fulfillment system directly.

  • DCall the fulfillment system synchronously from a Lambda function with reserved concurrency of 2,000

    Incorrect

    Reserved concurrency of 2,000 allows 2,000 parallel calls into a system that handles 50, and synchronous calls make customers wait for fulfillment.

A queue in front of a fixed-capacity system turns a burst into a backlog; the consumer's concurrency limit sets the drain rate.

Question 2 · choose 1

A company will run 60 Lambda functions in one account and Region. A payment function must always be able to scale to 400 concurrent executions, even when batch functions in the same account consume all the concurrency they can get. Load tests show that total demand at peak will exceed the account's current concurrency quota. Which actions should a solutions architect take?

  1. ARaise the memory of the payment function so that each invocation finishes sooner and needs less concurrency
  2. BAdd a dead-letter queue to the payment function so that throttled invocations can be retried later when capacity is free
  3. CDeploy the payment function in a second Availability Zone so that it has a separate concurrency pool
  4. DReserve 400 concurrency for the payment function and request a higher account concurrency quota in Service Quotas
Show the answer and why
  • ARaise the memory of the payment function so that each invocation finishes sooner and needs less concurrency

    Incorrect

    More memory can shorten invocations, but it guarantees nothing when other functions take all of the account's unreserved concurrency.

  • BAdd a dead-letter queue to the payment function so that throttled invocations can be retried later when capacity is free

    Incorrect

    A dead-letter queue keeps failed asynchronous events. Payments would still be throttled at the moment they are needed.

  • CDeploy the payment function in a second Availability Zone so that it has a separate concurrency pool

    Incorrect

    Lambda concurrency quotas apply per AWS Region, not per Availability Zone, so there is no separate pool to move the function into.

  • DReserve 400 concurrency for the payment function and request a higher account concurrency quota in Service Quotas

    Correct

    Reserved concurrency sets aside that much of the account's concurrency for one function so others cannot use it, and the Regional concurrency quota can be raised through a quota increase request.

Quotas are a design input: reserve what a critical function needs, and raise the shared quota before peak demand meets it.

Question 3 · choose 1

A web application will run in eu-central-1, us-east-1 and ap-southeast-1. For legal reasons, users in Germany must always be served from eu-central-1. All other users must be served from the Region with the lowest latency for them. Which Route 53 configuration meets these requirements?

  1. ALatency records for the three Regions, each with a health check so that unhealthy Regions are skipped automatically
  2. BGeoproximity records with a positive bias on eu-central-1 large enough to cover Germany
  3. CA geolocation record for Germany to eu-central-1, plus a default geolocation record that aliases latency records
  4. DWeighted records that send a share of traffic to each Region in proportion to its share of users
Show the answer and why
  • ALatency records for the three Regions, each with a health check so that unhealthy Regions are skipped automatically

    Incorrect

    Latency routing picks the Region with the lowest latency, which for some German users might not be eu-central-1, so the legal requirement is not guaranteed.

  • BGeoproximity records with a positive bias on eu-central-1 large enough to cover Germany

    Incorrect

    Geoproximity routing shifts traffic by distance and bias. It does not guarantee that every user in one country goes to one Region.

  • CA geolocation record for Germany to eu-central-1, plus a default geolocation record that aliases latency records

    Correct

    Geolocation routing answers by the user's location, and the default record catches every other location. Records with different routing policies can be combined by aliasing one to the other.

  • DWeighted records that send a share of traffic to each Region in proportion to its share of users

    Incorrect

    Weighted routing distributes traffic by weight regardless of the user's location or latency.

Hard location rules use geolocation; "everyone else by latency" sits behind the default geolocation record as a latency record set.

Question 4 · choose 2

A new booking application will use Amazon RDS for MySQL for bookings and an Amazon ElastiCache (Redis OSS) cluster for session data. The architecture review found that each data store would run in a single Availability Zone. The application must keep working, with automatic recovery, if one Availability Zone fails. Which changes meet this requirement? (Choose TWO.)

  1. ADeploy the database as a Multi-AZ deployment so that RDS fails over to a standby in another Availability Zone
  2. BTake hourly snapshots of the database and copy each snapshot to a second Availability Zone
  3. CUse a Redis OSS replication group with replicas in other zones and Multi-AZ automatic failover turned on
  4. DLaunch the database and cache nodes in a cluster placement group to get the lowest latency between them
  5. EMove both data stores to larger node types so that a single node can handle the full load
Show the answer and why
  • ADeploy the database as a Multi-AZ deployment so that RDS fails over to a standby in another Availability Zone

    Correct

    In a Multi-AZ deployment RDS keeps a synchronous standby in a different Availability Zone and fails over to it automatically if the primary has a problem.

  • BTake hourly snapshots of the database and copy each snapshot to a second Availability Zone

    Incorrect

    Snapshots allow a manual restore later. Restoring is not automatic and loses the changes made since the last snapshot.

  • CUse a Redis OSS replication group with replicas in other zones and Multi-AZ automatic failover turned on

    Correct

    With Multi-AZ turned on, ElastiCache promotes a read replica in another Availability Zone when the primary node fails.

  • DLaunch the database and cache nodes in a cluster placement group to get the lowest latency between them

    Incorrect

    A cluster placement group packs instances close together in one Availability Zone, which is the opposite of surviving a zone failure.

  • EMove both data stores to larger node types so that a single node can handle the full load

    Incorrect

    Larger nodes add capacity but are still single points of failure in one Availability Zone.

Surviving the loss of a zone needs a second copy in another zone and automatic failover: Multi-AZ for RDS and Multi-AZ replication groups for ElastiCache.

Question 5 · choose 1

In a new event-driven design, a consumer service had a bug for six hours and processed events incorrectly from an EventBridge event bus. After the fix, the team wants to send those six hours of events again. What should the design include?

  1. AA longer retry policy on the rule's target so delivery is attempted again
  2. BCloudWatch Logs of the consumer's output
  3. CA dead-letter queue on the rule
  4. DAn EventBridge archive on the event bus, replayed after the fix
Show the answer and why
  • AA longer retry policy on the rule's target so delivery is attempted again

    Incorrect

    Retries apply to failed deliveries, not to events that were delivered and processed incorrectly.

  • BCloudWatch Logs of the consumer's output

    Incorrect

    Logs of output cannot resend the original events.

  • CA dead-letter queue on the rule

    Incorrect

    A dead-letter queue holds failed deliveries, not successfully delivered events.

  • DAn EventBridge archive on the event bus, replayed after the fix

    Correct

    EventBridge can archive events so you can replay them to the event bus that originally received them.

Resending past events is EventBridge archive and replay.

Question 6 · choose 1

A new Aurora PostgreSQL database will serve a workload that is quiet most of the day and spikes sharply at unpredictable times. The team wants capacity to follow demand automatically and to pay only for what it uses. Which configuration fits?

  1. AA provisioned instance resized by a scheduled job
  2. BAurora Serverless capacity for the DB instances
  3. CThe largest provisioned instance class sized for the peak
  4. DMore read replicas added by hand during spikes
Show the answer and why
  • AA provisioned instance resized by a scheduled job

    Incorrect

    Schedules cannot follow spikes that come at unpredictable times.

  • BAurora Serverless capacity for the DB instances

    Correct

    Aurora Serverless adjusts capacity automatically based on application demand, and you are charged only for the resources consumed.

  • CThe largest provisioned instance class sized for the peak

    Incorrect

    Sizing for the peak pays for unused capacity most of the day.

  • DMore read replicas added by hand during spikes

    Incorrect

    Manual changes are slow and replicas do not add write capacity.

Unpredictable database load fits Aurora Serverless.

Question 7 · choose 1

A new ECS service behind an Application Load Balancer handles requests whose cost is roughly the same. The team wants the number of tasks to follow traffic automatically and keep a steady number of requests per task. Which scaling setup fits?

  1. ATarget tracking on ALBRequestCountPerTarget for the service
  2. BA daily scheduled scale-out at 9:00
  3. CA fixed desired count set for the busiest hour of the week
  4. DScaling on the memory use of the cluster's EC2 instances
Show the answer and why
  • ATarget tracking on ALBRequestCountPerTarget for the service

    Correct

    Service Auto Scaling target tracking can use the ALBRequestCountPerTarget metric to keep requests per task near a target.

  • BA daily scheduled scale-out at 9:00

    Incorrect

    Schedules do not follow actual traffic.

  • CA fixed desired count set for the busiest hour of the week

    Incorrect

    A fixed count wastes capacity off-peak and cannot follow changes.

  • DScaling on the memory use of the cluster's EC2 instances

    Incorrect

    Instance memory does not track requests per task.

Request-driven ECS scaling uses target tracking on requests per target.

Question 8 · choose 1

A new UDP game lobby service runs on eight EC2 instances with public IP addresses and no load balancer. Clients look up the service by DNS name. The team wants DNS to return several addresses and leave out instances that fail health checks. Which Route 53 routing policy fits?

  1. AFailover routing with one primary instance
  2. BSimple routing with all eight addresses in one record for the service
  3. CGeolocation routing by continent
  4. DMultivalue answer routing with a health check per record
Show the answer and why
  • AFailover routing with one primary instance

    Incorrect

    Failover routing sends all traffic to one primary, not to several healthy instances.

  • BSimple routing with all eight addresses in one record for the service

    Incorrect

    Simple routing cannot use health checks to drop failed addresses.

  • CGeolocation routing by continent

    Incorrect

    Geolocation routes by user location, not by health of several addresses.

  • DMultivalue answer routing with a health check per record

    Correct

    Multivalue answer routing returns multiple values and returns only values for healthy resources.

Several health-checked answers come from multivalue answer routing.

Question 9 · choose 1

A new web application keeps user sessions in the memory of each web server, so users are signed out whenever an instance is replaced or scaled in. The team wants stateless web servers with fast session access. Which change fits?

  1. AUse larger instances so they are replaced less often
  2. BTurn on sticky sessions on the load balancer only
  3. CStore sessions on each instance's EBS volume
  4. DStore sessions in Amazon ElastiCache
Show the answer and why
  • AUse larger instances so they are replaced less often

    Incorrect

    Instances are still replaced and scaled in.

  • BTurn on sticky sessions on the load balancer only

    Incorrect

    Stickiness keeps a user on one server, but the session is still lost when that server goes away.

  • CStore sessions on each instance's EBS volume

    Incorrect

    Sessions stay tied to one instance.

  • DStore sessions in Amazon ElastiCache

    Correct

    ElastiCache provides caching primitives for session state, keeping it outside the web servers.

Stateless web tiers keep session state in a shared store such as ElastiCache.

Practise domain 2 →Practise all domains →