Skip to content
BytePatterns

SAA-C03 · Domain 2: Design Resilient Architectures · 26% of the exam

Task 2.1: Design scalable and loosely coupled architectures.

Splitting a system so each part scales and fails on its own: queues and topics between tiers, event-driven and serverless designs, containers, API Gateway, load balancers, caching, and stateless services that scale out instead of up.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An online store's web tier sends each order synchronously to an order-processing tier on EC2 instances. During sales the processing tier falls behind, requests time out, and orders are lost. The company wants every order kept until it is processed and the processing tier to scale with the amount of waiting work. What should a solutions architect do?

  1. APublish orders to an Amazon SNS topic, and subscribe the processing instances to it through HTTPS endpoints
  2. BMove the processing tier to larger instance types and raise the load balancer's idle timeout for slow requests
  3. CSend orders to an Amazon SQS queue, and scale the processing Auto Scaling group on the backlog per instance
  4. DPut the processing tier behind an Application Load Balancer with sticky sessions turned on for the web tier
Show the answer and why
  • APublish orders to an Amazon SNS topic, and subscribe the processing instances to it through HTTPS endpoints

    Incorrect

    SNS pushes each message to subscribers instead of storing work for them to pull, and its retries to an HTTPS endpoint stop after at most an hour, so a busy tier can still miss orders.

  • BMove the processing tier to larger instance types and raise the load balancer's idle timeout for slow requests

    Incorrect

    A bigger instance raises the ceiling but keeps the tiers synchronously coupled. A large enough spike still overwhelms it and loses orders, because nothing stores the work.

  • CSend orders to an Amazon SQS queue, and scale the processing Auto Scaling group on the backlog per instance

    Correct

    The queue holds each order until a consumer processes and deletes it, so neither tier has to be available at the same moment. Backlog per instance is the documented metric for scaling consumers of a queue.

  • DPut the processing tier behind an Application Load Balancer with sticky sessions turned on for the web tier

    Incorrect

    A load balancer spreads requests across healthy targets, and stickiness pins a client to one target. Neither stores requests that the targets cannot handle yet.

"Keep every order until processed" asks for a buffer between the tiers. A queue decouples them, and scaling on backlog per instance matches capacity to the waiting work.

Question 2 · choose 1

When an order is placed, three independent services (billing, shipping and analytics) must each receive a copy of the order event and process it at their own pace. The analytics service can be offline for up to two days for maintenance and must not lose any events. Which design meets these requirements?

  1. APublish each event to an Amazon SNS topic, with a separate Amazon SQS queue per service subscribed
  2. BSend each event to one Amazon SQS standard queue that all three services poll for new messages
  3. CSend each event to an Amazon SQS FIFO queue and give each of the three services its own message group ID
  4. DPublish each event to an Amazon SNS topic, with each service's Lambda function subscribed directly
Show the answer and why
  • APublish each event to an Amazon SNS topic, with a separate Amazon SQS queue per service subscribed

    Correct

    SNS-to-SQS fanout gives every service its own copy in its own queue. A queue keeps messages for 4 days by default and up to 14, so analytics catches up after maintenance.

  • BSend each event to one Amazon SQS standard queue that all three services poll for new messages

    Incorrect

    A message received by one consumer is hidden from the others and then deleted, so each event would reach only one of the three services.

  • CSend each event to an Amazon SQS FIFO queue and give each of the three services its own message group ID

    Incorrect

    Message group IDs keep related messages in order; they do not copy a message to several consumers. Each event is still delivered for processing once.

  • DPublish each event to an Amazon SNS topic, with each service's Lambda function subscribed directly

    Incorrect

    SNS invokes Lambda asynchronously. By default Lambda keeps retrying an event for at most 6 hours (a function error gets only two more attempts), so events from a two-day outage would be dropped.

One event, many independent consumers, and long outages: fan out with SNS into one SQS queue per consumer. The queues absorb each consumer's downtime separately.

Question 3 · choose 1

A loan application workflow calls several services in sequence, retries failed steps, and must pause for up to 5 days while a human underwriter approves each application. The company wants a managed, auditable workflow with a visual history of every execution. What should a solutions architect recommend?

  1. AAn AWS Step Functions Express workflow that waits for the approval through a callback task token
  2. BA single AWS Lambda function that runs every step and polls for the approval in a loop until granted
  3. CAn Amazon SQS delay queue that keeps each application hidden until the approval period has ended
  4. DAn AWS Step Functions Standard workflow that waits for the approval through a callback task token
Show the answer and why
  • AAn AWS Step Functions Express workflow that waits for the approval through a callback task token

    Incorrect

    Express workflows run for at most five minutes and support only the request-response pattern, not waiting for a callback.

  • BA single AWS Lambda function that runs every step and polls for the approval in a loop until granted

    Incorrect

    A Lambda function can run for at most 15 minutes, far short of a 5-day wait, and the team would code retries and history itself.

  • CAn Amazon SQS delay queue that keeps each application hidden until the approval period has ended

    Incorrect

    The maximum delay of an SQS delay queue is 15 minutes, and a queue does not orchestrate steps or retries.

  • DAn AWS Step Functions Standard workflow that waits for the approval through a callback task token

    Correct

    Standard workflows run for up to one year, keep an auditable execution history, and support waiting for a callback with a task token.

Long human pauses, retries and an audit trail describe a Standard workflow. Express workflows, Lambda functions and delay queues all have limits measured in minutes.

Question 4 · choose 1

A company is launching a public API for partners. Traffic is unpredictable and close to zero at night. Each partner needs its own API key with a monthly request quota and per-client throttling, and the company does not want to manage any servers. Which solution meets these requirements?

  1. AAn Amazon API Gateway HTTP API with a JWT authorizer, backed by AWS Lambda functions
  2. BAn Amazon API Gateway REST API with usage plans and API keys, in front of AWS Lambda
  3. CAn Application Load Balancer that routes requests to Lambda functions registered as targets
  4. DAmazon ECS on EC2 instances running an API proxy that tracks each partner's key and quota
Show the answer and why
  • AAn Amazon API Gateway HTTP API with a JWT authorizer, backed by AWS Lambda functions

    Incorrect

    HTTP APIs are the lower-cost, minimal option and do not support API keys, per-client rate limiting or usage throttling.

  • BAn Amazon API Gateway REST API with usage plans and API keys, in front of AWS Lambda

    Correct

    Usage plans for REST APIs attach API keys to throttling and quota limits per client, and Lambda scales with demand without servers.

  • CAn Application Load Balancer that routes requests to Lambda functions registered as targets

    Incorrect

    An ALB can invoke Lambda functions, but it has no API keys, quotas or per-client throttling.

  • DAmazon ECS on EC2 instances running an API proxy that tracks each partner's key and quota

    Incorrect

    With EC2 capacity the company manages the underlying instances itself, which the requirement rules out, and it would build the key and quota logic too.

API keys, quotas and per-client throttling are REST API features in API Gateway. HTTP APIs trade them away for a lower price.

Question 5 · choose 1

A company wants to move a stateless Java web application from on-premises virtual machines into containers on AWS. The team has no Kubernetes experience and does not want to provision, patch or scale any servers or clusters of instances. Which solution meets these requirements with the LEAST operational overhead?

  1. AAmazon ECS services that run the application's containers on AWS Fargate
  2. BAmazon EKS with a node group of EC2 instances for the application's pods
  3. CDocker on EC2 instances in an Auto Scaling group behind a load balancer
  4. DAmazon ECS with an Auto Scaling group of EC2 instances as its capacity
Show the answer and why
  • AAmazon ECS services that run the application's containers on AWS Fargate

    Correct

    With Fargate, ECS runs containers without servers or clusters of EC2 instances to provision, configure, patch or scale, and ECS avoids Kubernetes.

  • BAmazon EKS with a node group of EC2 instances for the application's pods

    Incorrect

    EKS runs Kubernetes, which the team does not know, and a node group of EC2 instances is still capacity the team chooses and maintains.

  • CDocker on EC2 instances in an Auto Scaling group behind a load balancer

    Incorrect

    The team would manage the instances, the container runtime and the deployments itself, which is the most work of all the options.

  • DAmazon ECS with an Auto Scaling group of EC2 instances as its capacity

    Incorrect

    With EC2 instances as capacity, the team manages instance selection, configuration and maintenance itself.

"No servers or clusters of instances" plus "no Kubernetes" points to ECS on Fargate.

Question 6 · choose 2

A web application runs on EC2 instances in an Auto Scaling group behind an Application Load Balancer. Users are logged out whenever the group scales in, because session data lives in each instance's memory. Which changes make the web tier stateless so that scaling no longer affects users? (Choose TWO.)

  1. ATurn on sticky sessions so that each user always returns to the same instance
  2. BMove the session files to the instance store volumes of each instance
  3. CKeep session data in an Amazon ElastiCache cluster that every instance uses
  4. DTurn on scale-in protection for every instance in the Auto Scaling group
  5. EKeep session data in an Amazon DynamoDB table with a TTL attribute for expiry
Show the answer and why
  • ATurn on sticky sessions so that each user always returns to the same instance

    Incorrect

    Stickiness sends a user back to the same target, but the session is still lost when that instance is terminated.

  • BMove the session files to the instance store volumes of each instance

    Incorrect

    Instance store data does not persist when an instance is stopped or terminated, and it is still local to one instance.

  • CKeep session data in an Amazon ElastiCache cluster that every instance uses

    Correct

    A shared in-memory store holds session state outside the instances, so any instance can serve any user.

  • DTurn on scale-in protection for every instance in the Auto Scaling group

    Incorrect

    Protected instances are never removed, which defeats scaling in, and sessions are still lost when an instance fails.

  • EKeep session data in an Amazon DynamoDB table with a TTL attribute for expiry

    Correct

    A table outside the instances survives any of them, and TTL deletes expired sessions without consuming write throughput.

Stateless means the state lives outside the instance. A shared cache or a table both work; stickiness and local disks keep the state on one machine.

Question 7 · choose 1

An order application uses an Amazon RDS for MySQL DB instance. Analysts run long reporting queries during business hours that slow down order processing. The reports can use data that is a few seconds old. What should a solutions architect do to take the reporting load off the primary database?

  1. AConvert the database to a Multi-AZ DB instance deployment and point the reports at the standby
  2. BScale the primary DB instance up to a larger instance class with more vCPUs and more memory
  3. CCreate a read replica of the database and point the reporting tool at the replica's endpoint
  4. DPut an Amazon ElastiCache cluster in front of the database and cache the reporting queries
Show the answer and why
  • AConvert the database to a Multi-AZ DB instance deployment and point the reports at the standby

    Incorrect

    The standby of a Multi-AZ DB instance deployment cannot serve read traffic; it exists for failover.

  • BScale the primary DB instance up to a larger instance class with more vCPUs and more memory

    Incorrect

    A bigger instance costs more and the reports still compete with orders on the same database.

  • CCreate a read replica of the database and point the reporting tool at the replica's endpoint

    Correct

    Read replicas are asynchronous, read-only copies meant for read-heavy work such as business reporting, which tolerates a little lag.

  • DPut an Amazon ElastiCache cluster in front of the database and cache the reporting queries

    Incorrect

    A cache helps when the same data is read again and again. Long ad hoc reports read different data each time, and the application would need caching logic.

Reporting that tolerates lag belongs on a read replica. A Multi-AZ standby is for failover and cannot be read.

Question 8 · choose 1

Business partners upload files every day to a company's SFTP server in its data center. The company is moving to AWS, wants the files to land in Amazon S3, and does not want to run or patch any SFTP servers. Partners must keep using their existing SFTP clients. Which solution meets these requirements?

  1. AAWS DataSync agents at each partner site that copy the files into the bucket on a schedule
  2. BAn AWS Transfer Family SFTP server that stores the uploaded files directly in the bucket
  3. CAn Amazon S3 File Gateway that the partners connect to with their existing SFTP clients
  4. DAn EC2 instance in an Auto Scaling group running an SFTP server that writes uploads to S3
Show the answer and why
  • AAWS DataSync agents at each partner site that copy the files into the bucket on a schedule

    Incorrect

    DataSync moves data between storage systems through agents you deploy. Partners would have to run agents instead of their SFTP clients.

  • BAn AWS Transfer Family SFTP server that stores the uploaded files directly in the bucket

    Correct

    Transfer Family is a fully managed SFTP service that writes into S3, and partners keep their existing client configuration.

  • CAn Amazon S3 File Gateway that the partners connect to with their existing SFTP clients

    Incorrect

    S3 File Gateway presents file shares over NFS and SMB, not SFTP.

  • DAn EC2 instance in an Auto Scaling group running an SFTP server that writes uploads to S3

    Incorrect

    It works, but the company would run and patch the SFTP servers, which it wants to avoid.

A managed SFTP endpoint in front of S3 is Transfer Family. File Gateway speaks NFS and SMB, and DataSync needs agents.

Question 9 · choose 1

Users upload up to 10,000 images per minute to an S3 bucket at peak. Each image must be sent to a third-party moderation API that allows at most 50 concurrent calls. If the API is down, the images must wait and be processed later, even after an outage of up to 3 days. Which design meets these requirements?

  1. AHave S3 invoke a Lambda function directly for every upload, and set its reserved concurrency to 50
  2. BRun a Lambda function every minute that lists the bucket and calls the API for each new object
  3. CSend the S3 event notifications to an Amazon SNS topic with the moderation API subscribed via HTTPS
  4. DSend S3 notifications to an SQS queue, and consume it with Lambda at a maximum concurrency of 50
Show the answer and why
  • AHave S3 invoke a Lambda function directly for every upload, and set its reserved concurrency to 50

    Incorrect

    S3 invokes Lambda asynchronously. By default throttled events are retried for at most 6 hours and a function error gets only two more attempts, so a 3-day outage would drop events.

  • BRun a Lambda function every minute that lists the bucket and calls the API for each new object

    Incorrect

    This rebuilds event handling by hand: listing a growing bucket, tracking what was sent and retrying failures are all custom code.

  • CSend the S3 event notifications to an Amazon SNS topic with the moderation API subscribed via HTTPS

    Incorrect

    SNS pushes every notification to the endpoint with no cap on concurrent calls, and its HTTPS retries stop after at most an hour.

  • DSend S3 notifications to an SQS queue, and consume it with Lambda at a maximum concurrency of 50

    Correct

    The queue keeps messages for up to 14 days, and the maximum concurrency setting of the SQS event source caps how many invocations run at once.

A queue absorbs both the burst and the outage, and the event source mapping's maximum concurrency protects the downstream API.

Question 10 · choose 1

A company's order, payment and shipping microservices publish events about their work. New consumers keep being added: a Lambda function for refunds over $500, an AWS Step Functions workflow for international shipments, and a partner's HTTPS API that accepts at most 10 requests per second. Publishers must not change when a consumer is added. Which solution meets these requirements with the LEAST custom code?

  1. AA central Lambda function that receives every event and invokes each consumer that the event concerns
  2. BAn Amazon EventBridge event bus, with one rule per consumer that matches on event content
  3. CAn Amazon Kinesis data stream that every consumer reads, with each consumer discarding the events it does not need
  4. DAn Amazon SQS standard queue that all consumers poll, each one deleting the messages that it handles
Show the answer and why
  • AA central Lambda function that receives every event and invokes each consumer that the event concerns

    Incorrect

    The router function is custom code that must know every consumer and change whenever one is added, and it would also have to throttle the calls to the partner's API itself.

  • BAn Amazon EventBridge event bus, with one rule per consumer that matches on event content

    Correct

    Rules filter events by their content and deliver them to targets such as Lambda functions and Step Functions state machines. An API destination calls an HTTPS endpoint with an invocation rate limit per second, so adding a consumer means adding a rule.

  • CAn Amazon Kinesis data stream that every consumer reads, with each consumer discarding the events it does not need

    Incorrect

    Every consumer needs code to read the stream and filter it, and a partner's HTTPS API cannot read a stream at all, so more custom code would have to push the events to it.

  • DAn Amazon SQS standard queue that all consumers poll, each one deleting the messages that it handles

    Incorrect

    A queue hands each message to one consumer at a time, and the consumer that processes it deletes it. Consumers compete for messages instead of each getting the events they need.

Content-based routing to many kinds of targets, with publishers unaware of consumers, is what an event bus with rules is for. EventBridge also calls external HTTPS APIs at a set rate through API destinations.

Question 11 · choose 1

A retail site caches product prices in Amazon ElastiCache with lazy loading and a 1-hour TTL. When the pricing service updates a price in the database, shoppers can see the old price for up to an hour. The company wants prices read from the cache and the new price returned right after an update. What should a solutions architect change?

  1. ALower the TTL to 1 minute so that stale prices expire from the cache sooner
  2. BAdd read replicas to the database and send every cache miss to one of the replicas
  3. CWrite each new price to the cache at the same time as it is written to the database
  4. DMove the cache cluster to larger node types so that it can keep many more prices in memory at once
Show the answer and why
  • ALower the TTL to 1 minute so that stale prices expire from the cache sooner

    Incorrect

    A shorter TTL shrinks the window but does not close it: a price can still be stale for up to a minute, and more reads miss the cache and go to the database.

  • BAdd read replicas to the database and send every cache miss to one of the replicas

    Incorrect

    Replicas take read load off the primary, but they do not change what the cache holds. Replication is also asynchronous, so a miss could even load an old price from a replica.

  • CWrite each new price to the cache at the same time as it is written to the database

    Correct

    This is the write-through strategy: the cache is updated whenever the database is, so the cached price is never older than the database.

  • DMove the cache cluster to larger node types so that it can keep many more prices in memory at once

    Incorrect

    More memory means fewer evictions, but staleness comes from how the cache is filled, not from its size.

Lazy loading only refreshes an item when it expires or is missing. Write-through updates the cache on every write, which keeps it current; a TTL can stay as a safety net.

Question 12 · choose 1

A company runs 40 microservices on a self-managed Kubernetes cluster in its data center and deploys them with Helm charts and kubectl. It is moving them to AWS, wants to keep its Kubernetes manifests and tooling, and no longer wants to operate the Kubernetes control plane. Which solution meets these requirements?

  1. AAmazon EKS, with the existing Helm charts and manifests deployed to the new cluster
  2. BAmazon ECS on AWS Fargate, with each Kubernetes deployment converted to an ECS task definition and service
  3. CA self-managed Kubernetes cluster on EC2 instances spread across three Availability Zones
  4. DAWS Elastic Beanstalk with its Docker platform, one environment for each microservice
Show the answer and why
  • AAmazon EKS, with the existing Helm charts and manifests deployed to the new cluster

    Correct

    EKS runs a conformant Kubernetes control plane for you across Availability Zones, so the existing manifests, Helm charts and kubectl keep working.

  • BAmazon ECS on AWS Fargate, with each Kubernetes deployment converted to an ECS task definition and service

    Incorrect

    ECS is a different orchestrator with its own API. Every manifest would have to be rewritten as task definitions and services, and the Kubernetes tooling would no longer apply.

  • CA self-managed Kubernetes cluster on EC2 instances spread across three Availability Zones

    Incorrect

    The manifests would work, but the company would still install, operate and upgrade the control plane itself, which it wants to stop doing.

  • DAWS Elastic Beanstalk with its Docker platform, one environment for each microservice

    Incorrect

    Elastic Beanstalk can run containers, but it does not use Kubernetes manifests or Helm, so the deployment tooling would be replaced.

"Keep Kubernetes, stop running the control plane" is the case for a managed Kubernetes service. ECS and Elastic Beanstalk run containers too, but not with Kubernetes APIs.

Question 13 · choose 1

A payments platform passes account transactions through a queue. Transactions for the same account must be processed in the order they were sent and exactly once, while transactions for different accounts should be processed in parallel by many consumers. Which solution meets these requirements?

  1. AAn Amazon SQS standard queue with a long visibility timeout and long polling turned on
  2. BAn Amazon SNS standard topic with an AWS Lambda function subscribed for each account
  3. CAn Amazon Kinesis data stream with the account ID as the partition key and one consumer per shard
  4. DAn Amazon SQS FIFO queue, with the account ID as the message group ID of each message
Show the answer and why
  • AAn Amazon SQS standard queue with a long visibility timeout and long polling turned on

    Incorrect

    A standard queue delivers each message at least once and orders messages on a best-effort basis, so duplicates and reordering are possible.

  • BAn Amazon SNS standard topic with an AWS Lambda function subscribed for each account

    Incorrect

    A standard topic does not guarantee the order of messages, and one subscription per account would not scale to many accounts.

  • CAn Amazon Kinesis data stream with the account ID as the partition key and one consumer per shard

    Incorrect

    The partition key keeps each account's records in order, but records can be delivered to the consumer more than once, for example after producer or consumer retries, so processing is not exactly once.

  • DAn Amazon SQS FIFO queue, with the account ID as the message group ID of each message

    Correct

    A FIFO queue keeps order within a message group and processes each message exactly once. Different message groups are handed to different consumers in parallel.

Ordering per key with parallelism across keys is the message group model of FIFO queues, and FIFO queues add exactly-once processing.

Question 14 · choose 2

A checkout API on Amazon API Gateway and AWS Lambda charges the card and then, in the same request, calls a third-party email service to send the receipt. When the email service is slow, checkouts time out. The company wants checkout to return as soon as payment succeeds, receipts to be retried if the email service fails, and receipts that keep failing to be kept for investigation. Which steps meet these requirements? (Choose TWO.)

  1. ARaise the API Gateway integration timeout and the function timeout so that the email call can finish
  2. BAfter payment, send a receipt message to an Amazon SQS queue that a separate Lambda function consumes
  3. CGive the checkout function more memory so that it calls the email service faster
  4. DTurn on provisioned concurrency for the checkout function so that more checkouts can run at once
  5. EConfigure a dead-letter queue for the receipt queue to receive messages that fail several times
Show the answer and why
  • ARaise the API Gateway integration timeout and the function timeout so that the email call can finish

    Incorrect

    Checkout would still wait for the email service, so a slow email provider still makes checkout slow, and a failed email still fails the request.

  • BAfter payment, send a receipt message to an Amazon SQS queue that a separate Lambda function consumes

    Correct

    The queue decouples checkout from the email service: checkout returns after payment, and a message that fails processing becomes visible again and is retried.

  • CGive the checkout function more memory so that it calls the email service faster

    Incorrect

    More memory adds CPU to the function, but the time is spent waiting for a remote service, which does not get faster.

  • DTurn on provisioned concurrency for the checkout function so that more checkouts can run at once

    Incorrect

    Provisioned concurrency keeps execution environments initialized. It does not shorten a request that waits on the email service.

  • EConfigure a dead-letter queue for the receipt queue to receive messages that fail several times

    Correct

    After the maximum number of receives, SQS moves a message to the dead-letter queue, where it is kept for investigation instead of being retried forever or lost.

Work that does not need to finish before the response belongs on a queue. The queue gives retries, and a dead-letter queue catches the messages that keep failing.

Practise domain 2 →Practise all domains →