Skip to content
BytePatterns

SAP-C02 · Domain 3: Continuous Improvement for Existing Solutions · 25% of the exam

Task 3.4: Determine a strategy to improve reliability.

Finding the weak spots in a running architecture: single points of failure, missing replication, quotas close to their limit, and the self-healing and elastic features that remove them.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

Business partners upload files to a company over SFTP. The SFTP server is a single EC2 instance that writes the files to its own EBS volume, and each instance failure stops all uploads until someone rebuilds the server. The company wants to remove this single point of failure with the least operational work, and the partners must keep using SFTP. What should a solutions architect recommend?

  1. AMove to an AWS Transfer Family SFTP server that stores the files in Amazon S3
  2. BTake hourly EBS snapshots of the volume so that the server can be rebuilt faster
  3. CPut the instance in an Auto Scaling group of one so that it is replaced after a failure
  4. DAsk partners to upload with the S3 console instead of SFTP
Show the answer and why
  • AMove to an AWS Transfer Family SFTP server that stores the files in Amazon S3

    Correct

    Transfer Family is a managed service for SFTP transfers into and out of Amazon S3 or Amazon EFS, so there is no server for the company to keep running.

  • BTake hourly EBS snapshots of the volume so that the server can be rebuilt faster

    Incorrect

    Snapshots shorten a rebuild but leave one instance as the single point of failure.

  • CPut the instance in an Auto Scaling group of one so that it is replaced after a failure

    Incorrect

    A replacement instance launches from the group's launch template, so files written to the old instance's volume are not on it, and uploads stop while it launches.

  • DAsk partners to upload with the S3 console instead of SFTP

    Incorrect

    The partners must keep using SFTP, so changing their tool is not allowed.

A managed, highly available protocol endpoint removes the server rather than protecting it.

Question 2 · choose 1

A DynamoDB table in provisioned mode with auto scaling serves a mobile app. As the app has grown, traffic now jumps to about twice its previous peak within a minute several times a day, and writes are throttled until auto scaling catches up. The team wants the throttling gone without predicting capacity. Which change meets these requirements?

  1. AAdd a DynamoDB Accelerator (DAX) cluster in front of the table
  2. BRaise the target utilization of the auto scaling policy to 90%
  3. CSwitch the table to on-demand capacity mode
  4. DAdd a replica of the table in a second Region as a global table
Show the answer and why
  • AAdd a DynamoDB Accelerator (DAX) cluster in front of the table

    Incorrect

    DAX caches reads. Writes still go to the table, so write throttling remains.

  • BRaise the target utilization of the auto scaling policy to 90%

    Incorrect

    A higher target leaves less headroom, so throttling during a sudden jump would get worse.

  • CSwitch the table to on-demand capacity mode

    Correct

    On-demand mode serves requests without capacity planning and instantly accommodates up to double the previous peak traffic on a table.

  • DAdd a replica of the table in a second Region as a global table

    Incorrect

    A replica in another Region does not add write capacity to the table that the app writes to, and replicated writes consume capacity there too.

Sudden jumps within twice the previous peak are what on-demand mode absorbs; provisioned auto scaling reacts after the fact.

Question 3 · choose 1

A content management system stores its shared files on Amazon EFS in eu-west-1. The company's new continuity plan requires a copy of the file system in eu-central-1 that is normally no more than about 15 minutes behind and that the application can fail over to during a Regional outage. Which solution meets these requirements with the least effort?

  1. ACreate an AWS Backup plan that copies daily EFS backups to eu-central-1
  2. BTurn on Amazon EFS replication to a destination file system in eu-central-1
  3. CSchedule an AWS DataSync task that copies the file system to eu-central-1 every hour
  4. DConfigure S3 Cross-Region Replication for the file system
Show the answer and why
  • ACreate an AWS Backup plan that copies daily EFS backups to eu-central-1

    Incorrect

    A daily backup copy can be up to a day behind, far beyond 15 minutes.

  • BTurn on Amazon EFS replication to a destination file system in eu-central-1

    Correct

    EFS replication keeps a destination file system in another Region in sync automatically, with an RPO of 15 minutes for most file systems, and you can fail over to the replica during a disaster.

  • CSchedule an AWS DataSync task that copies the file system to eu-central-1 every hour

    Incorrect

    An hourly task can leave the copy up to an hour behind, and the team would manage the schedule and the failover itself.

  • DConfigure S3 Cross-Region Replication for the file system

    Incorrect

    S3 replication works on S3 buckets. An EFS file system is replicated with EFS replication.

A Regional copy of EFS that stays within minutes of the source is built in: EFS replication.

Question 4 · choose 1

EC2 instances in an Auto Scaling group sit behind an Application Load Balancer. Sometimes the application process hangs: the load balancer marks the target unhealthy and stops sending it traffic, but the instance keeps running for days until someone replaces it, and the group runs short of capacity. Which change makes the system heal itself?

  1. ATurn on instance scale-in protection so that the group keeps its instances
  2. BAdd a CloudWatch alarm on the CPUUtilization metric that notifies the team
  3. CIncrease the desired capacity by one so that there is always a spare instance
  4. DTurn on Elastic Load Balancing health checks for the Auto Scaling group
Show the answer and why
  • ATurn on instance scale-in protection so that the group keeps its instances

    Incorrect

    Scale-in protection stops the group from terminating instances during scale-in, which keeps the hung instance even longer.

  • BAdd a CloudWatch alarm on the CPUUtilization metric that notifies the team

    Incorrect

    A notification still needs a person to act, and a hung process does not reliably show up in CPU use.

  • CIncrease the desired capacity by one so that there is always a spare instance

    Incorrect

    A spare instance hides the lost capacity but leaves unhealthy instances running and paid for.

  • DTurn on Elastic Load Balancing health checks for the Auto Scaling group

    Correct

    With ELB health checks turned on, Auto Scaling also uses the load balancer's health check results and replaces instances that fail them.

By default, Auto Scaling looks only at EC2 status checks; adding the load balancer's health checks lets it replace instances whose application failed.

Question 5 · choose 1

Fourteen VPCs attach to a transit gateway in us-east-1. The company's data center reaches them over a single 1 Gbps dedicated AWS Direct Connect connection, with a transit virtual interface to a Direct Connect gateway that is associated with the transit gateway. Last month a fiber cut at the Direct Connect location stopped all hybrid traffic for six hours. There is no budget for another Direct Connect connection this year. During a Direct Connect failure, traffic may cross the internet at lower bandwidth if it is encrypted, and failover and failback must happen without manual steps. Which change meets these requirements?

  1. ATurn on SiteLink for the transit virtual interface so traffic can reach AWS through other Direct Connect locations
  2. BAdd a private virtual interface on the existing connection, attached to the same Direct Connect gateway as a standby path
  3. CBundle the connection with a second 1 Gbps connection at the same location into a link aggregation group
  4. DAdd a Site-to-Site VPN attachment with dynamic BGP routing to the transit gateway as a backup for the Direct Connect path
Show the answer and why
  • ATurn on SiteLink for the transit virtual interface so traffic can reach AWS through other Direct Connect locations

    Incorrect

    SiteLink connects on-premises sites to each other between Direct Connect points of presence over the AWS network. It does not give the data center a path when its only connection fails.

  • BAdd a private virtual interface on the existing connection, attached to the same Direct Connect gateway as a standby path

    Incorrect

    Virtual interfaces are what you create on a Direct Connect connection to use it. A second one still runs over the same physical connection and location, so the next fiber cut takes both down.

  • CBundle the connection with a second 1 Gbps connection at the same location into a link aggregation group

    Incorrect

    A LAG needs another dedicated connection, which the budget does not allow, and all connections in a LAG terminate at the same Direct Connect endpoint, so a failure at that location still cuts both.

  • DAdd a Site-to-Site VPN attachment with dynamic BGP routing to the transit gateway as a backup for the Direct Connect path

    Correct

    AWS describes a Site-to-Site VPN terminating on a transit gateway as a lower-cost backup for Direct Connect. For the same prefix, the transit gateway prefers routes propagated from the Direct Connect gateway over routes propagated from a VPN, so the VPN carries traffic only while Direct Connect is down.

The single point of failure is the one physical connection at one location, and the budget rules out a second connection. An IPsec VPN over the internet to the same transit gateway is the lower-cost backup AWS suggests for connections up to 1 Gbps. Because the transit gateway ranks Direct Connect gateway routes above VPN-propagated routes, traffic returns to Direct Connect on its own once it recovers. Use BGP rather than static VPN routes, because static routes would win over Direct Connect.

Question 6 · choose 1

A nightly batch job runs on 200 container workers that write to one provisioned DynamoDB table through an AWS SDK. The launch script of each worker exports AWS_MAX_ATTEMPTS=1, so short bursts of throttling fail items. In a test with a fixed one-second delay between retries, the workers retried in lockstep and were throttled again. The team must ride out brief throttling without failing items, avoid synchronized retries, keep the table's capacity mode and cost unchanged, add no new infrastructure, and change only configuration, not code. Which solution meets these requirements?

  1. AKeep the one-second retries but give each worker a different fixed delay derived from its worker number
  2. BSwitch the table to on-demand capacity mode so that DynamoDB absorbs the bursts without throttling the workers
  3. CSend the writes to an SQS queue and have a separate consumer write them to the table at a controlled rate
  4. DRemove the launch script's override and set the SDK's retry mode to standard with a higher max attempts value
Show the answer and why
  • AKeep the one-second retries but give each worker a different fixed delay derived from its worker number

    Incorrect

    Different fixed delays would spread the first retry, but the delays never grow while throttling lasts, and computing them per worker is a code change. The SDK's standard mode already backs off with jitter.

  • BSwitch the table to on-demand capacity mode so that DynamoDB absorbs the bursts without throttling the workers

    Incorrect

    On-demand mode handles throughput management and scaling for you, which can remove the throttling, but it changes the table's capacity mode and how it is billed, which the team must keep as they are.

  • CSend the writes to an SQS queue and have a separate consumer write them to the table at a controlled rate

    Incorrect

    A queue buffers requests and decouples the workers from the table, which smooths bursts, but it is new infrastructure and new code, both ruled out.

  • DRemove the launch script's override and set the SDK's retry mode to standard with a higher max attempts value

    Correct

    Standard mode retries failed requests with exponential backoff and jitter, using longer delays for throttling errors, so workers stop retrying in lockstep. Retry mode and max attempts can be set through environment variables or the shared config file, without code changes.

The constraints are no failed items, no synchronized retries, unchanged table capacity, no new infrastructure and configuration-only changes. On-demand mode changes billing, a queue adds infrastructure, and staggered fixed delays need code and never back off. The SDK's standard retry mode, set through configuration, adds exponential backoff with jitter.

Question 7 · choose 1

A Lambda function reads a Kinesis data stream. One malformed record makes the whole batch fail again and again, and the shard stops making progress. Which configuration helps isolate the bad record?

  1. AIncrease the batch size
  2. BIncrease the number of shards
  3. CTurn on BisectBatchOnFunctionError
  4. DLower the stream's retention period
Show the answer and why
  • AIncrease the batch size

    Incorrect

    Larger batches still fail as a whole.

  • BIncrease the number of shards

    Incorrect

    More shards do not remove the bad record from its batch.

  • CTurn on BisectBatchOnFunctionError

    Correct

    When an invocation fails and bisecting is on, the batch is split so the failing record can be isolated.

  • DLower the stream's retention period

    Incorrect

    Retention controls how long records stay, not how failures are handled.

Splitting failed Kinesis batches isolates poison records.

Question 8 · choose 1

A pipeline deploys an Amazon ECS service on Fargate by calling UpdateService with a new task definition; the service uses the rolling update deployment type. Two recent releases went wrong: in one, new tasks kept failing health checks while the deployment kept retrying for hours; in the other, the tasks were healthy but the service's 5xx CloudWatch alarm fired. The team wants ECS to fail the deployment and return to the last working version automatically in both cases, while keeping rolling updates and the current load balancer setup, with no extra infrastructure. Which solution meets these requirements?

  1. AMove the service into a CloudFormation stack and add the 5xx alarm as a rollback trigger for every stack update
  2. BLengthen the health check grace period so that new tasks have more time to pass their load balancer health checks
  3. CSet the minimum healthy percent to 100 and the maximum percent to 200 so that old tasks keep running until new tasks are ready
  4. DTurn on the deployment circuit breaker with rollback and add the 5xx alarm to the service's deployment alarms with rollback
Show the answer and why
  • AMove the service into a CloudFormation stack and add the 5xx alarm as a rollback trigger for every stack update

    Incorrect

    Rollback triggers make CloudFormation roll back a stack create or update when an alarm breaches, which suits stack-managed releases. This pipeline deploys with UpdateService, and the trigger does nothing about tasks that never reach a steady state.

  • BLengthen the health check grace period so that new tasks have more time to pass their load balancer health checks

    Incorrect

    A grace period suits applications that start slowly, because it delays health check evaluation. It neither fails nor rolls back a deployment whose tasks are broken or whose error alarm fires.

  • CSet the minimum healthy percent to 100 and the maximum percent to 200 so that old tasks keep running until new tasks are ready

    Incorrect

    These settings bound how many tasks run during a rolling update, which protects capacity. They do not detect a failed deployment or return the service to the previous version.

  • DTurn on the deployment circuit breaker with rollback and add the 5xx alarm to the service's deployment alarms with rollback

    Correct

    The circuit breaker detects tasks that do not reach a steady state and can roll back to the deployment in the COMPLETED state. Deployment alarms mark a deployment as failed when a CloudWatch alarm enters ALARM and can also roll it back, all within rolling updates.

The constraints are automatic failure detection for both unhealthy tasks and a breaching alarm, automatic return to the last good version, and rolling updates with no new infrastructure. Grace periods and task percentages tune a rollout but never roll back, and CloudFormation rollback triggers apply only to stack updates. The circuit breaker plus deployment alarms, both with rollback, cover both failure modes.

Question 9 · choose 1

Workers that process an SQS queue sometimes hang without crashing, and orders silently wait for hours before anyone notices. Which alarm would detect this condition quickly?

  1. AAn alarm on the queue's message size
  2. BAn alarm on the CPU of the worker instances
  3. CAn alarm on the number of messages sent to the queue
  4. DAn alarm on the queue's ApproximateAgeOfOldestMessage
Show the answer and why
  • AAn alarm on the queue's message size

    Incorrect

    Message size does not change when consumers hang.

  • BAn alarm on the CPU of the worker instances

    Incorrect

    Hung workers can show normal or low CPU.

  • CAn alarm on the number of messages sent to the queue

    Incorrect

    Producers keep sending normally while consumers hang.

  • DAn alarm on the queue's ApproximateAgeOfOldestMessage

    Correct

    This metric reports the age of the oldest unprocessed message, which grows when consumers stop processing.

Stalled consumers show up as a growing age of the oldest message.

Practise domain 3 →Practise all domains →