Skip to content
BytePatterns

DOP-C02 · Domain 3: Resilient Cloud Solutions · 15% of the exam

Task 3.1: Implement highly available solutions to meet resilience and business requirements.

Turning an availability target into a design: Multi-AZ and multi-Region data and compute, replication for stateful services, removing single points of failure, and load balancing across Availability Zones without downtime.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 2

An order service runs on Amazon EC2 instances in an Auto Scaling group that spans three Availability Zones behind an Application Load Balancer. It uses a single-AZ Amazon RDS for PostgreSQL DB instance, and all private subnets send traffic to a payment provider through one zonal NAT gateway in the first Availability Zone. The service must keep working if any one Availability Zone fails, with the least ongoing operational effort. Which changes should the DevOps engineer make? (Choose TWO.)

  1. AConvert the DB instance to a Multi-AZ DB instance deployment
  2. BCreate a read replica in a second Availability Zone and promote it by hand if the primary DB instance becomes unavailable
  3. CReplace the zonal NAT gateway with a regional NAT gateway and route the private subnets in every Availability Zone to it
  4. DAdd a second zonal NAT gateway in another Availability Zone and keep all private route tables pointed at the first NAT gateway
  5. EMove the DB instance to a larger instance class with Provisioned IOPS storage to absorb the extra load after a zone failure
Show the answer and why
  • AConvert the DB instance to a Multi-AZ DB instance deployment

    Correct

    RDS then keeps a synchronous standby in a different Availability Zone and fails over to it automatically when the primary or its Availability Zone fails.

  • BCreate a read replica in a second Availability Zone and promote it by hand if the primary DB instance becomes unavailable

    Incorrect

    Promotion is a manual step that turns the replica into a standalone instance, so it adds operational work during the outage.

  • CReplace the zonal NAT gateway with a regional NAT gateway and route the private subnets in every Availability Zone to it

    Correct

    A regional NAT gateway expands across the Availability Zones where the workload runs, so outbound traffic no longer depends on one zone.

  • DAdd a second zonal NAT gateway in another Availability Zone and keep all private route tables pointed at the first NAT gateway

    Incorrect

    Traffic still goes only through the first NAT gateway, so losing its Availability Zone still cuts off the other zones.

  • EMove the DB instance to a larger instance class with Provisioned IOPS storage to absorb the extra load after a zone failure

    Incorrect

    More capacity in one Availability Zone does not help when that zone is the one that fails.

The two single points of failure are the single-AZ database and the NAT gateway in one zone. Multi-AZ gives the database automatic failover, and a regional NAT gateway removes the shared dependency on one Availability Zone.

Question 2 · choose 1

A payments ledger is moving to an Amazon DynamoDB table and must run in us-east-1 and us-east-2. Either Region must accept writes and keep serving if the other is impaired, no acknowledged write may be lost, and a strongly consistent read in either Region must always return the latest version of an item. The team accepts a third Region for availability but does not want to run the application or store a full readable copy of the data there. What should the DevOps engineer do?

  1. ACreate a separate table in each Region and have the application write every item to both tables in one TransactWriteItems call
  2. BCreate an MRSC global table with replicas in us-east-1 and us-east-2 only, and no component in any other Region
  3. CCreate an MRSC global table with replicas in us-east-1 and us-east-2 and a witness in a third supported Region
  4. DKeep a single-Region table with point-in-time recovery and copy its backups to us-east-2 with AWS Backup every hour
Show the answer and why
  • ACreate a separate table in each Region and have the application write every item to both tables in one TransactWriteItems call

    Incorrect

    A TransactWriteItems call can only target tables in the same AWS account and Region, so it cannot write to tables in two Regions atomically.

  • BCreate an MRSC global table with replicas in us-east-1 and us-east-2 only, and no component in any other Region

    Incorrect

    An MRSC global table must be deployed in exactly three Regions, as three replicas or as two replicas and a witness.

  • CCreate an MRSC global table with replicas in us-east-1 and us-east-2 and a witness in a third supported Region

    Correct

    MRSC replicates each write synchronously to another Region before it succeeds, strongly consistent reads return the latest version, and a witness holds no readable copy in the third Region.

  • DKeep a single-Region table with point-in-time recovery and copy its backups to us-east-2 with AWS Backup every hour

    Incorrect

    Restoring a backup copy loses the writes made since the last copy, and us-east-2 would not serve traffic until the restore finishes.

Multi-Region strong consistency (MRSC) replicates each write synchronously to at least one other Region before it succeeds, so no acknowledged write is lost and a strongly consistent read in any replica Region returns the latest version. It needs three Regions, and one of them can be a witness that takes part in the replication without holding a readable copy of the data.

Question 3 · choose 1

A TCP service runs behind a Network Load Balancer in two Availability Zones. After a partial capacity shortage, the target group has 2 healthy targets in the first zone and 8 in the second. Clients receive both load balancer node addresses from DNS about equally, and the 2 targets in the first zone are now overloaded while the others are mostly idle. Traffic must be spread evenly across all healthy targets. What should the DevOps engineer do?

  1. ATurn on cross-zone load balancing for the Network Load Balancer or for its target group
  2. BRaise the target group's deregistration delay so that the busy targets finish their in-flight connections before new ones arrive
  3. CTurn on stickiness for the target group so that each client keeps using the same target for the length of its session
  4. DLower the DNS TTL of the load balancer's record so that clients resolve the node addresses again more often
Show the answer and why
  • ATurn on cross-zone load balancing for the Network Load Balancer or for its target group

    Correct

    With cross-zone load balancing each node spreads traffic across targets in all enabled zones. It is off by default for Network Load Balancers.

  • BRaise the target group's deregistration delay so that the busy targets finish their in-flight connections before new ones arrive

    Incorrect

    The deregistration delay only controls how long a deregistering target stays in the draining state. It does not change where new traffic goes.

  • CTurn on stickiness for the target group so that each client keeps using the same target for the length of its session

    Incorrect

    Stickiness keeps clients on targets they already use, which does not move load from the first zone to the second.

  • DLower the DNS TTL of the load balancer's record so that clients resolve the node addresses again more often

    Incorrect

    Each node already gets about half of the traffic. The imbalance comes from each node routing only within its own zone.

When cross-zone load balancing is off, each load balancer node routes only to targets in its own Availability Zone, so 2 targets share half the traffic while 8 share the other half. Application Load Balancers always have it on at the load balancer level; Network Load Balancers do not by default.

Question 4 · choose 1

A regulator requires that every new object written to an S3 bucket in eu-west-1 be copied to a bucket in eu-central-1. The company must be able to show a predictable replication time, with most objects copied in seconds and 99.9% within 15 minutes, and the operations team must be notified about any object that takes longer than that. Which solution meets these requirements with the least custom code?

  1. AAn S3 Event Notification on each PUT that invokes a Lambda function to copy the object and to publish an alert if the copy takes longer than 15 minutes
  2. BA Cross-Region Replication rule with S3 Replication Time Control, and an event notification for OperationMissedThreshold events
  3. CA Cross-Region Replication rule without Replication Time Control, and a CloudWatch alarm on the bucket's BucketSizeBytes metric in eu-central-1
  4. DAn S3 Batch Replication job that an EventBridge Scheduler schedule starts every 15 minutes for objects that have not been replicated
Show the answer and why
  • AAn S3 Event Notification on each PUT that invokes a Lambda function to copy the object and to publish an alert if the copy takes longer than 15 minutes

    Incorrect

    This is custom code that the team must run and monitor, and it gives no AWS-defined replication time.

  • BA Cross-Region Replication rule with S3 Replication Time Control, and an event notification for OperationMissedThreshold events

    Correct

    S3 RTC copies most objects in seconds and 99.9% within 15 minutes, and it emits an OperationMissedThreshold event for objects that miss the threshold.

  • CA Cross-Region Replication rule without Replication Time Control, and a CloudWatch alarm on the bucket's BucketSizeBytes metric in eu-central-1

    Incorrect

    Without RTC there is no defined replication time, and a storage size metric cannot tell which objects are late.

  • DAn S3 Batch Replication job that an EventBridge Scheduler schedule starts every 15 minutes for objects that have not been replicated

    Incorrect

    Batch Replication is for existing objects. Running it on a schedule delays new copies and still needs custom tracking to prove timeliness.

S3 Replication Time Control adds a replication time target and visibility to replication: S3 Replication metrics for pending and late objects, and event notifications when an object misses or later meets the 15-minute threshold.

Question 5 · choose 1

Behind an Application Load Balancer, one or two instances at a time start returning a high share of HTTP 500 errors because of a slow dependency on their host, yet they still pass the target group health checks. The team wants the load balancer to send less traffic to such targets automatically until they recover. What should the DevOps engineer configure on the target group?

  1. AThe least outstanding requests algorithm with slow start mode turned on
  2. BA health check with a lower unhealthy threshold and a shorter interval
  3. CThe weighted random algorithm with anomaly mitigation turned on for the group
  4. DSticky sessions so that each client keeps using the same target over time
Show the answer and why
  • AThe least outstanding requests algorithm with slow start mode turned on

    Incorrect

    Slow start ramps up new targets; it does not react to targets that return errors.

  • BA health check with a lower unhealthy threshold and a shorter interval

    Incorrect

    The targets pass health checks, so stricter checks may still not catch them and would remove them fully.

  • CThe weighted random algorithm with anomaly mitigation turned on for the group

    Correct

    Anomaly mitigation routes traffic away from targets whose error rates deviate from others, and it needs weighted random routing.

  • DSticky sessions so that each client keeps using the same target over time

    Incorrect

    Stickiness keeps clients on a target even when it returns errors.

Automatic Target Weights detects targets with significantly higher error rates than others in the target group. With weighted random routing, mitigation reduces their traffic until they recover.

Question 6 · choose 1

An application runs actively in us-east-1, eu-west-1 and ap-southeast-1, and all three Regions must serve users at all times. Each user should reach the Region that gives the lowest latency, and if the stack in one Region fails its health checks, its users must move to the next best Region automatically, without anyone changing DNS. What should the DevOps engineer configure in Route 53?

  1. AWeighted records with equal weights for the three Regions, each associated with a health check
  2. BLatency records for the three Regions, each associated with a health check
  3. CLatency records for the three Regions without health checks, with a low TTL on each record
  4. DFailover records with us-east-1 as primary and the other two Regions as secondary
Show the answer and why
  • AWeighted records with equal weights for the three Regions, each associated with a health check

    Incorrect

    Health checks remove a failed Region, but equal weights spread users without regard to their latency to each Region.

  • BLatency records for the three Regions, each associated with a health check

    Correct

    Latency routing picks the Region with the lowest latency, and when that record is unhealthy Route 53 chooses another healthy record by the same criteria.

  • CLatency records for the three Regions without health checks, with a low TTL on each record

    Incorrect

    Users get the lowest-latency Region, but records without a health check are always considered healthy, so users keep going to a failed Region.

  • DFailover records with us-east-1 as primary and the other two Regions as secondary

    Incorrect

    Failover is active-passive: it sends all users to one Region at a time and leaves the others idle.

Latency-based routing picks the Region with the lowest latency for each user. Health checks on every record let Route 53 skip an unhealthy Region and answer with the next healthy one.

Question 7 · choose 1

An application uses an Amazon MQ for ActiveMQ single-instance broker. When its Availability Zone had problems, messaging stopped until the zone recovered. The team wants the broker to keep working through the loss of one zone, with durable storage across zones and no application redesign. What should the DevOps engineer do?

  1. AMove the broker to a larger instance type in the same Availability Zone
  2. BReplace the broker with an SQS queue and change the application to use the SQS API
  3. CTake daily backups of the broker configuration so that a new broker can be created after a failure
  4. DUse an active/standby broker deployment with two brokers in different Availability Zones
Show the answer and why
  • AMove the broker to a larger instance type in the same Availability Zone

    Incorrect

    A larger broker in one zone still fails with that zone.

  • BReplace the broker with an SQS queue and change the application to use the SQS API

    Incorrect

    This is an application redesign, which the team wants to avoid.

  • CTake daily backups of the broker configuration so that a new broker can be created after a failure

    Incorrect

    Rebuilding a broker after a failure still stops messaging for a while.

  • DUse an active/standby broker deployment with two brokers in different Availability Zones

    Correct

    Active/standby brokers run a redundant pair across two zones with storage that is redundant across zones.

A single-instance broker runs in one zone. An active/standby deployment places two brokers in different zones as a redundant pair, with one active and one on standby.

Question 8 · choose 1

A shared file system for a web fleet in three Availability Zones was created as an EFS One Zone file system to save cost. Instances in the other two zones lost access to their content during a zonal outage. The content must stay available even if a zone is unavailable. What should the DevOps engineer do?

  1. AMove the content to an EFS Regional file system with mount targets in all three zones
  2. BKeep One Zone and turn on its automatic backups so that content can be restored after an outage
  3. CAdd mount targets for the One Zone file system in the other two zones
  4. DCopy the content to an EBS volume on each instance every night
Show the answer and why
  • AMove the content to an EFS Regional file system with mount targets in all three zones

    Correct

    Regional file systems store data redundantly across multiple zones and stay available when a zone is unavailable.

  • BKeep One Zone and turn on its automatic backups so that content can be restored after an outage

    Incorrect

    Backups help recovery but do not keep the content available during the outage.

  • CAdd mount targets for the One Zone file system in the other two zones

    Incorrect

    The data would still be stored in one zone.

  • DCopy the content to an EBS volume on each instance every night

    Incorrect

    Nightly copies on each instance are stale and add manual work.

EFS offers Regional and One Zone file systems. Regional, the recommended type, keeps data available across zones; One Zone data lives in a single zone.

Question 9 · choose 1

A training cluster reads curated datasets from S3 Express One Zone in its own Availability Zone for the lowest latency. The curated datasets exist only there and take weeks to rebuild. The business requires that losing an Availability Zone must not lose the datasets, while training keeps its low latency. What should the DevOps engineer recommend?

  1. AKeep the datasets only in S3 Express One Zone, because it stores data on multiple devices
  2. BMove the cluster's reads to S3 Standard and stop using S3 Express One Zone
  3. CKeep the source of truth in a general purpose bucket and stage working data in Express
  4. DCreate a second directory bucket in the same Availability Zone as a backup copy
Show the answer and why
  • AKeep the datasets only in S3 Express One Zone, because it stores data on multiple devices

    Incorrect

    Its devices are within a single zone, so losing the zone can affect the data.

  • BMove the cluster's reads to S3 Standard and stop using S3 Express One Zone

    Incorrect

    This gives up the low latency that training needs.

  • CKeep the source of truth in a general purpose bucket and stage working data in Express

    Correct

    General purpose buckets store data across zones, while the directory bucket keeps serving low-latency reads.

  • DCreate a second directory bucket in the same Availability Zone as a backup copy

    Incorrect

    Both copies would be in the same zone.

S3 Express One Zone stores data in directory buckets within one zone for high performance. Data that must survive a zone loss belongs in a storage class that spans zones.

Practise domain 3 →Practise all domains →