Skip to content
BytePatterns

SAP-C02 · Domain 2: Design for New Solutions · 29% of the exam

Task 2.2: Design a solution to ensure business continuity

Continuity built into a new design: cross-Region replication of data and databases, automated backups across zones and Regions, DNS failover, disaster recovery tests, and monitoring that triggers recovery before users notice.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A gaming company is building a player profile service on Amazon DynamoDB in us-east-1 and eu-west-1. Players in each Region must write with single-digit millisecond latency to their nearest Region, and if one Region fails, the other must already have the data and keep accepting writes. The company accepts that the last write wins when the same item is changed in both Regions at once. Which solution meets these requirements?

  1. ACreate a DynamoDB global table with replicas in both Regions in multi-Region eventual consistency mode
  2. BCopy backups of the table to the other Region every hour with AWS Backup and restore them if a Region fails
  3. CWrite to a table in us-east-1 only and add DynamoDB Accelerator clusters in eu-west-1 for European players
  4. DMigrate the profiles to an Aurora global database with its primary cluster in us-east-1
Show the answer and why
  • ACreate a DynamoDB global table with replicas in both Regions in multi-Region eventual consistency mode

    Correct

    Global tables replicate writes between Regions in which every replica accepts reads and writes. In multi-Region eventual consistency mode, conflicting concurrent writes are resolved with last writer wins.

  • BCopy backups of the table to the other Region every hour with AWS Backup and restore them if a Region fails

    Incorrect

    Restored backups lose the changes made since the last backup and take time to restore, so the other Region would not already have the data.

  • CWrite to a table in us-east-1 only and add DynamoDB Accelerator clusters in eu-west-1 for European players

    Incorrect

    DAX is an in-memory cache for a table in its own Region. European writes would still travel to us-east-1, and nothing would take writes if that Region failed.

  • DMigrate the profiles to an Aurora global database with its primary cluster in us-east-1

    Incorrect

    An Aurora global database accepts writes in its primary Region. Players near the secondary Region would not get local writes.

Local writes in two Regions with last-writer-wins conflict handling is the default mode of DynamoDB global tables.

Question 2 · choose 1

A media company stores new video assets in an S3 bucket in us-west-2. For continuity, a copy of every new object must exist in an S3 bucket in us-east-2, and the company's contract requires a predictable replication time backed by an SLA, with alerts when objects take longer. Which solution meets these requirements?

  1. ASchedule an AWS DataSync task that copies new objects from the source bucket to the destination bucket every 15 minutes
  2. BRun S3 Batch Replication every night to copy the objects that were created since the previous run
  3. CConfigure S3 Cross-Region Replication with S3 Replication Time Control and replication metrics and events
  4. DConfigure S3 Same-Region Replication to a second bucket and copy that bucket to us-east-2 with a scheduled job
Show the answer and why
  • ASchedule an AWS DataSync task that copies new objects from the source bucket to the destination bucket every 15 minutes

    Incorrect

    DataSync can move data between buckets on a schedule, but it offers no replication-time SLA for each new object.

  • BRun S3 Batch Replication every night to copy the objects that were created since the previous run

    Incorrect

    Batch Replication copies existing objects on demand. A nightly run is not continuous replication of new objects with a time target.

  • CConfigure S3 Cross-Region Replication with S3 Replication Time Control and replication metrics and events

    Correct

    S3 Replication Time Control replicates most new objects in seconds, is backed by a service level agreement for replication within 15 minutes, and provides metrics and event notifications for objects that miss the threshold.

  • DConfigure S3 Same-Region Replication to a second bucket and copy that bucket to us-east-2 with a scheduled job

    Incorrect

    Same-Region Replication keeps the copy in us-west-2. The extra scheduled copy adds delay and has no SLA.

"Predictable time with an SLA" is the wording for S3 Replication Time Control on top of Cross-Region Replication.

Question 3 · choose 2

A company protects 30 on-premises servers with AWS Elastic Disaster Recovery and backs up its Amazon RDS databases with AWS Backup. Auditors want proof every quarter that both can be recovered, and the tests must not interrupt production or stop data replication. Which actions meet these requirements? (Choose TWO.)

  1. AFail over the production Aurora cluster to its secondary Region every quarter and fail it back afterward
  2. BUse CloudFormation drift detection on the recovery stacks every quarter as evidence that recovery works
  3. CDetach the replication agents from the source servers during the quarterly test so that the drill uses a fixed copy
  4. DLaunch drill instances in Elastic Disaster Recovery, check the application, and terminate them afterward
  5. ECreate AWS Backup restore testing plans that restore recovery points on a schedule and validate the results
Show the answer and why
  • AFail over the production Aurora cluster to its secondary Region every quarter and fail it back afterward

    Incorrect

    A failover moves production to the other Region and is meant for real outages, so it interrupts production rather than testing beside it.

  • BUse CloudFormation drift detection on the recovery stacks every quarter as evidence that recovery works

    Incorrect

    Drift detection compares deployed resources with their templates. It does not restore or launch anything, so it proves nothing about recovery.

  • CDetach the replication agents from the source servers during the quarterly test so that the drill uses a fixed copy

    Incorrect

    Stopping replication is exactly what the auditors ruled out, and it would leave the servers unprotected during the test.

  • DLaunch drill instances in Elastic Disaster Recovery, check the application, and terminate them afterward

    Correct

    Elastic Disaster Recovery supports non-disruptive recovery drills. Performing a drill does not stop replication and does not affect the source servers.

  • ECreate AWS Backup restore testing plans that restore recovery points on a schedule and validate the results

    Correct

    Restore testing in AWS Backup runs restore jobs from chosen recovery points on a schedule and monitors them, and its results can be used to show auditors that restores succeed.

Both services have a built-in way to prove recovery without touching production: drills in Elastic Disaster Recovery and restore testing in AWS Backup.

Question 4 · choose 1

An internal API runs in a primary and a standby Region behind internal Application Load Balancers in private subnets. Clients resolve api.corp.internal from a Route 53 private hosted zone. When the primary Region's error-rate CloudWatch alarm goes off, DNS must send clients to the standby Region automatically. What should a solutions architect configure?

  1. AFailover records whose primary has a Route 53 health check that probes the private IP addresses of the primary load balancer
  2. BLoad balancer health checks on the primary targets, with the standby load balancer registered as an extra target
  3. CMultivalue answer records for both load balancers so that clients retry the other address when one fails
  4. DFailover records whose primary is associated with a Route 53 health check that monitors the state of the CloudWatch alarm
Show the answer and why
  • AFailover records whose primary has a Route 53 health check that probes the private IP addresses of the primary load balancer

    Incorrect

    Route 53 health checkers run outside your VPCs and cannot reach private IP addresses, so a health check on a private endpoint cannot report its state.

  • BLoad balancer health checks on the primary targets, with the standby load balancer registered as an extra target

    Incorrect

    Load balancer health checks decide which targets of that load balancer receive traffic. They do not change DNS answers to another Region.

  • CMultivalue answer records for both load balancers so that clients retry the other address when one fails

    Incorrect

    Without health checks, multivalue answers keep returning the failed Region, and nothing reacts to the CloudWatch alarm.

  • DFailover records whose primary is associated with a Route 53 health check that monitors the state of the CloudWatch alarm

    Correct

    For resources Route 53 cannot reach, such as private endpoints, a health check can be based on a CloudWatch alarm, and failover routing then answers with the secondary record while that check is unhealthy.

Private endpoints cannot be probed by Route 53 health checkers, so the health check watches a CloudWatch alarm instead and drives failover routing.

Question 5 · choose 1

A new product catalog runs on an Amazon DocumentDB instance-based cluster in us-east-1 and is accessed with MongoDB drivers. Shoppers in Europe need low-latency reads served from eu-west-1. If us-east-1 fails, eu-west-1 must take over writes within minutes and lose no more than a few seconds of data. The team will not build or monitor its own replication pipelines and will change nothing in the application except connection strings. Which solution meets these requirements?

  1. ARun an AWS DMS task with ongoing replication from the cluster to a separate DocumentDB cluster in eu-west-1
  2. BCopy cluster snapshots to eu-west-1 every hour and restore the latest one in eu-west-1 if us-east-1 fails
  3. CAdd eu-west-1 as a secondary Region of a DocumentDB global cluster and promote it if us-east-1 fails
  4. DAdd replica instances in two more Availability Zones of us-east-1 and send the European reads to the cluster's reader endpoint
Show the answer and why
  • ARun an AWS DMS task with ongoing replication from the cluster to a separate DocumentDB cluster in eu-west-1

    Incorrect

    DMS supports DocumentDB as a source with change data capture, so it can keep a readable copy in eu-west-1. But the team would build, size and monitor the replication tasks itself, which it refuses to do.

  • BCopy cluster snapshots to eu-west-1 every hour and restore the latest one in eu-west-1 if us-east-1 fails

    Incorrect

    DocumentDB can copy snapshots to another Region, but an hourly copy can lose up to an hour of writes, a new cluster must be restored after the failure, and nothing serves European reads in the meantime.

  • CAdd eu-west-1 as a secondary Region of a DocumentDB global cluster and promote it if us-east-1 fails

    Correct

    A global cluster replicates from the primary Region to read-only secondary Regions on dedicated infrastructure with latency typically under a second. A secondary cluster serves local reads and can be promoted to primary within minutes, with an RPO typically measured in seconds.

  • DAdd replica instances in two more Availability Zones of us-east-1 and send the European reads to the cluster's reader endpoint

    Incorrect

    Replicas in more zones raise read capacity and availability within us-east-1, but every read still crosses the Atlantic, and a Regional failure takes all of them down together.

The constraints are local reads in Europe, Regional takeover within minutes with seconds of data loss, and no replication pipeline to run. Snapshot copies lose too much data and recover too slowly, more replicas stay in one Region, and DMS works but is a pipeline the team would operate. A DocumentDB global cluster gives a read-only secondary in eu-west-1 that can be promoted.

Question 6 · choose 1

A new application will run active-passive in us-east-1 and us-west-2 on an Aurora PostgreSQL global database. Its database credentials are a Secrets Manager secret in us-east-1 that rotates every 30 days. If us-east-1 is down, the application in us-west-2 must read the current credentials from its own Region, the credentials must stay identical in both Regions after every rotation, and the team will not write or run code that copies secrets. Which solution meets these requirements?

  1. AKeep a SecureString parameter in us-west-2 and update it from a Lambda function that runs after each rotation in us-east-1
  2. BReplicate the secret to us-west-2 and keep the 30-day rotation configured on the primary secret in us-east-1
  3. CCreate a second secret in us-west-2 for the same database user with its own 30-day rotation schedule
  4. DLet the us-west-2 application read the secret from the us-east-1 endpoint through an IAM policy that allows the cross-Region call
Show the answer and why
  • AKeep a SecureString parameter in us-west-2 and update it from a Lambda function that runs after each rotation in us-east-1

    Incorrect

    A parameter in us-west-2 would be readable during the outage, but the Lambda function is copy code the team refuses to run, and AWS recommends Secrets Manager rather than Parameter Store for database credentials.

  • BReplicate the secret to us-west-2 and keep the 30-day rotation configured on the primary secret in us-east-1

    Correct

    Replica secrets hold the same secret data in other Regions. Secrets Manager rotates the primary secret and propagates the new value to every replica, so rotation is not managed per Region, and a replica can be promoted to a standalone secret if needed.

  • CCreate a second secret in us-west-2 for the same database user with its own 30-day rotation schedule

    Incorrect

    A local secret is readable during the outage, but two independent rotations would each set their own password for the same database user, so the values in the two Regions would not stay identical.

  • DLet the us-west-2 application read the secret from the us-east-1 endpoint through an IAM policy that allows the cross-Region call

    Incorrect

    A cross-Region read works only while us-east-1 is reachable. During the outage the application could not fetch the credentials at all.

Three constraints decide it: a local copy during the outage, identical values after rotation, and no copy code. A cross-Region read fails in the outage, a separately rotated secret drifts, and a parameter kept in sync by a function is the code the team will not run. A replica secret receives every rotated value from the primary automatically.

Question 7 · choose 1

A new IoT platform stores device state in an Amazon Keyspaces table in us-east-1. The platform must accept writes in both us-east-1 and eu-west-1, have every change from either Region reach the other within seconds, and keep reading and writing in either Region alone if the other fails, without a promotion or failover step in the database. Last-writer-wins handling of concurrent updates is acceptable, and the team will not add replication or conflict logic to its services. Which solution meets these requirements?

  1. AKeep the keyspace in us-east-1, which already stores its data in three Availability Zones, and add capacity for the European traffic
  2. BCreate a second table in eu-west-1 and have each service write every change to both Regions' tables through a shared library
  3. CTurn on CDC streams for the table and have a consumer apply each change event to a table in eu-west-1
  4. DAdd eu-west-1 to the keyspace so that it becomes a multi-Region keyspace with Keyspaces multi-Region replication
Show the answer and why
  • AKeep the keyspace in us-east-1, which already stores its data in three Availability Zones, and add capacity for the European traffic

    Incorrect

    By default, Keyspaces replicates data across three Availability Zones in one Region, which protects against a zone failure. Nothing would accept writes in eu-west-1 or keep working if us-east-1 failed.

  • BCreate a second table in eu-west-1 and have each service write every change to both Regions' tables through a shared library

    Incorrect

    Dual writes would put the data in both Regions, but the services would own the replication, retries and the reconciliation of concurrent updates, which is the conflict logic the team refuses to add.

  • CTurn on CDC streams for the table and have a consumer apply each change event to a table in eu-west-1

    Incorrect

    Keyspaces CDC records time-ordered, row-level change events that a consumer can apply elsewhere, but only in one direction. Writes made in eu-west-1 would never reach us-east-1 unless the team built a reverse pipeline and its own handling of conflicting updates.

  • DAdd eu-west-1 to the keyspace so that it becomes a multi-Region keyspace with Keyspaces multi-Region replication

    Correct

    A Region can be added to a single-Region keyspace once client-side timestamps are on. Replication is then fully managed and active-active: each Region reads and writes in isolation, changes typically arrive in under a second, and concurrent updates are reconciled by last writer wins without application involvement.

The constraints are writes in two Regions, changes flowing both ways within seconds, independent survival of each Region, and no replication or conflict logic in the services. Three-zone storage covers only one Region, dual writes push replication into the services, and a change stream is a one-way feed. A multi-Region keyspace replicates in both directions with managed last-writer-wins reconciliation and no promotion step.

Question 8 · choose 1

A CI pipeline in a tools account pushes container images to Amazon ECR in us-east-1. Production runs on Amazon ECS in a separate production account in us-east-1 and fails over to us-west-2 in the same account. During an outage of us-east-1, ECS in us-west-2 must be able to pull every production image, including images it has never pulled before. Only repositories whose names start with prod/ may leave the tools account, and the team will not change the pipeline or write code that copies images. Which solution meets these requirements?

  1. AAdd a step to the pipeline that also pushes each production image to a registry in us-west-2 in the production account
  2. BConfigure ECR replication of the prod/ prefix from the tools account to the production account's registry in us-west-2
  3. CCreate an ECR pull through cache rule in us-west-2 in the production account with the tools account registry as its upstream
  4. DTurn on tag immutability for the prod/ repositories so that the image tags that ECS uses can never be overwritten
Show the answer and why
  • AAdd a step to the pipeline that also pushes each production image to a registry in us-west-2 in the production account

    Incorrect

    Pushing twice would place every image in us-west-2 before an outage, but it is the pipeline change that the team has ruled out.

  • BConfigure ECR replication of the prod/ prefix from the tools account to the production account's registry in us-west-2

    Correct

    ECR private image replication works across Regions and accounts once the destination registry's permissions policy allows it, and a repository prefix filter limits which repositories are replicated. Each new image is copied when it is pushed, without code.

  • CCreate an ECR pull through cache rule in us-west-2 in the production account with the tools account registry as its upstream

    Incorrect

    Pull through cache supports an ECR upstream, but an image is fetched from the upstream the first time it is pulled and served from the cache only after that. Images never pulled before the outage would be unavailable.

  • DTurn on tag immutability for the prod/ repositories so that the image tags that ECS uses can never be overwritten

    Incorrect

    Tag immutability prevents image tags from being overwritten. It makes deployments more predictable, but every image still exists only in us-east-1.

The constraints are images available in us-west-2 before any pull, only the prod/ repositories crossing accounts, and no pipeline or code changes. A second push changes the pipeline, a pull through cache fills only on demand, and tag immutability copies nothing. ECR replication with a prefix filter to the production account in us-west-2 meets all of them.

Question 9 · choose 2

A new web tier runs in three Availability Zones. The team wants it to keep serving full traffic if one zone fails, without having to launch instances or change configuration during the event. Which design choices support this? (Choose TWO.)

  1. ARun the fleet in a single zone and rely on Auto Scaling to launch elsewhere on failure
  2. BRun all of the capacity in each zone on Spot Instances to save cost
  3. CLaunch the instances in a cluster placement group for lower latency
  4. DPre-provision enough capacity in each zone that the remaining two can carry the full load
  5. EDesign recovery so it does not depend on control plane changes during the event
Show the answer and why
  • ARun the fleet in a single zone and rely on Auto Scaling to launch elsewhere on failure

    Incorrect

    Launching new instances during the failure depends on control plane actions and is the bimodal behavior to avoid.

  • BRun all of the capacity in each zone on Spot Instances to save cost

    Incorrect

    Spot Instances are returned when Amazon EC2 needs the capacity back, so the headroom the design relies on could be interrupted at any time, including during the zone event.

  • CLaunch the instances in a cluster placement group for lower latency

    Incorrect

    A cluster placement group keeps instances within a single Availability Zone for low-latency networking, which is the opposite of spreading capacity across zones.

  • DPre-provision enough capacity in each zone that the remaining two can carry the full load

    Correct

    Having the extra capacity pre-provisioned, instead of launching new instances during a zone failure, helps achieve static stability.

  • EDesign recovery so it does not depend on control plane changes during the event

    Correct

    Removing control plane dependencies from the recovery path helps produce more resilient, statically stable workloads.

Static stability means enough capacity is already in place and nothing has to change during a failure.

Practise domain 2 →Practise all domains →