Skip to content
BytePatterns

DOP-C02 · Domain 3: Resilient Cloud Solutions · 15% of the exam

Task 3.2: Implement solutions that are scalable to meet business requirements.

Scaling on the right signal: Auto Scaling policies and metrics, queues that decouple tiers, caching, containers on ECS and EKS, serverless with API Gateway, Lambda and Fargate, and serving users from several Regions.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

Workers on Amazon EC2 instances in an Auto Scaling group process video jobs from an Amazon SQS queue. Each job takes about 2 minutes, and a job must not wait in the queue for more than about 10 minutes. The group scales on average CPU utilization, but CPU stays near 60% whether 50 or 5,000 jobs are waiting, so backlogs build up during peaks. Which scaling approach should the DevOps engineer use?

  1. AA target tracking policy on ApproximateNumberOfMessagesVisible for the queue with a target value of 100 messages
  2. BA scheduled scaling action that adds instances during the hours when peaks have usually happened over the past month
  3. CA target tracking policy on average CPU utilization with a lower target value of 30% so that the group scales out earlier
  4. DA target tracking policy that uses metric math for backlog per instance, with a target of about 5 messages per instance
Show the answer and why
  • AA target tracking policy on ApproximateNumberOfMessagesVisible for the queue with a target value of 100 messages

    Incorrect

    The number of messages does not change in proportion to the size of the group, so a fixed queue length target does not track how much work each instance has.

  • BA scheduled scaling action that adds instances during the hours when peaks have usually happened over the past month

    Incorrect

    A schedule reacts to past patterns, not to the actual backlog, so unexpected peaks still build up.

  • CA target tracking policy on average CPU utilization with a lower target value of 30% so that the group scales out earlier

    Incorrect

    CPU stays flat whatever the backlog is, so a lower CPU target does not relate the group's size to the waiting jobs.

  • DA target tracking policy that uses metric math for backlog per instance, with a target of about 5 messages per instance

    Correct

    Backlog per instance divides the queue length by the InService instances. A target equal to the acceptable latency divided by the time per message, here 10 / 2 = 5, keeps the wait within limits.

For queue workers, the right scaling metric is backlog per instance: queue length divided by running capacity, compared with the backlog an instance can clear within the acceptable latency. Metric math lets a target tracking policy compute it without publishing a custom metric.

Question 2 · choose 1

An Amazon ECS cluster runs services on Amazon EC2 container instances from an Auto Scaling group. Service auto scaling adds tasks during traffic peaks, but many new tasks stay in the PROVISIONING or PENDING state because no container instance has room for them, while the group's CPU-based scaling policy reacts slowly. The team wants the number of instances to follow what the tasks need, with no custom code. What should the DevOps engineer do?

  1. ACreate an Auto Scaling group capacity provider for the group with managed scaling turned on and use it as the cluster's capacity provider
  2. BLower the target of the service auto scaling policy so that services start fewer tasks during peaks and the cluster is not overloaded
  3. CAdd a step scaling policy to the Auto Scaling group that adds instances when the cluster's memory reservation stays above 80% for 10 minutes
  4. DWrite a Lambda function that counts pending tasks every minute and sets the group's desired capacity to match them
Show the answer and why
  • ACreate an Auto Scaling group capacity provider for the group with managed scaling turned on and use it as the cluster's capacity provider

    Correct

    With managed scaling, ECS creates the metrics, alarms and target tracking policy for the group and scales instances based on what the tasks use, reaching the targetCapacity you set.

  • BLower the target of the service auto scaling policy so that services start fewer tasks during peaks and the cluster is not overloaded

    Incorrect

    This throttles the application instead of adding capacity, and service auto scaling changes the number of tasks, not the number of instances.

  • CAdd a step scaling policy to the Auto Scaling group that adds instances when the cluster's memory reservation stays above 80% for 10 minutes

    Incorrect

    A hand-tuned step policy on one reservation metric still reacts late and must be maintained by the team, while ECS can manage the group's scaling itself.

  • DWrite a Lambda function that counts pending tasks every minute and sets the group's desired capacity to match them

    Incorrect

    This is custom code, which the team wants to avoid, and ECS cluster auto scaling already does this work.

ECS cluster auto scaling ties the Auto Scaling group to the tasks: once managed scaling is on for the capacity provider, ECS manages scale-out and scale-in so that tasks waiting for capacity get instances.

Question 3 · choose 1

A Node.js AWS Lambda function behind Amazon API Gateway serves a mobile app. Every weekday at 08:00 traffic jumps from a few requests to about 400 concurrent requests, and users see slow first responses caused by cold starts. Latency must stay low during the morning peak, and the team does not want to pay for idle capacity overnight. What should the DevOps engineer do?

  1. ASet reserved concurrency of 400 on the function so that this capacity is always kept for it and no other function can use it
  2. BTurn on Lambda SnapStart for the function's published versions and point the API at an alias of the latest version
  3. CConfigure provisioned concurrency on an alias and schedule it with Application Auto Scaling around the morning peak
  4. DIncrease the function's memory to the maximum so that each cold start finishes faster
Show the answer and why
  • ASet reserved concurrency of 400 on the function so that this capacity is always kept for it and no other function can use it

    Incorrect

    Reserved concurrency sets the maximum and minimum concurrency for the function but does not pre-initialize execution environments, so cold starts remain.

  • BTurn on Lambda SnapStart for the function's published versions and point the API at an alias of the latest version

    Incorrect

    SnapStart does not support Node.js managed runtimes, so it cannot be turned on for this function.

  • CConfigure provisioned concurrency on an alias and schedule it with Application Auto Scaling around the morning peak

    Correct

    Provisioned concurrency keeps pre-initialized execution environments ready to respond immediately, and Application Auto Scaling can manage the amount on a schedule.

  • DIncrease the function's memory to the maximum so that each cold start finishes faster

    Incorrect

    Memory sets the resources of each environment, but environments are still created on demand when the peak starts, so users still wait for cold starts.

Cold starts at a predictable peak call for pre-initialized environments exactly when they are needed. Provisioned concurrency on a version or alias, scheduled with Application Auto Scaling, does that and can be lowered again after the peak.

Question 4 · choose 1

An Aurora MySQL cluster serves a read-heavy API through its reader endpoint. On some days, with no fixed schedule, read traffic triples for a few hours. The number of connections stays steady, but the two Aurora Replicas then run at 95 percent CPU. The team wants read capacity added and removed automatically and does not want to pay for idle replicas between peaks. What should the DevOps engineer configure?

  1. AAurora Auto Scaling with a target tracking policy on the average CPU utilization of the readers
  2. BScheduled actions in Application Auto Scaling that add replicas every weekday morning and remove them in the evening
  3. CAn RDS Proxy read-only endpoint in front of the reader endpoint
  4. DEight Aurora Replicas running at all times so that the largest peak always fits
Show the answer and why
  • AAurora Auto Scaling with a target tracking policy on the average CPU utilization of the readers

    Correct

    Aurora Auto Scaling adds Aurora Replicas when the reader metric rises and removes the ones it created when load drops, and the reader endpoint spreads connections to them.

  • BScheduled actions in Application Auto Scaling that add replicas every weekday morning and remove them in the evening

    Incorrect

    Scheduled scaling suits peaks with a known time, but these peaks have no fixed schedule.

  • CAn RDS Proxy read-only endpoint in front of the reader endpoint

    Incorrect

    The problem is reader CPU while the number of connections stays steady, and a proxy adds no read capacity.

  • DEight Aurora Replicas running at all times so that the largest peak always fits

    Incorrect

    This covers the peaks without any scaling, but the extra replicas cost money between peaks.

Aurora Auto Scaling adjusts the number of Aurora Replicas with a scaling policy that tracks a CloudWatch metric, such as the average CPU of the readers, and removes only replicas it created. The reader endpoint lets applications use new replicas.

Question 5 · choose 1

A REST API in API Gateway is used by 300 partner companies, which authenticate through a Lambda authorizer. One partner's batch jobs sometimes send so many requests that other partners are throttled. Each partner must get its own request rate and its own monthly request quota, premium partners must get higher limits than standard partners, and the API team wants to manage this as a product offering in API Gateway. What should the DevOps engineer configure?

  1. AA stage-level throttling limit that is high enough for the combined traffic of every partner
  2. BAWS WAF rate-based rules that aggregate requests by each partner's identifying header
  3. CUsage plans for each tier with throttling and monthly quotas, and an API key for each partner in its tier's plan
  4. DAPI caching on the stage so that repeated requests from the batch jobs are answered from the cache
Show the answer and why
  • AA stage-level throttling limit that is high enough for the combined traffic of every partner

    Incorrect

    A stage limit protects the backend as a whole, but one shared limit still lets a single partner use most of it.

  • BAWS WAF rate-based rules that aggregate requests by each partner's identifying header

    Incorrect

    Custom aggregation keys can limit each partner separately, but WAF evaluates windows of at most 10 minutes and offers no monthly quota.

  • CUsage plans for each tier with throttling and monthly quotas, and an API key for each partner in its tier's plan

    Correct

    A usage plan sets throttling and quota limits that apply to each API key, and keys identify clients, while the authorizer still controls access.

  • DAPI caching on the stage so that repeated requests from the batch jobs are answered from the cache

    Incorrect

    Caching helps repeated reads but does not limit any partner's usage.

Usage plans make an API available as a product offering, with throttling and quota limits per API key. API keys identify clients but should not be used for authorization; an authorizer, IAM or Amazon Cognito controls access.

Question 6 · choose 1

A Kinesis data stream in provisioned mode receives clickstream data. During marketing campaigns, write volume rises sharply at unpredictable times, to at most about 1.5 times the highest peak of the previous month, and is low the rest of the time. Producers are throttled during spikes, and engineers spend hours resharding. The team wants no capacity planning, no scaling code to maintain, and no payment for idle shards between campaigns. What should the DevOps engineer do?

  1. ASwitch the stream to on-demand capacity mode and let Kinesis Data Streams manage the shards
  2. BKeep provisioned mode and permanently double the number of shards so that every spike fits
  3. CHave a Lambda function call UpdateShardCount when CloudWatch alarms report throttled writes on the stream
  4. DRegister the consumers for enhanced fan-out so that producers are no longer throttled during campaigns
Show the answer and why
  • ASwitch the stream to on-demand capacity mode and let Kinesis Data Streams manage the shards

    Correct

    On-demand streams need no capacity planning, are billed per GB written and read, and accommodate up to double the peak write throughput of the previous 30 days.

  • BKeep provisioned mode and permanently double the number of shards so that every spike fits

    Incorrect

    More shards absorb the spikes without code, but the size still has to be planned and the extra shards are paid for between campaigns.

  • CHave a Lambda function call UpdateShardCount when CloudWatch alarms report throttled writes on the stream

    Incorrect

    This automates provisioned scaling, but the team maintains the code, and each call can at most double the shard count, ten times a day.

  • DRegister the consumers for enhanced fan-out so that producers are no longer throttled during campaigns

    Incorrect

    Enhanced fan-out changes how consumers read from the stream; it does not add write capacity for producers.

On-demand mode manages shards for the needed throughput and suits unpredictable traffic. Writes are throttled only if traffic grows beyond double the previous peak within 15 minutes, which these campaigns do not reach.

Question 7 · choose 1

An ECS service on Fargate processes interruption-tolerant image jobs and scales from 4 to 200 tasks. The team wants to cut cost by running most of the extra tasks on Fargate Spot, while always keeping at least four tasks on regular Fargate. What should the DevOps engineer configure?

  1. AA capacity provider strategy with FARGATE base 4 and weight 1, and FARGATE_SPOT weight 3
  2. BA capacity provider strategy with only FARGATE_SPOT and base 4
  3. CTwo separate services, one on Fargate and one on Fargate Spot, scaled by different alarms
  4. DThe FARGATE launch type with a scheduled action that moves tasks to Spot at night
Show the answer and why
  • AA capacity provider strategy with FARGATE base 4 and weight 1, and FARGATE_SPOT weight 3

    Correct

    The base places the first tasks on FARGATE, and the weights split the remaining tasks between the providers.

  • BA capacity provider strategy with only FARGATE_SPOT and base 4

    Incorrect

    All tasks, including the four that must stay on regular Fargate, would run on Spot.

  • CTwo separate services, one on Fargate and one on Fargate Spot, scaled by different alarms

    Incorrect

    This splits one workload into two services instead of using a strategy.

  • DThe FARGATE launch type with a scheduled action that moves tasks to Spot at night

    Incorrect

    Launch types do not move tasks to Spot; capacity providers place them.

A capacity provider strategy uses base to set a minimum number of tasks on a provider and weight to share the rest. Fargate Spot tasks get a two-minute warning before they are interrupted.

Question 8 · choose 1

A new service needs a Valkey-compatible cache that is highly available across Availability Zones. Its load at launch is unknown and may grow quickly in bursts. The team has no data to choose node types, shard counts or minimum and maximum capacity, and it does not want to patch or replace cache nodes itself. What should the DevOps engineer use?

  1. AA single-node ElastiCache cluster sized for the expected first-month load
  2. BA self-managed Valkey cluster on EC2 instances in an Auto Scaling group across three Availability Zones
  3. CA node-based ElastiCache cluster with cluster mode on and ElastiCache auto scaling of shards and replicas
  4. DAn ElastiCache Serverless cache for Valkey, created with its default settings
Show the answer and why
  • AA single-node ElastiCache cluster sized for the expected first-month load

    Incorrect

    ElastiCache would manage the node, but one node is not highly available and its size must still be chosen.

  • BA self-managed Valkey cluster on EC2 instances in an Auto Scaling group across three Availability Zones

    Incorrect

    The group adds instances for load and spans zones, but the team would run, patch and replace the Valkey nodes itself.

  • CA node-based ElastiCache cluster with cluster mode on and ElastiCache auto scaling of shards and replicas

    Incorrect

    Auto scaling adjusts shards and replicas, but the team still chooses a supported node type and the minimum and maximum capacity.

  • DAn ElastiCache Serverless cache for Valkey, created with its default settings

    Correct

    ElastiCache Serverless creates a highly available cache without choosing nodes, scales with the application's use, and handles patching and node replacement.

ElastiCache runs as serverless caches or node-based clusters. Serverless monitors memory, compute and network use and scales to meet them, while node-based clusters, even with auto scaling, need node type and capacity choices.

Practise domain 3 →Practise all domains →