Skip to content
BytePatterns

DOP-C02 · Domain 5: Incident and Event Response · 14% of the exam

Task 5.1: Manage event sources to process, notify, and take action in response to events.

Wiring events to work: AWS Health, CloudTrail and EventBridge as sources, and fan-out, queuing and streaming workflows with SNS, SQS, Kinesis, Lambda and Step Functions.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company has 80 accounts in AWS Organizations and runs workloads in us-east-1 and eu-west-1. The central operations team wants an Amazon SNS notification whenever AWS Health reports an event, such as scheduled EC2 maintenance, that affects any account in the organization. The team wants to manage this from its own operations account, without creating rules in each member account and without polling. What should the DevOps engineer do?

  1. ADeploy an EventBridge rule for aws.health events to every member account with a StackSet, each rule sending to a topic in the operations account
  2. BTurn on Health organizational view with the operations account as delegated administrator, and create EventBridge rules there in both Regions
  3. CSchedule a Lambda function in the operations account that calls the AWS Health API for the organization every five minutes and publishes new events
  4. DCreate one EventBridge rule for aws.health events in the operations account without turning on organizational view
Show the answer and why
  • ADeploy an EventBridge rule for aws.health events to every member account with a StackSet, each rule sending to a topic in the operations account

    Incorrect

    This works only by creating a rule in every account, which the team wants to avoid. Organizational view provides one feed instead.

  • BTurn on Health organizational view with the operations account as delegated administrator, and create EventBridge rules there in both Regions

    Correct

    With organizational view, the management account or a delegated administrator receives a single feed of Health events from all accounts, and rules are created for each Region to cover.

  • CSchedule a Lambda function in the operations account that calls the AWS Health API for the organization every five minutes and publishes new events

    Incorrect

    This is the polling the team wants to avoid, while EventBridge delivers Health events as they arrive.

  • DCreate one EventBridge rule for aws.health events in the operations account without turning on organizational view

    Incorrect

    Without organizational view, a rule receives Health events only for its own account, not for the other accounts in the organization.

AWS Health events reach EventBridge per account and per Region. Organizational view, best used with a delegated administrator, turns them into one feed for the whole organization, and a rule in each Region of interest routes them.

Question 2 · choose 1

An order service publishes an event for every order. Three teams process the events independently: billing, shipping and analytics. Each team must receive every event, process at its own pace, retry failures without affecting the other teams, and keep events for up to four days if its consumers are down. Which design meets these requirements?

  1. ASend the events to one Amazon SQS queue with a dead-letter queue, and have the consumers of all three teams poll that queue
  2. BHave a Step Functions workflow call the billing, shipping and analytics services one after another for each event
  3. CPublish the events to an Amazon SNS topic that invokes each team's Lambda function directly as a subscriber, with retries for failed calls
  4. DPublish the events to an Amazon SNS topic with one Amazon SQS queue per team subscribed, each with its own dead-letter queue
Show the answer and why
  • ASend the events to one Amazon SQS queue with a dead-letter queue, and have the consumers of all three teams poll that queue

    Incorrect

    Consumers of one queue compete for messages, so each event reaches only one team instead of all three.

  • BHave a Step Functions workflow call the billing, shipping and analytics services one after another for each event

    Incorrect

    A sequential workflow couples the teams: a slow or failing service holds up the others.

  • CPublish the events to an Amazon SNS topic that invokes each team's Lambda function directly as a subscriber, with retries for failed calls

    Incorrect

    Nothing buffers the events for a team whose consumers are down for days; the requirement calls for a durable queue per team.

  • DPublish the events to an Amazon SNS topic with one Amazon SQS queue per team subscribed, each with its own dead-letter queue

    Correct

    SNS fans each message out to every subscribed queue, and each queue lets its team process asynchronously at its own pace with its own retries.

Fan-out to SQS gives every consumer group its own durable copy of each message. SQS can retain messages for up to 14 days, so four days of downtime is covered, and a dead-letter queue per team isolates poison messages.

Question 3 · choose 2

A Lambda function processes clickstream records from an Amazon Kinesis data stream through an event source mapping with default settings. A malformed record makes the function throw an error, and processing on that shard stops for days while other shards continue. The team wants the shard to keep moving, wants as few good records as possible to be retried, and needs the failed records kept for later analysis. The function code cannot change this quarter. Which actions should the DevOps engineer take? (Choose TWO.)

  1. AIncrease the number of shards in the stream so that the malformed record is processed by more function instances
  2. BMap the function to an enhanced fan-out consumer so that it gets dedicated read throughput for each shard
  3. CTurn on BisectBatchOnFunctionError and set MaximumRetryAttempts to a small number on the event source mapping
  4. DAdd an on-failure destination, such as an Amazon SQS queue, to the event source mapping
  5. EAdd a dead-letter queue to the function's asynchronous invocation configuration
Show the answer and why
  • AIncrease the number of shards in the stream so that the malformed record is processed by more function instances

    Incorrect

    Resharding spreads load across more shards, but the batch with the bad record still fails and blocks its shard.

  • BMap the function to an enhanced fan-out consumer so that it gets dedicated read throughput for each shard

    Incorrect

    Enhanced fan-out changes how records are read, not how a failing batch is retried, so the shard stays blocked.

  • CTurn on BisectBatchOnFunctionError and set MaximumRetryAttempts to a small number on the event source mapping

    Correct

    Bisecting splits a failed batch to isolate the bad record without using the retry quota, and a retry limit stops Lambda from retrying the record until it expires.

  • DAdd an on-failure destination, such as an Amazon SQS queue, to the event source mapping

    Correct

    When retries are exhausted, Lambda sends a record about the discarded batch to the destination, so the failure can be analyzed later.

  • EAdd a dead-letter queue to the function's asynchronous invocation configuration

    Incorrect

    For stream event sources, failed batches are kept by adding a destination to the event source mapping, not through the function's settings for asynchronous invocation.

With default settings, Lambda retries a failing stream batch until the records expire, which can block a shard for up to a week. Bisecting, a retry limit or maximum record age, and an on-failure destination keep the stream moving and keep evidence of what was skipped.

Question 4 · choose 1

An EventBridge rule in a central account sends order events to a Lambda function in another team's account. While that account's permissions were being changed, deliveries failed and events were silently lost. The team wants to keep only the events that could not be delivered, together with the error code for each failure, so that it can analyze them and send them again once the cause is fixed. What should the DevOps engineer configure?

  1. AA higher MaximumEventAgeInSeconds and MaximumRetryAttempts in the target's retry policy
  2. BAn archive on the event bus with no event pattern, replayed after the cause is fixed
  3. CA CloudWatch alarm on the rule's FailedInvocations metric that notifies the team
  4. DA dead-letter queue, a standard SQS queue in the rule's Region, for the rule's target
Show the answer and why
  • AA higher MaximumEventAgeInSeconds and MaximumRetryAttempts in the target's retry policy

    Incorrect

    Longer retries help with transient errors, but failures such as missing permissions are not retried, and events are dropped when retries end.

  • BAn archive on the event bus with no event pattern, replayed after the cause is fixed

    Incorrect

    An archive keeps every event for replay, but it does not record which deliveries failed or why.

  • CA CloudWatch alarm on the rule's FailedInvocations metric that notifies the team

    Incorrect

    The alarm reports that deliveries failed, but it does not keep the failed events.

  • DA dead-letter queue, a standard SQS queue in the rule's Region, for the rule's target

    Correct

    EventBridge sends events it cannot deliver to the target's dead-letter queue, with attributes such as the error code, error message and target ARN.

Target dead-letter queues capture events that could not be delivered after retries or because of errors that are not retried. Message attributes show why delivery failed, and the events can be resent after the cause is fixed.

Question 5 · choose 1

An ECS service consumes an SQS standard queue and writes to a database. A bug sent 40,000 messages to the queue's dead-letter queue, and the bug is now fixed. The messages must be processed again by the same service without writing a script to copy them and without changing or redeploying the service. The database can absorb only about 50 extra messages per second on top of normal traffic. What should the DevOps engineer do?

  1. AAdd a redrive allow policy to the dead-letter queue that names the source queue as an allowed queue
  2. BIncrease the maxReceiveCount in the source queue's redrive policy so that messages get more processing attempts
  3. CPoint the service's queue URL at the dead-letter queue until it is empty, then point it back
  4. DStart a dead-letter queue redrive to the source queue with a custom maximum velocity of 50 messages per second
Show the answer and why
  • AAdd a redrive allow policy to the dead-letter queue that names the source queue as an allowed queue

    Incorrect

    A redrive allow policy controls which source queues may use the dead-letter queue; it does not move any messages.

  • BIncrease the maxReceiveCount in the source queue's redrive policy so that messages get more processing attempts

    Incorrect

    This allows more receives before future failures are moved; messages already in the dead-letter queue are not affected.

  • CPoint the service's queue URL at the dead-letter queue until it is empty, then point it back

    Incorrect

    This needs no copy script, but it changes and redeploys the service and does not limit how fast the backlog reaches the database.

  • DStart a dead-letter queue redrive to the source queue with a custom maximum velocity of 50 messages per second

    Correct

    Redrive moves the messages back to the source queue without code, and a custom maximum velocity limits how many messages per second it moves.

Dead-letter queue redrive moves messages to the source queue or to another queue of the same type, so the regular consumer processes them again. Velocity control can be system optimized or a custom maximum of up to 500 messages per second.

Question 6 · choose 1

A Lambda function processes a Kinesis data stream with 20 shards. Each batch waits on a slow external API, and the IteratorAge metric keeps growing. Records with the same partition key must stay in order, and the team cannot add shards right now. What should the DevOps engineer change?

  1. ARaise ParallelizationFactor on the event source mapping, for example to 5
  2. BSwitch the event source mapping to read from TRIM_HORIZON
  3. CLower the batch size to 1 so that each invocation finishes faster
  4. DSet reserved concurrency of 20 on the function to guarantee capacity
Show the answer and why
  • ARaise ParallelizationFactor on the event source mapping, for example to 5

    Correct

    More concurrent batches per shard raise throughput while order is kept per partition key.

  • BSwitch the event source mapping to read from TRIM_HORIZON

    Incorrect

    The starting position does not change how fast batches are processed.

  • CLower the batch size to 1 so that each invocation finishes faster

    Incorrect

    Smaller batches add invocations per shard without adding concurrency.

  • DSet reserved concurrency of 20 on the function to guarantee capacity

    Incorrect

    This caps the function at one batch per shard instead of adding parallelism.

ParallelizationFactor, from 1 to 10, lets Lambda process a shard with several concurrent batches. Lambda still keeps in-order processing at the partition-key level.

Question 7 · choose 1

An SNS standard topic delivers alerts to a partner's HTTPS endpoint and to several internal subscribers. The partner's endpoint is sometimes down for several hours, and the alerts sent during those outages are lost. The team wants every alert that could not be delivered to the partner kept for up to 14 days, separate from those that were delivered, so that it can be analyzed and resent. The other subscribers must not be affected. What should the DevOps engineer configure?

  1. AA custom delivery policy on the partner's subscription with 100 retries and a maximum delay of 3,600 seconds
  2. BAn SQS dead-letter queue in the same account and Region, with 14-day retention, on the partner's subscription
  3. CAn additional SQS queue subscribed to the topic that keeps a copy of every alert for 14 days
  4. DDelivery status logging for HTTPS on the topic, with logs sent to CloudWatch Logs
Show the answer and why
  • AA custom delivery policy on the partner's subscription with 100 retries and a maximum delay of 3,600 seconds

    Incorrect

    A custom policy helps with short outages, but the total retry time for an HTTPS endpoint is limited to 3,600 seconds, after which messages are discarded.

  • BAn SQS dead-letter queue in the same account and Region, with 14-day retention, on the partner's subscription

    Correct

    Messages that cannot be delivered after the retry policy are held in the subscription's dead-letter queue, and only that subscription is affected.

  • CAn additional SQS queue subscribed to the topic that keeps a copy of every alert for 14 days

    Incorrect

    Fan-out to a queue keeps a copy of all alerts, but it does not show which ones failed to reach the partner.

  • DDelivery status logging for HTTPS on the topic, with logs sent to CloudWatch Logs

    Incorrect

    Delivery status logs show whether each delivery succeeded and the endpoint's response, but they do not keep the alerts for resending.

When the delivery policy is exhausted, SNS discards a message unless a dead-letter queue is attached to the subscription. The queue must be in the same account and Region, and a 14-day retention gives time to analyze and resend the messages.

Question 8 · choose 1

Order events on an EventBridge bus must be sent to a SaaS ticketing system's HTTPS webhook that accepts at most 10 requests per second and requires an API key header. The team wants no custom code between EventBridge and the webhook. What should the DevOps engineer configure?

  1. AA Lambda function target that calls the webhook and sleeps between calls
  2. BAn SNS topic target with an HTTPS subscription to the webhook
  3. CAn API destination whose connection holds the API key, with a rate of 10 per second
  4. DAn SQS queue target with a consumer on EC2 that forwards messages to the webhook
Show the answer and why
  • AA Lambda function target that calls the webhook and sleeps between calls

    Incorrect

    This is custom code that API destinations replace.

  • BAn SNS topic target with an HTTPS subscription to the webhook

    Incorrect

    SNS does not limit the request rate to the webhook as needed here.

  • CAn API destination whose connection holds the API key, with a rate of 10 per second

    Correct

    API destinations invoke HTTPS endpoints with authorization from a connection and a configurable invocation rate.

  • DAn SQS queue target with a consumer on EC2 that forwards messages to the webhook

    Incorrect

    This adds a consumer service to build and run.

EventBridge API destinations call HTTPS endpoints as rule or pipe targets. Connections hold authorization, and the invocation rate limits calls per second, with events backing up when the rate is lower than the event rate.

Practise domain 5 →Practise all domains →