Skip to content
BytePatterns

DOP-C02 · Domain 5: Incident and Event Response · 14% of the exam

Task 5.2: Implement configuration changes in response to events.

Fixing state when something happens: Systems Manager and Auto Scaling acting on a fleet, AWS Config finding drift from the desired state, and automated changes to infrastructure that bring it back.

Study it

  • Changing configuration on events: Systems Manager Automation and AWS Config remediation

    Lesson coming

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A company runs EBS-backed Amazon EC2 instances that are managed by Systems Manager. When AWS Health schedules an instance for retirement, the operations team stops and starts the instance by hand at a quiet moment so that it moves to a new host. The team wants this to happen automatically as soon as the retirement notice arrives. What should the DevOps engineer do?

  1. ACreate an EventBridge rule for the Health retirement event that starts the AWS-RestartEC2Instance Automation runbook for the instance
  2. BCreate an EventBridge rule for the Health retirement event that reboots the affected instance with the AWS-RebootEC2Instance runbook
  3. CDo nothing and let the scheduled stop happen during the maintenance window, because EC2 then starts the instance again on a new host
  4. DAdd a CloudWatch alarm with an EC2 recover action on the StatusCheckFailed_System metric of every instance
Show the answer and why
  • ACreate an EventBridge rule for the Health retirement event that starts the AWS-RestartEC2Instance Automation runbook for the instance

    Correct

    AWS Health sends AWS_EC2_PERSISTENT_INSTANCE_RETIREMENT_SCHEDULED for EBS-backed instances, and an EventBridge rule can run a Systems Manager Automation document that stops and starts the instance, which migrates it to a new host.

  • BCreate an EventBridge rule for the Health retirement event that reboots the affected instance with the AWS-RebootEC2Instance runbook

    Incorrect

    A reboot keeps the instance on the same host. Stopping and starting is what moves it, and it is not the same as rebooting.

  • CDo nothing and let the scheduled stop happen during the maintenance window, because EC2 then starts the instance again on a new host

    Incorrect

    The scheduled event stops the instance; the documented ways to move it are a manual or automated stop and start.

  • DAdd a CloudWatch alarm with an EC2 recover action on the StatusCheckFailed_System metric of every instance

    Incorrect

    The recover action reacts to failed system status checks. A scheduled retirement notice arrives before anything fails, so the alarm does not act on it.

Scheduled retirement of an EBS-backed instance can be handled ahead of time. AWS documents an EventBridge rule on the Health event that runs AWS-RestartEC2Instance through Systems Manager Automation, with an IAM role that lets EventBridge start the automation.

Question 2 · choose 1

Instances in an Auto Scaling group write application logs to local disk and upload them to Amazon S3 every hour. When the group scales in, up to an hour of logs is lost with each terminated instance. The team wants every instance's remaining logs uploaded before it is terminated, without keeping instances running longer than needed. What should the DevOps engineer do?

  1. ATurn on scale-in protection for all instances so that the group cannot terminate any instance that still holds logs on its local disk
  2. BRaise the health check grace period so that instances stay in service longer before the group terminates them
  3. CAdd a termination lifecycle hook, upload the logs when its EventBridge event arrives, then complete the lifecycle action
  4. DAdd a scheduled action that raises the minimum size every evening so that fewer instances are terminated at night
Show the answer and why
  • ATurn on scale-in protection for all instances so that the group cannot terminate any instance that still holds logs on its local disk

    Incorrect

    Protected instances are not terminated on scale in, so the group cannot shrink and capacity is wasted.

  • BRaise the health check grace period so that instances stay in service longer before the group terminates them

    Incorrect

    The grace period applies to health checks of new instances. It does not delay a scale-in termination.

  • CAdd a termination lifecycle hook, upload the logs when its EventBridge event arrives, then complete the lifecycle action

    Correct

    A termination hook puts the instance in a wait state, during which logs can be downloaded or uploaded; completing the lifecycle action lets the termination continue.

  • DAdd a scheduled action that raises the minimum size every evening so that fewer instances are terminated at night

    Incorrect

    Fewer terminations still lose logs when they happen, and the group keeps more instances than it needs.

Lifecycle hooks pause an instance at launch or termination for a set time, one hour by default. A termination hook is the hook point for last-minute work such as shipping logs, followed by complete-lifecycle-action.

Question 3 · choose 3

Amazon GuardDuty runs in a production account. When it reports a CryptoCurrency finding for an Amazon EC2 instance, the security team wants the instance isolated from the network within minutes, automatically, and kept intact for forensics. The team wants to use an AWS-provided Systems Manager runbook instead of writing code. Which actions should the DevOps engineer take? (Choose THREE.)

  1. ACreate an EventBridge rule on the default event bus that matches source aws.guardduty, detail-type GuardDuty Finding and the CryptoCurrency finding types
  2. BWrite a Lambda function that replaces the instance's security groups and call it from the rule
  3. CSet the rule's target to the AWSSupport-ContainEC2Instance Automation runbook with Action set to Contain and an input transformer that passes the instance ID
  4. DAdd an EC2 terminate action to an alarm on the instance so that the compromised instance is removed
  5. EGive the rule an IAM role that EventBridge can assume and that may start the Automation execution
Show the answer and why
  • ACreate an EventBridge rule on the default event bus that matches source aws.guardduty, detail-type GuardDuty Finding and the CryptoCurrency finding types

    Correct

    GuardDuty publishes each finding to EventBridge as an event with this source and detail type, so a rule can select the finding types of interest.

  • BWrite a Lambda function that replaces the instance's security groups and call it from the rule

    Incorrect

    This is custom code, which the team wants to avoid when an AWS-provided runbook does the containment.

  • CSet the rule's target to the AWSSupport-ContainEC2Instance Automation runbook with Action set to Contain and an input transformer that passes the instance ID

    Correct

    The runbook performs network containment of an instance and keeps its original security group configuration for a later restore, and input transformation passes values from the event to the target.

  • DAdd an EC2 terminate action to an alarm on the instance so that the compromised instance is removed

    Incorrect

    Terminating the instance destroys the evidence that the team wants to keep for forensics.

  • EGive the rule an IAM role that EventBridge can assume and that may start the Automation execution

    Correct

    To relay events to targets such as Systems Manager Automation, EventBridge needs an IAM role with the permissions for that target.

Event-driven remediation with Systems Manager has three parts: a rule that matches the event, an Automation runbook as the target with the right input, and a role that lets EventBridge start the automation.

Question 4 · choose 1

An ECS cluster runs part of its EC2 container instances on Spot. When a Spot instance is interrupted, its tasks are killed abruptly and requests fail. The team wants ECS to react to the two-minute interruption notice by moving service tasks off the instance before it is reclaimed. What should the DevOps engineer configure?

  1. ATurn on Amazon ECS Spot Instance draining on the Linux container instances
  2. BRaise the ECS service's health check grace period to five minutes
  3. CTurn on Capacity Rebalancing on the Auto Scaling group and rely on it to notify ECS
  4. DUse the hibernate interruption behavior so that tasks pause until the instance returns
Show the answer and why
  • ATurn on Amazon ECS Spot Instance draining on the Linux container instances

    Correct

    With Spot draining on, ECS receives the interruption notice and sets the instance to DRAINING, so replacement tasks start elsewhere.

  • BRaise the ECS service's health check grace period to five minutes

    Incorrect

    The grace period affects new tasks' health checks, not interrupted instances.

  • CTurn on Capacity Rebalancing on the Auto Scaling group and rely on it to notify ECS

    Incorrect

    ECS does not receive a notice when Capacity Rebalancing removes instances.

  • DUse the hibernate interruption behavior so that tasks pause until the instance returns

    Incorrect

    Hibernation gives no two-minute notice, and tasks would not move.

Spot draining places an interrupted container instance into DRAINING. ECS stops scheduling new tasks there and starts replacements for service tasks on other instances.

Question 5 · choose 1

Worker tasks in an ECS service read jobs from an SQS queue. Customers complain when jobs wait more than five minutes, but CPU of the workers stays low because they wait on I/O. The team wants the service to add tasks in steps when jobs start waiting too long and remove them when the wait drops. What should the DevOps engineer configure?

  1. ATarget tracking on the service's average CPU utilization at 50 percent
  2. BA scheduled action that doubles the tasks during business hours
  3. CA longer visibility timeout on the queue so that jobs wait less
  4. DA step scaling policy driven by an alarm on ApproximateAgeOfOldestMessage
Show the answer and why
  • ATarget tracking on the service's average CPU utilization at 50 percent

    Incorrect

    CPU stays low while jobs wait, so it does not track the backlog.

  • BA scheduled action that doubles the tasks during business hours

    Incorrect

    Schedules do not react to how long jobs actually wait.

  • CA longer visibility timeout on the queue so that jobs wait less

    Incorrect

    The visibility timeout does not add workers or speed up processing.

  • DA step scaling policy driven by an alarm on ApproximateAgeOfOldestMessage

    Correct

    Step scaling changes capacity by steps based on the alarm, and the age of the oldest message reflects how long jobs wait.

Application Auto Scaling step scaling adjusts capacity in steps based on CloudWatch alarm breaches. ApproximateAgeOfOldestMessage measures how long the oldest message has waited in the queue.

Question 6 · choose 1

An engineer must troubleshoot one instance in an Auto Scaling group behind an Application Load Balancer. Each time the engineer restarts the application, the group marks the instance unhealthy and replaces it. The instance must stay in the group and receive no traffic during the work, the group must keep the same number of instances serving traffic, and every other instance must still be replaced if it fails its health checks. What should the DevOps engineer do?

  1. ASuspend the HealthCheck and ReplaceUnhealthy processes for the group until the work is done
  2. BTurn on instance scale-in protection for the instance being troubleshot
  3. CPut the instance into Standby without decrementing the desired capacity, then exit Standby afterwards
  4. DDetach the instance from the group and attach it again when the work is done
Show the answer and why
  • ASuspend the HealthCheck and ReplaceUnhealthy processes for the group until the work is done

    Incorrect

    Suspension stops the replacement, but it applies to the whole group, and the instance keeps receiving load balancer traffic.

  • BTurn on instance scale-in protection for the instance being troubleshot

    Incorrect

    Scale-in protection guards against scale-in, but not against replacement after a failed health check, and traffic still arrives.

  • CPut the instance into Standby without decrementing the desired capacity, then exit Standby afterwards

    Correct

    Standby deregisters the instance from the load balancer and skips its health checks, and keeping the desired capacity makes the group launch a replacement.

  • DDetach the instance from the group and attach it again when the work is done

    Incorrect

    Detaching removes it from the load balancer, but the instance is no longer part of the group, which must keep it.

Standby removes an instance from service temporarily so that it can be updated or troubleshot without being terminated by health checks or scale-in. Choosing not to decrement the desired capacity keeps the number of instances in service.

Question 7 · choose 1

GuardDuty findings show instances behind an Application Load Balancer being probed by known malicious IP addresses. Security wants each such IP address blocked at the load balancer's web ACL within minutes of a finding, without people in the loop. What should the DevOps engineer build?

  1. AAn AWS Firewall Manager AWS WAF policy that applies the same web ACL to every load balancer in the organization
  2. BAn EventBridge rule for GuardDuty findings that runs a function to add the IP to a WAF IP set
  3. CA network ACL rule for each address in the load balancer subnets, added by the security team
  4. DA GuardDuty threat list that contains the addresses from earlier findings, uploaded to S3 and activated
Show the answer and why
  • AAn AWS Firewall Manager AWS WAF policy that applies the same web ACL to every load balancer in the organization

    Incorrect

    Firewall Manager keeps web ACLs consistent across accounts, but it does not react to GuardDuty findings or add the reported addresses to a rule.

  • BAn EventBridge rule for GuardDuty findings that runs a function to add the IP to a WAF IP set

    Correct

    GuardDuty publishes findings to EventBridge, and an IP set can be used in a blocking rule of the web ACL.

  • CA network ACL rule for each address in the load balancer subnets, added by the security team

    Incorrect

    This is manual and limited by the number of network ACL rules.

  • DA GuardDuty threat list that contains the addresses from earlier findings, uploaded to S3 and activated

    Incorrect

    Threat lists make GuardDuty generate findings for activity involving the listed addresses; they do not block any traffic.

GuardDuty sends findings to EventBridge in near real time. An automated target can update an AWS WAF IP set that a rule in the web ACL blocks.

Question 8 · choose 1

New instances in an Auto Scaling group behind an Application Load Balancer must download a large dataset and register with a configuration service, which takes 5 to 10 minutes, before they receive traffic. Today they are registered with the load balancer early and fail requests for several minutes. The team also wants an instance whose setup fails to be terminated and replaced automatically instead of being put in service. What should the DevOps engineer configure?

  1. AA default instance warmup of 600 seconds for the group so that new instances can settle first
  2. BA launch lifecycle hook that holds instances in Pending:Wait until the setup script reports CONTINUE or ABANDON
  3. CA health check grace period of 15 minutes so that new instances are not replaced while they are still setting up
  4. DA slow start duration of 600 seconds on the target group so that new instances ramp up to a full share of requests
Show the answer and why
  • AA default instance warmup of 600 seconds for the group so that new instances can settle first

    Incorrect

    The warmup keeps new instances out of the aggregated scaling metrics for a while, but they are still registered with the load balancer and receive traffic.

  • BA launch lifecycle hook that holds instances in Pending:Wait until the setup script reports CONTINUE or ABANDON

    Correct

    Instances wait before load balancer registration until the action completes, and an ABANDON result makes the group terminate and replace the instance.

  • CA health check grace period of 15 minutes so that new instances are not replaced while they are still setting up

    Incorrect

    The grace period only delays Auto Scaling health checks; the instance is already registered with the load balancer and receiving requests.

  • DA slow start duration of 600 seconds on the target group so that new instances ramp up to a full share of requests

    Incorrect

    Slow start only ramps up requests to a target that is registered and healthy, so the instance still receives requests during setup, and a failed setup is not replaced.

Launch lifecycle hooks put instances in Pending:Wait, by default for up to an hour, so that custom actions finish before the instance is registered with the load balancer and becomes InService. An ABANDON result tells the group to terminate and replace the instance.

Practise domain 5 →Practise all domains →