Skip to content
BytePatterns

DOP-C02 · Domain 5: Incident and Event Response · 14% of the exam

Task 5.3: Troubleshoot system and application failures.

Root cause analysis: reading failed pipelines, builds, deployments and stack operations, synthetic canaries, and failures in scaling, ECS and EKS with CloudWatch, X-Ray, AWS Health and OpsCenter.

Study it

  • Troubleshooting failed pipelines, deployments and stacks

    Lesson coming

  • Troubleshooting scaling, ECS and EKS failures with CloudWatch, X-Ray and OpsCenter

    Partly covered by: CloudWatch, Alarms & X-Ray, Tracing a Request

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A stack update failed, and AWS CloudFormation started to roll it back. The rollback then failed too, and the stack is now in UPDATE_ROLLBACK_FAILED. The events show that an engineer had deleted a Lambda permission outside of CloudFormation, so CloudFormation cannot restore it. The team must get the stack back to a state where it can be updated, without deleting the stack or its other resources. What should the DevOps engineer do?

  1. ARun an update on the stack with a corrected template that recreates the deleted Lambda permission resource
  2. BContinue rolling back the update, skipping the logical ID of the Lambda permission that cannot be restored
  3. CDelete the stack with the option to retain all resources, and create a new stack that imports them
  4. DRun drift detection on the stack and accept the detected differences so that the stack becomes updatable
Show the answer and why
  • ARun an update on the stack with a corrected template that recreates the deleted Lambda permission resource

    Incorrect

    A stack in UPDATE_ROLLBACK_FAILED cannot be updated. It must first be rolled back to a working state.

  • BContinue rolling back the update, skipping the logical ID of the Lambda permission that cannot be restored

    Correct

    Continuing the rollback returns the stack to UPDATE_ROLLBACK_COMPLETE, and resources that cannot roll back can be skipped by logical ID.

  • CDelete the stack with the option to retain all resources, and create a new stack that imports them

    Incorrect

    This throws away the stack the team wants to keep and turns a rollback fix into a migration.

  • DRun drift detection on the stack and accept the detected differences so that the stack becomes updatable

    Incorrect

    Drift detection reports differences between the template and the resources; it does not change the stack's state.

UPDATE_ROLLBACK_FAILED is fixed by continuing the rollback: either repair the cause by hand, or skip the resources that cannot roll back. After the stack reaches UPDATE_ROLLBACK_COMPLETE, it can be updated again, and the skipped resources should be brought back in line.

Question 2 · choose 1

A new Amazon ECS service on AWS Fargate runs in private subnets of a VPC that has no NAT gateway, because security policy forbids internet access from these subnets. Tasks stop with CannotPullContainerError and a timeout while pulling the image from Amazon ECR in the same Region. The task execution role has the required ECR permissions. What should the DevOps engineer do?

  1. AAdd a NAT gateway in a public subnet and a default route to it from the private subnets
  2. BTurn on auto-assign public IP for the tasks so that they can reach the public ECR endpoints from the private subnets
  3. CCreate only an interface VPC endpoint for com.amazonaws.region.ecr.dkr in the private subnets
  4. DCreate interface endpoints for ecr.api and ecr.dkr and an Amazon S3 gateway endpoint
Show the answer and why
  • AAdd a NAT gateway in a public subnet and a default route to it from the private subnets

    Incorrect

    This would fix the pull by giving the subnets internet access, which the security policy forbids.

  • BTurn on auto-assign public IP for the tasks so that they can reach the public ECR endpoints from the private subnets

    Incorrect

    A public IP in a private subnet without an internet route does not give access, and the policy forbids internet access anyway.

  • CCreate only an interface VPC endpoint for com.amazonaws.region.ecr.dkr in the private subnets

    Incorrect

    Fargate tasks also need the ecr.api endpoint and the Amazon S3 gateway endpoint, because image layers are stored in Amazon S3.

  • DCreate interface endpoints for ecr.api and ecr.dkr and an Amazon S3 gateway endpoint

    Correct

    Tasks pull images over PrivateLink through the two ECR interface endpoints and need the S3 gateway endpoint for the image layers.

A timeout in CannotPullContainerError points to a network path problem to the ECR endpoint. In VPCs without internet access, ECR is reached through its interface endpoints plus the S3 gateway endpoint; a CloudWatch Logs endpoint is also needed if the tasks use the awslogs driver.

Question 3 · choose 1

During a product launch, an Auto Scaling group that uses a single instance type in one Availability Zone fails to scale out. Its activity history shows "We currently do not have sufficient capacity in the Availability Zone you requested... Launching EC2 instance failed." The application can run on several similar instance types. How should the DevOps engineer make future scale-outs more reliable?

  1. AAdd more Availability Zones to the group and use a mixed instances policy with several suitable instance types
  2. BRequest an increase of the account's On-Demand instance quota for the instance family in the Region
  3. CRaise the group's maximum size so that the next scale-out requests more instances of the same type at once
  4. DTurn off the group's health checks during the launch so that new instances are not replaced while they start
Show the answer and why
  • AAdd more Availability Zones to the group and use a mixed instances policy with several suitable instance types

    Correct

    AWS recommends expanding the group to more Availability Zones and using a diverse set of instance types so that the group does not rely on one instance type.

  • BRequest an increase of the account's On-Demand instance quota for the instance family in the Region

    Incorrect

    The error is about EC2 capacity for that instance type in that Availability Zone, not about the account's quota.

  • CRaise the group's maximum size so that the next scale-out requests more instances of the same type at once

    Incorrect

    Asking for more of the same instance type in the same zone does not help when that zone has no capacity for it.

  • DTurn off the group's health checks during the launch so that new instances are not replaced while they start

    Incorrect

    The instances are never launched, so health checks play no part in the failure.

Insufficient capacity errors are tied to an instance type in an Availability Zone. Spreading the group across zones and instance types gives Amazon EC2 Auto Scaling more pools to launch from.

Question 4 · choose 1

A new AWS CodeBuild project builds a container image with docker build and pushes it to Amazon ECR. The build fails in the build phase with the message "Cannot connect to the Docker daemon at unix:/var/run/docker.sock. Is the docker daemon running?" The service role already has the ECR permissions it needs. What should the DevOps engineer do?

  1. AAdd ecr:GetAuthorizationToken and the ECR push actions to the CodeBuild service role
  2. BIncrease the compute type of the build project so that the Docker daemon has enough memory to start
  3. CTurn on privileged mode for the build project's environment so that the build can use the Docker daemon
  4. DTurn on local Docker layer caching for the project so that the build reuses the daemon from earlier builds
Show the answer and why
  • AAdd ecr:GetAuthorizationToken and the ECR push actions to the CodeBuild service role

    Incorrect

    Missing ECR permissions cause authorization errors when logging in or pushing, not a failure to reach the Docker daemon.

  • BIncrease the compute type of the build project so that the Docker daemon has enough memory to start

    Incorrect

    The documented cause of this error is that the build is not running in privileged mode, not a lack of memory.

  • CTurn on privileged mode for the build project's environment so that the build can use the Docker daemon

    Correct

    CodeBuild documents this error as caused by not running the build in privileged mode and fixes it by turning privileged mode on.

  • DTurn on local Docker layer caching for the project so that the build reuses the daemon from earlier builds

    Incorrect

    Docker layer caching itself requires privileged mode, and it speeds up builds rather than providing a daemon.

Building Docker images inside CodeBuild needs privileged mode on the project's environment. The error message about the Docker socket is the typical sign that it is off.

Question 5 · choose 1

Deleting an old CloudFormation stack ends in DELETE_FAILED because one of its S3 buckets still contains objects. The business wants to keep that bucket and its objects, but the stack itself must be removed. What should the DevOps engineer do?

  1. ARetry the deletion with RetainResources set to the bucket's logical ID
  2. BEmpty the bucket of all objects and then delete the stack again
  3. CTurn on termination protection and delete the stack again
  4. DAdd a stack policy that denies Update:Delete on the bucket
Show the answer and why
  • ARetry the deletion with RetainResources set to the bucket's logical ID

    Correct

    Rerunning the deletion with RetainResources deletes the stack without deleting the retained resource.

  • BEmpty the bucket of all objects and then delete the stack again

    Incorrect

    This deletes the objects that the business wants to keep.

  • CTurn on termination protection and delete the stack again

    Incorrect

    Termination protection prevents stack deletion.

  • DAdd a stack policy that denies Update:Delete on the bucket

    Incorrect

    Stack policies protect resources from stack updates, not from stack deletion.

Some resources, such as S3 buckets with objects, cannot be deleted. When a stack is in DELETE_FAILED, the deletion can be retried while retaining the resource that CloudFormation cannot delete.

Question 6 · choose 1

An ECS service uses CodeDeploy blue/green deployments with the all-at-once configuration. Twice this month CodeDeploy reported success, but the new task set then failed its load balancer health checks and the application went down. What should the DevOps engineer change so that bad releases fail before customers are affected?

  1. ASwitch to a canary or linear configuration so health checks run before traffic shifts
  2. BLower the target group's healthy threshold so that new tasks are marked healthy faster
  3. CAdd a CloudWatch alarm on unhealthy hosts to the deployment group and roll back automatically when it goes into ALARM
  4. DTurn on the ECS deployment circuit breaker with rollback for the service
Show the answer and why
  • ASwitch to a canary or linear configuration so health checks run before traffic shifts

    Correct

    With canary or linear configurations, health checks start on the replacement tasks before the AllowTraffic event.

  • BLower the target group's healthy threshold so that new tasks are marked healthy faster

    Incorrect

    Faster healthy marks do not stop broken code from receiving traffic.

  • CAdd a CloudWatch alarm on unhealthy hosts to the deployment group and roll back automatically when it goes into ALARM

    Incorrect

    Alarm-based rollback reacts after the threshold is met, and with all-at-once the health checks fail only after production traffic has already shifted.

  • DTurn on the ECS deployment circuit breaker with rollback for the service

    Incorrect

    The circuit breaker is supported only for services that use the rolling update (ECS) deployment controller, not CodeDeploy blue/green.

With all-at-once, health checks run on the green task set only after traffic has shifted. Canary and linear configurations let health checks run during installation, so failures stop the deployment first.

Question 7 · choose 1

An ECS service on EC2 container instances behind a Classic Load Balancer keeps replacing tasks because they are marked unhealthy, although the application answers on port 80 when tested from inside the instance. The load balancer health check pings port 80. What should the DevOps engineer check first?

  1. AWhether the task definition reserves enough CPU and memory for the app
  2. BWhether the ECR image was scanned for vulnerabilities when it was pushed
  3. CWhether the instance security group allows port 80 from the load balancer
  4. DWhether the container instances run the latest ECS-optimized AMI version
Show the answer and why
  • AWhether the task definition reserves enough CPU and memory for the app

    Incorrect

    The application already answers on the port, so resource sizing is not the first suspect.

  • BWhether the ECR image was scanned for vulnerabilities when it was pushed

    Incorrect

    Image scanning does not affect health checks.

  • CWhether the instance security group allows port 80 from the load balancer

    Correct

    If the container maps to port 80, the instance security group must allow it for the health checks to pass.

  • DWhether the container instances run the latest ECS-optimized AMI version

    Incorrect

    The AMI version does not explain blocked health check traffic.

Health checks fail when the load balancer cannot reach the mapped port, or when health check settings are too strict or point to the wrong port.

Question 8 · choose 1

A Linux EC2 instance with an unencrypted root volume became unreachable by SSH and Session Manager after a network configuration change. The team wants an AWS-provided automated way to diagnose and, where possible, repair common connectivity problems on the instance. What should the DevOps engineer run?

  1. AThe AWS-RestartEC2Instance runbook on a schedule until the instance responds
  2. BThe AWSSupport-ExecuteEC2Rescue runbook against the instance
  3. CAn EC2 instance recovery alarm on StatusCheckFailed_System
  4. DRun Command with a script that resets the network settings
Show the answer and why
  • AThe AWS-RestartEC2Instance runbook on a schedule until the instance responds

    Incorrect

    Restarting does not diagnose or repair the configuration.

  • BThe AWSSupport-ExecuteEC2Rescue runbook against the instance

    Correct

    This runbook uses EC2Rescue to troubleshoot and repair common connectivity issues for Linux or Windows instances.

  • CAn EC2 instance recovery alarm on StatusCheckFailed_System

    Incorrect

    Recovery moves an instance to new hardware; it does not fix the network configuration.

  • DRun Command with a script that resets the network settings

    Incorrect

    Run Command needs a reachable managed node, which this instance is not.

AWSSupport-ExecuteEC2Rescue runs EC2Rescue to diagnose and repair common connectivity issues. Instances with encrypted root volumes are not supported.

Question 9 · choose 1

A partner account uploads files to a company's old S3 bucket, which uses the Object writer ownership setting. The company's processing role, which the bucket policy allows, gets 403 Access Denied when it reads those files, but not for files the company uploads itself. What should the DevOps engineer change?

  1. AAdd s3:GetObjectAcl to the processing role's IAM policy for the bucket
  2. BAsk the partner to make the files public so that the role can read them
  3. CTurn on default encryption with SSE-S3 for the bucket
  4. DSet Object Ownership to Bucket owner enforced, which disables ACLs
Show the answer and why
  • AAdd s3:GetObjectAcl to the processing role's IAM policy for the bucket

    Incorrect

    The objects are owned by the partner, so the bucket owner's policies do not control them.

  • BAsk the partner to make the files public so that the role can read them

    Incorrect

    Public objects expose company data to everyone.

  • CTurn on default encryption with SSE-S3 for the bucket

    Incorrect

    Encryption settings do not change object ownership.

  • DSet Object Ownership to Bucket owner enforced, which disables ACLs

    Correct

    With Bucket owner enforced, the bucket owner owns all objects and access is controlled by policies.

With Object writer, the uploading account owns its objects. Bucket owner enforced, the default for new buckets, disables ACLs and makes the bucket owner the owner of every object.

Practise domain 5 →Practise all domains →