Skip to content
BytePatterns

DVA-C02 · Domain 4: Troubleshooting and Optimization · 18% of the exam

Task 4.1: Assist in a root cause analysis

Finding why something broke: reading metrics, logs and traces, querying logs with CloudWatch Logs Insights, custom metrics with the embedded metric format, dashboards, and the service logs that explain a failed deployment or a broken integration.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

Users report that an API backed by an AWS Lambda function was slow during the last hour. In the function's log group, the developer wants the 20 slowest invocations of that hour, with their request IDs, to look them up one by one. Which CloudWatch Logs Insights query should the developer run?

  1. Afields @requestId, @duration | filter @type = "REPORT" | sort @duration desc | limit 20
  2. Bfilter @type = "REPORT" | stats max(@duration) by bin(5m)
  3. Cfields @requestId, @duration | filter @type = "REPORT" | sort @duration asc | limit 20
  4. Dfields @requestId, @message | filter @message like /ERROR/ | sort @timestamp desc | limit 20
Show the answer and why
  • Afields @requestId, @duration | filter @type = "REPORT" | sort @duration desc | limit 20

    Correct

    Lambda REPORT lines carry the discovered @duration and @requestId fields. Sorting by duration in descending order and limiting to 20 returns the slowest invocations.

  • Bfilter @type = "REPORT" | stats max(@duration) by bin(5m)

    Incorrect

    stats aggregates into one value per 5-minute bin, so individual invocations and their request IDs are lost.

  • Cfields @requestId, @duration | filter @type = "REPORT" | sort @duration asc | limit 20

    Incorrect

    Ascending order returns the shortest durations first, so this lists the 20 fastest invocations, not the slowest.

  • Dfields @requestId, @message | filter @message like /ERROR/ | sort @timestamp desc | limit 20

    Incorrect

    This returns the 20 most recent log events that contain ERROR. Slow invocations do not have to log errors.

Lambda REPORT log lines expose @duration, @billedDuration and @maxMemoryUsed; filter, sort and limit turn them into a top-N list.

Question 2 · choose 1

A busy AWS Lambda function records business metrics such as OrdersPlaced and CartValue, with a PaymentMethod dimension, by calling the CloudWatch PutMetricData API during every invocation. These calls add latency to each request. The developer wants the same custom metrics without making a separate metrics API call in the function. What should the developer do?

  1. ASend each data point to an SQS queue for another function that calls PutMetricData
  2. BWrite the values as embedded metric format log lines to the function's logs
  3. CTurn on CloudWatch Lambda Insights for the function
  4. DPublish the metrics with PutMetricData as high-resolution metrics
Show the answer and why
  • ASend each data point to an SQS queue for another function that calls PutMetricData

    Incorrect

    The function would still make an API call per invocation, to SQS instead of CloudWatch, and a second function would be added to run.

  • BWrite the values as embedded metric format log lines to the function's logs

    Correct

    The embedded metric format generates custom metrics asynchronously from structured logs written to CloudWatch Logs; CloudWatch extracts the metrics, and no PutMetricData call is needed.

  • CTurn on CloudWatch Lambda Insights for the function

    Incorrect

    Lambda Insights collects system-level metrics such as CPU time, memory and cold starts. It does not create business metrics such as OrdersPlaced.

  • DPublish the metrics with PutMetricData as high-resolution metrics

    Incorrect

    High resolution changes the granularity to one second. The function still calls PutMetricData on every invocation.

EMF is the low-overhead way to emit custom metrics from Lambda and containers: log a structured JSON line, and CloudWatch creates the metric.

Question 3 · choose 1

A first deployment of an AWS SAM application creates an AWS CloudFormation stack, which fails and ends in ROLLBACK_COMPLETE. The stack's event list is long: dozens of CREATE_IN_PROGRESS, CREATE_COMPLETE and DELETE_COMPLETE events. The developer wants to find the event that most likely caused the failure, and its explanation, as quickly as possible. What should the developer do?

  1. ARead the status reason of the stack's final ROLLBACK_COMPLETE event
  2. BRead the DELETE_COMPLETE events that the rollback produced
  3. CRun drift detection on the stack and review the drifted resources
  4. DRun Detect root cause in the console and read the event it labels
Show the answer and why
  • ARead the status reason of the stack's final ROLLBACK_COMPLETE event

    Incorrect

    ROLLBACK_COMPLETE only reports that the resources created during the failed creation were removed. It is the end state, not the cause.

  • BRead the DELETE_COMPLETE events that the rollback produced

    Incorrect

    Those events record the cleanup after the failure. They show what was removed, not why the creation failed.

  • CRun drift detection on the stack and review the drifted resources

    Incorrect

    Drift detection compares existing resources with the stack template. It does not explain why a stack creation failed.

  • DRun Detect root cause in the console and read the event it labels

    Correct

    CloudFormation analyzes the failure and labels the event that is likely the root cause; its Status reason, and sometimes a linked CloudTrail event, gives the details.

For a failed stack, the useful event is the failure that started it all; Detect root cause finds it among the cancellations and rollback events.

Question 4 · choose 1

A REST API in Amazon API Gateway uses a Lambda proxy integration. The Python handler ends with return {"status": 200, "data": order}. The function's logs show that every invocation succeeds, yet every client receives a 502 Bad Gateway response. What should the developer change?

  1. AReturn {"statusCode": 200, "body": body}, with body as the order in JSON
  2. BTurn on API caching for the method so that responses come from the cache
  3. CIncrease the function's memory so that it returns its result faster
  4. DReturn the original dictionary serialized into one JSON string instead
Show the answer and why
  • AReturn {"statusCode": 200, "body": body}, with body as the order in JSON

    Correct

    A proxy integration expects an object with statusCode and a string body (plus optional headers). Any other output format makes API Gateway return 502 Bad Gateway.

  • BTurn on API caching for the method so that responses come from the cache

    Incorrect

    Caching stores endpoint responses for a TTL. A malformed response is still malformed, and the first request still reaches the function.

  • CIncrease the function's memory so that it returns its result faster

    Incorrect

    Memory adds CPU to the function. The invocations already succeed; the problem is the shape of what the function returns.

  • DReturn the original dictionary serialized into one JSON string instead

    Incorrect

    A plain string is still not the output format a proxy integration expects, so API Gateway keeps returning 502.

With Lambda proxy integrations the function owns the HTTP response: statusCode, headers and a string body, or the client sees 502.

Question 5 · choose 1

A newly created AWS Lambda function calls a partner API that usually answers in four to six seconds. Every invocation fails, and the logs show "Task timed out after 3.00 seconds". The developer did not change any settings after creating the function. What should the developer do?

  1. ARaise the function's memory to 10,240 MB, the largest available setting
  2. BSet reserved concurrency on the function
  3. CTurn on AWS X-Ray active tracing
  4. DRaise the function's timeout above the API's response time
Show the answer and why
  • ARaise the function's memory to 10,240 MB, the largest available setting

    Incorrect

    More memory speeds up computation, but the function is waiting on the partner API, so it still stops at the timeout.

  • BSet reserved concurrency on the function

    Incorrect

    Concurrency settings do not change how long one invocation may run.

  • CTurn on AWS X-Ray active tracing

    Incorrect

    Tracing shows where time is spent but does not stop the timeout.

  • DRaise the function's timeout above the API's response time

    Correct

    The default timeout is 3 seconds. Raising it above the partner API's response time lets the call finish.

New functions start with a 3-second timeout; set it to match the slowest expected calls.

Question 6 · choose 1

Tasks of a new Amazon ECS service on AWS Fargate fail to start with a CannotPullContainerError for an image in Amazon ECR. The tasks run in private subnets without public IP addresses. The error says the connection timed out. What is the most likely cause?

  1. AThe container's health check command is wrong
  2. BThe task definition requests too little memory for the container image
  3. CThe subnets have no network route to Amazon ECR
  4. DThe service's desired count is set too high
Show the answer and why
  • AThe container's health check command is wrong

    Incorrect

    Health checks run after the container starts; this task never pulled its image.

  • BThe task definition requests too little memory for the container image

    Incorrect

    Memory limits affect running containers, not the network path to the registry.

  • CThe subnets have no network route to Amazon ECR

    Correct

    A timeout means there is no route; tasks in private subnets need a NAT gateway or VPC endpoints to reach the registry.

  • DThe service's desired count is set too high

    Incorrect

    The number of tasks does not cause an image pull to time out.

CannotPullContainer timeouts point to networking; permission errors point to the task execution role.

Question 7 · choose 1

An AWS CodeBuild project compiles and tests successfully, but the build fails in the UPLOAD_ARTIFACTS phase with an AccessDenied error for the Amazon S3 output bucket. The bucket was recently changed to a new name. What should the developer check first?

  1. AThe source repository credentials that the project uses to download code
  2. BThe service role's S3 permissions for the new bucket
  3. CThe compute type of the build environment
  4. DThe buildspec's install phase runtime versions
Show the answer and why
  • AThe source repository credentials that the project uses to download code

    Incorrect

    The source was already downloaded and built; the failure is when writing artifacts.

  • BThe service role's S3 permissions for the new bucket

    Correct

    CodeBuild uses its service role to call other services, including s3:PutObject for the artifact bucket.

  • CThe compute type of the build environment

    Incorrect

    Compute size does not affect permissions to write to S3.

  • DThe buildspec's install phase runtime versions

    Incorrect

    Runtime versions affect the build, which already succeeded.

AccessDenied in CodeBuild usually points to the service role; check that it covers renamed resources.

Question 8 · choose 1

A CloudFormation stack update failed, and the automatic rollback also failed because a database instance it needed had been deleted outside CloudFormation. The stack is now in UPDATE_ROLLBACK_FAILED and cannot be updated. What should the developer do?

  1. ADelete the whole stack and recreate it later from the original template file
  2. BRun another stack update with the previous template
  3. CContinue the rollback and skip the resource that cannot roll back
  4. DTurn on termination protection and wait for the stack to recover
Show the answer and why
  • ADelete the whole stack and recreate it later from the original template file

    Incorrect

    This destroys the other resources in the stack when the stack can be returned to a working state instead.

  • BRun another stack update with the previous template

    Incorrect

    A stack in UPDATE_ROLLBACK_FAILED cannot be updated until the rollback is completed.

  • CContinue the rollback and skip the resource that cannot roll back

    Correct

    Continuing the rollback, if needed with the failing resources skipped, returns the stack to UPDATE_ROLLBACK_COMPLETE.

  • DTurn on termination protection and wait for the stack to recover

    Incorrect

    Termination protection prevents deletion; the stack does not recover on its own.

Fix the cause or skip the failing resources, then continue the rollback so that the stack can be updated again.

Question 9 · choose 1

In the AWS X-Ray trace map, a downstream service node shows a high rate of responses marked as errors, while faults stay at zero. A developer must decide where to look first. What do these error marks indicate?

  1. AServer faults, meaning 500 series errors in the downstream service
  2. BThrottling, meaning 429 Too Many Requests responses
  3. CClient errors, meaning 400 series responses
  4. DNetwork timeouts between the two services
Show the answer and why
  • AServer faults, meaning 500 series errors in the downstream service

    Incorrect

    X-Ray records 500 series errors as faults, and faults are at zero here.

  • BThrottling, meaning 429 Too Many Requests responses

    Incorrect

    X-Ray marks 429 responses as throttles, a separate category.

  • CClient errors, meaning 400 series responses

    Correct

    X-Ray marks client errors, the 400 series, as errors, so the requests sent to the service are the first place to look.

  • DNetwork timeouts between the two services

    Incorrect

    Errors in X-Ray refer to the response class, not to timeouts between nodes.

In X-Ray, errors are 4xx, faults are 5xx, and throttles are 429.

Practise domain 4 →Practise all domains →