Skip to content
BytePatterns

DVA-C02 · Domain 4: Troubleshooting and Optimization · 18% of the exam

Task 4.2: Instrument code for observability

Making the application explain itself: logging versus monitoring versus observability, structured logs, custom metrics from code, traces with annotations, alarms and notifications for important events, and health checks.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

A checkout service sends traces to AWS X-Ray. Support engineers want to find every trace for a given customer ID with a filter expression in the console. They also want the full order object, a nested JSON structure, stored on the trace so that they can read it in the trace details, but they never search by it. How should the developer record these two values?

  1. AThe customer ID as metadata and the order object as an annotation
  2. BThe customer ID as an annotation and the order object as metadata
  3. CBoth values as metadata on the segment
  4. DBoth values as annotations on the segment
Show the answer and why
  • AThe customer ID as metadata and the order object as an annotation

    Incorrect

    Metadata is not indexed, so the customer ID could not be searched, and annotations hold simple values, not nested objects.

  • BThe customer ID as an annotation and the order object as metadata

    Correct

    Annotations are simple key-value pairs indexed for filter expressions. Metadata can hold values of any type, including objects, and is stored on the trace without being indexed.

  • CBoth values as metadata on the segment

    Incorrect

    Metadata is not indexed for search, so filter expressions could not find traces by customer ID.

  • DBoth values as annotations on the segment

    Incorrect

    Annotation values are Booleans, numbers or strings, so the nested order object does not fit, and X-Ray indexes up to 50 annotations per trace.

Annotate what you search by; attach as metadata what you only want to read. With OpenTelemetry, attributes become metadata unless their keys are listed in aws.xray.annotations.

Question 2 · choose 1

A Python AWS Lambda function logs with the standard logging module, and its logs reach CloudWatch Logs as plain text. The team wants each log event written as JSON key-value pairs so that fields can be searched, and wants DEBUG output kept out of CloudWatch Logs in production while still being able to switch it on later without changing code. What should the developer do?

  1. AAdd a metric filter to the log group that turns each line into JSON
  2. BSend the function's logs to a custom log group shared with the other functions
  3. CMove the log group to the Infrequent Access log class
  4. DSet the function's log format to JSON and its application log level to INFO
Show the answer and why
  • AAdd a metric filter to the log group that turns each line into JSON

    Incorrect

    A metric filter turns matching log data into numeric CloudWatch metrics. It does not change the format of the log events.

  • BSend the function's logs to a custom log group shared with the other functions

    Incorrect

    Several functions can share a custom log group, but that changes where logs go, not their format or which levels are sent.

  • CMove the log group to the Infrequent Access log class

    Incorrect

    The Infrequent Access class lowers ingestion cost with a subset of features. It neither formats logs as JSON nor filters them by level.

  • DSet the function's log format to JSON and its application log level to INFO

    Correct

    With the JSON log format, supported logging methods are captured as structured JSON, and log-level filtering sends only the chosen level and less detailed levels; changing the level later needs no code change.

Lambda's log format and log-level settings give structured, filtered logs without code changes for supported runtimes and logging methods.

Question 3 · choose 2

A team deploys a service to production with AWS CodeDeploy. The developers want an email every time a deployment to the production deployment group succeeds or fails, and they do not want to write code for it. Which TWO solutions would each meet this requirement? (Choose TWO.)

  1. AAn AWS CloudTrail trail that delivers CodeDeploy API activity to Amazon S3
  2. BA CloudWatch alarm on the service function's Errors metric with an SNS action
  3. CA deployment group trigger on success and failure events that publishes to SNS
  4. DAn EventBridge rule on CodeDeploy deployment state changes that targets SNS
  5. EAn AWS X-Ray group that filters traces for the deployment's version
Show the answer and why
  • AAn AWS CloudTrail trail that delivers CodeDeploy API activity to Amazon S3

    Incorrect

    CloudTrail records actions taken in the account as events for auditing. Delivering them to S3 sends no one an email when a deployment finishes.

  • BA CloudWatch alarm on the service function's Errors metric with an SNS action

    Incorrect

    The Errors metric counts failed invocations of the function. It says nothing about whether a deployment succeeded.

  • CA deployment group trigger on success and failure events that publishes to SNS

    Correct

    CodeDeploy triggers send notifications for deployment events to an SNS topic, and an email subscription on the topic delivers them.

  • DAn EventBridge rule on CodeDeploy deployment state changes that targets SNS

    Correct

    Event rules can react when a CodeDeploy deployment changes state and invoke targets such as SNS topics, which can email subscribers.

  • EAn AWS X-Ray group that filters traces for the deployment's version

    Incorrect

    X-Ray collects data about requests the application serves. It does not report deployment outcomes.

Deployment notifications come from the deployment service itself: CodeDeploy triggers or EventBridge rules on its state-change events, delivered by SNS.

Question 4 · choose 1

An Amazon ECS service on AWS Fargate runs a Java API behind an Application Load Balancer. Each new task needs about 120 seconds to warm up before its /health endpoint returns 200. During deployments, new tasks fail the load balancer health checks while warming up, ECS stops them as unhealthy, and the deployment never settles. The application itself is fine once it is warm. What should the developer change?

  1. ASet a health check grace period on the ECS service above the warm-up time
  2. BIncrease the target group's deregistration delay to 300 seconds
  3. CLower the target group's health check interval to 5 seconds
  4. DRemove the container health check from the task definition and redeploy
Show the answer and why
  • ASet a health check grace period on the ECS service above the warm-up time

    Correct

    During the grace period, the ECS service scheduler ignores unhealthy load balancer and container health checks for a newly started task, so it is not stopped while it warms up.

  • BIncrease the target group's deregistration delay to 300 seconds

    Incorrect

    The deregistration delay only controls how long the load balancer waits for in-flight requests while a target is deregistering.

  • CLower the target group's health check interval to 5 seconds

    Incorrect

    A shorter interval checks more often, so a warming task reaches the unhealthy threshold even sooner.

  • DRemove the container health check from the task definition and redeploy

    Incorrect

    The tasks are failing the load balancer's health checks, which ECS also acts on; removing the container check does not stop that.

Give slow-starting tasks a health check grace period instead of loosening or removing the checks that protect steady-state traffic.

Question 5 · choose 1

A developer publishes a custom CloudWatch metric, CheckoutLatency, and adds a UserId dimension so that latency can be broken down by user. The app has two million users, and CloudWatch costs rise sharply. Dashboards only need latency by Region and by payment method. What should the developer change?

  1. AKeep the UserId dimension and publish the metric less often
  2. BReplace UserId with Region and PaymentMethod dimensions
  3. CSwitch the metric to high resolution
  4. DMove the metric to a new namespace
Show the answer and why
  • AKeep the UserId dimension and publish the metric less often

    Incorrect

    Each user ID still creates its own variation of the metric, so the number of metrics stays the same.

  • BReplace UserId with Region and PaymentMethod dimensions

    Correct

    Every unique name/value pair creates a new variation of a metric, so low-cardinality dimensions that match the dashboards keep the metric count small.

  • CSwitch the metric to high resolution

    Incorrect

    High resolution adds detail in time and higher cost; it does not reduce the number of metric variations.

  • DMove the metric to a new namespace

    Incorrect

    The namespace does not change how many dimension combinations exist.

Choose metric dimensions with low cardinality; put per-user detail in logs instead of metric dimensions.

Question 6 · choose 1

A nightly job publishes a custom metric, JobCompleted, with a value of 1 when it finishes. The team creates an alarm for values below 1, but on a night when the job crashed and published nothing, the alarm stayed quiet. What should the developer change on the alarm?

  1. ATreat missing data as breaching
  2. BTreat missing data as not breaching
  3. CRaise the threshold to 2
  4. DShorten the alarm period to 10 seconds
Show the answer and why
  • ATreat missing data as breaching

    Correct

    With breaching, missing data points count as bad and breaching the threshold, so a job that publishes nothing raises the alarm.

  • BTreat missing data as not breaching

    Incorrect

    This treats missing points as good, which is why the alarm stayed quiet.

  • CRaise the threshold to 2

    Incorrect

    The threshold is compared with data points; with no data there is nothing to compare.

  • DShorten the alarm period to 10 seconds

    Incorrect

    A shorter period on a metric that sends no data still sees no data.

For heartbeat-style metrics, missing data is itself the failure signal.

Question 7 · choose 1

A web service runs on EC2 instances behind an Application Load Balancer. Sometimes an instance's application process hangs while the instance itself stays up, and the load balancer keeps sending it requests. The developer adds a /health endpoint that returns 200 only when the application can serve requests. What else should the developer configure?

  1. AA target group health check that uses the /health path
  2. BA CloudWatch alarm on the instances' CPUUtilization metric
  3. CSticky sessions on the target group
  4. DAccess logs for the load balancer
Show the answer and why
  • AA target group health check that uses the /health path

    Correct

    Load balancer nodes route requests only to healthy targets, based on the health check settings of the target group.

  • BA CloudWatch alarm on the instances' CPUUtilization metric

    Incorrect

    A hung process may use little CPU, and an alarm alone does not stop routing to the instance.

  • CSticky sessions on the target group

    Incorrect

    Stickiness keeps sending a client to the same target, which makes a hung target worse.

  • DAccess logs for the load balancer

    Incorrect

    Access logs record requests; they do not change where requests are routed.

Application-level health endpoints work only when the load balancer's health checks use them.

Question 8 · choose 1

Several AWS Lambda functions write JSON logs to CloudWatch Logs. The team wants every log event with "level":"ERROR" delivered to another Lambda function within seconds, which posts a summary to the team's chat tool. What should the developer configure on the log groups?

  1. AA metric filter that counts ERROR events
  2. BA daily export of the log groups to an Amazon S3 bucket for the summary job
  3. CA longer retention period on each log group
  4. DA subscription filter for ERROR events to the function
Show the answer and why
  • AA metric filter that counts ERROR events

    Incorrect

    A metric filter produces a count, not the log events themselves for the summary.

  • BA daily export of the log groups to an Amazon S3 bucket for the summary job

    Incorrect

    A daily export is far too slow for delivery within seconds.

  • CA longer retention period on each log group

    Incorrect

    Retention controls how long logs are kept, not where they are sent.

  • DA subscription filter for ERROR events to the function

    Correct

    Subscriptions deliver a real-time feed of matching log events to destinations such as a Lambda function.

Subscription filters stream matching log events to Lambda, Kinesis or Firehose for real-time processing.

Question 9 · choose 1

A traced service shows in AWS X-Ray that a request handler takes 2 seconds, but the trace has no detail about which part of the handler is slow. The developer suspects a local image-resizing step. How can the developer see that step's duration in each trace?

  1. AAdd an annotation with the request's user ID
  2. BRecord a custom subsegment around the image-resizing code
  3. CIncrease the sampling rate of the default rule
  4. DTurn on CloudWatch Logs Insights queries for the function's log group
Show the answer and why
  • AAdd an annotation with the request's user ID

    Incorrect

    Annotations add searchable values to a trace; they do not time a block of code.

  • BRecord a custom subsegment around the image-resizing code

    Correct

    A segment can break down the work done into subsegments, which provide more granular timing for each part of a request.

  • CIncrease the sampling rate of the default rule

    Incorrect

    More traces show the same missing detail more often.

  • DTurn on CloudWatch Logs Insights queries for the function's log group

    Incorrect

    Log queries analyze log text; they do not add timing to traces.

Wrap important local work in subsegments (spans in OpenTelemetry) so that traces show where time is spent.

Practise domain 4 →Practise all domains →