Skip to content
BytePatterns

DOP-C02 · Domain 4: Monitoring and Logging · 15% of the exam

Task 4.2: Audit, monitor, and analyze logs and metrics to detect issues.

Finding problems in the data: alarms on standard, custom and anomaly detection metrics, dashboards, X-Ray traces, Logs Insights and Athena queries, CloudTrail events, AWS Config rules and Amazon Inspector findings.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An order API's request count follows a strong daily and weekly pattern: heavy during business hours, low at night and at weekends. A static CloudWatch alarm on low request counts fires every night, and a static alarm on high counts misses daytime drops that still sit above the night-time level. The team wants one alarm that fires when the request count leaves its normal range for that time of day or week, in either direction, without hand-tuned thresholds. What should the DevOps engineer configure?

  1. AA composite alarm that combines the existing low-count and high-count alarms with an OR rule
  2. BA metric math expression that averages the request count over the last 24 hours, with a static alarm on the result of that expression
  3. CAn anomaly detection alarm on the request count metric that fires when the value is outside the band in either direction
  4. DTwo scheduled EventBridge Scheduler jobs that change the static alarm threshold at 08:00 and 20:00 on weekdays
Show the answer and why
  • AA composite alarm that combines the existing low-count and high-count alarms with an OR rule

    Incorrect

    A composite alarm combines the states of other alarms. The underlying static thresholds still ignore the time of day.

  • BA metric math expression that averages the request count over the last 24 hours, with a static alarm on the result of that expression

    Incorrect

    A 24-hour average flattens the daily pattern and still needs a fixed threshold, so short drops during busy hours are hidden.

  • CAn anomaly detection alarm on the request count metric that fires when the value is outside the band in either direction

    Correct

    Anomaly detection models account for hourly, daily and weekly seasonality, and the alarm can fire above the band, below it, or both.

  • DTwo scheduled EventBridge Scheduler jobs that change the static alarm threshold at 08:00 and 20:00 on weekdays

    Incorrect

    Switching thresholds by schedule is the hand tuning the team wants to avoid, and it does not cover weekends or gradual changes.

Seasonal metrics need a moving baseline. A CloudWatch anomaly detection alarm compares each data point with the expected range from the model instead of a static threshold.

Question 2 · choose 1

Requests to an Amazon API Gateway REST API invoke a Lambda function, which calls a service on Amazon ECS with AWS Fargate, which reads from Amazon DynamoDB. Some requests take over 3 seconds, and the team cannot tell which hop adds the time. The team wants an end-to-end view of each traced request across all of these services and prefers the instrumentation that AWS recommends for new work. What should the DevOps engineer do?

  1. ATurn on access logging for the API stage and use CloudWatch Logs Insights to compare the integration latency with the total latency
  2. BTurn on VPC Flow Logs for the Fargate tasks' subnets and look for slow connections between the tasks and DynamoDB
  3. CInstall the X-Ray daemon on each Fargate task host and use the X-Ray SDK in the ECS service, leaving tracing turned off in API Gateway and in the Lambda function
  4. DTurn on X-Ray tracing for the API stage and the Lambda function, and run the AWS Distro for OpenTelemetry collector as a sidecar in the ECS tasks
Show the answer and why
  • ATurn on access logging for the API stage and use CloudWatch Logs Insights to compare the integration latency with the total latency

    Incorrect

    Access logs show latency at the API only. They cannot split the time among Lambda, the ECS service and DynamoDB.

  • BTurn on VPC Flow Logs for the Fargate tasks' subnets and look for slow connections between the tasks and DynamoDB

    Incorrect

    Flow logs record IP traffic metadata, not request timings inside applications, so they cannot attribute latency to a hop.

  • CInstall the X-Ray daemon on each Fargate task host and use the X-Ray SDK in the ECS service, leaving tracing turned off in API Gateway and in the Lambda function

    Incorrect

    With Fargate there are no servers to manage, the X-Ray SDKs and daemon are in maintenance mode, and leaving tracing off upstream breaks the end-to-end view.

  • DTurn on X-Ray tracing for the API stage and the Lambda function, and run the AWS Distro for OpenTelemetry collector as a sidecar in the ECS tasks

    Correct

    API Gateway and Lambda can send traces to X-Ray, and ECS uses an ADOT sidecar container to collect and route trace data to X-Ray. AWS recommends OpenTelemetry for new instrumentation.

X-Ray shows the latency of an entire request and of each downstream service it traced. With the X-Ray SDKs and daemon in maintenance mode since February 2026, OpenTelemetry through ADOT is the recommended way to instrument containers.

Question 3 · choose 1

Application Load Balancer access logs for the last year are stored in Amazon S3 under the default date-based prefixes. An Amazon Athena table over the logs has no partitions, so every investigation scans the whole year, and queries are slow and costly. Engineers usually query one or two days. The team wants queries to read only the days they ask for, including days added in the future, without running jobs or code to maintain partitions. What should the DevOps engineer do?

  1. APartition the table by day and add a scheduled Lambda function that runs ALTER TABLE ADD PARTITION for each new day
  2. BRecreate the table with partition projection on the date, using the known prefix structure of ALB access logs
  3. CKeep the table as it is and add a WHERE condition on the request time column to every query
  4. DSend the access logs to CloudWatch Logs instead and query them with CloudWatch Logs Insights
Show the answer and why
  • APartition the table by day and add a scheduled Lambda function that runs ALTER TABLE ADD PARTITION for each new day

    Incorrect

    This prunes partitions, but it is code that must run for every new day, which the team wants to avoid.

  • BRecreate the table with partition projection on the date, using the known prefix structure of ALB access logs

    Correct

    Because ALB logs have a known partition scheme, projection lets Athena compute partitions as new data arrives, with no need for ALTER TABLE ADD PARTITION.

  • CKeep the table as it is and add a WHERE condition on the request time column to every query

    Incorrect

    Partitions are what restrict the data a query scans. Without them, a condition on a column does not keep Athena from reading the whole year.

  • DSend the access logs to CloudWatch Logs instead and query them with CloudWatch Logs Insights

    Incorrect

    ALB access logs are delivered to Amazon S3. Moving them into another service needs a separate pipeline and does not use the existing year of data.

Partition projection calculates partition values and locations from table properties instead of reading them from the Data Catalog. For logs with a predictable date layout, such as ALB access logs, it also removes partition maintenance.

Question 4 · choose 1

After a release, a service logs about two million error lines an hour with request IDs, user IDs and timestamps that differ in every line. Engineers want to see quickly which few kinds of error messages make up most of the volume across all matching events. What should the DevOps engineer use?

  1. AA metric filter for each error message text so that counts appear as metrics
  2. BLive Tail on the log group so that engineers can read the errors as they arrive
  3. CA Logs Insights query that filters the error lines and appends the pattern command
  4. DA subscription filter that sends the errors to a Lambda function that counts identical lines
Show the answer and why
  • AA metric filter for each error message text so that counts appear as metrics

    Incorrect

    Writing a filter for every message requires knowing the messages first.

  • BLive Tail on the log group so that engineers can read the errors as they arrive

    Incorrect

    Reading events as they arrive does not group millions of them.

  • CA Logs Insights query that filters the error lines and appends the pattern command

    Correct

    The pattern command finds shared text structures across the matching events, compressing many events into a few patterns.

  • DA subscription filter that sends the errors to a Lambda function that counts identical lines

    Incorrect

    Lines differ by IDs and timestamps, so identical-line counts miss the common message structure.

Logs Insights pattern analysis groups log events with shared structure and variable parts, which makes large error sets easier to understand.

Question 5 · choose 1

An application runs in three accounts and two Regions. The operations team works from a central monitoring account and wants one CloudWatch dashboard that graphs key metrics and alarms from all six account-Region combinations, without signing in and out of accounts or switching Regions. The metrics must stay in CloudWatch in their own accounts rather than be copied to another store. What should the DevOps engineer create?

  1. AA separate dashboard in each account and Region, linked from a wiki page that the operators use as a start page
  2. BA cross-account cross-Region dashboard in the monitoring account, with the other accounts set up as sharing accounts
  3. CA metric stream in each account that sends metrics to an S3 bucket in the monitoring account for Athena queries
  4. DA composite alarm in each account that combines the alarms for the key metrics
Show the answer and why
  • AA separate dashboard in each account and Region, linked from a wiki page that the operators use as a start page

    Incorrect

    Each team keeps its own dashboard without any sharing setup, but operators still switch accounts and Regions to see them.

  • BA cross-account cross-Region dashboard in the monitoring account, with the other accounts set up as sharing accounts

    Correct

    These dashboards summarize CloudWatch data from multiple accounts and Regions in one place, with widgets that read each account's data where it is.

  • CA metric stream in each account that sends metrics to an S3 bucket in the monitoring account for Athena queries

    Incorrect

    Metric streams centralize the data for analysis, but they copy it out of CloudWatch and Athena queries are not a live dashboard.

  • DA composite alarm in each account that combines the alarms for the key metrics

    Incorrect

    Composite alarms reduce alarm noise by combining alarm states; they do not graph metrics from several accounts.

CloudWatch cross-account cross-Region dashboards give a single view of data from many accounts and Regions, with drill-down into more specific dashboards, once sharing and monitoring accounts are set up.

Question 6 · choose 1

VPC flow logs from 50 VPCs are delivered to an S3 bucket. A few times a month a security analyst needs to find, with SQL, which source and destination address pairs sent the most rejected traffic during a given week. The company wants nothing to run or pay for between these queries, and the flow logs must stay in S3 rather than be copied into another data store. What should the DevOps engineer set up?

  1. AAn Athena table over the flow logs in S3, queried with SQL for rejected traffic
  2. BAn Amazon OpenSearch Service domain that Firehose loads with the flow logs, searched from OpenSearch Dashboards
  3. CAn Amazon Redshift provisioned cluster that loads the flow logs every night with COPY
  4. DReachability Analyzer runs between every pair of addresses
Show the answer and why
  • AAn Athena table over the flow logs in S3, queried with SQL for rejected traffic

    Correct

    Athena runs standard SQL directly on data in S3, is serverless and charges only for the queries that run.

  • BAn Amazon OpenSearch Service domain that Firehose loads with the flow logs, searched from OpenSearch Dashboards

    Incorrect

    OpenSearch Service suits continuous log analytics, but the domain runs and is paid for by the hour, and the logs are copied into it.

  • CAn Amazon Redshift provisioned cluster that loads the flow logs every night with COPY

    Incorrect

    Redshift suits a busy data warehouse, but its nodes are billed while the cluster runs, and the logs are loaded into the cluster.

  • DReachability Analyzer runs between every pair of addresses

    Incorrect

    Reachability Analyzer checks configured paths; it does not analyze recorded traffic.

Athena can create tables over VPC flow logs stored in S3, including partitioned and Parquet layouts, and run SQL on them on demand, for example to find top talkers or rejected traffic.

Question 7 · choose 1

Users of an internet-facing application in a few countries report slow responses, but the application's own metrics and load balancer metrics look normal. The team suspects problems with internet service providers between users and AWS, and wants to see performance and availability by location and network. What should the DevOps engineer use?

  1. ACloudWatch Internet Monitor for the application's resources
  2. BVPC Flow Logs on the load balancer subnets
  3. CX-Ray tracing on the backend services
  4. DA CloudWatch alarm on the load balancer's TargetResponseTime metric
Show the answer and why
  • ACloudWatch Internet Monitor for the application's resources

    Correct

    Internet Monitor shows how internet issues affect performance and availability between AWS-hosted applications and end users.

  • BVPC Flow Logs on the load balancer subnets

    Incorrect

    Flow logs show traffic inside the VPC, not provider problems on the internet path.

  • CX-Ray tracing on the backend services

    Incorrect

    Backend traces do not show internet path issues.

  • DA CloudWatch alarm on the load balancer's TargetResponseTime metric

    Incorrect

    Target response time is measured inside AWS and already looks normal.

Internet Monitor uses connectivity data from AWS's global network to set a baseline and raises health events when end users in certain locations and networks are affected.

Question 8 · choose 1

An API behind an Application Load Balancer has an average latency near 80 ms, yet about 2 percent of requests take more than 3 seconds and customers complain. An existing alarm on the average TargetResponseTime never fires. The team wants an alarm that fires when the slowest few percent of requests are slow, but not when only a single request among the tens of thousands in a period is slow. What should the DevOps engineer change?

  1. ALower the threshold of the average latency alarm to 50 ms
  2. BAlarm on the Maximum statistic of TargetResponseTime above 3 seconds
  3. CAlarm on the trimmed mean tm99 of TargetResponseTime above 1 second
  4. DAlarm on the p99 statistic of TargetResponseTime above 3 seconds
Show the answer and why
  • ALower the threshold of the average latency alarm to 50 ms

    Incorrect

    A lower average threshold fires on normal traffic and still hides the tail.

  • BAlarm on the Maximum statistic of TargetResponseTime above 3 seconds

    Incorrect

    Maximum shows the slowest request, but a single slow request in the period would set off the alarm.

  • CAlarm on the trimmed mean tm99 of TargetResponseTime above 1 second

    Incorrect

    A trimmed mean is a robust view of the typical request, but tm99 drops the slowest 1 percent and averages the rest, so the tail barely moves it.

  • DAlarm on the p99 statistic of TargetResponseTime above 3 seconds

    Correct

    p99 is the value that 99 percent of requests stay below, so it rises above 3 seconds when more than 1 percent of requests are that slow, and not for one outlier.

Averages hide outliers, and Maximum reacts to any single one. A percentile such as p99 shows how slow the worst experiences are while ignoring the extreme 1 percent, and ALB metrics support percentile statistics.

Practise domain 4 →Practise all domains →