Skip to content
BytePatterns

DOP-C02 · Domain 4: Monitoring and Logging · 15% of the exam

Task 4.1: Configure the collection, aggregation, and storage of logs and metrics.

Getting telemetry in and keeping it safe: the CloudWatch agent and custom metrics, metric filters and metric streams, subscriptions to Kinesis, Lambda and OpenSearch Service, retention and lifecycle, and KMS encryption of log data.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 2

A compliance review of an application's Amazon CloudWatch Logs log groups finds two gaps. The log data must be encrypted with a customer managed KMS key that the security team controls and can disable, and log events older than 400 days must be deleted automatically. Today the log groups use the default encryption and never expire. Which actions should the DevOps engineer take? (Choose TWO.)

  1. AAssociate the customer managed KMS key with each log group, after allowing the CloudWatch Logs service principal to use the key in its key policy
  2. BSet the retention of each log group to 400 days
  3. CExport the log groups to an S3 bucket every day and turn on SSE-KMS with the customer managed key for the bucket
  4. DAttach a data protection policy to each log group so that log events are protected with the security team's KMS key
  5. EMove the log groups to the Infrequent Access log class so that old log events expire after 400 days
Show the answer and why
  • AAssociate the customer managed KMS key with each log group, after allowing the CloudWatch Logs service principal to use the key in its key policy

    Correct

    Encryption with AWS KMS is turned on per log group by associating a KMS key, and the key policy must give the CloudWatch Logs service principal permission to use the key.

  • BSet the retention of each log group to 400 days

    Correct

    400 days is one of the allowed retention values, and data older than the retention setting is deleted automatically.

  • CExport the log groups to an S3 bucket every day and turn on SSE-KMS with the customer managed key for the bucket

    Incorrect

    This encrypts the exported copies in S3, not the log data stored in CloudWatch Logs, and the log groups still never expire.

  • DAttach a data protection policy to each log group so that log events are protected with the security team's KMS key

    Incorrect

    A data protection policy does not change the KMS key that encrypts the log group's data.

  • EMove the log groups to the Infrequent Access log class so that old log events expire after 400 days

    Incorrect

    The log class does not decide when log events expire. Expiry is set by the retention setting.

CloudWatch Logs always encrypts data at rest; associating a customer managed KMS key with a log group puts the encryption under the key owner's control. Retention is a separate per-log-group setting with fixed allowed values.

Question 2 · choose 1

A company's central data platform must receive Amazon CloudWatch metrics from the AWS/EC2, AWS/RDS and AWS/ApplicationELB namespaces continuously, with low latency, and store them in Amazon S3 in Parquet format for analysis with Amazon Athena. The team wants no polling code to maintain. What should the DevOps engineer do?

  1. ASchedule a Lambda function every minute that calls GetMetricData for the three namespaces and writes the results to S3 as Parquet files with a date prefix
  2. BCreate a CloudWatch metric stream for the three namespaces to an Amazon Data Firehose stream that converts the records to Parquet in S3
  3. CAdd a CloudWatch Logs subscription filter for the three namespaces that sends the metric data points to the Firehose stream for S3
  4. DBuild a CloudWatch dashboard with the three namespaces and export its widgets to S3 as CSV files on a schedule
Show the answer and why
  • ASchedule a Lambda function every minute that calls GetMetricData for the three namespaces and writes the results to S3 as Parquet files with a date prefix

    Incorrect

    This is the polling code the team wants to avoid, and its latency and completeness depend on the schedule and the code.

  • BCreate a CloudWatch metric stream for the three namespaces to an Amazon Data Firehose stream that converts the records to Parquet in S3

    Correct

    Metric streams deliver metrics continually with low latency to Firehose, where a transformation can convert the data to formats such as Parquet before it lands in S3.

  • CAdd a CloudWatch Logs subscription filter for the three namespaces that sends the metric data points to the Firehose stream for S3

    Incorrect

    Subscription filters stream log events from log groups. Metrics in the AWS/EC2, AWS/RDS and AWS/ApplicationELB namespaces are not log events.

  • DBuild a CloudWatch dashboard with the three namespaces and export its widgets to S3 as CSV files on a schedule

    Incorrect

    Dashboards display metrics. They are not a continuous delivery mechanism for metric data to a data lake.

Metric streams are the push path out of CloudWatch. With the custom Firehose setup you choose the namespaces, and Firehose can transform the native JSON or OpenTelemetry records into Parquet for Athena.

Question 3 · choose 1

A team runs 300 Amazon EC2 Linux instances. It needs memory utilization for each instance, the CPU used by the java process on each instance, and the files under /var/log/app/ in Amazon CloudWatch Logs. The collection settings must live in one place, and changes must reach the whole fleet without anyone logging in to instances or writing collection code. What should the DevOps engineer do?

  1. ATurn on detailed monitoring for every instance so that EC2 publishes its metrics at one-minute intervals
  2. BTurn on Container Insights for the account so that CloudWatch collects memory and process metrics from the instances
  3. CAdd a cron job on each instance that reads memory and process CPU every minute and calls PutMetricData, and ship the log files to CloudWatch Logs with a script
  4. DKeep one CloudWatch agent configuration with procstat and the log files in Parameter Store, and roll the agent out with Systems Manager
Show the answer and why
  • ATurn on detailed monitoring for every instance so that EC2 publishes its metrics at one-minute intervals

    Incorrect

    Detailed monitoring changes the period of the metrics EC2 already sends. It adds neither memory nor per-process metrics, nor log files.

  • BTurn on Container Insights for the account so that CloudWatch collects memory and process metrics from the instances

    Incorrect

    Container Insights collects metrics and logs from containerized applications on ECS and EKS. These are plain EC2 instances.

  • CAdd a cron job on each instance that reads memory and process CPU every minute and calls PutMetricData, and ship the log files to CloudWatch Logs with a script

    Incorrect

    This is collection code installed on every instance, and changing it means touching every instance again.

  • DKeep one CloudWatch agent configuration with procstat and the log files in Parameter Store, and roll the agent out with Systems Manager

    Correct

    The agent collects memory, per-process metrics through procstat, and log files. Systems Manager installs it fleet-wide and applies the configuration that is kept in Parameter Store.

The CloudWatch agent covers in-guest metrics, process metrics and log files. Keeping its configuration in Parameter Store and rolling it out with Systems Manager gives one source of truth and no logins to individual servers.

Question 4 · choose 1

An API writes JSON access logs to Amazon CloudWatch Logs. Each event has the fields status and clientIp. During an incident, the on-call engineer needs the 10 client IP addresses that received the most HTTP 5xx responses in the last hour. Which CloudWatch Logs Insights query returns this result?

  1. Afilter status >= 500 | stats count(*) as errors by clientIp | sort errors desc | limit 10
  2. Bstats count(*) as errors by clientIp | filter status >= 500 | sort errors desc | limit 10
  3. Cfilter status >= 500 | sort errors desc | stats count(*) as errors by clientIp | limit 10
  4. Dfilter status >= 500 | sort @timestamp desc | limit 10
Show the answer and why
  • Afilter status >= 500 | stats count(*) as errors by clientIp | sort errors desc | limit 10

    Correct

    The filter keeps only 5xx events, stats counts them per IP, and sort with limit after the last stats returns the top 10.

  • Bstats count(*) as errors by clientIp | filter status >= 500 | sort errors desc | limit 10

    Incorrect

    After a stats command only the fields it defines can be referenced, so status is no longer available to filter on.

  • Cfilter status >= 500 | sort errors desc | stats count(*) as errors by clientIp | limit 10

    Incorrect

    A sort or limit command must come after the last stats command; otherwise the query is not valid.

  • Dfilter status >= 500 | sort @timestamp desc | limit 10

    Incorrect

    This returns the 10 most recent 5xx events, not the IP addresses with the most errors.

Logs Insights commands run in order, separated by pipes. Filter raw events first, aggregate with stats, then use sort and limit for a top-N result.

Question 5 · choose 1

A Java application writes stack traces that span many lines to a log file. The CloudWatch agent sends each line as a separate log event, so one error becomes dozens of events that are hard to search. Every log message starts with a timestamp. What should the DevOps engineer change in the agent configuration?

  1. ARaise the agent's metrics collection interval so that lines are batched together
  2. BSet multi_line_start_pattern to match the timestamp at the start of each message
  3. CSet the log stream name to the instance ID so that related lines are kept together
  4. DCreate a metric filter on the log group that merges lines with the same timestamp
Show the answer and why
  • ARaise the agent's metrics collection interval so that lines are batched together

    Incorrect

    The metrics interval controls metric collection, not how log lines are grouped.

  • BSet multi_line_start_pattern to match the timestamp at the start of each message

    Correct

    A message is a line that matches the pattern plus the following lines that do not, so each stack trace becomes one event.

  • CSet the log stream name to the instance ID so that related lines are kept together

    Incorrect

    The stream name does not combine lines into one event.

  • DCreate a metric filter on the log group that merges lines with the same timestamp

    Incorrect

    Metric filters create metrics; they do not merge log events.

The CloudWatch agent identifies the start of each log message with multi_line_start_pattern, so multi-line output such as stack traces is sent as single events.

Question 6 · choose 1

A Firehose stream delivers application logs to an OpenSearch Service domain. During monthly domain maintenance of up to 90 minutes, Firehose sends records to its S3 error bucket, and engineers must redrive them to OpenSearch afterwards. The team wants Firehose to keep retrying through the maintenance instead. What should the DevOps engineer change?

  1. ARaise the buffer size of the stream to the maximum value
  2. BTurn off the S3 error bucket so that no records leave the stream
  3. CSet the OpenSearch retry duration above 90 minutes, such as 7200 seconds
  4. DAdd index rotation by day so that the domain accepts writes during maintenance
Show the answer and why
  • ARaise the buffer size of the stream to the maximum value

    Incorrect

    Larger buffers do not keep Firehose retrying a failing destination.

  • BTurn off the S3 error bucket so that no records leave the stream

    Incorrect

    After retries end, records go to the error bucket rather than being held by the stream.

  • CSet the OpenSearch retry duration above 90 minutes, such as 7200 seconds

    Correct

    Firehose retries with backoff until the retry duration expires, and a higher value avoids delivering to the error bucket during downtime.

  • DAdd index rotation by day so that the domain accepts writes during maintenance

    Incorrect

    Index rotation names indexes by time; it does not help when the domain is unavailable.

For OpenSearch destinations, Firehose retries failed index requests for the retry duration, from 0 to 7200 seconds, and then sends data to the S3 error bucket, from which it must be redriven.

Question 7 · choose 1

A data team wants to analyze the individual requests made to a large S3 bucket each week, including the requester, operation, object key and HTTP status of every request, at the lowest possible cost. Records may arrive hours late, and the team does not need security-grade guarantees such as log file integrity validation. The security team already audits bucket-level actions with CloudTrail. What should the DevOps engineer turn on?

  1. ACloudTrail data events for all objects in the bucket on a new trail
  2. BS3 server access logging delivered to a separate logging bucket
  3. CS3 Storage Lens advanced metrics with a daily metrics export for the bucket
  4. DCloudTrail Insights on the existing trail
Show the answer and why
  • ACloudTrail data events for all objects in the bucket on a new trail

    Incorrect

    Data events record each object-level request within minutes, but they incur a fee per event, which this analysis does not need to pay.

  • BS3 server access logging delivered to a separate logging bucket

    Correct

    Server access logs record each request, cost nothing beyond log storage when delivered to S3, and arrive within a few hours.

  • CS3 Storage Lens advanced metrics with a daily metrics export for the bucket

    Incorrect

    Storage Lens shows aggregated usage and activity trends at low effort, but it does not record individual requests.

  • DCloudTrail Insights on the existing trail

    Incorrect

    Insights flags unusual API call rates; it does not record each request.

S3 offers server access logging and CloudTrail logging. AWS recommends CloudTrail for auditing, with integrity validation and faster delivery; server access logs are a cheaper record of each request, delivered within a few hours.

Question 8 · choose 1

A trading service publishes a custom queue-depth metric to CloudWatch once a minute. Spikes that last 10 to 20 seconds cause failed trades but do not show in the one-minute data. The team wants to see these spikes on a graph and to be alarmed within about 30 seconds of one starting, and it accepts higher CloudWatch charges for this one metric only. What should the DevOps engineer do?

  1. APublish it as a high-resolution metric every second and alarm on it with a 10-second period
  2. BTurn on detailed monitoring for the EC2 instances that run the service
  3. CPublish it every 10 seconds at standard resolution and alarm on the Maximum statistic over one minute
  4. DCreate an anomaly detection alarm on the existing one-minute metric
Show the answer and why
  • APublish it as a high-resolution metric every second and alarm on it with a 10-second period

    Correct

    High-resolution metrics are stored at 1-second resolution, can be read with periods down to 1 second, and support alarms with 10- or 30-second periods at a higher charge.

  • BTurn on detailed monitoring for the EC2 instances that run the service

    Incorrect

    Detailed monitoring sends EC2 metrics every minute; it changes neither the custom metric nor its resolution.

  • CPublish it every 10 seconds at standard resolution and alarm on the Maximum statistic over one minute

    Incorrect

    The one-minute Maximum would include the spike's value, but alarms on standard-resolution data use periods of at least 60 seconds, which is too slow.

  • DCreate an anomaly detection alarm on the existing one-minute metric

    Incorrect

    Anomaly detection replaces a static threshold with a learned band, but it still evaluates the one-minute values that hide the spikes.

Custom metrics can be standard resolution, with one-minute granularity, or high resolution, with one-second granularity. Only high-resolution metrics allow high-resolution alarms with 10- or 30-second periods, which cost more.

Question 9 · choose 1

An audit found hundreds of Lambda log groups that keep data forever, because Lambda creates each log group automatically the first time a function runs. Functions are deployed with CloudFormation through a pipeline. Logs must be kept for 90 days and no longer, every new function must comply from its first invocation, and the platform team does not want a separate job or remediation process to maintain. What should the DevOps engineer do?

  1. ADeploy the AWS Config managed rule cw-loggroup-retention-period-check with MinRetentionTime set to 90
  2. BExport all log groups to S3 every day and then delete the exported log events
  3. CTurn on log-level filtering at WARN for every function in the templates
  4. DDeclare each function's log group in its template with RetentionInDays set to 90
Show the answer and why
  • ADeploy the AWS Config managed rule cw-loggroup-retention-period-check with MinRetentionTime set to 90

    Incorrect

    The rule checks for a minimum retention, so it treats Never expire as compliant, and it only reports; it does not set retention.

  • BExport all log groups to S3 every day and then delete the exported log events

    Incorrect

    This adds a daily job to maintain instead of setting retention on the log groups.

  • CTurn on log-level filtering at WARN for every function in the templates

    Incorrect

    Filtering sends fewer log lines to CloudWatch Logs, but whatever is sent is still kept forever.

  • DDeclare each function's log group in its template with RetentionInDays set to 90

    Correct

    Defining the log group in the template creates it with its retention setting when the function is deployed, before the first invocation.

By default, log data is stored indefinitely, and Lambda creates a log group on first use. Declaring each log group with a retention period in infrastructure as code applies the setting from the start.

Practise domain 4 →Practise all domains →