Skip to content
BytePatterns

DEA-C01 · Domain 3: Data Operations and Support · 22% of the exam

Task 3.3: Maintain and monitor data pipelines

Keeping pipelines healthy and traceable: CloudWatch metrics, logs and alarms, CloudTrail for API calls, finding why a Glue or EMR job is slow or failing, and searching logs with Logs Insights, Athena and OpenSearch Service.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

Auditors require a record of which IAM principal read each object in a sensitive S3 bucket. The records must appear in the company's existing CloudTrail trail, which logs management events only. What should a data engineer do?

  1. AAdd data events for the bucket's objects to the trail
  2. BTurn on CloudTrail Insights events for the trail
  3. CTurn on S3 Versioning for the sensitive bucket
  4. DAdd an AWS Config rule that checks the bucket's logging settings
Show the answer and why
  • AAdd data events for the bucket's objects to the trail

    Correct

    Trails do not log data events by default. S3 object-level activity such as GetObject is a data event, and logging it costs extra.

  • BTurn on CloudTrail Insights events for the trail

    Incorrect

    Insights events flag unusual API call rates or error rates against a baseline. They do not record each object read.

  • CTurn on S3 Versioning for the sensitive bucket

    Incorrect

    Versioning keeps every version of an object. It records nothing about who read them.

  • DAdd an AWS Config rule that checks the bucket's logging settings

    Incorrect

    Config rules evaluate resource configurations. They do not record requests made against the bucket.

CloudTrail splits activity into management events (logged by default) and data events such as S3 object reads (opt-in, per resource).

Question 2 · choose 1

An AWS Glue Spark job has become slow. A data engineer suspects that a few tasks in one stage run far longer than the others and wants to see the stages, tasks and executors of each run without changing the script. What should the engineer turn on?

  1. AJob bookmarks, so that each run reads only the new data
  2. BCustom log group settings for the job's continuous logging
  3. CThe Spark UI for the job, with event logs stored in Amazon S3
  4. DData events in AWS CloudTrail for the job's S3 bucket
Show the answer and why
  • AJob bookmarks, so that each run reads only the new data

    Incorrect

    Bookmarks track processed data between runs. They give no view of how the tasks of a run performed.

  • BCustom log group settings for the job's continuous logging

    Incorrect

    These settings decide where driver and executor log lines go. Log lines are not a per-stage breakdown of task times.

  • CThe Spark UI for the job, with event logs stored in Amazon S3

    Correct

    With the Spark UI turned on, Glue writes Spark event logs to S3 and the console shows stages, tasks and executors, which is where uneven task times are visible.

  • DData events in AWS CloudTrail for the job's S3 bucket

    Incorrect

    Data events record object-level API calls. They do not describe Spark stages or tasks.

For Spark performance questions the Spark UI is the first tool: the Stage tab's event timeline and summary metrics show skew and stragglers. Glue observability metrics add CloudWatch metrics such as skewness.

Question 3 · choose 1

During an incident, a data engineer must find the 10 most frequent error messages across 30 Lambda log groups for the last 24 hours. The question is one-off, and nothing new should be built or deployed. What should the engineer use?

  1. AA metric filter on each log group for the word ERROR
  2. BA subscription filter that streams the logs to Amazon OpenSearch Service
  3. CAn export of the log groups to Amazon S3 queried later with Athena
  4. DA CloudWatch Logs Insights query across the log groups
Show the answer and why
  • AA metric filter on each log group for the word ERROR

    Incorrect

    Metric filters turn new log events into metrics from the time they are created. They do not look back at the last 24 hours or group messages.

  • BA subscription filter that streams the logs to Amazon OpenSearch Service

    Incorrect

    Subscriptions feed new log events to other services in real time. That builds a pipeline and does not cover the past day.

  • CAn export of the log groups to Amazon S3 queried later with Athena

    Incorrect

    This works but means exporting, defining a table and querying, which is more than a one-off question needs.

  • DA CloudWatch Logs Insights query across the log groups

    Correct

    Logs Insights interactively searches and analyzes log data already in CloudWatch Logs, across several log groups, with aggregations such as counts.

Ad hoc questions over logs already in CloudWatch belong to Logs Insights. Metric filters and subscriptions act on new events going forward.

Question 4 · choose 2

A Lambda function in a pipeline writes the line RECORD_REJECTED to its CloudWatch log group (Standard log class) for every record it rejects. The team must be notified by email when more than 100 records are rejected within 5 minutes. Which actions should a data engineer take? (Choose TWO.)

  1. ACreate a subscription filter that sends the log events to Amazon Data Firehose
  2. BCreate a metric filter on the log group that counts RECORD_REJECTED lines
  3. CCreate a CloudWatch alarm on that metric with an action that notifies an SNS topic
  4. DSave a CloudWatch Logs Insights query that counts the lines and add it to a dashboard
  5. ETurn on CloudTrail data events for the Lambda function
Show the answer and why
  • ACreate a subscription filter that sends the log events to Amazon Data Firehose

    Incorrect

    A subscription delivers the events to another service for processing. It does not count them or alert anyone.

  • BCreate a metric filter on the log group that counts RECORD_REJECTED lines

    Correct

    A metric filter turns matching log events into a CloudWatch metric that can be graphed or alarmed on.

  • CCreate a CloudWatch alarm on that metric with an action that notifies an SNS topic

    Correct

    The alarm compares the 5-minute sum with the threshold of 100, and its action publishes to an SNS topic with an email subscription.

  • DSave a CloudWatch Logs Insights query that counts the lines and add it to a dashboard

    Incorrect

    A dashboard shows the count when someone looks at it. It sends no notification.

  • ETurn on CloudTrail data events for the Lambda function

    Incorrect

    Lambda data events record Invoke API calls. They do not read the function's own log lines.

Log line to alert is a two-step pattern: a metric filter makes a metric from the log, and an alarm on the metric notifies through SNS.

Question 5 · choose 1

A data engineer must list every query that failed in the last 24 hours on an Amazon Redshift Serverless workgroup, together with the error text. Which source should the engineer query?

  1. AThe SVV_TABLE_INFO view, filtered on the tables that the queries used
  2. BThe SYS_QUERY_HISTORY view, filtered on status = 'failed'
  3. CThe connection log from Redshift database audit logging
  4. DAWS CloudTrail event history for the Redshift workgroup
Show the answer and why
  • AThe SVV_TABLE_INFO view, filtered on the tables that the queries used

    Incorrect

    SVV_TABLE_INFO summarizes tables for diagnosing table design issues. It has no query history.

  • BThe SYS_QUERY_HISTORY view, filtered on status = 'failed'

    Correct

    SYS_QUERY_HISTORY holds a row per user query, with a status that includes failed and an error_message column.

  • CThe connection log from Redshift database audit logging

    Incorrect

    The connection log records connection activity, not the outcome of each query.

  • DAWS CloudTrail event history for the Redshift workgroup

    Incorrect

    Event history records management events for AWS API calls, not the result of SQL queries inside the database.

SYS monitoring views are the place to look for query outcomes in Redshift; CloudTrail covers API calls around the warehouse, not the SQL inside it.

Question 6 · choose 1

Developers launch Amazon EMR clusters for ad hoc Spark work and often forget to terminate them, so idle clusters run for days. What should a data engineer configure so that idle clusters shut down on their own?

  1. ATermination protection on every new cluster
  2. BEMR managed scaling with a small minimum number of units
  3. CSpot Instances for all of the core nodes
  4. DAn auto-termination policy with an idle timeout
Show the answer and why
  • ATermination protection on every new cluster

    Incorrect

    Termination protection guards against accidental termination, which is the opposite of what is needed.

  • BEMR managed scaling with a small minimum number of units

    Incorrect

    Managed scaling changes the number of instances. The cluster itself keeps running while idle.

  • CSpot Instances for all of the core nodes

    Incorrect

    Spot Instances change the price of the capacity. An idle cluster still keeps running.

  • DAn auto-termination policy with an idle timeout

    Correct

    An auto-termination policy shuts a cluster down after it has been idle for the time you specify, without manual monitoring.

Idle clusters cost money for nothing; an auto-termination policy is the built-in cleanup for clusters that people forget.

Question 7 · choose 1

During a test run, a data engineer wants to watch new log events from three Lambda log groups as they arrive and show only the lines that contain one request ID. What should the engineer use?

  1. ACloudWatch Logs Live Tail
  2. BA metric filter on each of the three log groups
  3. CA subscription filter that sends the logs to Firehose
  4. DA scheduled CloudWatch Logs Insights query every hour
Show the answer and why
  • ACloudWatch Logs Live Tail

    Correct

    Live Tail shows a streaming list of new log events as they are ingested, and you can filter and highlight them by terms.

  • BA metric filter on each of the three log groups

    Incorrect

    Metric filters turn matching log data into numeric CloudWatch metrics. They do not display the log lines.

  • CA subscription filter that sends the logs to Firehose

    Incorrect

    Subscriptions deliver a real-time feed of log events to other services for processing, not to a person's screen.

  • DA scheduled CloudWatch Logs Insights query every hour

    Incorrect

    Logs Insights searches and analyzes stored log data. An hourly query does not show events as they arrive.

For live, interactive troubleshooting, Live Tail streams and filters new log events; Logs Insights is for querying what was already stored.

Question 8 · choose 1

A consumer group reads a topic on an Amazon MSK cluster. The team wants a CloudWatch alarm when the consumers fall more than 10 minutes behind the latest data. Which metric should the alarm use?

  1. AMaxOffsetLag across the partitions for the consumer group
  2. BEstimatedMaxTimeLag for the consumer group and topic
  3. CBytesInPerSec for the topic on each broker
  4. DCpuUser for each broker in the cluster
Show the answer and why
  • AMaxOffsetLag across the partitions for the consumer group

    Incorrect

    MaxOffsetLag counts offsets across partitions. It measures lag in messages, not in minutes.

  • BEstimatedMaxTimeLag for the consumer group and topic

    Correct

    EstimatedMaxTimeLag is a time estimate, in seconds, to drain the maximum offset lag, so it maps directly to a 10-minute threshold.

  • CBytesInPerSec for the topic on each broker

    Incorrect

    BytesInPerSec is the rate at which brokers receive data from clients. It says nothing about how far consumers lag.

  • DCpuUser for each broker in the cluster

    Incorrect

    CpuUser is the broker's CPU use in user space. It does not measure how far consumers lag.

Consumer lag metrics quantify how far readers trail the newest data; pick the time-based one when the requirement is stated in time.

Question 9 · choose 2

A pipeline stage reads from an Amazon SQS standard queue. Sometimes the consumer stops processing, and messages pile up unnoticed for hours. Which TWO CloudWatch metrics should a data engineer alarm on to detect a growing backlog? (Choose TWO.)

  1. ANumberOfMessagesSent
  2. BApproximateNumberOfMessagesNotVisible
  3. CApproximateAgeOfOldestMessage
  4. DApproximateNumberOfMessagesDelayed
  5. EApproximateNumberOfMessagesVisible
Show the answer and why
  • ANumberOfMessagesSent

    Incorrect

    This counts messages added to the queue. It reflects the producer, not whether anyone is consuming.

  • BApproximateNumberOfMessagesNotVisible

    Incorrect

    This counts in-flight messages that were received but not yet deleted. A stopped consumer receives nothing, so the count does not grow.

  • CApproximateAgeOfOldestMessage

    Correct

    This is the age of the oldest unprocessed message in the queue, which keeps rising while the consumer is stopped.

  • DApproximateNumberOfMessagesDelayed

    Incorrect

    This counts messages that are delayed and not yet available. A stopped consumer does not change it.

  • EApproximateNumberOfMessagesVisible

    Correct

    This is the number of messages available for retrieval, which reflects the current processing backlog.

Backlog shows up as more visible messages and an older oldest message; producer-side metrics stay normal while consumers are stuck.

Practise domain 3 →Practise all domains →