Skip to content
BytePatterns

SOA-C03 · Domain 1: Monitoring, Logging, Analysis, Remediation, and Performance Optimization · 22% of the exam

Task 1.1: Implement metrics, alarms, and filters by using AWS monitoring and logging services.

Seeing what a workload does: CloudWatch metrics and the CloudWatch agent on instances and container clusters, log groups and metric filters, alarms and composite alarms with their actions, dashboards shared across accounts and Regions, CloudTrail, and notifications through SNS.

Study it

Sample questions

Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.

Question 1 · choose 1

An operations team runs a fleet of Amazon EC2 Linux instances. The team must alarm when memory utilization or used disk space on any instance passes 85%. Neither value appears among the EC2 metrics in Amazon CloudWatch. What should the team do?

  1. AInstall the CloudWatch agent on the instances and configure it to collect memory and disk metrics
  2. BTurn on detailed monitoring for every instance so that EC2 sends its metrics to CloudWatch in 1-minute periods
  3. CCreate VPC flow logs for the subnets and add a metric filter to the flow log group
  4. DCreate an Amazon EventBridge rule that matches EC2 instance state-change events
Show the answer and why
  • AInstall the CloudWatch agent on the instances and configure it to collect memory and disk metrics

    Correct

    The agent collects in-guest system metrics such as mem_used_percent and disk_used_percent and publishes them to CloudWatch, where alarms can watch them.

  • BTurn on detailed monitoring for every instance so that EC2 sends its metrics to CloudWatch in 1-minute periods

    Incorrect

    Detailed monitoring only shortens the period of the metrics EC2 already sends. It does not add memory or file system metrics.

  • CCreate VPC flow logs for the subnets and add a metric filter to the flow log group

    Incorrect

    Flow logs record IP traffic to and from network interfaces. They carry no memory or disk usage data.

  • DCreate an Amazon EventBridge rule that matches EC2 instance state-change events

    Incorrect

    State-change events report transitions such as running or stopped, not resource utilization inside the instance.

Memory and disk usage live inside the guest operating system, so EC2 cannot see them from the hypervisor. The CloudWatch agent collects them as custom metrics in the CWAgent namespace.

Question 2 · choose 1

An on-call engineer is paged many times a night by two CloudWatch alarms: one on high CPU utilization and one on high memory utilization for the same web tier. Either condition alone is harmless. The team wants a single page only when both alarms are in the ALARM state at the same time, and it wants to keep both alarms. What should the team do?

  1. ARaise the number of evaluation periods on each alarm so that short spikes no longer reach the ALARM state
  2. BCreate a composite alarm with the rule ALARM(cpu) AND ALARM(memory), and set the notification only on it
  3. CSend both alarms to one Amazon SNS topic and add a subscription filter policy on the pager endpoint
  4. DAdd both alarms to a CloudWatch dashboard as alarm status widgets
Show the answer and why
  • ARaise the number of evaluation periods on each alarm so that short spikes no longer reach the ALARM state

    Incorrect

    Evaluation periods only set how many recent data points each alarm weighs. Each alarm still notifies on its own when its own condition holds.

  • BCreate a composite alarm with the rule ALARM(cpu) AND ALARM(memory), and set the notification only on it

    Correct

    A composite alarm combines the states of other alarms with Boolean logic. With actions only at the composite level, the team is notified only when both alarms are in ALARM.

  • CSend both alarms to one Amazon SNS topic and add a subscription filter policy on the pager endpoint

    Incorrect

    A filter policy decides which messages a subscriber receives, one message at a time. It cannot wait for a second alarm to fire.

  • DAdd both alarms to a CloudWatch dashboard as alarm status widgets

    Incorrect

    A dashboard displays alarm states. It does not change when or how anyone is notified.

Composite alarms exist to cut alarm noise: the underlying alarms keep their states without actions, and only the combined condition notifies anyone.

Question 3 · choose 1

An application writes its logs to an Amazon CloudWatch Logs log group. The operations team must be notified within minutes when more than 20 log lines that contain the word ERROR arrive in 5 minutes. The team wants the solution with the least ongoing effort. What should the team do?

  1. AExport the log group to Amazon S3 every hour and count the ERROR lines with Amazon Athena queries
  2. BTurn on CloudTrail Insights events for the account and alarm on the Insights events
  3. CCreate a metric filter on the log group for ERROR and a CloudWatch alarm on the new metric
  4. DAdd a subscription filter that streams the log group to an Amazon Data Firehose stream that writes to S3
Show the answer and why
  • AExport the log group to Amazon S3 every hour and count the ERROR lines with Amazon Athena queries

    Incorrect

    Log data can take up to 12 hours to become available for export, and AWS does not recommend export for continuous processing. It cannot notify within minutes.

  • BTurn on CloudTrail Insights events for the account and alarm on the Insights events

    Incorrect

    CloudTrail Insights looks for unusual API call rates and API error rates. It does not read application log lines.

  • CCreate a metric filter on the log group for ERROR and a CloudWatch alarm on the new metric

    Correct

    A metric filter turns matching log events into a CloudWatch metric as they arrive, and an alarm on that metric can notify through Amazon SNS.

  • DAdd a subscription filter that streams the log group to an Amazon Data Firehose stream that writes to S3

    Incorrect

    A subscription delivers log events to another service, but delivery alone counts nothing and raises no alarm.

Counting a pattern in logs and alarming on the count is the job metric filters were built for: no code, no extra pipeline.

Question 4 · choose 1

A company runs its workloads in six AWS accounts and in two AWS Regions. The operations team works from a central monitoring account and wants one CloudWatch dashboard that shows metrics from all six accounts and both Regions, without signing in to each account. What should the team set up?

  1. AShare each account's dashboard by email with the operations engineers
  2. BAn AWS CloudTrail organization trail that delivers events from every account to the monitoring account
  3. CA CloudWatch Logs subscription filter in each account that sends log events to the monitoring account
  4. DCross-account cross-Region sharing to the monitoring account, and one dashboard there
Show the answer and why
  • AShare each account's dashboard by email with the operations engineers

    Incorrect

    Dashboard sharing lets people outside the account view one account's dashboard. The team would still have six dashboards, not one.

  • BAn AWS CloudTrail organization trail that delivers events from every account to the monitoring account

    Incorrect

    CloudTrail records API activity. It does not deliver CloudWatch metrics or build dashboards.

  • CA CloudWatch Logs subscription filter in each account that sends log events to the monitoring account

    Incorrect

    A subscription filter forwards log events, not metrics, and does not make a cross-account dashboard.

  • DCross-account cross-Region sharing to the monitoring account, and one dashboard there

    Correct

    After each account enables sharing with the monitoring account, a single dashboard there can summarize metrics from many accounts and Regions.

CloudWatch dashboards can mix Regions, and cross-account sharing lets the monitoring account read the other accounts' metrics. Dashboard sharing is for people without account access, and CloudTrail and log subscriptions move other data.

Question 5 · choose 2

A licensing server runs on a single Amazon EC2 instance that cannot be placed in an Auto Scaling group. A CloudWatch alarm watches the instance's StatusCheckFailed_System metric. When the alarm fires, the instance must move to healthy hardware automatically and the on-call team must get an email. Which actions should the alarm have? (Choose TWO.)

  1. AAn EC2 recover action for the instance
  2. BAn EC2 reboot action for the instance
  3. CA Systems Manager Run Command action that restarts the service
  4. DA notification to an Amazon SNS topic that has an email subscription
  5. EAn action that starts an AWS Step Functions state machine
Show the answer and why
  • AAn EC2 recover action for the instance

    Correct

    The recover action is made for system status check failures: the instance is migrated to new hardware and keeps its instance ID, private IP addresses and Elastic IP addresses.

  • BAn EC2 reboot action for the instance

    Incorrect

    AWS recommends the reboot action for instance status check failures. A system check failure points at the underlying host, which the recover action addresses.

  • CA Systems Manager Run Command action that restarts the service

    Incorrect

    Run Command is not among the actions an alarm can take. The supported Systems Manager actions create OpsItems or incidents.

  • DA notification to an Amazon SNS topic that has an email subscription

    Correct

    An alarm can notify an SNS topic, and the topic delivers the message to its email subscribers, including the status of the recovery attempt.

  • EAn action that starts an AWS Step Functions state machine

    Incorrect

    Starting a state machine is not a CloudWatch alarm action; it would need an EventBridge rule on the alarm's state change.

For StatusCheckFailed_System the alarm's recover action moves the instance off the impaired host, and an SNS action tells people about it. Alarms act through SNS, Lambda, EC2, Auto Scaling, Systems Manager OpsItems and incidents, and investigations; anything else goes through EventBridge.

Question 6 · choose 1

An application publishes a custom CloudWatch metric, PaymentErrors, only when a payment fails. An alarm fires when the sum is above 5 in 5 minutes. Most of the day the alarm shows INSUFFICIENT_DATA, which confuses the on-call team. The alarm should show OK whenever no errors are being reported. What should the team change?

  1. AConfigure the alarm to treat missing data as notBreaching
  2. BConfigure the alarm to treat missing data as breaching
  3. CConfigure the alarm to treat missing data as ignore
  4. DPublish PaymentErrors as a high-resolution metric with a granularity of one second
Show the answer and why
  • AConfigure the alarm to treat missing data as notBreaching

    Correct

    notBreaching counts missing data points as within the threshold. AWS recommends it for metrics that report data points only when an error occurs.

  • BConfigure the alarm to treat missing data as breaching

    Incorrect

    Breaching counts every missing data point as bad, so the alarm would go to ALARM whenever no payment fails.

  • CConfigure the alarm to treat missing data as ignore

    Incorrect

    Ignore keeps the current alarm state, so the alarm would not move to OK just because errors stopped being reported.

  • DPublish PaymentErrors as a high-resolution metric with a granularity of one second

    Incorrect

    Resolution changes how fine the data points are when they exist. With no errors there are still no data points at all.

The default treatment, missing, moves the alarm to INSUFFICIENT_DATA when the whole evaluation range is empty. For an errors-only metric, silence is good news, so notBreaching is the right choice.

Practise domain 1 →Practise all domains →