SOA-C03 · Domain 1: Monitoring, Logging, Analysis, Remediation, and Performance Optimization · 22% of the exam
Task 1.1: Implement metrics, alarms, and filters by using AWS monitoring and logging services.
Seeing what a workload does: CloudWatch metrics and the CloudWatch agent on instances and container clusters, log groups and metric filters, alarms and composite alarms with their actions, dashboards shared across accounts and Regions, CloudTrail, and notifications through SNS.
Study it
CloudWatch metrics, alarms, composite alarms and cross-account dashboards
Try each one before opening the answer. Every option is explained, with the AWS documentation page that proves it.
Question 1 · choose 1
An operations team runs a fleet of Amazon EC2 Linux instances. The team must alarm when memory utilization or used disk space on any instance passes 85%. Neither value appears among the EC2 metrics in Amazon CloudWatch. What should the team do?
AInstall the CloudWatch agent on the instances and configure it to collect memory and disk metrics
BTurn on detailed monitoring for every instance so that EC2 sends its metrics to CloudWatch in 1-minute periods
CCreate VPC flow logs for the subnets and add a metric filter to the flow log group
DCreate an Amazon EventBridge rule that matches EC2 instance state-change events
Show the answer and why
AInstall the CloudWatch agent on the instances and configure it to collect memory and disk metrics
Correct
The agent collects in-guest system metrics such as mem_used_percent and disk_used_percent and publishes them to CloudWatch, where alarms can watch them.
BTurn on detailed monitoring for every instance so that EC2 sends its metrics to CloudWatch in 1-minute periods
Incorrect
Detailed monitoring only shortens the period of the metrics EC2 already sends. It does not add memory or file system metrics.
CCreate VPC flow logs for the subnets and add a metric filter to the flow log group
Incorrect
Flow logs record IP traffic to and from network interfaces. They carry no memory or disk usage data.
DCreate an Amazon EventBridge rule that matches EC2 instance state-change events
Incorrect
State-change events report transitions such as running or stopped, not resource utilization inside the instance.
Memory and disk usage live inside the guest operating system, so EC2 cannot see them from the hypervisor. The CloudWatch agent collects them as custom metrics in the CWAgent namespace.
An on-call engineer is paged many times a night by two CloudWatch alarms: one on high CPU utilization and one on high memory utilization for the same web tier. Either condition alone is harmless. The team wants a single page only when both alarms are in the ALARM state at the same time, and it wants to keep both alarms. What should the team do?
ARaise the number of evaluation periods on each alarm so that short spikes no longer reach the ALARM state
BCreate a composite alarm with the rule ALARM(cpu) AND ALARM(memory), and set the notification only on it
CSend both alarms to one Amazon SNS topic and add a subscription filter policy on the pager endpoint
DAdd both alarms to a CloudWatch dashboard as alarm status widgets
Show the answer and why
ARaise the number of evaluation periods on each alarm so that short spikes no longer reach the ALARM state
Incorrect
Evaluation periods only set how many recent data points each alarm weighs. Each alarm still notifies on its own when its own condition holds.
BCreate a composite alarm with the rule ALARM(cpu) AND ALARM(memory), and set the notification only on it
Correct
A composite alarm combines the states of other alarms with Boolean logic. With actions only at the composite level, the team is notified only when both alarms are in ALARM.
CSend both alarms to one Amazon SNS topic and add a subscription filter policy on the pager endpoint
Incorrect
A filter policy decides which messages a subscriber receives, one message at a time. It cannot wait for a second alarm to fire.
DAdd both alarms to a CloudWatch dashboard as alarm status widgets
Incorrect
A dashboard displays alarm states. It does not change when or how anyone is notified.
Composite alarms exist to cut alarm noise: the underlying alarms keep their states without actions, and only the combined condition notifies anyone.
An application writes its logs to an Amazon CloudWatch Logs log group. The operations team must be notified within minutes when more than 20 log lines that contain the word ERROR arrive in 5 minutes. The team wants the solution with the least ongoing effort. What should the team do?
AExport the log group to Amazon S3 every hour and count the ERROR lines with Amazon Athena queries
BTurn on CloudTrail Insights events for the account and alarm on the Insights events
CCreate a metric filter on the log group for ERROR and a CloudWatch alarm on the new metric
DAdd a subscription filter that streams the log group to an Amazon Data Firehose stream that writes to S3
Show the answer and why
AExport the log group to Amazon S3 every hour and count the ERROR lines with Amazon Athena queries
Incorrect
Log data can take up to 12 hours to become available for export, and AWS does not recommend export for continuous processing. It cannot notify within minutes.
BTurn on CloudTrail Insights events for the account and alarm on the Insights events
Incorrect
CloudTrail Insights looks for unusual API call rates and API error rates. It does not read application log lines.
CCreate a metric filter on the log group for ERROR and a CloudWatch alarm on the new metric
Correct
A metric filter turns matching log events into a CloudWatch metric as they arrive, and an alarm on that metric can notify through Amazon SNS.
DAdd a subscription filter that streams the log group to an Amazon Data Firehose stream that writes to S3
Incorrect
A subscription delivers log events to another service, but delivery alone counts nothing and raises no alarm.
Counting a pattern in logs and alarming on the count is the job metric filters were built for: no code, no extra pipeline.
A company runs its workloads in six AWS accounts and in two AWS Regions. The operations team works from a central monitoring account and wants one CloudWatch dashboard that shows metrics from all six accounts and both Regions, without signing in to each account. What should the team set up?
AShare each account's dashboard by email with the operations engineers
BAn AWS CloudTrail organization trail that delivers events from every account to the monitoring account
CA CloudWatch Logs subscription filter in each account that sends log events to the monitoring account
DCross-account cross-Region sharing to the monitoring account, and one dashboard there
Show the answer and why
AShare each account's dashboard by email with the operations engineers
Incorrect
Dashboard sharing lets people outside the account view one account's dashboard. The team would still have six dashboards, not one.
BAn AWS CloudTrail organization trail that delivers events from every account to the monitoring account
Incorrect
CloudTrail records API activity. It does not deliver CloudWatch metrics or build dashboards.
CA CloudWatch Logs subscription filter in each account that sends log events to the monitoring account
Incorrect
A subscription filter forwards log events, not metrics, and does not make a cross-account dashboard.
DCross-account cross-Region sharing to the monitoring account, and one dashboard there
Correct
After each account enables sharing with the monitoring account, a single dashboard there can summarize metrics from many accounts and Regions.
CloudWatch dashboards can mix Regions, and cross-account sharing lets the monitoring account read the other accounts' metrics. Dashboard sharing is for people without account access, and CloudTrail and log subscriptions move other data.
A licensing server runs on a single Amazon EC2 instance that cannot be placed in an Auto Scaling group. A CloudWatch alarm watches the instance's StatusCheckFailed_System metric. When the alarm fires, the instance must move to healthy hardware automatically and the on-call team must get an email. Which actions should the alarm have? (Choose TWO.)
AAn EC2 recover action for the instance
BAn EC2 reboot action for the instance
CA Systems Manager Run Command action that restarts the service
DA notification to an Amazon SNS topic that has an email subscription
EAn action that starts an AWS Step Functions state machine
Show the answer and why
AAn EC2 recover action for the instance
Correct
The recover action is made for system status check failures: the instance is migrated to new hardware and keeps its instance ID, private IP addresses and Elastic IP addresses.
BAn EC2 reboot action for the instance
Incorrect
AWS recommends the reboot action for instance status check failures. A system check failure points at the underlying host, which the recover action addresses.
CA Systems Manager Run Command action that restarts the service
Incorrect
Run Command is not among the actions an alarm can take. The supported Systems Manager actions create OpsItems or incidents.
DA notification to an Amazon SNS topic that has an email subscription
Correct
An alarm can notify an SNS topic, and the topic delivers the message to its email subscribers, including the status of the recovery attempt.
EAn action that starts an AWS Step Functions state machine
Incorrect
Starting a state machine is not a CloudWatch alarm action; it would need an EventBridge rule on the alarm's state change.
For StatusCheckFailed_System the alarm's recover action moves the instance off the impaired host, and an SNS action tells people about it. Alarms act through SNS, Lambda, EC2, Auto Scaling, Systems Manager OpsItems and incidents, and investigations; anything else goes through EventBridge.
An application publishes a custom CloudWatch metric, PaymentErrors, only when a payment fails. An alarm fires when the sum is above 5 in 5 minutes. Most of the day the alarm shows INSUFFICIENT_DATA, which confuses the on-call team. The alarm should show OK whenever no errors are being reported. What should the team change?
AConfigure the alarm to treat missing data as notBreaching
BConfigure the alarm to treat missing data as breaching
CConfigure the alarm to treat missing data as ignore
DPublish PaymentErrors as a high-resolution metric with a granularity of one second
Show the answer and why
AConfigure the alarm to treat missing data as notBreaching
Correct
notBreaching counts missing data points as within the threshold. AWS recommends it for metrics that report data points only when an error occurs.
BConfigure the alarm to treat missing data as breaching
Incorrect
Breaching counts every missing data point as bad, so the alarm would go to ALARM whenever no payment fails.
CConfigure the alarm to treat missing data as ignore
Incorrect
Ignore keeps the current alarm state, so the alarm would not move to OK just because errors stopped being reported.
DPublish PaymentErrors as a high-resolution metric with a granularity of one second
Incorrect
Resolution changes how fine the data points are when they exist. With no errors there are still no data points at all.
The default treatment, missing, moves the alarm to INSUFFICIENT_DATA when the whole evaluation range is empty. For an errors-only metric, silence is good news, so notBreaching is the right choice.