CloudWatch Metrics, Alarms, Logs and X-Ray Explained
9 min readBytePatterns
CloudWatch alarms, Logs and X-Ray in one incident: M out of N alarm evaluation, log groups and metric filters, trace segments, sampling, and the SDK status.
"How would you know your service is broken, and how would you find out why?" Weak answers list services. The strong answer is an order of questions. Metrics tell you that something is wrong, traces tell you where, and logs tell you why. On AWS those are CloudWatch metrics and alarms, X-Ray traces, and CloudWatch Logs. Everything below comes from the AWS documentation pages listed at the end, as of September 2026.
The problem it solves
A request goes through a load balancer, an application and a database. After a deploy, some requests start failing. You need detection that does not wait for a customer and does not page anyone over one noisy minute, then the location of the failing hop, then the cause: the actual error message. No single tool answers all three well, which is why they come as a set.
The intuition
Metrics are numbers per period, named by a namespace and dimensions, such as a load balancer's 5XX count.
An alarm watches one metric over several periods. It is defined by three settings. The Period is the length of each datapoint in seconds. Evaluation Periods, N, is how many of the most recent datapoints to look at. Datapoints to Alarm, M, is how many of those must breach the threshold. The breaching datapoints do not have to be consecutive; they only have to fall within the last N. An alarm is in one of three states, OK, ALARM or INSUFFICIENT_DATA, and its actions can notify an SNS topic or scale a group. High-resolution alarms can use 10, 20 or 30 second periods, at a higher charge.
Logs are organised by source. A log event is a timestamp and a raw message; a log stream is the events from one source, such as one instance; a log group collects streams that share retention, monitoring and access control settings. Expired events are deleted automatically. Metric filters turn matching log events into metric datapoints, so an error line can feed an alarm. Logs Insights runs queries over log groups.
Traces follow one request. Each service sends X-Ray a segment describing its work, and a segment can contain subsegments for downstream calls such as a database query. Segments that share a trace ID form a trace. The ID travels between services in the X-Amzn-Trace-Id header. By default the X-Ray SDK records the first request each second and five percent of any additional requests, so a trace is a sample, not a record of everything. Annotations are indexed key-value pairs you can filter traces by, up to 50 per trace; metadata is stored but not indexed.
A note for new designs: the X-Ray SDKs and daemon entered maintenance mode on February 25, 2026, receiving security fixes only, and AWS recommends OpenTelemetry for instrumenting applications and sending traces to X-Ray.
Watch it run
The animation walks one incident, asking in order: is it broken, where, and why. The load balancer publishes its 5XX count to CloudWatch each period, and the baseline is 0. After a bad deploy, errors climb to 14: one of three datapoints over a threshold of 10. The alarm looks at several periods, so one noisy minute is not an incident; the count reaches 22, two of three. At 31, the third period in breach, the alarm goes from OK to ALARM, and its action publishes to an SNS topic that pages whoever is on call. The metric says that. A trace says where: X-Ray stitches one request's segments across every service it touched, and the time is going into a database call that now times out. The logs say why: a Logs Insights query over the app's log group, filtered to that trace, finds "pool exhausted". Roll back, the next periods are clean, and the alarm returns to OK on its own.
CloudWatch, Alarms & X-Ray
Step 1 of 11
Every incident asks three things in order: is it broken, where, and why. Each has its own tool.
The same interactive animation as the lesson — step through it with the controls.
The code
The lesson's alarm, as the CLI call that creates it:
# illustrative — load balancer name and topic ARN are placeholders
aws cloudwatch put-metric-alarm --alarm-name api-5xx \
--namespace AWS/ApplicationELB --metric-name HTTPCode_Target_5XX_Count \
--dimensions Name=LoadBalancer,Value=app/web/1234567890abcdef \
--statistic Sum --period 60 --evaluation-periods 3 --threshold 10 \
--comparison-operator GreaterThanOrEqualToThreshold \
--alarm-actions arn:aws:sns:us-east-1:123456789012:oncall
A toy model of M out of N evaluation, not the CloudWatch service. It ignores missing data and evaluation timing, and treats fewer than N datapoints as INSUFFICIENT_DATA:
from collections import deque
class ToyAlarm:
"""Toy model of CloudWatch metric alarm evaluation, as documented in September 2026."""
def __init__(self, threshold, datapoints_to_alarm, evaluation_periods):
self.threshold = threshold
self.m, self.n = datapoints_to_alarm, evaluation_periods
self.window = deque(maxlen=evaluation_periods) # the last N datapoints
self.breaching = 0
self.state = "INSUFFICIENT_DATA"
self.actions = []
def add(self, value):
if len(self.window) == self.n:
self.breaching -= self.window[0] >= self.threshold # oldest leaves
self.window.append(value)
self.breaching += value >= self.threshold
if len(self.window) < self.n:
new = "INSUFFICIENT_DATA"
elif self.breaching >= self.m:
new = "ALARM"
else:
new = "OK"
if new != self.state:
self.actions.append(f"{self.state}->{new}") # actions fire on a change
self.state = new
return new
alarm = ToyAlarm(threshold=10, datapoints_to_alarm=3, evaluation_periods=3)
five_xx = [0, 0, 0, 14, 22, 31, 0, 0, 0]
print([alarm.add(v)[0] for v in five_xx])
# ['I', 'I', 'O', 'O', 'O', 'A', 'O', 'O', 'O']
print(alarm.actions)
# ['INSUFFICIENT_DATA->OK', 'OK->ALARM', 'ALARM->OK']
noisy = [0, 12, 0, 0, 15, 0, 11, 0]
for m in (3, 2):
a = ToyAlarm(10, m, 3)
print(m, [a.add(v)[0] for v in noisy])
# 3 ['I', 'I', 'O', 'O', 'O', 'O', 'O', 'O']
# 2 ['I', 'I', 'O', 'O', 'O', 'O', 'A', 'O']
With 2 out of 3, 15 and 11 alarm even with a clean period between them. Two more toys: "where did the time go?" as each segment's duration minus its children's, and the default sampling rule as expected traces per second:
trace = [ # name, parent, start ms, end ms
("alb", None, 0, 3020),
("api", "alb", 5, 3015),
("auth", "api", 10, 40),
("db query", "api", 45, 3010),
]
self_ms = {name: end - start for name, _, start, end in trace}
for name, parent, start, end in trace:
if parent:
self_ms[parent] -= end - start
print(max(self_ms, key=self_ms.get)) # db query
print([round(1 + 0.05 * (rps - 1), 2) for rps in (1, 20, 200)])
# [1.0, 1.95, 10.95]
The incremental window against a direct recount of the last N datapoints, on 2,000 random series and settings:
import random
random.seed(18)
ok = True
for _ in range(2000):
n = random.randint(1, 5)
m = random.randint(1, n)
t = random.randint(0, 20)
a = ToyAlarm(t, m, n)
series = [random.randint(0, 25) for _ in range(random.randint(0, 30))]
for i, v in enumerate(series):
last = series[max(0, i + 1 - n):i + 1]
if len(last) < n:
want = "INSUFFICIENT_DATA"
else:
want = "ALARM" if sum(x >= t for x in last) >= m else "OK"
ok &= a.add(v) == want
print(ok) # True
The complexity
The costs are delay, noise and coverage:
- Detection delay is roughly the period times the number of breaching datapoints needed: with 3 out of 3 one-minute periods, the alarm cannot fire before the third breaching minute.
- Noise versus speed. A larger M ignores more spikes and reacts later.
- Sampling keeps tracing cheap, but the failing request you care about may not have a trace.
Where it goes wrong
- Alarming on causes, not symptoms. CPU at 80 percent may be fine; errors and latency users feel are not.
- An alarm on one datapoint. Every spike pages someone.
- Expecting logs to answer where. Without a trace ID in each log line, joining them across services is guesswork.
- Using CloudWatch to find who changed a resource. API activity is recorded by CloudTrail.
- New instrumentation on the X-Ray SDK. It is in maintenance mode; use OpenTelemetry.
When it shows up in interviews
It appears in AWS interviews as "how do you monitor this?" and in system design as the operations section at the end. Expect follow-ups on what triggers an alarm, how to avoid alert fatigue, and how you would find the slow service. Tying the three signals to the three questions separates a structured answer from a list of product names.
How to say it in an interview
"Metrics tell me that something is wrong, traces tell me where, logs tell me why. I would alarm on user-facing symptoms, like the load balancer's 5XX count, with an M out of N rule so one noisy minute does not page anyone, and send the alarm to SNS. For location I would trace requests with OpenTelemetry into X-Ray and look at which segment holds the time. For the cause I would query that service's log group, filtered by the trace ID. And for who changed what, CloudTrail."
Scaling on the same metrics is covered in EC2 Auto Scaling, and the SNS side of the alarm in SQS vs SNS vs EventBridge.
Sources
- Alarm evaluation — Amazon CloudWatch User Guide
- Amazon CloudWatch Logs concepts — CloudWatch Logs User Guide
- AWS X-Ray concepts — AWS X-Ray Developer Guide
- X-Ray SDK and daemon support timeline — AWS X-Ray Developer Guide