Skip to content
BytePatterns

Observability Explained: Logs, Metrics and Traces

8 min readBytePatterns

Observability basics: what logs, metrics and traces each answer, why you alert on percentiles, label cardinality, span self time and trace sampling, in code.

A request fails somewhere in one of twelve services, at three in the morning. Observability is whether you can find out where and why from the outside, without shipping new code to ask. It rests on three kinds of telemetry, metrics, logs and traces, and knowing which question goes to which is most of the skill.

The problem it solves

A monolith can be debugged with a stack trace and a log file. In a distributed system one request crosses several services, each with its own logs. You need to answer, in order:

  1. Is something broken? Something must notice before users do.
  2. Where? Which service, which endpoint, which hop inside the request.
  3. Why? The specific error, input or resource that failed.

Each signal is cheap for one question and poor at the others.

The intuition

  • Metrics are numbers aggregated over time: request rate, error rate, latency percentiles, CPU, queue depth. They are tiny, fast to query and cheap to keep for months, which makes them the right place for alerts and dashboards. They cannot tell you about one specific request.
  • Logs are timestamped records of individual events with context: the error message, the user, the parameters. They answer "why", but volume makes them expensive. Structured logs (key-value fields rather than free text) make them queryable.
  • Traces follow one request across every service it touched. Each service records a span with a start time and a duration, linked by a shared trace id and a parent span id passed along in request headers. A trace is the only view that shows which hop cost the time.

Two rules keep the bill sane. Alert on percentiles, not averages: a mean hides the slow tail that users feel. And keep metric labels low-cardinality: every distinct combination of label values is a separate time series, so adding a user id label multiplies storage by the number of users. Traces, the most expensive signal, are sampled.

As of October 2026, OpenTelemetry is the widely used vendor-neutral standard for emitting all three signals, and the W3C Trace Context traceparent header is the usual way to propagate trace ids between services (from memory). For the AWS versions of these tools, see CloudWatch alarms, logs and X-Ray.

Watch it run

The animation walks one incident through the three views. Metrics are cheap numbers over time, continuous, aggregated and small enough to keep for a year; errors sit at 0.6 to 0.7%. Then the error rate crosses its 2% threshold at 4.1% and an alert fires. Alerting belongs on metrics precisely because a threshold on one number is easy to trust. Watch percentiles, not averages, because a handful of ten-second requests barely move a mean, here 42 ms against a p99 of 1.9 seconds. The dashboard rolls the same metric up per service, and catalog, at 4.9%, is the source. But no metric can say which hop inside a request cost the time, so it opens one request's trace, 214 ms. Each span is one service's slice of that single request, and there it is: 186 of 214 ms sat inside catalog, only 8 ms of that in the database. Now it reads that one service's logs around the failing requests: its connection pool was exhausted, 25 of 25 busy. Logs carry the context the other two views threw away. The last frames are the costs: keep labels low-cardinality, service and route, because a label per user id will bankrupt you, and sample traces, 1 in 100, instead of recording every request.

Observability Basics

Step 1 of 13

Metrics are cheap numbers over time: request rate, error rate, latency percentiles.

The same interactive animation as the lesson — step through it with the controls.

The code

A toy model of the three ideas that cause most confusion. First, why a mean is a poor alert. A calm window, then an incident where 1.2% of requests get stuck behind a pool:

import math, random
from fractions import Fraction

def percentile(xs, p):
    """Nearest rank: the smallest sample with at least p% of samples at or below it."""
    s = sorted(xs)
    k = math.ceil(Fraction(str(p)) * len(s) / 100)
    return s[max(k, 1) - 1]

rng = random.Random(35)
calm = [rng.gauss(40, 8) for _ in range(10_000)]
incident = calm[:9_880] + [rng.uniform(800, 3_000) for _ in range(120)]  # 1.2% stuck
for window in (calm, incident):
    mean = sum(window) / len(window)
    print(round(mean), round(percentile(window, 50)), round(percentile(window, 99)))
# 40 40 59
# 62 40 1143

labels = {"service": 12, "route": 40, "status": 5}
print(math.prod(labels.values()))                     # 2400 time series
labels["user_id"] = 2_000_000
print(f"{math.prod(labels.values()):,}")              # 4,800,000,000

The median did not move, the mean drifted from 40 to 62 ms (under a typical 100 ms alert), and p99 went from 59 ms to over a second. One user id label turns 2,400 series into 4.8 billion.

Next, a trace. A span's self time is its duration minus the time covered by its children, merged so that parallel children are not double-counted. The span with the most self time is where to look:

SPANS = {  # name: (parent, start_ms, duration_ms) -- one request's trace
    "gateway": (None, 0, 214), "auth": ("gateway", 6, 12),
    "catalog": ("gateway", 20, 186), "db": ("catalog", 26, 8),
}

def self_time(spans):
    """Time a span spent itself: its duration minus the union of its children."""
    out = {}
    for name, (_, start, dur) in spans.items():
        kids = sorted((s, s + d) for p, s, d in spans.values() if p == name)
        covered, reach = 0, start
        for s, e in kids:
            s, e = max(s, reach), min(e, start + dur)
            if e > s:
                covered += e - s
                reach = e
        out[name] = dur - covered
    return out

st = self_time(SPANS)
print(st, max(st, key=st.get))
# {'gateway': 16, 'auth': 12, 'catalog': 178, 'db': 8} catalog

Catalog spent 178 ms itself and only 8 ms waiting on the database, which points at something inside catalog, the exhausted pool. Finally, sampling. Deciding from the trace id means every service makes the same choice, so kept traces are complete; independent coin flips at three services almost never keep all three spans of a trace. Then both functions are checked against brute force on 300 seeded cases: percentiles against the definition, self time by counting every millisecond no child covers, with overlapping children:

def keep(trace_id, one_in=100):
    """Head sampling: decide from the trace id, so every service decides alike."""
    return int(trace_id[-8:], 16) % one_in == 0

ids = ["%032x" % rng.getrandbits(128) for _ in range(100_000)]
kept = [t for t in ids if keep(t)]
coin = sum(all(rng.random() < 0.01 for _ in range(3)) for _ in ids)  # 3 services flip coins
print(len(kept), coin)                                                     # 1045 0

ok = True
for _ in range(300):
    n = rng.randint(1, 60)
    xs = [rng.randint(0, 50) for _ in range(n)]
    p = rng.choice([1, 10, 50, 90, 95, 99, 99.9, 100])
    brute = min(v for v in xs if sum(x <= v for x in xs) * 100 >= Fraction(str(p)) * n)
    ok &= percentile(xs, p) == brute
    spans = {"s0": (None, 0, rng.randint(1, 80))}
    for i in range(1, rng.randint(1, 8)):
        parent = "s%d" % rng.randrange(i)
        _, ps, pd = spans[parent]
        s = rng.randint(ps, ps + pd - 1)
        spans["s%d" % i] = (parent, s, rng.randint(1, ps + pd - s))   # kids may overlap
    for name, t in self_time(spans).items():
        _, s, d = spans[name]
        kids = [(ks, ks + kd) for p_, ks, kd in spans.values() if p_ == name]
        ok &= t == sum(not any(a <= ms < b for a, b in kids) for ms in range(s, s + d))
print(ok)                                                                 # True

The complexity

  • Metrics: storage grows with the number of series, the product of label value counts, not with traffic.
  • Logs: grow with traffic times verbosity; retention and sampling of debug logs are the levers.
  • Traces: grow with traffic times spans per request, divided by the sampling rate.

Where it goes wrong

  • Alerting on averages or on causes. Alert on user-facing symptoms, error rate and latency percentiles, and use dashboards to find causes.
  • High-cardinality labels. User ids, request ids and raw URLs belong in logs and traces, not metric labels.
  • Per-service random sampling. It produces traces with missing spans. Sample by trace id, or sample after the trace completes ("tail sampling") to keep every error and slow trace, at the cost of buffering.
  • Logs without a trace id. Put the trace id in every log line, so a slow trace leads straight to its logs.

When it shows up in interviews

Inside almost every system design question as "how would you monitor this?", and on its own as "how would you debug a latency spike across microservices?". Interviewers listen for the order: an alert on a symptom metric, a dashboard to find the service, a trace to find the hop, logs to find the cause. With message queues, queue depth and consumer lag are the metrics that matter.

How to say it in an interview

"I use three signals for three questions. Metrics tell me something is broken: I alert on error rate and latency percentiles, never averages, and keep labels low-cardinality. A dashboard per service tells me where. A distributed trace, propagated by trace id through every hop, shows which span spent the time. Then that service's logs, tagged with the same trace id, tell me why. Traces are expensive, so I sample by trace id to keep them complete, and keep every error trace if I can afford tail sampling."