Skip to content
BytePatterns

The Outage Story

Behavioral: lesson 10 of 12

What you did in the first ten minutes, and what you changed after.

Lesson 10 of 12 · 5 min

The Outage Story

Step 1 of 10

"The deploy took checkout down for forty minutes" is a headline, not a story. The story is the order you did things.

The Idea

An outage story is judged on sequence, not drama. Stop the bleeding, then find out why. Interviewers are listening for someone who reached for a rollback before a debugger, kept people informed while working, and left behind a change that makes the same failure boring.

Real-World Example

"The deploy took checkout down for forty minutes." Better: "Error rate alerted at 14:02. I rolled back at 14:06 and posted in the channel. The cause was a migration that dropped a column still read by the old pods; we now gate drops behind a two-release rule."

The Tradeoff

Own your part without performing guilt — an interviewer reading contrition as instability is a real risk. And never end on a person. The strongest version ends on the mechanism: a check, an alert, a rule that no longer depends on anyone remembering.

Your turn

Put the steps in the right order.

  1. Find the cause once customers are served again
  2. Restore service by the fastest safe route
  3. Land a mechanism that makes the failure impossible or boring
  4. Say how it was detected, and how long that took

Mini quiz

1 / 3

The first priority during an incident is:

New lessons land every few weeks

Leave an address and we will tell you when the next one is up. That is the only reason we will use it.

One address, stored so we can email you. Nothing else, ever.