Back to Frameworks
TPL-04Reliability

Incident Review Template

A good post-incident review makes the system stronger, not the engineer smaller. The structure below is blameless by design — every section asks about the system, not the individual.

Summary

Two to three sentences anyone in the company can understand. What happened, who was affected, how long it lasted, how it was resolved.

Severity & impact

  • Severity tier and why it was classified that way
  • Users / requests / revenue affected
  • Duration: detection → mitigation → full resolution

Timeline

A minute-by-minute log in UTC. Detection, escalation, key decisions, mitigation, all-clear. Include the things that took longer than they should have — that's where the learning lives.

Contributing factors

Plural, on purpose. Most incidents have three to five contributing factors, not a single root cause. Look at code, configuration, process, and human factors.

What went well

Reinforce the behaviours you want more of — fast escalation, clear comms, good instinct on a roll-back. This is not filler; it's how good incident response becomes a habit.

Action items

Each item has an owner, a due date, and a priority. Anything without all three is a wish, not an action. Track them to completion in your normal backlog.

Common pitfalls

  • Naming individuals as the cause. The system let them make that mistake.
  • Generating 30 action items no one ever ships. Five is plenty if they're real.
  • Skipping the review because "we already know what happened." The review is for the org, not just the team.

Leadership in your inbox.

Leadership lessons, frameworks, and field notes for modern engineering teams. No motivational fluff.

Trouble seeing the form?Subscribe on Beehiiv instead

Free. One email a week. Unsubscribe anytime.