Back to Frameworks
TPL-04Reliability

Incident Review Template

A good post-incident review makes the system stronger, not the engineer smaller. The structure below is blameless by design — every section asks about the system, not the individual.

Summary

Two to three sentences anyone in the company can understand. What happened, who was affected, how long it lasted, how it was resolved.

Severity & impact

  • Severity tier and why it was classified that way
  • Users / requests / revenue affected
  • Duration: detection → mitigation → full resolution

Timeline

A minute-by-minute log in UTC. Detection, escalation, key decisions, mitigation, all-clear. Include the things that took longer than they should have — that's where the learning lives.

Contributing factors

Plural, on purpose. Most incidents have three to five contributing factors, not a single root cause. Look at code, configuration, process, and human factors.

What went well

Reinforce the behaviours you want more of — fast escalation, clear comms, good instinct on a roll-back. This is not filler; it's how good incident response becomes a habit.

Action items

Each item has an owner, a due date, and a priority. Anything without all three is a wish, not an action. Track them to completion in your normal backlog.

Common pitfalls

  • — Naming individuals as the cause. The system let them make that mistake.
  • — Generating 30 action items no one ever ships. Five is plenty if they're real.
  • — Skipping the review because "we already know what happened." The review is for the org, not just the team.

One practical engineering leadership lesson every week.

Get one useful idea you can apply with your team — a framework, conversation technique, template, or lesson from real engineering leadership.

Trouble seeing the form?Get the Next Issue →

Free. Useful. No motivational fluff.