Summary
Two to three sentences anyone in the company can understand. What happened, who was affected, how long it lasted, how it was resolved.
Severity & impact
- Severity tier and why it was classified that way
- Users / requests / revenue affected
- Duration: detection → mitigation → full resolution
Timeline
A minute-by-minute log in UTC. Detection, escalation, key decisions, mitigation, all-clear. Include the things that took longer than they should have — that's where the learning lives.
Contributing factors
Plural, on purpose. Most incidents have three to five contributing factors, not a single root cause. Look at code, configuration, process, and human factors.
What went well
Reinforce the behaviours you want more of — fast escalation, clear comms, good instinct on a roll-back. This is not filler; it's how good incident response becomes a habit.
Action items
Each item has an owner, a due date, and a priority. Anything without all three is a wish, not an action. Track them to completion in your normal backlog.
Common pitfalls
- — Naming individuals as the cause. The system let them make that mistake.
- — Generating 30 action items no one ever ships. Five is plenty if they're real.
- — Skipping the review because "we already know what happened." The review is for the org, not just the team.