Operations

Your Incident Report Says "Better Processes." Here's the Template That Actually Works.

By Felix Maru · August 8, 2026 · 8 min read

We had the same outage twice in three months. Both times, an unannounced maintenance window took down the same system. The first post-incident report concluded: "Improve communication around planned maintenance." Nobody owned a change. Nobody had a deadline. The document was filed, and three months later, a different engineer made the exact same call without telling anyone.

I have seen this pattern at every organisation I have worked in. The incident report proves the review happened. It does not prove anything changed.

Here is the framework I use now, and why each section exists.

What Most Incident Reports Actually Contain

A typical post-incident report has three things: a timeline of events, a description of how the incident was resolved, and a "lessons learned" section. The lessons learned section is where things go wrong.

If you open most incident reports and look at the action items, you will find language like:

None of those are action items. They are aspirations. There is no system named, no owner, no deadline. Nobody knows what "improve" means by next Tuesday. The report gets filed, the incident fades, and the conditions that caused it remain exactly in place.

Documentation is not improvement. Filing the report proves you reviewed the incident. Specific actions with named owners are what actually change behavior.

The Five-Section Framework

Every post-incident review I write now has exactly five sections. No more, no fewer. The discipline of the structure is the point: it forces specificity in the places where vagueness is most tempting.

1
Timeline
What happened, in what order, with approximate timestamps. Factual only, no interpretation yet.
2
Root Cause
Keep asking "why" until you reach a system or process gap, not a person's mistake.
3
Impact
Customers or users affected, ticket volume generated, resolution time, any CSAT or SLA hit.
4
Specific Actions
Each action names the change, the system or process it lives in, and a completion date.
5
Owner and Deadline
One named person per action. One calendar date. Not "the team" and not "soon."
The five sections that turn a post-incident review from a record into a change. Steps 1 to 4 are analytical; step 5 is where accountability lands on a person.

Writing Root Cause Without Blame

The root cause section is the hardest to write well. Most teams write it by describing what someone did or failed to do. That is blame, not root cause, and it produces action items that are unenforceable (you cannot write "do not forget things" as a policy).

Good root cause lives in systems, processes, documentation, and training, not in individuals. The question is not who failed but why the conditions existed for failure to happen.

Compare these two framings of the same event:

The first version cannot generate a useful action item. The second version already points at its own fix: add a notification checkpoint to the maintenance runbook, with a named reviewer, before the window opens.

When I am working through root cause, I ask "why" four times. The first answer almost always names a person or an action. By the fourth, I am usually pointing at a missing document, an undefined process, or a system that does not surface the right information at the right moment. That is where the fix belongs.

The 48-Hour Rule

Write the post-incident review within 48 hours of the incident being closed. Not 72 hours. Not "later this week." Forty-eight hours.

After 48 hours, memories blur. People reconstruct what they think happened rather than what they actually observed. The ticket queue fills back up and the incident stops feeling urgent. By the end of the week, the window for a good review has passed, and what gets written is a polished narrative rather than an honest one.

My process is a 15-minute call with everyone who touched the incident. I open the five-section document before the call. We fill in the timeline together, agree on root cause, and assign owners to each action item before we hang up. The document goes out to the team the same day.

Fifteen minutes is not enough time to assign blame or to have a long defensive conversation. It is exactly enough time to get the facts on paper and decide what changes. The constraint is the point.

What a Specific Action Actually Looks Like

Section four is where most teams return to vague language after the discipline of root cause. Here is the test I apply to every action item before I let it into the document: it must answer four questions.

If you cannot answer all four, it is not an action item yet. It is a note to think about later, which is another way of saying it will not happen.

Here is the same action item written two ways:

Aspirational: "Improve the communication process for planned maintenance windows."

Operational: "Add a mandatory 'Support team notified?' checkbox to the maintenance runbook in Notion, with a 2-hour pre-window SLA. [Name] owns the runbook update by August 15."

The operational version is auditable. Either the checkbox is in the runbook by August 15 or it is not. The aspirational version can be claimed as complete without anything changing.

Storing and Following Up

A post-incident review that nobody can find later is only half as useful. I keep all PIRs in a labeled folder in the same place the team already lives, whether that is Notion, a Google Drive operations folder, or a pinned Help Scout note on the account. What matters is that someone facing a similar incident two years from now can search for it and find what changed the last time.

I also put the action items on the next team meeting agenda. Not the next incident, the next regular sync. This closes the loop: the review is written, the actions are assigned, and the team confirms completion before the topic goes quiet. It takes five minutes. It is the step that most teams skip, and skipping it is why the same incidents keep happening.

The first post-incident review you run this way will take 45 minutes. By the fifth, you are done in 20. The template does most of the work. The discipline is in running it every time, not only after the dramatic incidents, but after the smaller ones too, the access provisioning error that a user caught before it became a problem, the ticket that bounced between three agents because the escalation path was not written down. The pattern matters as much as the severity.

Start with the Next One

You do not need to retrofit every past incident. Start with the next one. Run the 15-minute call, fill in the five sections, assign one owner per action, and get the document out before 48 hours pass. Do it again for the one after that. Within a few months you will have a body of reviews that actually shows you where your processes are weak, and a team that trusts the process enough to be honest in it.

If you are working through a recurring incident pattern and want a second pair of eyes on what your current review is missing, drop me a note here. Happy to look at a real example and point out where the template usually falls short.

Share 𝕏 in

Comments