Post-Mortem Template
Instructions
Copy this template for each post-mortem. Fill in all sections. Conduct the post-mortem meeting within 48 hours of the incident for SEV1/SEV2, within 1 week for SEV3.
Post-Mortem: [Incident Title]
Date: [YYYY-MM-DD] Severity: [SEV1/SEV2/SEV3/SEV4] Duration: [Total duration] Author: [Name] Incident Commander: [Name] Attendees: [List of post-mortem participants]
1. Summary
One paragraph describing what happened, when, and the impact on users.
Example: On January 15, 2024, from 14:30 to 15:15 UTC (45 minutes), the backend API returned 503 errors for approximately 80% of requests. The root cause was database connection pool exhaustion triggered by a new endpoint that failed to release connections. An estimated 2,400 users were affected during the incident window.
2. Impact
| Metric | Value |
|---|---|
| Duration | minutes/hours |
| Users affected | count or percentage |
| Requests failed | count or percentage |
| Revenue impact | if applicable |
| SLA impact | if applicable |
| Data loss | yes/no, describe if yes |
3. Timeline (UTC)
| Time | Event |
|---|---|
| HH:MM | First sign of impact (from metrics/logs) |
| HH:MM | Alert fired / issue detected |
| HH:MM | Incident declared, severity assigned |
| HH:MM | Incident commander designated |
| HH:MM | Investigation started |
| HH:MM | Root cause identified |
| HH:MM | Mitigation applied |
| HH:MM | Service recovered |
| HH:MM | Incident declared resolved |
4. Root Cause
Describe the fundamental reason the incident occurred. Use the Five Whys technique.
Five Whys:
- Why did [symptom]? Because [cause 1].
- Why did [cause 1]? Because [cause 2].
- Why did [cause 2]? Because [cause 3].
- Why did [cause 3]? Because [cause 4].
- Why did [cause 4]? Because [root cause].
Root cause: One sentence describing the fundamental issue.
5. Contributing Factors
List all conditions that contributed to the incident occurring or worsening.
- Factor 1: Description
- Factor 2: Description
- Factor 3: Description
6. Detection
| Question | Answer |
|---|---|
| How was the incident detected? | Alert / user report / manual check |
| Time from impact to detection | minutes |
| Was the right alert in place? | yes / no |
| Did the alert fire promptly? | yes / no / N/A |
7. Response
| Question | Answer |
|---|---|
| Time from detection to response | minutes |
| Were the right people paged? | yes / no |
| Was the runbook useful? | yes / no / no runbook existed |
| Time from response to mitigation | minutes |
| Was communication clear and timely? | yes / no |
8. What Went Well
- List things that worked effectively during the incident
- Example: Alert fired within 2 minutes of impact
- Example: Rollback procedure completed in under 5 minutes
9. What Could Be Improved
- List gaps or problems in the response
- Example: No alert for connection pool saturation
- Example: Runbook did not mention this failure mode
- Example: It took 15 minutes to find the right dashboard
10. Action Items
| # | Action | Owner | Priority | Due Date | Tracking |
|---|---|---|---|---|---|
| 1 | Description | Name | P1/P2/P3 | Date | Ticket link |
| 2 | Description | Name | P1/P2/P3 | Date | Ticket link |
| 3 | Description | Name | P1/P2/P3 | Date | Ticket link |
Action item categories:
- Prevent: Changes to prevent this class of incident from recurring
- Detect: Improvements to detect similar issues faster
- Mitigate: Changes to reduce time-to-recovery
- Process: Improvements to incident response procedures
11. Lessons Learned
Key insights from this incident that the broader team should understand.
- Lesson 1
- Lesson 2
- Lesson 3
Reviewed and approved by: [Engineering Manager Name], [Date]