Blameless Postmortem Template
Incident metadata
- Incident ID:
@@INCIDENT_ID@@ - Severity:
@@SEVERITY@@ - Duration:
@@DURATION@@ - Affected services:
@@AFFECTED_SERVICES@@ - User / customer impact:
@@USER_IMPACT@@ - Status: Draft — pending team review
Executive summary
[2–3 sentences: what happened, impact, how it was resolved.]
Impact
- Services affected:
- User impact:
- SLA / SLO impact:
- Business impact (if known):
Timeline (UTC)
| Time | Event |
|---|---|
| HH:MM | First anomaly detected in metrics |
| HH:MM | Alert fired |
| HH:MM | On-call engineer acknowledged |
| HH:MM | Investigation started |
| HH:MM | Root cause identified |
| HH:MM | Mitigation applied |
| HH:MM | Service fully recovered |
Root cause
[Technical explanation, evidence-backed.]
Five Whys
- Why did [symptom]? → [reason 1]
- Why [reason 1]? → [reason 2]
- Why [reason 2]? → [reason 3]
- Why [reason 3]? → [reason 4]
- Why [reason 4]? → [root systemic cause]
Detection & response assessment
| Metric | Value | Target | Assessment |
|---|---|---|---|
| Time to detect | Xm | <5m | ✅/⚠️/❌ |
| Time to acknowledge | Xm | <15m | ✅/⚠️/❌ |
| Time to mitigate | Xm | <30m | ✅/⚠️/❌ |
| Time to resolve | Xm | <2h | ✅/⚠️/❌ |
What went well
- [Positive aspects of the detection and response.]
What could be improved
- [Areas for systemic improvement.]
Action items
| ID | Action | Category | Owner | Priority | Due | Status |
|---|---|---|---|---|---|---|
| AI-1 | Prevent recurrence | High | +7d | Open | ||
| AI-2 | Improve detection | Medium | +14d | Open |
Categories: Prevent recurrence / Improve detection / Improve response / Improve resilience
Lessons learned
- [Key systemic takeaways for the team.]
This postmortem was auto-generated by Azure SRE Agent and must be reviewed by the incident team before publishing. Focus on systems and processes ; never individuals.