All skills
hieutrtr avatar

/incident-response

@4afa41d

Production incident response procedures for Python/React applications. Use when responding to production outages, investigating error spikes, diagnosing performance degradation, or conducting post-mortems. Covers severity classification (SEV1-SEV4), incident commander role, communication templates, diagnostic commands for FastAPI/ PostgreSQL/Redis, rollback procedures, and blameless post-mortem process. Does NOT cover monitoring setup (use monitoring-setup) or deployment procedures (use deployment-pipeline).

Use this Skill: https://skilld.dev/gh/hieutrtr/ai1-skills/incident-response

This session only. Nothing lands on disk.

referencespost-mortem-template.md

≈1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Post-Mortem Template

Instructions

Copy this template for each post-mortem. Fill in all sections. Conduct the post-mortem meeting within 48 hours of the incident for SEV1/SEV2, within 1 week for SEV3.


Post-Mortem: [Incident Title]

Date: [YYYY-MM-DD] Severity: [SEV1/SEV2/SEV3/SEV4] Duration: [Total duration] Author: [Name] Incident Commander: [Name] Attendees: [List of post-mortem participants]

1. Summary

One paragraph describing what happened, when, and the impact on users.

Example: On January 15, 2024, from 14:30 to 15:15 UTC (45 minutes), the backend API returned 503 errors for approximately 80% of requests. The root cause was database connection pool exhaustion triggered by a new endpoint that failed to release connections. An estimated 2,400 users were affected during the incident window.

2. Impact

Metric Value
Duration minutes/hours
Users affected count or percentage
Requests failed count or percentage
Revenue impact if applicable
SLA impact if applicable
Data loss yes/no, describe if yes

3. Timeline (UTC)

Time Event
HH:MM First sign of impact (from metrics/logs)
HH:MM Alert fired / issue detected
HH:MM Incident declared, severity assigned
HH:MM Incident commander designated
HH:MM Investigation started
HH:MM Root cause identified
HH:MM Mitigation applied
HH:MM Service recovered
HH:MM Incident declared resolved

4. Root Cause

Describe the fundamental reason the incident occurred. Use the Five Whys technique.

Five Whys:

  1. Why did [symptom]? Because [cause 1].
  2. Why did [cause 1]? Because [cause 2].
  3. Why did [cause 2]? Because [cause 3].
  4. Why did [cause 3]? Because [cause 4].
  5. Why did [cause 4]? Because [root cause].

Root cause: One sentence describing the fundamental issue.

5. Contributing Factors

List all conditions that contributed to the incident occurring or worsening.

  • Factor 1: Description
  • Factor 2: Description
  • Factor 3: Description

6. Detection

Question Answer
How was the incident detected? Alert / user report / manual check
Time from impact to detection minutes
Was the right alert in place? yes / no
Did the alert fire promptly? yes / no / N/A

7. Response

Question Answer
Time from detection to response minutes
Were the right people paged? yes / no
Was the runbook useful? yes / no / no runbook existed
Time from response to mitigation minutes
Was communication clear and timely? yes / no

8. What Went Well

  • List things that worked effectively during the incident
  • Example: Alert fired within 2 minutes of impact
  • Example: Rollback procedure completed in under 5 minutes

9. What Could Be Improved

  • List gaps or problems in the response
  • Example: No alert for connection pool saturation
  • Example: Runbook did not mention this failure mode
  • Example: It took 15 minutes to find the right dashboard

10. Action Items

# Action Owner Priority Due Date Tracking
1 Description Name P1/P2/P3 Date Ticket link
2 Description Name P1/P2/P3 Date Ticket link
3 Description Name P1/P2/P3 Date Ticket link

Action item categories:

  • Prevent: Changes to prevent this class of incident from recurring
  • Detect: Improvements to detect similar issues faster
  • Mitigate: Changes to reduce time-to-recovery
  • Process: Improvements to incident response procedures

11. Lessons Learned

Key insights from this incident that the broader team should understand.

  1. Lesson 1
  2. Lesson 2
  3. Lesson 3

Reviewed and approved by: [Engineering Manager Name], [Date]

Source: SKILL.md on GitHub

1 warning13d4 checks · Risk SAFE
  • Gen Agent Trust Hub13d

    This skill provides a structured framework for production incident response, including log analysis and report generation. The primary security considerations are a surface for indirect prompt injection via the processing of untrusted logs and a configuration mismatch between the allowed tools and the commands used in the instructions.

  • Socket13d

    No alerts

  • Snyk13d

    Risk: LOW · No issues

  • Runlayer7mo

    6/6 files flagged

Signed by skilld at 4afa41d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Dormantupdated 8 months ago
compatibility
Any backend/frontend stack
context
fork
All 1 allowed tools
Read Grep Glob Write Bash(curl:*) Bash(jq:*)
Other metadata
metadata
{
  "author": "platform-team",
  "version": "1.0.0",
  "sdlc-phase": "operations"
}

README badge

README badge for hieutrtr/ai1-skills/incident-response