All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

bundlesknowledge-lifecycletemplatesblameless-postmortem.md

≈519 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Blameless Postmortem Template

Incident metadata

  • Incident ID: @@INCIDENT_ID@@
  • Severity: @@SEVERITY@@
  • Duration: @@DURATION@@
  • Affected services: @@AFFECTED_SERVICES@@
  • User / customer impact: @@USER_IMPACT@@
  • Status: Draft — pending team review

Executive summary

[2–3 sentences: what happened, impact, how it was resolved.]

Impact

  • Services affected:
  • User impact:
  • SLA / SLO impact:
  • Business impact (if known):

Timeline (UTC)

Time Event
HH:MM First anomaly detected in metrics
HH:MM Alert fired
HH:MM On-call engineer acknowledged
HH:MM Investigation started
HH:MM Root cause identified
HH:MM Mitigation applied
HH:MM Service fully recovered

Root cause

[Technical explanation, evidence-backed.]

Five Whys

  1. Why did [symptom]? → [reason 1]
  2. Why [reason 1]? → [reason 2]
  3. Why [reason 2]? → [reason 3]
  4. Why [reason 3]? → [reason 4]
  5. Why [reason 4]? → [root systemic cause]

Detection & response assessment

Metric Value Target Assessment
Time to detect Xm <5m ✅/⚠️/❌
Time to acknowledge Xm <15m ✅/⚠️/❌
Time to mitigate Xm <30m ✅/⚠️/❌
Time to resolve Xm <2h ✅/⚠️/❌

What went well

  • [Positive aspects of the detection and response.]

What could be improved

  • [Areas for systemic improvement.]

Action items

ID Action Category Owner Priority Due Status
AI-1 Prevent recurrence High +7d Open
AI-2 Improve detection Medium +14d Open

Categories: Prevent recurrence / Improve detection / Improve response / Improve resilience

Lessons learned

  • [Key systemic takeaways for the team.]

This postmortem was auto-generated by Azure SRE Agent and must be reviewed by the incident team before publishing. Focus on systems and processes ; never individuals.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent