All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referenceskt-methodology.md

≈1.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Kepner-Tregoe (KT) Operations Overlay

Use KT to improve incident discipline without adding unnecessary runtime or paperwork. The default is proportional KT: use the smallest KT structure that makes the decision safer and clearer.

KT Depth Policy

Scenario Required depth Output target
P1/P2 incident Full KT: SA -> PA -> DA -> PPA Structured incident response and closure artifact
P3/P4 incident Lightweight KT: brief SA plus targeted PA or DA Concise evidence and next action
Production write action Minimum DA + PPA; use full KT if P1/P2, broad blast radius, or irreversible Approval-ready remediation rationale
Read-only scheduled health/reporting task ADAC only Short operational summary
Routine low-risk recommendation No formal KT unless uncertainty or impact warrants it Evidence, confidence, next step

Do not expand a low-risk read-only task into a full KT worksheet. Prefer compact bullets and only include sections that change the decision.

ADAC Lightweight Reliability Pattern

Use ADAC when full KT is unnecessary:

  1. Auto-Detect: what signal, scope, and failure boundary were identified.
  2. Auto-Declare: what can fail, expected degradation, blast radius, confidence, and owner.
  3. Auto-Communicate: how to validate, who is notified, and what action is recommended.

ADAC is enough for routine health checks, anomaly summaries, and low-risk scheduled tasks.

KT Flow Mapping

SA - Situation Appraisal

Purpose: classify concerns, set priorities, and choose the next analysis.

Use when:

  • severity or blast radius is unclear
  • multiple symptoms compete for attention
  • the agent needs to choose between PA, DA, or PPA

Minimum output:

  • concern/deviation
  • impact and urgency
  • known/unknown
  • next analysis path

PA - Problem Analysis

Purpose: identify the most likely cause using IS / IS NOT logic.

Use when:

  • a deviation exists and root cause is unknown
  • rollback/remediation depends on knowing the cause
  • symptoms could come from platform, workload, dependency, or release change

Minimum output:

  • problem statement
  • key IS / IS NOT distinctions by WHAT, WHERE, WHEN, EXTENT
  • likely causes and confidence
  • verification method

DA - Decision Analysis

Purpose: choose the best option using explicit objectives.

Use when:

  • selecting between rollback, scale, restart, patch, failover, or wait-and-observe
  • proposing production write actions
  • human approval is required

Minimum output:

  • decision to be made
  • must/want criteria
  • options considered
  • selected option and rejected options
  • rationale and approval need

PPA - Potential Problem Analysis

Purpose: protect execution of the chosen action.

Use when:

  • a production action can make things worse
  • rollback/roll-forward must be ready
  • action success requires validation thresholds

Minimum output:

  • potential problems
  • preventive actions
  • contingent actions
  • rollback trigger and validation window

Output Shapes

Full KT Incident Output

Use for P1/P2 and high-risk incidents:

  1. Situation Appraisal
  2. Problem Analysis
  3. Decision Analysis
  4. Potential Problem Analysis
  5. Recommended Action
  6. Risk and Rollback
  7. Owners and Timeboxed Next Steps
  8. Validation Evidence

Lightweight KT Output

Use for P3/P4 or targeted analysis:

  1. Situation Summary
  2. Evidence and Confidence
  3. Chosen Analysis (PA, DA, or PPA)
  4. Recommendation
  5. Validation / Escalation Trigger

ADAC Output

Use for routine read-only automation:

  1. Auto-Detect
  2. Auto-Declare
  3. Auto-Communicate

Domain Guidance

AKS

  • SA: split cluster, node pool, namespace, workload, and dependency concerns.
  • PA: compare affected vs unaffected nodes, pods, namespaces, versions, and deployment windows.
  • DA: compare restart, scale, rollback, config fix, node cordon/drain, or wait-and-observe.
  • PPA: define failure triggers and safe rollback before action.

Container Apps

  • SA: classify impact by app, environment, revision, and ingress path.
  • PA: compare current revision against known-good revision and recent configuration changes.
  • DA: choose traffic shift, rollback, scale, config patch, or hold.
  • PPA: protect rollout with threshold triggers and abort plan.

Drasi on AKS

  • SA: separate ingestion lag, query staleness, runtime faults, platform faults, and external dependency faults.
  • PA: isolate Drasi runtime vs AKS platform vs upstream/downstream dependency.
  • DA: choose scale, rollback, configuration correction, restart, or dependency escalation.
  • PPA: define fallback, validation windows, and post-mitigation probes.

Governance Hook Policy

Use hooks to enforce only the minimum required KT depth:

  1. P1/P2: reject outputs missing meaningful SA, PA, DA, and PPA.
  2. Production write action: reject outputs missing meaningful DA and PPA; require full KT only when severity or blast radius warrants it.
  3. P3/P4 read-only: do not block for missing full KT; require concise evidence, confidence, and next step.

See kt-templates.md and hooks-governance.md.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent