All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencesdeployment-patterns.md

≈1.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Deployment Patterns from Official Samples

This guide captures production-grade patterns seen in official microsoft/sre-agent sample deployments and scripts.

Pattern 1: Idempotent Post-Provision Setup

Post-provision automation should be safe to re-run.

Use upsert semantics for:

  1. Knowledge base uploads
  2. Custom agents
  3. Incident response plans
  4. Scheduled tasks
  5. Connectors

Design goal: "fix-forward" without redeploying infrastructure.

Pattern 2: Eventual Consistency Retries

Some operations are not immediately available after provisioning.

Use bounded retry loops for:

  • incident platform availability
  • response plan creation/update
  • connector readiness

Recommended defaults:

  1. 3 attempts
  2. 10-30 second delay
  3. clear failure message with operator next step

Pattern 3: Quickstart Cleanup

If quickstart plans are auto-created, remove overlap before custom routes.

Automation step:

  1. list active plans
  2. detect quickstart handler
  3. disable/delete when custom equivalent exists

Pattern 4: Full Verification Pass

After setup, verify all control-plane objects:

  1. knowledge files indexed
  2. custom agents present
  3. incident platform connected
  4. connectors healthy
  5. response plans active
  6. scheduled tasks active

Return a single readiness summary.

Pattern 5: Memory-First Investigation

Investigation agents should search memory before fresh diagnostics.

Default flow:

  1. memory lookup for similar incidents
  2. runbook lookup
  3. logs + metrics evidence
  4. root-cause hypothesis
  5. action proposal

Pattern 6: Structured Incident Report Templates

Use a fixed report schema across teams:

  1. Summary
  2. Impact
  3. Timeline (UTC)
  4. Evidence
  5. Root Cause
  6. Remediation
  7. Action Items
  8. References

This supports reliable handoffs and postmortems.

Pattern 7: Scheduled Issue Triage Workflows

For code-facing operations:

  1. scheduled task triggers issue triage agent
  2. triage agent classifies and labels issues
  3. triage agent posts standardized response comment
  4. escalation for P0/P1 style issues

Pattern 8: KT-Governed Major Incident Handling

For high-severity incidents and production write paths:

  1. perform Situation Appraisal before deep investigation
  2. produce Problem Analysis with explicit IS / IS NOT logic
  3. evaluate remediation options with Decision Analysis (MUST/WANT)
  4. protect execution plan with Potential Problem Analysis

Use hooks to reject incomplete major-incident responses.

Pattern 9: Starter Lab Deployment (azd-based)

The official starter lab at microsoft/sre-agent/labs/starter-lab uses azd up for deployment:

  1. Deploy: SRE Agent, sample app (Grubify), Log Analytics, App Insights, Alert, Container Registry, Managed Identity.
  2. Three scenario tracks: IT Operations, Developer (GitHub), Workflow Automation.
  3. Estimated time: ~40 minutes.

Use this as a reference for azd-based SRE Agent provisioning patterns. See https://github.com/microsoft/sre-agent/tree/main/labs/starter-lab.

Pattern 10: Event-driven Terraform Drift via HTTP Triggers

Use HTTP triggers to process drift notifications from Terraform Cloud or other webhook-capable drift tools.

Reference architecture:

  1. Drift system emits webhook payload.
  2. Logic App / Azure Functions / APIM validates source and acquires Azure token.
  3. Authenticated POST invokes SRE Agent HTTP trigger.
  4. Trigger routes to a drift-analysis custom agent in Review mode first.

Use this for event-driven operations. Keep scheduled tasks for periodic checks.

Pattern 11: Drift Investigation Evidence Contract

For each drift event, gather and correlate:

  1. Terraform desired state and actual Azure resource state.
  2. Activity Log actor and timestamp for each drifted change.
  3. Incident impact signals from Application Insights or Azure Monitor.
  4. Recent code or deployment changes that explain the drift.

This prevents unsafe "diff-only" remediation.

Pattern 12: Severity-Gated Remediation for Drift

Use severity classes with explicit actions:

  1. Benign (tags/metadata) -> safe rollback path.
  2. Risky (security downgrade) -> prioritize rollback with security evidence.
  3. Critical (incident mitigation change, such as emergency scale-up) -> do not auto-revert until root cause is fixed and rollback is safe.

Hard rule:

  1. Never auto-revert critical drift that is actively mitigating a live incident.
  2. Enforce approval + rollback evidence before executing critical write actions.

Pattern 13: Replay, Verification, and Learning Loop

Before enabling autonomous drift remediation:

  1. Validate with historical incident payloads and simulated webhook runs.
  2. Verify trigger execution history and report quality.
  3. Run execution review and update drift skill instructions when gaps are found.
  4. Track recurring drift classes to evolve bundle templates and governance hooks.

AKS and Container Apps Production Pattern

Split specialist custom agents

  1. aks-diagnostics for cluster/pod/network diagnostics
  2. containerapps-revision-analyzer for revision health and rollback evidence
  3. incident-notifier for outbound notifications and stakeholder updates

Keep responsibilities separate

  1. Diagnose agents collect evidence.
  2. Remediation agents execute approved changes.
  3. Notifier agents report outcomes.

Drasi on AKS Pattern

Use this when SRE Agent supports Drasi workloads on AKS:

  1. Detect symptom from alert:
    • event processing lag
    • query staleness
    • connector ingestion faults
  2. Gather evidence:
    • AKS node/pod health
    • Drasi operator/controller logs
    • backing data plane dependencies
  3. Correlate:
    • deployment/revision changes
    • config drift
    • recent cluster/network events
  4. Route:
    • infrastructure issue -> AKS remediator path
    • app/workflow issue -> Drasi workflow owner path
  5. Report:
    • service impact
    • likely failure domain
    • recommended action and risk

Use drasi-aks-playbook.md for a ready-to-use Drasi-specific operating template.

Sources

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent