All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencesproduction-patterns-guide.md

≈4.4k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Production Patterns Guide for Azure SRE Agent

This guide consolidates 10 high-value patterns extracted from the core SKILL.md to keep the main skill document focused. All patterns are source-grounded in microsoft/sre-agent public repo.

Last updated: 2026-06-25 | Version: 2.3.3


Recipe Design Framework

Five-dimension model for composing agents:

Dimension Examples Intent
Platform AKS, Container Apps, App Service, Functions Where agent runs
Connectors GitHub, ADO, Azure DevOps, Jira, Slack, Splunk External systems to orchestrate
Skills diagnostics, remediation, RCA, cost optimization Agent capabilities
Knowledge runbooks, incident templates, SOP, troubleshooting guides Domain guidance (KT)
Response autonomous, review-gated, escalation, handoff chains Decision mode and flow

Example: Edge Incident Agent

Platform: AKS cluster per region
Connectors: GitHub (for PR proposals), Splunk (telemetry)
Skills: connectivity-diagnosis, firewall-rule-repair, rollback
Knowledge: NSG-troubleshooting runbook
Response: Review-gated for writes; escalate if > 3 retries

Response Plan Patterns & Real-World Examples

Severity-Routed Response Plans

Different agents handle Sev0 (autonomous, fast) vs Sev2 (review, thorough):

# Response Plan: azmon-sev01-edge-diagnosis
metadata:
  name: azmon-sev01-edge-network
  handlingAgent: edge-connectivity-agent
spec:
  priorities: [Sev0, Sev1] # Only highest severity
  agentMode: Autonomous # Fast remediation for network drops
  customInstructions: |
    Step 1: Check App Gateway NSG rules (priority-100 legacy DENY often shadows ALLOW).
    Step 2: Query LAW for gateway → backend connectivity metrics.
    Step 3: If DENY rule found, comment noting removal; do NOT execute until human confirms.
  maxAttempts: 3

---
# Response Plan: azmon-sev2-app-perf
metadata:
  name: azmon-sev2-app-degradation
  handlingAgent: perf-investigation-agent
spec:
  priorities: [Sev2, Sev3]
  agentMode: Review # Requires human approval before any action
  customInstructions: |
    Step 1: Diagnose from AppInsights (dependencies, exceptions, traces).
    Step 2: Check for recent deployments or traffic spikes.
    Step 3: Propose remediation (scaling, code fix, or config change).
    Step 4: Wait for human review before executing.

Handoff Chains: Multi-Agent Orchestration

Complex incidents route through specialist subagents in sequence:

handlingAgent: incident-orchestrator
subagents:
  - triage-agent # Classifies symptom (edge vs app tier)
  - investigation-agent # Deep diagnostics
  - remediation-agent # Proposes safe fixes
  - rca-agent # Root-cause analysis + evidence
  - pr-creation-agent # GitHub PR for durable fix

Each agent runs with limited scope; orchestrator collects artifacts and routes to next phase.

Knowledge Injection via customInstructions

Step-by-step guidance dramatically improves agent accuracy:

customInstructions: |
  # CosmosDB Incident Response SOP

  **Triage (2 min):**
  - Check Cosmos request units (RU) from Azure Portal metrics
  - Query activity log for recent replication failures
  - Look for throttling errors (429) in application logs

  **Investigation (5 min):**
  - Run: SELECT * FROM c WHERE c.error = "429" in past 15 min
  - Count throttled requests by partition key
  - Check write-region vs read-region latency

  **Remediation (approval required):**
  - Propose: increase RU from X to Y (cost delta: $Z/month)
  - OR: archive old data to reduce hot set
  - OR: rebalance partition key strategy

  **Evidence:**
  - Attach before/after metrics snapshot
  - Document assumption: throttling is root cause

Multi-Backend Deployment Patterns

IaC deployment is owned by the upstream microsoft/sre-agent repository under sreagent-templates/, not by this skill. This skill supplies bundle content (response plans, agents, hooks, scheduled tasks) that you drop into the config directory those templates generate. Do not hand-author agent ARM templates when the upstream templates already cover the resource shape.

The agent resource type is Microsoft.App/agents (verified 2026-08-09 against deploy-iac).

Generate a config directory from a recipe, then deploy it with one of four backends:

git clone https://github.com/microsoft/sre-agent.git
cd sre-agent/sreagent-templates
bash bin/check-prerequisites.sh

./bin/new-agent.sh --recipe azmon-lawappinsights --non-interactive \
  --set agentName=my-agent \
  --set resourceGroup=rg-my-agent \
  --set location=australiaeast \
  --set targetRGs=rg-my-workload \
  -o my-agent/

./bin/deploy.sh my-agent/
Backend Command Use when
Bicep ./bin/deploy.sh my-agent/ Default; runs az deployment sub create
Terraform ./bin/deploy-tf.sh my-agent/ Terraform-managed infrastructure
PowerShell .\bin\ps\Deploy-Agent.ps1 -InputPath .\my-agent\ Windows / PowerShell 7 environments
azd cd my-agent/ && azd up azd-based workflows

All backends support --what-if / --dry-run for validation without deploying. For azd specifically, the documented preview flag is azd provision --preview, there is no --no-exec flag.

Module status (verified 2026-08-10): no official Azure Verified Module (Bicep or Azure/avm-* Terraform) exists for Microsoft.App/agents, so the upstream hand-rolled templates remain best practice. A community Terraform AVM module (kewalaka/terraform-azure-avm-res-app-agent) deploys the agent with connectors, model settings, identity, and monitoring; treat it as a community reference, re-check its maintenance status before use.

Reference implementations worth reading before designing your own:

  • yortch/agentic-devops-demo — end-to-end Copilot + SRE Agent loop: subagent YAML (incident-handler, code-analyzer), knowledge-base docs, scheduled tasks, chaos-engineering workflow.
  • microsoft.github.io/frontier-sre-agent-rvas — official RVAS reference architecture with named specialist subagents, response plans, and autonomous remediation chapters.
  • matthansen0/azure-sre-agent-sandbox — automated AKS break-fix sandbox for testing loops safely.

Recipes are prebuilt starting points (Azure Monitor, PagerDuty, Dynatrace and others) and follow the same <platform>-<key-connectors> naming this skill uses. Browse sreagent-templates/recipes rather than assuming a recipe name.

Two deployment phases

Deployment is split, and the split determines what your IaC pipeline can and cannot guarantee.

Phase Mechanism What it covers
Phase 1 — ARM Bicep / Terraform / PowerShell / azd Resource group, user-assigned managed identity, Log Analytics workspace, Application Insights, the Microsoft.App/agents resource, RBAC role assignments (Reader, Monitoring Reader, Log Analytics Reader, SRE Agent Administrator), plus connectors, skills, subagents and tools as ARM sub-resources
Phase 2 — data plane apply-extras.sh (runs automatically after Phase 1) Code repositories (need Git PAT/OAuth), hooks (no ARM sub-resource yet), HTTP triggers (URLs generated server-side), knowledge files (binary upload), plugin configuration

If the data-plane token is unavailable, restricted CI/CD, no interactive login, apply-extras.sh prints what it skipped instead of failing silently. Treat that output as a deployment gate: a Phase 1 success is not a deployed agent. When wiring this into azd, Phase 2 belongs in a postprovision hook, and the hook must fail closed if extras were skipped.

Day-2 operations

The upstream templates ship scripts for ongoing management: export (config directory from a running agent), clone (./bin/clone-agent.sh, export plus deploy under a new name and resource group), diff (local config against the live agent), and verify (a 22-point check across connectors, skills, subagents and hooks). Prefer export-then-diff over reading the portal when auditing drift.


Data-Plane Configuration Upload (Post-Deployment)

Response plans, custom agents, knowledge, and hooks don't deploy via ARM, they upload after infrastructure:

# 1. Package extras as tarball
tar -czf extras.tar.gz -C config/skills . -C config/agents .

# 2. Get auth token
TOKEN=$(az account get-access-token --resource https://azuresre.ai --query accessToken -o tsv)

# 3. Upload to data-plane
curl -X POST "https://${AGENT_ENDPOINT}/api/v2/extendedAgent/skills" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/gzip" \
  --data-binary @extras.tar.gz

# 4. Upload knowledge with RAG indexing
curl -X POST "${AGENT_ENDPOINT}/api/v1/WorkspaceMemory/synthesized-knowledge?triggerIndexing=true" \
  -F "files=@incident-template.md" \
  -F "files=@troubleshooting-guide.md" \
  -H "Authorization: Bearer $TOKEN"

Knowledge uploaded with ?triggerIndexing=true is immediately available to agents for RAG grounding.


Connector Auth Orchestration

Multi-path auth for different platforms; choose the one your team already uses:

GitHub Auth Paths

# Path 1: OAuth (no env vars, browser sign-in)
export GITHUB_OAUTH_ENABLED=true
# User is prompted to visit: https://github.com/login/device
# Code displays in terminal; user enters on device

# Path 2: PAT (Headless / CI)
export GITHUB_PAT=ghp_xxxxxxxxxxxxxxxxxxxx

# Path 3: GitHub App (Multi-tenant)
export GITHUB_APP_ID=123456
export GITHUB_APP_PRIVATE_KEY=$(cat private-key.pem)

ADO Auth Paths

# Path 1: Managed Identity (Recommended for Azure)
export ADO_USE_MI=true
# Agent automatically uses its system-assigned managed identity

# Path 2: Azure AD (Interactive)
export ADO_USE_AAD=true
# User is prompted to authenticate in browser

# Path 3: PAT (Legacy fallback)
export ADO_PAT=<pat_token_here>
export ADO_ORG=https://dev.azure.com/myorg

Logic Apps / Functions Bridge for Private Endpoints

When SRE Agent cannot reach Git or CI/CD directly:

HTTP Trigger → Azure Functions (with network access)
↓
[Transform & auth check]
↓
GitHub/ADO internal endpoint
↓
[Return to SRE Agent]

Incident Response Runbook Pattern

Six-step workflow embedded in knowledge; agent follows sequentially, collecting evidence at each step:

# Incident Response Runbook: Production API Degradation

## Step 1: Acknowledge (0–2 min)

- Send ack to PagerDuty (stops escalation)
- Create incident summary in Teams channel
- Assign investigator

## Step 2: Triage (2–5 min)

- Determine if issue is **edge** (networking, DNS, gateway) or **app** (compute, memory, dependencies)
- Check geography (global vs region-specific)
- Identify suspect component from alert context

## Step 3: Investigate (5–15 min)

- Query AppInsights for exception rates, latencies, dependencies
- Check infrastructure metrics: CPU, memory, disk, network
- Look for recent deployments or config changes
- Gather logs from target system

## Step 4: Mitigate (15–30 min)

- Apply smallest safe fix: scale, circuit-break, rollback, or temporary config
- Document decision and assumptions
- Monitor metrics for recovery confirmation
- Log action as incident note (audit trail)

## Step 5: RCA (30–60 min)

- Map facts: what happened, when, why, blast radius
- Link to root cause with evidence (logs, metrics, code)
- Identify whether fix is temporary or permanent

## Step 6: Close & Recommend (60+ min)

- Prepare PR to durable fix (code, IaC, policy, or KT)
- Add alert tuning if false positive rate is high
- Schedule follow-up to prevent recurrence

Export → Clone → Verify Workflow

Full agent lifecycle: export current → validate → clone to new → verify all configs:

# 1. Export running agent to portable format
./bin/export-agent.sh \
  --subscription "$SUB" \
  --resource-group "$RG" \
  --agent-name "$AGENT" \
  --output exported-config.json

# 2. Validate schema
./bin/verify-agent.sh expected-config.json

# 3. Clone to new environment
./bin/deploy.sh \
  --source exported-config.json \
  --target-rg rg-new-agent \
  --target-location swedencentral

# 4. Run 22-point verification checklist
./bin/verify-agent.sh exported-config.json
# Checks: skills, subagents, hooks, knowledge, response plans, connectors, tasks, etc.

Export output includes: skills, subagents, hooks, common prompts, response plans, scheduled tasks, connectors, incident platform, model provider, access level, run mode, and upgrade channel.


Common Prompt Safety Rules

Share across all custom agents to prevent dangerous patterns:

# Safety Rules (inject into all agent prompts)

1. **Never delete production resources without explicit approval.**
   - Log all delete operations as incident notes before execution.
   - Require human confirmation in a tracked approval message.

2. **Log all actions for audit and compliance.**
   - Record what changed, who authorized it, timestamp, and evidence.
   - Use Application Insights customEvents or Azure audit logs.

3. **Rollback on error; do not retry indefinitely.**
   - If 3 attempts fail, escalate and stop.
   - Preserve error state for investigation.

4. **Alert on cost anomalies.**
   - If proposed action increases monthly spend > 10%, notify owner.
   - Require explicit approval for high-cost changes.

5. **Test in non-prod first.**
   - Never run write operations in production on first attempt.
   - Validate logic and output in staging.

Webhook Bridge for Third-Party Platforms

When Dynatrace, Grafana, or Splunk webhook directly to SRE Agent via HTTP trigger:

Dynatrace/Grafana/Splunk
         ↓ (webhook)
    Azure Functions (auth bridge)
         ↓
    SRE Agent HTTP Trigger
         ↓
    Investigation + response plan

Function bridge responsibilities:

  • Validate webhook signature
  • Transform vendor JSON to SRE Agent payload schema
  • Inject correlation ID and tenant context
  • Forward to HTTP trigger with Bearer token

HTTP trigger authentication:

# SRE Agent response plan tied to HTTP trigger
metadata:
  name: external-webhook-receiver
spec:
  httpTrigger:
    enabled: true
    authLevel: Function # Requires API key in Function auth settings
    methods: [POST]
    route: "incidents/external"

Post-Deployment Verification Checklist

22-point validation to confirm agent is production-ready:

✅ Connectivity & Identity
  [ ] Agent identity (MI or SP) has correct RBAC on target subscriptions
  [ ] Connectors can reach GitHub, ADO, observability APIs
  [ ] MCP servers (if configured) are reachable and authenticated

✅ Configuration & Scope
  [ ] Response plans exist and have correct incident platform routing
  [ ] Scheduled tasks are enabled and have correct cron expressions
  [ ] HTTP triggers are disabled until readiness testing complete
  [ ] Skill and custom-agent catalog loaded without errors

✅ Observability & Cost
  [ ] AMBA alerts firing correctly; no alert storms detected
  [ ] AAU budget set and monitored in Azure Monitor
  [ ] Incident telemetry flowing to Application Insights

✅ Governance
  [ ] Approval hooks configured for write actions
  [ ] Run-mode posture matches blast radius (ReadOnly/Review/Autonomous)
  [ ] Access level is Low unless High explicitly justified

✅ Knowledge & Runbooks
  [ ] At least one runbook or incident template loaded
  [ ] Knowledge indexing completed (check Application Insights customEvents)

✅ Rollout Safety
  [ ] Rollback procedure documented and tested
  [ ] On-call team notified and trained
  [ ] Cost and AAU forecast reviewed

SLI-to-Incident Bridge

Azure SRE Agent consumes Azure Monitor SLI breaches as incident triggers. To design SLIs and set alert thresholds:

  1. Define SLIs & SLOs using the observability-monitoring skill, this covers Azure Monitor Service Groups, burn-rate targets, and SLO tracking.
  2. Create AMBA alerts on fast-burn and slow-burn conditions that translate SLO targets into alert rules.
  3. Route alerts to Azure SRE Agent — the agent ingests alert payloads, matches them to response plans, and dispatches runbooks.

The division of responsibility is clear: Observability skill owns SLI/SLO definition; Azure SRE Agent owns incident response to SLI breaches. Together they form the end-to-end incident management lifecycle.


Summary

All 10 patterns are production-tested and source-grounded in the microsoft/sre-agent repository. Use this guide as a starting point; customize frameworks, runbooks, and safety rules for your operational environment.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent