Production Patterns Guide for Azure SRE Agent
This guide consolidates 10 high-value patterns extracted from the core SKILL.md to keep the main skill document focused. All patterns are source-grounded in microsoft/sre-agent public repo.
Last updated: 2026-06-25 | Version: 2.3.3
Recipe Design Framework
Five-dimension model for composing agents:
| Dimension | Examples | Intent |
|---|---|---|
| Platform | AKS, Container Apps, App Service, Functions | Where agent runs |
| Connectors | GitHub, ADO, Azure DevOps, Jira, Slack, Splunk | External systems to orchestrate |
| Skills | diagnostics, remediation, RCA, cost optimization | Agent capabilities |
| Knowledge | runbooks, incident templates, SOP, troubleshooting guides | Domain guidance (KT) |
| Response | autonomous, review-gated, escalation, handoff chains | Decision mode and flow |
Example: Edge Incident Agent
Platform: AKS cluster per region
Connectors: GitHub (for PR proposals), Splunk (telemetry)
Skills: connectivity-diagnosis, firewall-rule-repair, rollback
Knowledge: NSG-troubleshooting runbook
Response: Review-gated for writes; escalate if > 3 retriesResponse Plan Patterns & Real-World Examples
Severity-Routed Response Plans
Different agents handle Sev0 (autonomous, fast) vs Sev2 (review, thorough):
# Response Plan: azmon-sev01-edge-diagnosis
metadata:
name: azmon-sev01-edge-network
handlingAgent: edge-connectivity-agent
spec:
priorities: [Sev0, Sev1] # Only highest severity
agentMode: Autonomous # Fast remediation for network drops
customInstructions: |
Step 1: Check App Gateway NSG rules (priority-100 legacy DENY often shadows ALLOW).
Step 2: Query LAW for gateway → backend connectivity metrics.
Step 3: If DENY rule found, comment noting removal; do NOT execute until human confirms.
maxAttempts: 3
---
# Response Plan: azmon-sev2-app-perf
metadata:
name: azmon-sev2-app-degradation
handlingAgent: perf-investigation-agent
spec:
priorities: [Sev2, Sev3]
agentMode: Review # Requires human approval before any action
customInstructions: |
Step 1: Diagnose from AppInsights (dependencies, exceptions, traces).
Step 2: Check for recent deployments or traffic spikes.
Step 3: Propose remediation (scaling, code fix, or config change).
Step 4: Wait for human review before executing.Handoff Chains: Multi-Agent Orchestration
Complex incidents route through specialist subagents in sequence:
handlingAgent: incident-orchestrator
subagents:
- triage-agent # Classifies symptom (edge vs app tier)
- investigation-agent # Deep diagnostics
- remediation-agent # Proposes safe fixes
- rca-agent # Root-cause analysis + evidence
- pr-creation-agent # GitHub PR for durable fixEach agent runs with limited scope; orchestrator collects artifacts and routes to next phase.
Knowledge Injection via customInstructions
Step-by-step guidance dramatically improves agent accuracy:
customInstructions: |
# CosmosDB Incident Response SOP
**Triage (2 min):**
- Check Cosmos request units (RU) from Azure Portal metrics
- Query activity log for recent replication failures
- Look for throttling errors (429) in application logs
**Investigation (5 min):**
- Run: SELECT * FROM c WHERE c.error = "429" in past 15 min
- Count throttled requests by partition key
- Check write-region vs read-region latency
**Remediation (approval required):**
- Propose: increase RU from X to Y (cost delta: $Z/month)
- OR: archive old data to reduce hot set
- OR: rebalance partition key strategy
**Evidence:**
- Attach before/after metrics snapshot
- Document assumption: throttling is root causeMulti-Backend Deployment Patterns
IaC deployment is owned by the upstream microsoft/sre-agent repository under sreagent-templates/, not by this skill. This skill supplies bundle content (response plans, agents, hooks, scheduled tasks) that you drop into the config directory those templates generate. Do not hand-author agent ARM templates when the upstream templates already cover the resource shape.
The agent resource type is Microsoft.App/agents (verified 2026-08-09 against deploy-iac).
Generate a config directory from a recipe, then deploy it with one of four backends:
git clone https://github.com/microsoft/sre-agent.git
cd sre-agent/sreagent-templates
bash bin/check-prerequisites.sh
./bin/new-agent.sh --recipe azmon-lawappinsights --non-interactive \
--set agentName=my-agent \
--set resourceGroup=rg-my-agent \
--set location=australiaeast \
--set targetRGs=rg-my-workload \
-o my-agent/
./bin/deploy.sh my-agent/| Backend | Command | Use when |
|---|---|---|
| Bicep | ./bin/deploy.sh my-agent/ |
Default; runs az deployment sub create |
| Terraform | ./bin/deploy-tf.sh my-agent/ |
Terraform-managed infrastructure |
| PowerShell | .\bin\ps\Deploy-Agent.ps1 -InputPath .\my-agent\ |
Windows / PowerShell 7 environments |
| azd | cd my-agent/ && azd up |
azd-based workflows |
All backends support --what-if / --dry-run for validation without deploying. For azd specifically, the documented preview flag is azd provision --preview, there is no --no-exec flag.
Module status (verified 2026-08-10): no official Azure Verified Module (Bicep or Azure/avm-* Terraform) exists for Microsoft.App/agents, so the upstream hand-rolled templates remain best practice. A community Terraform AVM module (kewalaka/terraform-azure-avm-res-app-agent) deploys the agent with connectors, model settings, identity, and monitoring; treat it as a community reference, re-check its maintenance status before use.
Reference implementations worth reading before designing your own:
yortch/agentic-devops-demo— end-to-end Copilot + SRE Agent loop: subagent YAML (incident-handler, code-analyzer), knowledge-base docs, scheduled tasks, chaos-engineering workflow.microsoft.github.io/frontier-sre-agent-rvas— official RVAS reference architecture with named specialist subagents, response plans, and autonomous remediation chapters.matthansen0/azure-sre-agent-sandbox— automated AKS break-fix sandbox for testing loops safely.
Recipes are prebuilt starting points (Azure Monitor, PagerDuty, Dynatrace and others) and follow the same <platform>-<key-connectors> naming this skill uses. Browse sreagent-templates/recipes rather than assuming a recipe name.
Two deployment phases
Deployment is split, and the split determines what your IaC pipeline can and cannot guarantee.
| Phase | Mechanism | What it covers |
|---|---|---|
| Phase 1 — ARM | Bicep / Terraform / PowerShell / azd | Resource group, user-assigned managed identity, Log Analytics workspace, Application Insights, the Microsoft.App/agents resource, RBAC role assignments (Reader, Monitoring Reader, Log Analytics Reader, SRE Agent Administrator), plus connectors, skills, subagents and tools as ARM sub-resources |
| Phase 2 — data plane | apply-extras.sh (runs automatically after Phase 1) |
Code repositories (need Git PAT/OAuth), hooks (no ARM sub-resource yet), HTTP triggers (URLs generated server-side), knowledge files (binary upload), plugin configuration |
If the data-plane token is unavailable, restricted CI/CD, no interactive login, apply-extras.sh prints what it skipped instead of failing silently. Treat that output as a deployment gate: a Phase 1 success is not a deployed agent. When wiring this into azd, Phase 2 belongs in a postprovision hook, and the hook must fail closed if extras were skipped.
Day-2 operations
The upstream templates ship scripts for ongoing management: export (config directory from a running agent), clone (./bin/clone-agent.sh, export plus deploy under a new name and resource group), diff (local config against the live agent), and verify (a 22-point check across connectors, skills, subagents and hooks). Prefer export-then-diff over reading the portal when auditing drift.
Data-Plane Configuration Upload (Post-Deployment)
Response plans, custom agents, knowledge, and hooks don't deploy via ARM, they upload after infrastructure:
# 1. Package extras as tarball
tar -czf extras.tar.gz -C config/skills . -C config/agents .
# 2. Get auth token
TOKEN=$(az account get-access-token --resource https://azuresre.ai --query accessToken -o tsv)
# 3. Upload to data-plane
curl -X POST "https://${AGENT_ENDPOINT}/api/v2/extendedAgent/skills" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/gzip" \
--data-binary @extras.tar.gz
# 4. Upload knowledge with RAG indexing
curl -X POST "${AGENT_ENDPOINT}/api/v1/WorkspaceMemory/synthesized-knowledge?triggerIndexing=true" \
-F "files=@incident-template.md" \
-F "files=@troubleshooting-guide.md" \
-H "Authorization: Bearer $TOKEN"Knowledge uploaded with ?triggerIndexing=true is immediately available to agents for RAG grounding.
Connector Auth Orchestration
Multi-path auth for different platforms; choose the one your team already uses:
GitHub Auth Paths
# Path 1: OAuth (no env vars, browser sign-in)
export GITHUB_OAUTH_ENABLED=true
# User is prompted to visit: https://github.com/login/device
# Code displays in terminal; user enters on device
# Path 2: PAT (Headless / CI)
export GITHUB_PAT=ghp_xxxxxxxxxxxxxxxxxxxx
# Path 3: GitHub App (Multi-tenant)
export GITHUB_APP_ID=123456
export GITHUB_APP_PRIVATE_KEY=$(cat private-key.pem)ADO Auth Paths
# Path 1: Managed Identity (Recommended for Azure)
export ADO_USE_MI=true
# Agent automatically uses its system-assigned managed identity
# Path 2: Azure AD (Interactive)
export ADO_USE_AAD=true
# User is prompted to authenticate in browser
# Path 3: PAT (Legacy fallback)
export ADO_PAT=<pat_token_here>
export ADO_ORG=https://dev.azure.com/myorgLogic Apps / Functions Bridge for Private Endpoints
When SRE Agent cannot reach Git or CI/CD directly:
HTTP Trigger → Azure Functions (with network access)
↓
[Transform & auth check]
↓
GitHub/ADO internal endpoint
↓
[Return to SRE Agent]Incident Response Runbook Pattern
Six-step workflow embedded in knowledge; agent follows sequentially, collecting evidence at each step:
# Incident Response Runbook: Production API Degradation
## Step 1: Acknowledge (0–2 min)
- Send ack to PagerDuty (stops escalation)
- Create incident summary in Teams channel
- Assign investigator
## Step 2: Triage (2–5 min)
- Determine if issue is **edge** (networking, DNS, gateway) or **app** (compute, memory, dependencies)
- Check geography (global vs region-specific)
- Identify suspect component from alert context
## Step 3: Investigate (5–15 min)
- Query AppInsights for exception rates, latencies, dependencies
- Check infrastructure metrics: CPU, memory, disk, network
- Look for recent deployments or config changes
- Gather logs from target system
## Step 4: Mitigate (15–30 min)
- Apply smallest safe fix: scale, circuit-break, rollback, or temporary config
- Document decision and assumptions
- Monitor metrics for recovery confirmation
- Log action as incident note (audit trail)
## Step 5: RCA (30–60 min)
- Map facts: what happened, when, why, blast radius
- Link to root cause with evidence (logs, metrics, code)
- Identify whether fix is temporary or permanent
## Step 6: Close & Recommend (60+ min)
- Prepare PR to durable fix (code, IaC, policy, or KT)
- Add alert tuning if false positive rate is high
- Schedule follow-up to prevent recurrenceExport → Clone → Verify Workflow
Full agent lifecycle: export current → validate → clone to new → verify all configs:
# 1. Export running agent to portable format
./bin/export-agent.sh \
--subscription "$SUB" \
--resource-group "$RG" \
--agent-name "$AGENT" \
--output exported-config.json
# 2. Validate schema
./bin/verify-agent.sh expected-config.json
# 3. Clone to new environment
./bin/deploy.sh \
--source exported-config.json \
--target-rg rg-new-agent \
--target-location swedencentral
# 4. Run 22-point verification checklist
./bin/verify-agent.sh exported-config.json
# Checks: skills, subagents, hooks, knowledge, response plans, connectors, tasks, etc.Export output includes: skills, subagents, hooks, common prompts, response plans, scheduled tasks, connectors, incident platform, model provider, access level, run mode, and upgrade channel.
Common Prompt Safety Rules
Share across all custom agents to prevent dangerous patterns:
# Safety Rules (inject into all agent prompts)
1. **Never delete production resources without explicit approval.**
- Log all delete operations as incident notes before execution.
- Require human confirmation in a tracked approval message.
2. **Log all actions for audit and compliance.**
- Record what changed, who authorized it, timestamp, and evidence.
- Use Application Insights customEvents or Azure audit logs.
3. **Rollback on error; do not retry indefinitely.**
- If 3 attempts fail, escalate and stop.
- Preserve error state for investigation.
4. **Alert on cost anomalies.**
- If proposed action increases monthly spend > 10%, notify owner.
- Require explicit approval for high-cost changes.
5. **Test in non-prod first.**
- Never run write operations in production on first attempt.
- Validate logic and output in staging.Webhook Bridge for Third-Party Platforms
When Dynatrace, Grafana, or Splunk webhook directly to SRE Agent via HTTP trigger:
Dynatrace/Grafana/Splunk
↓ (webhook)
Azure Functions (auth bridge)
↓
SRE Agent HTTP Trigger
↓
Investigation + response planFunction bridge responsibilities:
- Validate webhook signature
- Transform vendor JSON to SRE Agent payload schema
- Inject correlation ID and tenant context
- Forward to HTTP trigger with Bearer token
HTTP trigger authentication:
# SRE Agent response plan tied to HTTP trigger
metadata:
name: external-webhook-receiver
spec:
httpTrigger:
enabled: true
authLevel: Function # Requires API key in Function auth settings
methods: [POST]
route: "incidents/external"Post-Deployment Verification Checklist
22-point validation to confirm agent is production-ready:
✅ Connectivity & Identity
[ ] Agent identity (MI or SP) has correct RBAC on target subscriptions
[ ] Connectors can reach GitHub, ADO, observability APIs
[ ] MCP servers (if configured) are reachable and authenticated
✅ Configuration & Scope
[ ] Response plans exist and have correct incident platform routing
[ ] Scheduled tasks are enabled and have correct cron expressions
[ ] HTTP triggers are disabled until readiness testing complete
[ ] Skill and custom-agent catalog loaded without errors
✅ Observability & Cost
[ ] AMBA alerts firing correctly; no alert storms detected
[ ] AAU budget set and monitored in Azure Monitor
[ ] Incident telemetry flowing to Application Insights
✅ Governance
[ ] Approval hooks configured for write actions
[ ] Run-mode posture matches blast radius (ReadOnly/Review/Autonomous)
[ ] Access level is Low unless High explicitly justified
✅ Knowledge & Runbooks
[ ] At least one runbook or incident template loaded
[ ] Knowledge indexing completed (check Application Insights customEvents)
✅ Rollout Safety
[ ] Rollback procedure documented and tested
[ ] On-call team notified and trained
[ ] Cost and AAU forecast reviewedSLI-to-Incident Bridge
Azure SRE Agent consumes Azure Monitor SLI breaches as incident triggers. To design SLIs and set alert thresholds:
- Define SLIs & SLOs using the
observability-monitoringskill, this covers Azure Monitor Service Groups, burn-rate targets, and SLO tracking. - Create AMBA alerts on fast-burn and slow-burn conditions that translate SLO targets into alert rules.
- Route alerts to Azure SRE Agent — the agent ingests alert payloads, matches them to response plans, and dispatches runbooks.
The division of responsibility is clear: Observability skill owns SLI/SLO definition; Azure SRE Agent owns incident response to SLI breaches. Together they form the end-to-end incident management lifecycle.
Summary
All 10 patterns are production-tested and source-grounded in the microsoft/sre-agent repository. Use this guide as a starting point; customize frameworks, runbooks, and safety rules for your operational environment.