Deployment Patterns from Official Samples
This guide captures production-grade patterns seen in official microsoft/sre-agent
sample deployments and scripts.
Pattern 1: Idempotent Post-Provision Setup
Post-provision automation should be safe to re-run.
Use upsert semantics for:
- Knowledge base uploads
- Custom agents
- Incident response plans
- Scheduled tasks
- Connectors
Design goal: "fix-forward" without redeploying infrastructure.
Pattern 2: Eventual Consistency Retries
Some operations are not immediately available after provisioning.
Use bounded retry loops for:
- incident platform availability
- response plan creation/update
- connector readiness
Recommended defaults:
- 3 attempts
- 10-30 second delay
- clear failure message with operator next step
Pattern 3: Quickstart Cleanup
If quickstart plans are auto-created, remove overlap before custom routes.
Automation step:
- list active plans
- detect quickstart handler
- disable/delete when custom equivalent exists
Pattern 4: Full Verification Pass
After setup, verify all control-plane objects:
- knowledge files indexed
- custom agents present
- incident platform connected
- connectors healthy
- response plans active
- scheduled tasks active
Return a single readiness summary.
Pattern 5: Memory-First Investigation
Investigation agents should search memory before fresh diagnostics.
Default flow:
- memory lookup for similar incidents
- runbook lookup
- logs + metrics evidence
- root-cause hypothesis
- action proposal
Pattern 6: Structured Incident Report Templates
Use a fixed report schema across teams:
- Summary
- Impact
- Timeline (UTC)
- Evidence
- Root Cause
- Remediation
- Action Items
- References
This supports reliable handoffs and postmortems.
Pattern 7: Scheduled Issue Triage Workflows
For code-facing operations:
- scheduled task triggers issue triage agent
- triage agent classifies and labels issues
- triage agent posts standardized response comment
- escalation for P0/P1 style issues
Pattern 8: KT-Governed Major Incident Handling
For high-severity incidents and production write paths:
- perform Situation Appraisal before deep investigation
- produce Problem Analysis with explicit
IS / IS NOTlogic - evaluate remediation options with Decision Analysis (
MUST/WANT) - protect execution plan with Potential Problem Analysis
Use hooks to reject incomplete major-incident responses.
Pattern 9: Starter Lab Deployment (azd-based)
The official starter lab at microsoft/sre-agent/labs/starter-lab uses azd up for deployment:
- Deploy: SRE Agent, sample app (Grubify), Log Analytics, App Insights, Alert, Container Registry, Managed Identity.
- Three scenario tracks: IT Operations, Developer (GitHub), Workflow Automation.
- Estimated time: ~40 minutes.
Use this as a reference for azd-based SRE Agent provisioning patterns. See https://github.com/microsoft/sre-agent/tree/main/labs/starter-lab.
Pattern 10: Event-driven Terraform Drift via HTTP Triggers
Use HTTP triggers to process drift notifications from Terraform Cloud or other webhook-capable drift tools.
Reference architecture:
- Drift system emits webhook payload.
- Logic App / Azure Functions / APIM validates source and acquires Azure token.
- Authenticated POST invokes SRE Agent HTTP trigger.
- Trigger routes to a drift-analysis custom agent in Review mode first.
Use this for event-driven operations. Keep scheduled tasks for periodic checks.
Pattern 11: Drift Investigation Evidence Contract
For each drift event, gather and correlate:
- Terraform desired state and actual Azure resource state.
- Activity Log actor and timestamp for each drifted change.
- Incident impact signals from Application Insights or Azure Monitor.
- Recent code or deployment changes that explain the drift.
This prevents unsafe "diff-only" remediation.
Pattern 12: Severity-Gated Remediation for Drift
Use severity classes with explicit actions:
- Benign (tags/metadata) -> safe rollback path.
- Risky (security downgrade) -> prioritize rollback with security evidence.
- Critical (incident mitigation change, such as emergency scale-up) -> do not auto-revert until root cause is fixed and rollback is safe.
Hard rule:
- Never auto-revert critical drift that is actively mitigating a live incident.
- Enforce approval + rollback evidence before executing critical write actions.
Pattern 13: Replay, Verification, and Learning Loop
Before enabling autonomous drift remediation:
- Validate with historical incident payloads and simulated webhook runs.
- Verify trigger execution history and report quality.
- Run execution review and update drift skill instructions when gaps are found.
- Track recurring drift classes to evolve bundle templates and governance hooks.
AKS and Container Apps Production Pattern
Split specialist custom agents
aks-diagnosticsfor cluster/pod/network diagnosticscontainerapps-revision-analyzerfor revision health and rollback evidenceincident-notifierfor outbound notifications and stakeholder updates
Keep responsibilities separate
- Diagnose agents collect evidence.
- Remediation agents execute approved changes.
- Notifier agents report outcomes.
Drasi on AKS Pattern
Use this when SRE Agent supports Drasi workloads on AKS:
- Detect symptom from alert:
- event processing lag
- query staleness
- connector ingestion faults
- Gather evidence:
- AKS node/pod health
- Drasi operator/controller logs
- backing data plane dependencies
- Correlate:
- deployment/revision changes
- config drift
- recent cluster/network events
- Route:
- infrastructure issue -> AKS remediator path
- app/workflow issue -> Drasi workflow owner path
- Report:
- service impact
- likely failure domain
- recommended action and risk
Use drasi-aks-playbook.md for a ready-to-use Drasi-specific operating template.
Sources
- Official samples: https://github.com/microsoft/sre-agent
- Hands-on lab patterns: https://github.com/microsoft/sre-agent/tree/main/labs
- Deployment compliance sample: https://github.com/microsoft/sre-agent/tree/main/labs/deployment-compliance
- HTTP triggers (official docs): https://learn.microsoft.com/azure/sre-agent/http-triggers
- HTTP trigger tutorial: https://learn.microsoft.com/azure/sre-agent/create-http-trigger
- Terraform drift sample: https://github.com/microsoft/sre-agent/tree/main/labs/terraform-drift-detection
- Event-driven Terraform drift blog: https://techcommunity.microsoft.com/blog/appsonazureblog/event-driven-iac-operations-with-azure-sre-agent-terraform-drift-detection-via-h/4512233