---
name: azure-sre-agent
description: >-
  Design, configure, review, and operate production-grade Azure SRE Agent
  capabilities: response plans, scheduled tasks, HTTP triggers, custom agents,
  autonomous and review workflows, approval guardrails, AMBA observability,
  source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry
  posture, Digital Native governance, postmortem generation, and KT discipline.
compatibility: Azure SRE Agent; GitHub Copilot agent skills; new projects only
metadata:
  last_verified: "2026-08-25"
  version: "2.23.3"
  risk: "critical"
  last_updated: "2026-08-25"
---

# Azure SRE Agent Skill

Use this skill for **new Azure projects** that need Azure SRE Agent designed, reviewed, or productionized. Route broad Azure incident diagnosis to sibling troubleshooting skills unless the work is explicitly about SRE Agent capabilities, workflows, or governance.

## Overview

This skill helps you design and review production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, connectors, source-aware RCA, governance hooks, and AMBA-aligned observability. It assumes greenfield or newly structured deployments, not legacy preservation.

**SLI-to-incident flow:** Azure SRE Agent **consumes SLI breaches as incident triggers**. For SLI/SLO design, Azure Monitor baselines, and target-setting, use the `observability-monitoring` skill; this skill owns response and remediation execution.

## Use when

| Request pattern                                                                     | Use this skill? | Route or output                                      |
| ----------------------------------------------------------------------------------- | --------------- | ---------------------------------------------------- |
| Design a new Azure SRE Agent capability, response plan, trigger, or custom agent    | Yes             | Bundle selection, workflow, governance, rollout plan |
| Map Azure Monitor or AMBA alerts into SRE Agent workflows                           | Yes             | AMBA baseline, alert routing, autonomy posture       |
| Add source RCA, PR remediation, or GitHub and Azure DevOps integration to SRE Agent | Yes             | Source-aware investigation design                    |
| Review connector security, approval hooks, or SRE Agent RBAC                        | Yes             | Least-privilege and governance review                |
| Diagnose an Azure outage with no SRE Agent in scope                                 | No              | Use `../azure-troubleshooting/SKILL.md`              |
| Design or operate Azure Monitor Health Models (state rollup, health-state alerts)    | No              | Use `../azure-health-models/SKILL.md`               |
| General observability architecture without SRE Agent workflows                      | No              | Use `../observability-monitoring/SKILL.md`           |
| Generic product-agnostic autonomous-agent loop outside the SRE Agent surface        | No              | Use `../autonomous-agent-loops/SKILL.md`             |

## Production defaults

| Decision area     | Default                                                                                                                                 | Avoid                                                          |
| ----------------- | --------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- |
| Run mode          | Start in `ReadOnly` when scope/ownership is unclear, otherwise `Review`; promote one workflow at a time. Both surfaces use **`Autonomous`**; the data-plane response-plan value is **`autonomous`** (`automatic` is rejected everywhere). Note: the platform's own default for new response plans and scheduled tasks is `Autonomous` — this package deliberately overrides that to `Review` for production hardening | Defaulting new production automation to autonomous |
| Access level      | Start agent `accessLevel` at `Low`; raise to `High` only with explicit blast-radius, RBAC, and rollback decisions                       | Provisioning `High` access by default                          |
| Model tier        | Default General Purpose per workflow; reserve higher tiers for deep RCA where AAU cost is accepted                                      | A single global high-cost tier for all triggers/tasks          |
| Alerting baseline | Start from AMBA, then tune to SLOs and alert-noise evidence                                                                             | Copying baseline alerts unchanged into production              |
| Connector scope   | Enable only required operations, tools, repos, and RBAC scopes; stay within the 80-tool/agent budget                                    | Catch-all agents and broad write-capable connectors            |
| KT depth          | Full `SA -> PA -> DA -> PPA` for P1/P2; lightweight KT for lower severity; minimum `DA + PPA` for write actions                         | Forcing full KT for low-risk read-only tasks                   |
| Cost posture      | Budget for **4 AAUs per agent-hour fixed cost** plus model-dependent active-flow AAUs                                                   | Treating AAU cost as fixed only                                |

## Autonomy Progression Model

Promote workflows through levels based on demonstrated success, not configuration changes alone. Gate each transition on evaluation data, not intent.

| Level | Run Mode | Gate Criteria | Human Role |
| --- | --- | --- | --- |
| L0 - Manual | `ReadOnly` | Initial onboarding or untrusted scope | Investigates, approves, actuates all actions |
| L1 - Assisted | `ReadOnly` | Agent provides investigation insights; human actuates | Reviews agent findings, executes mitigations |
| L2 - Review | `Review` | Agent proposes mitigations; human approves each | Approves or rejects agent-proposed actions |
| L3 - Autonomous (bounded) | ARM `Autonomous` + plan `autonomous` | Sustained high success rate against Gold evaluation data for a specific, narrow workflow | Monitors, handles novel situations |
| L4 - Autonomous (multi-step) | ARM `Autonomous` + plan `autonomous` | Agent handles end-to-end incident lifecycle for well-defined scenarios; adapts strategy based on outcomes | Architectural governance, safety guardrail design |

**Progression rules:** promote one workflow at a time; L2→L3 requires demonstrated success against Gold-quality evaluation data (`references/knowledge-lifecycle.md`); L3→L4 requires multi-step resolution (diagnose → mitigate → verify → adapt); demote on any sustained increase in false positives or missed detections.

## Autonomy safety rules (live-verified)

Each rule below was reproduced against a live deployment; full evidence, error messages, and worked examples in `references/live-verified-operations.md`. Read that file before promoting any workflow or writing guardrails.

1. **Plan-level `review` does NOT restrain an agent-level `Autonomous`.** ARM `actionConfiguration.mode` is the governing control; supervised operation means ARM `mode: Review`, whatever plans say.
2. **Per-plan promotion works directly** — an earlier "raise ARM first" reading misread an invalid-enum rejection. `automatic` is invalid everywhere; set `agentMode: autonomous` directly, and enumerate every plan's mode explicitly before changing the agent-level default.
3. **Graduated autonomy needs two agents with two identities** (supervised + autonomous), never shared identities, a custom no-delete AKS role for the autonomous one, and no self-administration. Watch the two custom-role deploy-blockers (`AssignableScopeMismatch`, `InvalidDataActionOrNotDataAction`); neither fails at compile time.
4. **An agent without `incidentManagementConfiguration` deploys `Succeeded` then rejects every plan create with a bare 405**; repair via redeploy, not `az resource update` (redacted connection string breaks the round-trip).
5. **Audit autonomous actions via the `kube-audit-admin` Log Analytics category**, not Activity Log (data-plane writes are invisible there) and not `IncidentActivitySnapshot` (proves investigation, not action, and misattributes the acting identity).
6. **Read every hook's actual prompt before trusting its name** — a hook named for approval shipped as a self-judged gate that approved everything plausible. Match hooks to PascalCase tool names including `Terminal`/`RunInTerminal`.
7. **Incident titles are Prometheus rule-GROUP names**, so `titleContainsAny` routing misfires; overlapping plans silently suppress investigation (zero tool calls). One rule group per routing intent; verify outcomes via `ResponsePlanId`.
8. **Control stack order is hooks → tool access policies → connector Ask → run mode** — and a user-defined hook returning allow overrides even a global policy deny. Use global deny policies for coarse blocks and `toolGlob(argGlob)` argument patterns for boundaries RBAC can't express; audit every hook-allow override (`references/hooks-governance.md`).

## Stop conditions

Stop and require human review when rollback is unclear, run mode or ownership is unknown, the design spans multi-subscription or multi-tenant write scope without explicit isolation, or schema and auth behavior cannot be verified from current docs.

## Model-output reset

Before producing output, identify the workload, owner, environment, incident source, and blast radius; prefer evidence, guardrails, and rollout steps over slogans; keep AMBA, connectors, hooks, and KT proportional to the use case; and mark freshness gaps instead of smoothing them away.

## Anti-hallucination rule

**Trigger:** before writing YAML, API payloads, hook settings, trigger auth guidance, AAU pricing guidance, or region-sensitive deployment advice.

**Verification methods:** verify current Azure SRE Agent schema, run modes, connector behavior, pricing, and supported regions against `references/source-map.md`, Microsoft Learn, and official `microsoft/sre-agent` artifacts.

**Forbidden shortcuts:** do not guess `azuresre.ai` schema fields, hook shapes, token audiences, operation names, pricing tiers, region lists, or plugin catalog behavior.

**Safe degraded output:** if a field, endpoint, auth audience, or behavior cannot be verified, mark it `[VERIFY]`, point to the source to check, and stop at a reviewable seam. Do not invent a fallback workflow or fake-safe production path.

## Live-verified operations

Load `references/live-verified-operations.md` when promoting to autonomous, writing/auditing hooks, debugging engagement, verifying incident routing, auditing autonomous writes, or working with the data-plane API. It carries the verified loop shape, hook matcher mechanics, RBAC-vs-ingestion gaps, the kube-audit-admin audit path, and the data-plane path table (audience `https://azuresre.dev`; PUT creates / POST updates).

## Bundle Routing

Open only the bundle and reference files needed for the current task. When populating `@@PLACEHOLDER@@` values in bundle templates, use `parameters.example.yaml` as the substitution checklist.

| User intent                                                      | Load                                                                                                               |
| ---------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| Minimal production baseline                                      | `bundles/base-core/`, `references/production-blueprints.md`                                                        |
| AMBA, Azure Monitor, alert baselines, alert tuning               | `bundles/observability-amba/`, `references/observability-amba.md`                                                  |
| HTTP/event-driven invocation from CI/CD or alerts                | `bundles/http-triggers-production/`, `references/http-triggers-production.md`                                      |
| HTTP trigger auth bridge via Functions, Logic Apps, or APIM      | `bundles/http-trigger-auth-bridges/`, `references/http-trigger-auth-bridges.md`                                    |
| Source-code RCA, deployment regression, PR remediation           | `bundles/source-rca-remediation/`, `references/source-rca-remediation.md`                                          |
| Knowledge, memory, runbook lifecycle                             | `bundles/knowledge-lifecycle/`, `references/knowledge-lifecycle.md`                                                |
| Identity, RBAC, OBO fallback, private network                    | `bundles/security-identity/`, `references/security-identity-rbac.md`                                               |
| Value tracking, autonomy promotion, AAU governance               | `bundles/operational-metrics/`, `references/value-promotion-scorecard.md`                                          |
| Common Azure workload response-plan seeds                        | `bundles/azure-workload-production/`, `references/workload-bundle-seeding.md`                                      |
| AKS incidents or health checks                                   | `bundles/aks-production/`, `references/aks-containerapps-production.md`                                            |
| Container Apps incidents or revisions                            | `bundles/containerapps-production/`, `references/aks-containerapps-production.md`                                  |
| Drasi on AKS                                                     | `bundles/drasi-aks-production/`, `references/drasi-aks-playbook.md`                                                |
| Run mode, access level, model tier, upgrade channel, tool budget | `references/run-posture-and-operational-levers.md`                                                                 |
| Live-verified run modes, hooks, routing, data-plane API          | `references/live-verified-operations.md`                                                                           |
| KT, approval, audit, write-action governance                     | `bundles/governance-kt/`, `references/kt-methodology.md`, `references/hooks-governance.md`                         |
| Observability or incident connectors                             | `bundles/connectors-observability/`, `references/connectors-and-mcp.md`, `references/connector-token-security.md`  |
| Teams/GitHub handoff                                             | `bundles/connectors-collab-handoff/`, `references/connectors-and-mcp.md`                                           |
| Bundle maintenance                                               | `bundles/catalog.yaml`, `bundles/README.md`, `references/bundles-operations.md`, `references/capability-matrix.md` |
| Audit diagnostics or 2am investigation                           | `references/audit-diagnostics-2am.md`                                                                              |
| Source verification                                              | `references/source-map.md`                                                                                         |
| Output shaping/examples                                          | `references/output-examples.md`                                                                                    |
| Project/resource isolation and placement                         | `references/project-boundary-resource-placement.md`                                                                |
| First deployment (Quickstart)                                    | `references/quickstart-first-deployment.md`                                                                        |
| IaC deployment, recipes, backends, day-2 export/clone/diff/verify | `references/production-patterns-guide.md#multi-backend-deployment-patterns`                                        |
| HTTP trigger auth troubleshooting (token audience conflict)      | `references/http-triggers-production.md#authentication`                                                            |

## Workflow

For new production projects, build capabilities in this order:

1. Establish owner, region, incident platform, identity model, baseline connectors, AAU limit, and governance hooks.
2. Verify the current deployment region. Supported regions last verified 2026-08-25 (18): Australia East, Canada Central, Central US, East Asia, East US 2, France Central, Italy North, Japan East, Korea Central, North Central US, South Africa North, Southeast Asia, Spain Central, Sweden Central, UK South, West Central US, West US 2, West US 3. An agent's region is fixed at creation and cannot be changed; it can still manage resources elsewhere where its identity has permissions. Confirm subscription-specific availability before cutover.
3. Deploy AMBA-aligned monitoring or map existing alerts to AMBA categories before creating response plans.
4. Classify alerts as `investigate`, `notify`, `digest`, `tune`, or `candidate-for-autonomy` before routing them to agents.
5. Create specialist custom agents by concern: diagnostics, source RCA, remediation review, notification, workload specialist, and knowledge capture. Six generic subagents ship built in (Explore, Plan, CodeReview, Bash, Verification, GeneralPurpose) alongside your custom ones.
6. Create response plans and HTTP triggers in `Review`; apply cooldown and noise controls before any high-volume trigger goes live.
7. Add source code and IaC context where deployment correlation matters, then add knowledge lifecycle tasks only after first useful investigations.
8. Test new response plans and HTTP triggers in `Review` long enough to capture Intent Met, false positives, approval outcomes, and AAU cost before autonomy promotion.
9. Promote only narrow, repeatedly validated, low-blast-radius workflows to Autonomous.

## Recipe Design Framework

Production SRE Agent capabilities are shaped by five intersecting dimensions:

| Dimension      | Options / Examples                                                                      | Decision driver                                            |
| -------------- | --------------------------------------------------------------------------------------- | ---------------------------------------------------------- |
| **Platform**   | AzMonitor, PagerDuty, Dynatrace, ServiceNow, GitHub Actions, Azure DevOps, Grafana      | Where incidents originate; native webhook support          |
| **Connectors** | AppInsights, LAW (Log Analytics), Kusto, MCP integrations, GitHub, Azure DevOps APIs    | Which observability/CI/CD systems feed investigation       |
| **Skills**     | Investigation playbooks by workload: VM, CosmosDB, AKS, Container Apps, HTTP errors     | The target system or error class you need to automate      |
| **Knowledge**  | Runbooks, incident templates, RCA patterns, troubleshooting guides, approval checklists | Context and procedures the agent needs to reach conclusion |
| **Response**   | Severity-based routing, agentMode (Autonomous/Review), customInstructions, max attempts | How the agent behaves and when it escalates or acts        |

Design in this order: platform → connectors → skills → knowledge → response mode. Test each layer independently before combining. Name recipes `<platform>-<key-connectors>[-workload-variant]` (good: `pagerduty-law-vmcosmos`; avoid: `agent-automation`).

See [Production Patterns Guide](references/production-patterns-guide.md) for recipe patterns, multi-backend deployment, connector auth orchestration, export→clone→verify workflows, and post-deployment verification checklists.

## Resource Boundary Defaults

Last reviewed: **2026-08-25**. The `Microsoft.App/agents` resource has no diagnostic-settings or private-endpoint support; audit via the agent's own Application Insights `customEvents`, and reach private workloads via VNet integration (preview, `vnetConfiguration.subnetResourceId`); see `references/security-identity-rbac.md` and `references/run-posture-and-operational-levers.md`.
