All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

SKILL.md

≈97 tokens always: the name and description. ≈4.9k when used: this file. ≈110k more on demand in 190 files.

Azure SRE Agent Skill

Use this skill for new Azure projects that need Azure SRE Agent designed, reviewed, or productionized. Route broad Azure incident diagnosis to sibling troubleshooting skills unless the work is explicitly about SRE Agent capabilities, workflows, or governance.

Overview

This skill helps you design and review production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, connectors, source-aware RCA, governance hooks, and AMBA-aligned observability. It assumes greenfield or newly structured deployments, not legacy preservation.

SLI-to-incident flow: Azure SRE Agent consumes SLI breaches as incident triggers. For SLI/SLO design, Azure Monitor baselines, and target-setting, use the observability-monitoring skill; this skill owns response and remediation execution.

Use when

Request pattern Use this skill? Route or output
Design a new Azure SRE Agent capability, response plan, trigger, or custom agent Yes Bundle selection, workflow, governance, rollout plan
Map Azure Monitor or AMBA alerts into SRE Agent workflows Yes AMBA baseline, alert routing, autonomy posture
Add source RCA, PR remediation, or GitHub and Azure DevOps integration to SRE Agent Yes Source-aware investigation design
Review connector security, approval hooks, or SRE Agent RBAC Yes Least-privilege and governance review
Diagnose an Azure outage with no SRE Agent in scope No Use ../azure-troubleshooting/SKILL.md
Design or operate Azure Monitor Health Models (state rollup, health-state alerts) No Use ../azure-health-models/SKILL.md
General observability architecture without SRE Agent workflows No Use ../observability-monitoring/SKILL.md
Generic product-agnostic autonomous-agent loop outside the SRE Agent surface No Use ../autonomous-agent-loops/SKILL.md

Production defaults

Decision area Default Avoid
Run mode Start in ReadOnly when scope/ownership is unclear, otherwise Review; promote one workflow at a time. Both surfaces use Autonomous; the data-plane response-plan value is autonomous (automatic is rejected everywhere). Note: the platform's own default for new response plans and scheduled tasks is Autonomous — this package deliberately overrides that to Review for production hardening Defaulting new production automation to autonomous
Access level Start agent accessLevel at Low; raise to High only with explicit blast-radius, RBAC, and rollback decisions Provisioning High access by default
Model tier Default General Purpose per workflow; reserve higher tiers for deep RCA where AAU cost is accepted A single global high-cost tier for all triggers/tasks
Alerting baseline Start from AMBA, then tune to SLOs and alert-noise evidence Copying baseline alerts unchanged into production
Connector scope Enable only required operations, tools, repos, and RBAC scopes; stay within the 80-tool/agent budget Catch-all agents and broad write-capable connectors
KT depth Full SA -> PA -> DA -> PPA for P1/P2; lightweight KT for lower severity; minimum DA + PPA for write actions Forcing full KT for low-risk read-only tasks
Cost posture Budget for 4 AAUs per agent-hour fixed cost plus model-dependent active-flow AAUs Treating AAU cost as fixed only

Autonomy Progression Model

Promote workflows through levels based on demonstrated success, not configuration changes alone. Gate each transition on evaluation data, not intent.

Level Run Mode Gate Criteria Human Role
L0 - Manual ReadOnly Initial onboarding or untrusted scope Investigates, approves, actuates all actions
L1 - Assisted ReadOnly Agent provides investigation insights; human actuates Reviews agent findings, executes mitigations
L2 - Review Review Agent proposes mitigations; human approves each Approves or rejects agent-proposed actions
L3 - Autonomous (bounded) ARM Autonomous + plan autonomous Sustained high success rate against Gold evaluation data for a specific, narrow workflow Monitors, handles novel situations
L4 - Autonomous (multi-step) ARM Autonomous + plan autonomous Agent handles end-to-end incident lifecycle for well-defined scenarios; adapts strategy based on outcomes Architectural governance, safety guardrail design

Progression rules: promote one workflow at a time; L2→L3 requires demonstrated success against Gold-quality evaluation data (references/knowledge-lifecycle.md); L3→L4 requires multi-step resolution (diagnose → mitigate → verify → adapt); demote on any sustained increase in false positives or missed detections.

Autonomy safety rules (live-verified)

Each rule below was reproduced against a live deployment; full evidence, error messages, and worked examples in references/live-verified-operations.md. Read that file before promoting any workflow or writing guardrails.

  1. Plan-level review does NOT restrain an agent-level Autonomous. ARM actionConfiguration.mode is the governing control; supervised operation means ARM mode: Review, whatever plans say.
  2. Per-plan promotion works directly — an earlier "raise ARM first" reading misread an invalid-enum rejection. automatic is invalid everywhere; set agentMode: autonomous directly, and enumerate every plan's mode explicitly before changing the agent-level default.
  3. Graduated autonomy needs two agents with two identities (supervised + autonomous), never shared identities, a custom no-delete AKS role for the autonomous one, and no self-administration. Watch the two custom-role deploy-blockers (AssignableScopeMismatch, InvalidDataActionOrNotDataAction); neither fails at compile time.
  4. An agent without incidentManagementConfiguration deploys Succeeded then rejects every plan create with a bare 405; repair via redeploy, not az resource update (redacted connection string breaks the round-trip).
  5. Audit autonomous actions via the kube-audit-admin Log Analytics category, not Activity Log (data-plane writes are invisible there) and not IncidentActivitySnapshot (proves investigation, not action, and misattributes the acting identity).
  6. Read every hook's actual prompt before trusting its name — a hook named for approval shipped as a self-judged gate that approved everything plausible. Match hooks to PascalCase tool names including Terminal/RunInTerminal.
  7. Incident titles are Prometheus rule-GROUP names, so titleContainsAny routing misfires; overlapping plans silently suppress investigation (zero tool calls). One rule group per routing intent; verify outcomes via ResponsePlanId.
  8. Control stack order is hooks → tool access policies → connector Ask → run mode — and a user-defined hook returning allow overrides even a global policy deny. Use global deny policies for coarse blocks and toolGlob(argGlob) argument patterns for boundaries RBAC can't express; audit every hook-allow override (references/hooks-governance.md).

Stop conditions

Stop and require human review when rollback is unclear, run mode or ownership is unknown, the design spans multi-subscription or multi-tenant write scope without explicit isolation, or schema and auth behavior cannot be verified from current docs.

Model-output reset

Before producing output, identify the workload, owner, environment, incident source, and blast radius; prefer evidence, guardrails, and rollout steps over slogans; keep AMBA, connectors, hooks, and KT proportional to the use case; and mark freshness gaps instead of smoothing them away.

Anti-hallucination rule

Trigger: before writing YAML, API payloads, hook settings, trigger auth guidance, AAU pricing guidance, or region-sensitive deployment advice.

Verification methods: verify current Azure SRE Agent schema, run modes, connector behavior, pricing, and supported regions against references/source-map.md, Microsoft Learn, and official microsoft/sre-agent artifacts.

Forbidden shortcuts: do not guess azuresre.ai schema fields, hook shapes, token audiences, operation names, pricing tiers, region lists, or plugin catalog behavior.

Safe degraded output: if a field, endpoint, auth audience, or behavior cannot be verified, mark it [VERIFY], point to the source to check, and stop at a reviewable seam. Do not invent a fallback workflow or fake-safe production path.

Live-verified operations

Load references/live-verified-operations.md when promoting to autonomous, writing/auditing hooks, debugging engagement, verifying incident routing, auditing autonomous writes, or working with the data-plane API. It carries the verified loop shape, hook matcher mechanics, RBAC-vs-ingestion gaps, the kube-audit-admin audit path, and the data-plane path table (audience https://azuresre.dev; PUT creates / POST updates).

Bundle Routing

Open only the bundle and reference files needed for the current task. When populating @@PLACEHOLDER@@ values in bundle templates, use parameters.example.yaml as the substitution checklist.

User intent Load
Minimal production baseline bundles/base-core/, references/production-blueprints.md
AMBA, Azure Monitor, alert baselines, alert tuning bundles/observability-amba/, references/observability-amba.md
HTTP/event-driven invocation from CI/CD or alerts bundles/http-triggers-production/, references/http-triggers-production.md
HTTP trigger auth bridge via Functions, Logic Apps, or APIM bundles/http-trigger-auth-bridges/, references/http-trigger-auth-bridges.md
Source-code RCA, deployment regression, PR remediation bundles/source-rca-remediation/, references/source-rca-remediation.md
Knowledge, memory, runbook lifecycle bundles/knowledge-lifecycle/, references/knowledge-lifecycle.md
Identity, RBAC, OBO fallback, private network bundles/security-identity/, references/security-identity-rbac.md
Value tracking, autonomy promotion, AAU governance bundles/operational-metrics/, references/value-promotion-scorecard.md
Common Azure workload response-plan seeds bundles/azure-workload-production/, references/workload-bundle-seeding.md
AKS incidents or health checks bundles/aks-production/, references/aks-containerapps-production.md
Container Apps incidents or revisions bundles/containerapps-production/, references/aks-containerapps-production.md
Drasi on AKS bundles/drasi-aks-production/, references/drasi-aks-playbook.md
Run mode, access level, model tier, upgrade channel, tool budget references/run-posture-and-operational-levers.md
Live-verified run modes, hooks, routing, data-plane API references/live-verified-operations.md
KT, approval, audit, write-action governance bundles/governance-kt/, references/kt-methodology.md, references/hooks-governance.md
Observability or incident connectors bundles/connectors-observability/, references/connectors-and-mcp.md, references/connector-token-security.md
Teams/GitHub handoff bundles/connectors-collab-handoff/, references/connectors-and-mcp.md
Bundle maintenance bundles/catalog.yaml, bundles/README.md, references/bundles-operations.md, references/capability-matrix.md
Audit diagnostics or 2am investigation references/audit-diagnostics-2am.md
Source verification references/source-map.md
Output shaping/examples references/output-examples.md
Project/resource isolation and placement references/project-boundary-resource-placement.md
First deployment (Quickstart) references/quickstart-first-deployment.md
IaC deployment, recipes, backends, day-2 export/clone/diff/verify references/production-patterns-guide.md#multi-backend-deployment-patterns
HTTP trigger auth troubleshooting (token audience conflict) references/http-triggers-production.md#authentication

Workflow

For new production projects, build capabilities in this order:

  1. Establish owner, region, incident platform, identity model, baseline connectors, AAU limit, and governance hooks.
  2. Verify the current deployment region. Supported regions last verified 2026-08-25 (18): Australia East, Canada Central, Central US, East Asia, East US 2, France Central, Italy North, Japan East, Korea Central, North Central US, South Africa North, Southeast Asia, Spain Central, Sweden Central, UK South, West Central US, West US 2, West US 3. An agent's region is fixed at creation and cannot be changed; it can still manage resources elsewhere where its identity has permissions. Confirm subscription-specific availability before cutover.
  3. Deploy AMBA-aligned monitoring or map existing alerts to AMBA categories before creating response plans.
  4. Classify alerts as investigate, notify, digest, tune, or candidate-for-autonomy before routing them to agents.
  5. Create specialist custom agents by concern: diagnostics, source RCA, remediation review, notification, workload specialist, and knowledge capture. Six generic subagents ship built in (Explore, Plan, CodeReview, Bash, Verification, GeneralPurpose) alongside your custom ones.
  6. Create response plans and HTTP triggers in Review; apply cooldown and noise controls before any high-volume trigger goes live.
  7. Add source code and IaC context where deployment correlation matters, then add knowledge lifecycle tasks only after first useful investigations.
  8. Test new response plans and HTTP triggers in Review long enough to capture Intent Met, false positives, approval outcomes, and AAU cost before autonomy promotion.
  9. Promote only narrow, repeatedly validated, low-blast-radius workflows to Autonomous.

Recipe Design Framework

Production SRE Agent capabilities are shaped by five intersecting dimensions:

Dimension Options / Examples Decision driver
Platform AzMonitor, PagerDuty, Dynatrace, ServiceNow, GitHub Actions, Azure DevOps, Grafana Where incidents originate; native webhook support
Connectors AppInsights, LAW (Log Analytics), Kusto, MCP integrations, GitHub, Azure DevOps APIs Which observability/CI/CD systems feed investigation
Skills Investigation playbooks by workload: VM, CosmosDB, AKS, Container Apps, HTTP errors The target system or error class you need to automate
Knowledge Runbooks, incident templates, RCA patterns, troubleshooting guides, approval checklists Context and procedures the agent needs to reach conclusion
Response Severity-based routing, agentMode (Autonomous/Review), customInstructions, max attempts How the agent behaves and when it escalates or acts

Design in this order: platform → connectors → skills → knowledge → response mode. Test each layer independently before combining. Name recipes <platform>-<key-connectors>[-workload-variant] (good: pagerduty-law-vmcosmos; avoid: agent-automation).

See Production Patterns Guide for recipe patterns, multi-backend deployment, connector auth orchestration, export→clone→verify workflows, and post-deployment verification checklists.

Resource Boundary Defaults

Last reviewed: 2026-08-25. The Microsoft.App/agents resource has no diagnostic-settings or private-endpoint support; audit via the agent's own Application Insights customEvents, and reach private workloads via VNet integration (preview, vnetConfiguration.subnetResourceId); see references/security-identity-rbac.md and references/run-posture-and-operational-levers.md.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent