All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencesquickstart-first-deployment.md

≈2.1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Quickstart: Deploy Your First Response Plan

Want to go from zero to a working agent in <1 hour? Start here.

Prerequisites

  • Azure subscription with Owner role
  • Azure SRE Agent already deployed in your target region (18 supported regions last verified 2026-08-25: Australia East, Canada Central, Central US, East Asia, East US 2, France Central, Italy North, Japan East, Korea Central, North Central US, South Africa North, Southeast Asia, Spain Central, Sweden Central, UK South, West Central US, West US 2, West US 3 ; see ../SKILL.md Workflow and Microsoft Learn supported-regions for the current list)
  • Azure portal or Azure CLI access
  • One alert or incident you want to automate

Step 1: Choose Your Starting Bundle

Pick the one closest to your need:

  • Just starting? → Use bundles/base-core/ (minimal production baseline)
  • Using Azure Monitor / AMBA? → Use bundles/observability-amba/
  • Triggered from CI/CD, GitHub, or Azure DevOps? → Use bundles/http-triggers-production/
  • Need to investigate code or deployments? → Use bundles/source-rca-remediation/

Step 2: Populate Your Configuration

Copy parameters.example.yaml to my-deployment-params.yaml:

cp parameters.example.yaml my-deployment-params.yaml

Fill in each @@PLACEHOLDER@@ with your environment values. See the comments in the YAML file for guidance.

Required fields:

  • @@MODEL_PROVIDER@@, your preferred model provider (default: Azure OpenAI)
  • @@SRE_AGENT_RESOURCE_GROUP@@, the resource group containing your deployed SRE Agent
  • @@SRE_AGENT_HOST@@, the agent's sre.azure.com host (used by HTTP-trigger workflows)

Step 3: Apply the Bundle (Choose One Path)

Path A: Portal (No CLI Required)

  1. Go to Azure Portal → search "SRE Agent"
  2. Open your SRE Agent resource
  3. Navigate to "Response Plans" → "Add"
  4. Open bundles/base-core/response-plans/core-high-severity.yaml in a text editor
  5. Copy the YAML and paste into the Portal
  6. Replace @@PLACEHOLDER@@ values with your my-deployment-params.yaml values
  7. Click "Save"

Success signal: Response plan appears in your list without errors.

Path B: Data-Plane API (Azure CLI + REST)

There is no az sre-agent CLI command group (verified 2026-08-10). Apply the response plan through the agent's data-plane API using a token for the https://azuresre.dev audience:

# Authenticate to your subscription
az account set --subscription <your-subscription-id>

# Resolve the data-plane endpoint and token (api-version 2026-01-01)
AGENT=$(az resource list --resource-type Microsoft.App/agents --query '[0].id' -o tsv)
ENDPOINT=$(az rest --method GET --url "https://management.azure.com${AGENT}?api-version=2026-01-01" --query properties.agentEndpoint -o tsv)
TOKEN=$(az account get-access-token --resource "https://azuresre.dev" --query accessToken -o tsv)

# PUT creates, POST updates an existing plan. The id MUST be in the PATH --
# writing to the collection path returns 405 with an empty body.
# The body is JSON in the flat plan schema -- do NOT pipe the YAML template file
# from bundles/ as-is; it uses a template envelope, not this API's body shape.
cat > /tmp/core-high-severity.json <<'JSON'
{
  "id": "core-high-severity",
  "incidentPlatform": "AzMonitor",
  "titleContains": "",
  "agentMode": "review",
  "maxAutomatedInvestigationAttempts": 3,
  "isEnabled": true
}
JSON

curl -s -X PUT "$ENDPOINT/api/v1/incidentPlayground/filters/core-high-severity" \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  --data @/tmp/core-high-severity.json

Field names follow the flat plan schema (id, incidentPlatform, agentMode, ...) ; verify against references/live-verified-operations.md and your agent before scripting; if a field is rejected, diff against an existing plan via GET /api/v1/incidentPlayground/filters.

Success signal: API returns the created plan without errors. Never DELETE a live plan as a step toward changing it ; use POST to update, because a rejected PUT after DELETE leaves the alerts the plan handled routing nowhere.

Path C: Bicep / Terraform / PowerShell / azd (IaC Recommended for Production)

Use the upstream microsoft/sre-agent templates under sreagent-templates/, generate a config directory from a recipe with ./bin/new-agent.sh, drop this skill's bundle content into it, then deploy with ./bin/deploy.sh (Bicep), ./bin/deploy-tf.sh (Terraform), .\bin\ps\Deploy-Agent.ps1 (PowerShell), or azd up.

Remember that ARM covers Phase 1. Per official deploy-iac docs (2026-08-25): Phase 1 ARM includes connectors, skills, subagents, and tools as ARM sub-resources; hooks ("not yet exposed as ARM sub-resources at deploy time"), code repositories, HTTP triggers, knowledge files, and plugin configuration are applied in the data-plane Phase 2 by apply-extras.sh, which reports what it skipped when a data-plane token is unavailable.

See production-patterns-guide.md#multi-backend-deployment-patterns for the full backend table, the two-phase model, and day-2 export/clone/diff/verify operations.

Step 4: Test in the Agent Playground Before Deploying

Before attaching any custom agent or response plan to production, test it in the Agent Playground (Agent Canvas → Test playground; verified 2026-07-30):

  1. Edit and test side by side — split-screen: form/YAML editor on the left, live chat on the right. Chat input is disabled until you Apply or Discard, so you never test a stale configuration.
  2. Evaluate agent quality — the Evaluation tab runs AI scoring (0–100) covering intent match, completeness, tool fit, prompt clarity, actionability, and safety.
  3. Apply quick fixes — "Review and apply" opens the quick-fixes dialog: select fixes, preview the YAML diff, then "Accept selected fixes" (continue editing or save immediately).
  4. Test tools in isolation — system tools (fill parameters, execute, inspect raw JSON) and Kusto tools (run queries against connected clusters) can be validated without a live conversation.

Use the playground for the acceptance tests in bundles/base-core/checklists/acceptance-tests.md; it is the safe pre-deployment loop that keeps edits out of production threads.

Step 5: Test with Your First Alert

  1. Trigger an alert or manual incident in your environment
  2. Verify the response plan is invoked (check SRE Agent activity log)
  3. Confirm the result appears in your incident tracking system

Success signals:

  • Activity log shows "Response plan invoked" status
  • No 401/403 errors in logs
  • Investigation result appears in your incident platform

Step 6: Review Before Autonomy

Once the first run succeeds:

  1. Check Approval Outcomes, did the agent make the right decision?
  2. Check Knowledge Capture, is the output useful?
  3. Check AAU Cost, does the cost match your expectations?

If all three are "yes", you're ready to promote from Review to Autonomous (data-plane agentMode: autonomous; automatic is not a valid value on either surface).

Troubleshooting First Deployment

Problem Cause Fix
Response plan fails to save Syntax error in YAML or invalid placeholder Validate YAML syntax; check parameters.example.yaml comments for required fields
401 Unauthorized error HTTP trigger token audience mismatch See http-triggers-production.md#authentication
Response plan runs but produces no output Model response is empty or connector disabled Check connector permissions; verify model provider is available in your region
AAU cost is higher than expected Autonomous mode with verbose KT or high-frequency triggers Switch to Review mode; reduce trigger frequency or KT depth; see value-promotion-scorecard.md

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent