All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referenceshttp-triggers-production.md

≈840 tokens on demand. Your agent reads this file only when SKILL.md points to it.

HTTP Triggers Production Reference

Use HTTP triggers for event-driven SRE Agent execution from CI/CD pipelines, monitoring tools, or external automation.

Guardrails

  1. Disable triggers until bearer-token authentication, payload schema, replay controls, and secret handling are tested.
  2. Keep prompts bounded to the triggering service, resource scope, deployment, commit, or incident.
  3. Do not send secrets or raw customer data in trigger payloads.
  4. Set maxTurns well below service limits and tune after real executions.
  5. Default to Review mode and use approval hooks for production write actions.
  6. Record idempotency keys or calling-system correlation IDs.
  7. Review execution history for prompt injection, unexpected tool use, and AAU cost.

Good Uses

  • post-deployment health validation
  • pipeline failure triage
  • IaC drift investigation
  • release readiness checks
  • compliance evidence gathering

Poor Uses

  • broad open-ended investigation without scope
  • high-frequency alert fan-out without cooldown/noise control
  • direct production remediation from untrusted payloads

Authentication

The Token Audience Conflict

The official Microsoft Learn docs contain conflicting guidance about HTTP trigger token audiences:

  • The HTTP Trigger API reference section shows --resource https://management.azure.com
  • The Troubleshooting section says the audience must be the SRE Agent app ID 59f0a04a-b322-4310-adc9-39ac41e9631e
  • The data-plane API documentation references audience https://azuresre.dev

Working Workaround

Until this is clarified by Microsoft, use this two-step pattern for CI/CD integration:

If your trigger is from GitHub Actions / Azure DevOps:

  • Use OIDC Federated Identity with the SRE Agent app ID: 59f0a04a-b322-4310-adc9-39ac41e9631e
  • This is confirmed working as of 2026-06-25

If your token comes from Azure CLI or custom code:

  • Request a token with scope https://management.azure.com
  • If you get a 401, retry with scope https://azuresre.dev
  • Log which scope worked for your environment

The caller needs Microsoft.App/agents/threads/write permission on the agent resource (service principal, managed identity, or user). A successful invocation returns HTTP 202 (Accepted) immediately with {"message", "executionTime", "threadId", "success"} and the agent processes the request asynchronously.

Debug Flowchart

Is your HTTP trigger returning 401?
├─ Check token is not expired → If expired, request new token
├─ Check token audience is correct → See "Working Workaround" above
├─ Check SRE Agent resource exists and is in your subscription → If not, provision first
└─ If still 401, file issue at https://github.com/microsoft/sre-agent/issues with your token audience value

Success Signal

Your HTTP trigger returns 202 (Accepted) — the agent processes asynchronously — and the execution appears in trigger history / the response plan is invoked.

[VERIFY]
Claim = HTTP trigger token audience requirements
WhereToCheck = https://learn.microsoft.com/en-us/azure/sre-agent/http-triggers
Status = Conflict documented; working workaround provided; awaiting official docs clarification
Action = Test your specific CI/CD platform (GitHub Actions, Azure DevOps, etc.) to confirm token scope before production cutover

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent