All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencessource-map.md

≈3.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Source Map (Authoritative Basis)

This map ties key guidance in this skill to authoritative sources.

GitHub Copilot Agent Skills

  1. Agent skill concept, supported locations, and skill package shape:
  2. Agent Skills specification, frontmatter constraints, progressive disclosure, and validation:
  3. GitHub CLI skill validation/publishing preview:

Platform and Routing

  1. Incident platform behavior, single active platform, quickstart behavior:
  2. Response plan lifecycle, turn off/on, testing mode, quickstart overlap warning:
  3. Run mode semantics and permission interaction:

Scheduled Automation

  1. Scheduled task controls (Draft the cron for me, Polish instructions, Max executions precedence):
  2. Workflow design practices, Run task now, custom agent trigger model:
  3. HTTP triggers, event-driven execution, auth model, and limits:
  4. HTTP trigger setup and pipeline integration tutorial:
  5. HTTP trigger behavior details used in this package (404 on disabled triggers, 250-turn cap, and auth troubleshooting):

Custom Agents and Skills

  1. Custom agent behavior, testing, tool assignment, knowledge base limits:
  2. Skill activation model and constraints:
  3. Default-agent versus custom-agent tool/skill inheritance and override behavior:
  4. Memory, session insights, synthesized knowledge, uploaded document formats, and knowledge base limits:
  5. Code Interpreter sandbox, tool limits, timeout limits, and file-session behavior:

Connectors and MCP

  1. Connector categories, health states, heartbeat behavior, wildcard syntax and version notes:
  2. MCP transport, partner connector behavior, authentication, preconfigured partners, and tool assignment:
  3. Managed connector operation governance, parameter locking, credential isolation, and Ask approval caveats:
  4. SRE Agent MCP server for external tools calling into SRE Agent:
  5. investigate_yolo auto-approval behavior and MCP safety guardrails:
  6. Official plugin catalog and connector-specific setup guidance:

Governance with Hooks

  1. Hook events (Stop, PostToolUse), formats, limits, best practices:
  2. Hook API workflow and v2 extended-agent examples:
  3. Tool access policies (global/custom-agent/thread scopes; allow/ask/deny; hook-allow-overrides-deny interaction; toolGlob(argGlob) patterns):

Real Deployment Patterns

  1. Official end-to-end samples and labs:
  2. Hands-on lab post-provision and verification patterns (the repo moved samples/ into labs/; the hands-on lab is the labs/ root, azd up):
  3. Deployment-compliance sample (hooks + idempotent setup + policy workflows):
  4. Terraform drift detection sample (event-driven webhook flow + remediation patterns):
  5. Terraform drift operations reference blog:
  6. Recipe templates used for incident-platform and PR deployment-guard parity:
  7. Recipe template used for VM + Cosmos workload parity:
  8. This skill's full-capability blueprint patterns:
    • references/production-blueprints.md
    • bundles/governance-kt/hooks/kt-completeness-gate.yaml
  9. Bundle-first extension artifacts:
    • bundles/catalog.yaml
    • bundles/README.md
    • references/bundles-operations.md
    • references/capability-matrix.md

Observability and AMBA

  1. Azure Monitor Baseline Alerts home and service-level alert recommendations:
  2. AMBA GitHub repository and releases:
  3. AMBA ALZ deployment and policy guidance:
  4. Example AMBA service pages used when seeding workload bundles:
  5. Azure SRE Agent action auditing in Application Insights customEvents:

Project Boundary and AI Resource Isolation

  1. Document Intelligence managed identities and secure storage/private endpoint patterns:
  2. Azure AI Search RBAC and managed identity guidance:

Source RCA, Security, and Value Tracking

  1. Source-code connection and RCA behavior:
  2. GitHub connector capabilities:
  3. Security overview and defense-in-depth model:
  4. Incident value tracking, success rate, average TTM, and Intent Met quality:

Pricing, Regions, and Model Provider

  1. AAU billing model, always-on cost, token-based active flow, monthly allocation limit:
  2. Supported regions - 18 as of 2026-08-25 (Australia East, Canada Central, Central US, East Asia, East US 2, France Central, Italy North, Japan East, Korea Central, North Central US, South Africa North, Southeast Asia, Spain Central, Sweden Central, UK South, West Central US, West US 2, West US 3):
  3. Model provider selection, EUDB guidance, and provider switching behavior:
  4. Billing model update announcement (token-based active flow, effective April 15, 2026):
  5. API reference, data-plane auth audience, RBAC roles, and agent endpoint shape:
    • https://learn.microsoft.com/azure/sre-agent/api-reference
    • Verified 2026-08-25: Microsoft.App/agents REST versions are 2026-01-01 (GA) and 2025-05-01-preview (provider default / page-pinned). Data-plane token audience is https://azuresre.dev; agent endpoint pattern https://{name}--{id}.{hash}.{region}.azuresre.ai. The HTTP-trigger Troubleshooting section says the audience must be the SRE Agent app ID 59f0a04a-b322-4310-adc9-39ac41e9631e (an ARM https://management.azure.com audience returns 401), while the API-reference section shows --resource https://management.azure.com; three conflicting audiences, keep marked [VERIFY]. actionConfiguration.mode is Review, Autonomous, or ReadOnly (Automatic is rejected; the API-reference page's Automatic enum is stale; live-verified 2026-08-10); actionConfiguration.accessLevel is Low or High; defaultModel.provider is Anthropic or MicrosoftFoundry; upgradeChannel is Stable or Preview. ARM also documents sub-resources for skills/subagents/tools/scheduledTasks/incidentFilters/hooks/commonPrompts (PUT/GET/DELETE; base64 envelope for non-connectors) plus data-plane threads/approvals/repos paths; see references/live-verified-operations.md path table. RBAC roles: SRE Agent Administrator/User/Reader. See references/run-posture-and-operational-levers.md.
  6. Network requirements, required domains, resources created during provisioning, and regional availability pointer:
  7. Infrastructure-as-code deployment phases and data-plane configuration gaps:

GA Announcement and New Capabilities

  1. GA announcement (March 10, 2026):
  2. What's new in GA release:
  3. Plugin Marketplace and Skills:
  4. MCP connectors (including pre-configured partners):
  5. Starter lab (azd-based deployment):
  6. Additional labs/recipes inventory (verified 2026-08-25; re-sweep on next run): labs/deployment-guard, labs/zava-learning, recipes/minimal, recipes/azuretoaws-sre-agent (AWS cross-cloud), top-level appendix/lab-instructions.md.

Build 2026 and Post-GA Capabilities

  1. Build 2026 announcement (five enterprise releases: VNet integration preview, Managed Connectors preview, granular permissions, Private Plugin Marketplace, GitHub Enterprise/BYO GitHub App):
  2. GitHub connector (OAuth, PAT, BYO GitHub App; GitHub Enterprise Cloud requires BYO App):
  3. Log Analytics and Application Insights connectors (native MCP-backed query tools):
  4. Execute mitigations (built-in safety: delete/remove refused, az keyvault blocked, management locks respected, subscription GUID validation; AgentAzCliExecution audit event):
  5. Network requirements (allow-list domains, provisioning resources, Zscaler warning):
  6. Agent Playground (YAML diff preview and Accept-selected-fixes workflow ; re-verify current UX before quoting):
  7. AWS DevOps Agent cross-cloud integration (out of scope for this Azure-greenfield skill; roadmap note):
  8. Terraform drift detection via HTTP triggers (reference workflow):

Official Resources

  1. Product documentation: https://aka.ms/sreagent/docs
  2. Self-paced labs: https://aka.ms/sreagent/lab
  3. Technical videos: https://aka.ms/sreagent/youtube
  4. Home page: https://www.azure.com/sreagent
  5. X (Twitter): https://x.com/azuresreagent

Notes

  1. Community references are optional context only.
  2. Community proactive-skill reference reviewed for cadence ideas (non-authoritative):
  3. If a claim cannot be tied to the sources above, treat it as a local assumption and label it.
  4. KT templates in this skill are derived from a user-provided workbook: KT_Templates.xlsx (SA QU/WORKSHEET, PA QU/WORKSHEET, DA QU/WORKSHEET, PPA QU/WORKSHEET).

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent