All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencesobservability-amba.md

≈692 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Observability and AMBA Reference

Use Azure Monitor Baseline Alerts (AMBA) as the default monitoring baseline for new Azure SRE Agent implementations unless the customer has a stronger existing standard.

Core Pattern

  1. Use AMBA to identify recommended Service Health, Resource Health, Activity Log, Metric, and Log Search alert categories.
  2. Classify each alert before automation: investigate, notify, digest, tune, suppress, or candidate for later autonomy.
  3. Create SRE Agent response plans only when the alert has a meaningful investigation path or safe remediation proposal.
  4. Tune thresholds to workload SLOs and business criticality.
  5. Use Review mode by default and enable reinvestigation cooldown for recurring Azure Monitor alerts.
  6. Review noisy alerts and AAU cost before scaling response-plan coverage.
  7. Query Application Insights customEvents for agent audit events before promoting AMBA-triggered workflows to Autonomous.

Policy-Deployed AMBA Checks

When AMBA is deployed by Azure Policy/ALZ patterns, verify:

  • initiative and assignment scope
  • assignment effect and disabled policies
  • managed identity and remediation permissions
  • action groups and alert processing rules
  • exclusions, overrides, and resource selectors
  • compliance/remediation state
  • whether bulk remediation is being proposed as a production write action

Terraform/AVM Guidance

When using AMBA Terraform/AVM patterns:

  • Keep customer-specific overrides outside upstream module source.
  • Review plan output for policy assignments, action groups, alert rules, managed identities, and notification routing.
  • Treat BYO managed identity, BYO notifications, telemetry settings, and exclusions as explicit design decisions.
  • Validate management-group versus subscription scope before deployment.

SRE Agent Mapping

  • Service Health/Resource Health: impact assessment and stakeholder digest.
  • Activity Log delete/update: change validation, blast radius, rollback recommendation, KT DA/PPA if production-impacting.
  • Metric alerts: resource triage, trend analysis, scaling/remediation proposal, and threshold tuning.
  • Log Search alerts: fleet investigation, KQL summary, pattern detection, and knowledge capture.

Agent Audit Checks

Before expanding AMBA-to-response-plan coverage, review these audit signals:

  1. AgentToolExecution volume by tool and custom agent.
  2. ModelGeneration token trend for AAU and context growth.
  3. IncidentActivitySnapshot counts for assisted versus mitigated incidents.
  4. ApprovalDecision outcomes for unsafe or low-confidence remediation proposals.

Use the audit results to identify noisy alerts, expensive investigations, weak response-plan prompts, and workflows that are not safe for Autonomous mode.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent