All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencesoutput-examples.md

≈969 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Output Examples

Use these examples to keep Azure SRE Agent answers consistent. Keep outputs concise, evidence-led, and explicit about what is confirmed versus assumed.

Example 1: Production capability design review

## Azure SRE Agent Design Review

### Scope
- Workload: customer-facing API on App Service + Azure SQL
- Environment: Production
- Incident platform: Azure Monitor
- Mode: Review for all initial plans/tasks/triggers

### Recommended Bundles
- `base-core`
- `observability-amba`
- `azure-workload-production`
- `security-identity`
- `operational-metrics`

### Findings
1. AMBA baseline exists for App Service and SQL, but alert-to-response-plan ownership is not documented.
2. The proposed workload triage agent has both read and write tools; split read-only diagnostics from write-capable remediation review.
3. No AAU guardrail is defined for high-volume 5xx alerts.

### Recommended Changes
- Create one Review-mode response plan for high-severity 5xx incidents with a 15-minute cooldown.
- Use `workload-triage-agent` for investigation and route write recommendations to `remediation-review`.
- Add `agent-consumption-review` scheduled task weekly.
- Add approval hook before enabling `RunAzCliWriteCommands`.

### Residual Risks
- SQL failover actions remain manual until repeated safe evidence exists.
- Alert quality must be measured for two to four weeks before autonomy promotion.

Example 2: Incident triage response

## Incident Triage

### Evidence
- Alert: App Service HTTP 5xx rate above baseline
- Window: 2026-05-11 08:10-08:25 UTC
- Impact: 14% of requests failed for `/checkout`
- Recent change: deployment `2026.05.11.3` at 08:02 UTC

### Hypothesis
The deployment introduced a checkout dependency configuration issue.
Confidence: medium-high. The timing and route-specific error pattern match, but dependency telemetry still needs confirmation.

### Recommended Action
Do not restart the app yet. Validate dependency configuration and compare current revision/settings with the previous successful deployment.

### Risk and Rollback
- Restart risk: may hide evidence and create brief downtime.
- Rollback option: return to deployment `2026.05.11.2` if configuration regression is confirmed.

### Validation
- Confirm 5xx rate returns to baseline for 15 minutes.
- Confirm `/checkout` dependency calls succeed.
- Capture known-error record only if the root cause is confirmed.

Example 3: Autonomy promotion decision

## Autonomy Promotion Scorecard

### Candidate
- Plan: staging App Service restart after sustained memory saturation
- Current mode: Review
- Proposed mode: Autonomous for staging only

### Evidence
- 12 Review-mode executions over 30 days
- 11 approved without modification
- 1 denied due to unrelated deployment window
- No customer-facing impact
- Rollback path is clear: restart is reversible and validation is automatic

### Decision
Promote to Autonomous for staging only. Do not promote production equivalent yet.

### Guardrails
- Scope to staging resource group only
- Maximum one execution per 60 minutes
- Notify owner after every action
- Revert to Review if two failures or one unexpected side effect occurs

Example 4: HTTP trigger readiness report

## HTTP Trigger Readiness

### Trigger
- Name: post-deploy-validation
- Source: Azure DevOps pipeline
- Initial state: disabled
- Initial mode: Review

### Required Before Enablement
- Store bearer token in pipeline secret store.
- Validate payload fields: deploymentId, serviceName, resourceScope, commitSha, changeSummary.
- Add replay protection using deployment ID and timestamp.
- Limit `maxTurns` to the smallest useful bound.
- Test with a non-production resource scope.

### Enablement Decision
Not ready. Payload validation and replay protection are not yet documented.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent