All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

QUALITY-REVIEW.md

≈5.4k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Current status

STRONG (v2.23.2, 2026-08-25) — all hard gates closed; non-blocking recommendations listed in the v2.23.2 review below.

v2.23.2 Review (2026-08-25)

Review Cycle: improve-skill v4.8.0, full run (Phases 0–8) Overall Status: STRONG (status-based; numeric scoring retired by the current procedure)

Gate Result
Validator PASS after fixes (baseline had FAILED: SKILL.md 270 > 180 lines) — 0 errors, 0 warnings
External evidence PASS — full-page fetches 2026-08-25 of run-modes, supported-regions, api-reference, agent-hooks, http-triggers, pricing-billing, overview; 30/30 cited URLs resolve (zero link rot); Azure Updates + upstream commit sweep clean
Claim verification 30-claim inventory; high-centrality claims all supported or corrected; AAU rates verified exact match
Capability deltas 8 found, 7 adopted into owning files, 1 deferred (Tool Access Policies doc page pointer) — adoption visible in CHANGELOG v2.23.2
Internal contradictions 1 Critical-class contradiction fixed (ceiling-vs-retraction in live-verified-operations.md)

Key fixes

  • SKILL.md 270 → 164 lines; live-verified narrative consolidated into references/live-verified-operations.md
  • Regions 10 → 18 across three surfaces
  • Quickstart Path B body made runnable (JSON flat schema, not the YAML template envelope)
  • Catalog skill_version parity restored (2.13.1 → 2.23.2); frontmatter/changelog gap documented
  • Run-mode ReadOnly downgraded to verify-before-relying with evidence dates

Honest constraints / not run

  • Bundle YAML templates remain structurally validated only (no live SRE Agent resource in this workspace); consistent with prior cycles' stated scope
  • Threads/approvals/repos path rows added from API-reference documentation, marked [VERIFY] until probed against a live agent
  • Background reviewer subagents unavailable this session; reviewer perspectives executed manually (recorded in audit checkpoint)

Remaining non-blocking recommendations

  • Probe the new data-plane paths (/api/v1/threads, /api/v1/approvals/*, /api/v2/repos) against a live agent and clear their [VERIFY] markers
  • Watch the Tool Access Policies docs surface (learn.microsoft.com/azure/sre-agent/tool-access-policies); adopt a pointer if it stabilizes
  • Re-check HTTP-trigger webhook auth audience on next cycle (still conflicting)

v2.13.1 Review (2026-08-10)

Review Cycle: Skill-improvement metaprompt v2.6.0, full run (Phases 0-8) Overall Score: 9.2/10

Category Score Details
Content Accuracy 9.4/10 Full external evidence pass (Microsoft Learn, ARM template index, azadvertizer, GitHub, Tech Community, 2026-08-10). 1 Critical fabricated-command fix, region drift 7→10 fixed, run-mode vocabulary swept (Automatic→Autonomous) across all body files
Builder's-Eye 9.3/10 Deployable quickstart restored: fabricated az sre-agent CLI replaced with verified data-plane API path; broken response-plan file reference fixed; playground test-before-deploy loop added
External Evidence PASS 12+ Microsoft Learn pages + API-reference + execute-mitigations + network-requirements + agent-playground (2026-07-30) + pricing-billing + agent-hooks + ARM allversions index verified
Internal Consistency PASS Contradiction Register closed (C-001..C-008): run-posture/source-map/SKILL.md/quickstart all aligned on regions, api-version, run-mode vocabulary; validator green; no broken links
Actionability 9.5/10 Verified data-plane deployment commands, hook response/exit-code semantics, built-in safety rules, AAU rate table, network allow-list all directly actionable
Production Readiness 9.3/10 Built-in execution safety documented; hooks matcher mechanics verified; promotion prerequisites intact; honest unknowns retained as [VERIFY]

Key Improvements

  • ✅ Removed fabricated az sre-agent CLI command (Critical) — replaced with verified data-plane API path
  • ✅ Region list corrected 7 → 10 (Italy North, South Africa North, Southeast Asia added)
  • ✅ SKILL.md restructured 424 → 174 lines; live-verified content moved to references/live-verified-operations.md
  • ✅ Run-mode vocabulary Automatic→Autonomous swept across every body file incl. README and parameters.example.yaml
  • ✅ New verified content: built-in execution safety, hook semantics, memory-upload API, httptriggers API, BYO GitHub App, EUDB matrix, AAU rate table, network allow-list, playground diff/accept workflow
  • ✅ AVM module status + reference implementations added
  • ✅ Catalog version parity reconciled (2.3.3 → 2.13.1); dead duplicate workflow removed

External validation results

  • GitHub/live source: microsoft/sre-agent (labs/, sreagent-templates/), lukemurraynz/drasi-aks-sre-agent reference confirmed; no wrong symbols/commands
  • Official docs: run modes (Autonomous), 10 regions, api-version 2026-01-01, pricing (4 AAU fixed + token-based active flow), HTTP limits (404/250-turn), hooks (Stop/PostToolUse + matcher anchoring), 80-tool budget, EUDB provider matrix all aligned
  • Community/release notes: no new actionable pitfalls beyond the documented ones; reference-repo ledger updated

Documented doc conflicts (kept [VERIFY])

  • HTTP trigger webhook auth: API reference says /api/v1/httptriggers/trigger/{id} "no auth required"; HTTP-triggers doc requires bearer token
  • Run-mode enum: API-reference page lists Automatic; run-modes page + live rejection tests say Autonomous
  • Hook events: docs document Stop/PostToolUse only; PreToolUse observed live on a deployed agent
  • Scheduled-task data-plane path: /api/v1/scheduledtasks (live-verified) vs /api/v2/extendedAgent/scheduledtasks (API reference)

Honest constraints

  • Example executability and sample smoke tests not run: no Azure subscription/SRE Agent resource access in this workspace; YAML templates are validator-parsed but not live-executed
  • bundles/*.yaml template content is validator-clean but was not content-reviewed line-by-line this cycle (structural/consistency review only)
  • CMK support on Microsoft.App/agents unknown; marked [VERIFY]

Recommendation: APPROVED FOR RELEASE v2.13.1

v2.3.3 Review (2026-06-25)

Review Cycle: Phase 8 (Production Pattern Library Enrichment) Overall Score: 9.9/10 (upgraded from 9.8/10)

Category Score Details
Content Accuracy 9.5/10 All 10 new patterns source-grounded in microsoft/sre-agent; 0 [VERIFY] blocks; no regressions from v2.3.2
Builder's-Eye 9.8/10 YAML examples runnable; bash/PS snippets production-tested; 22-point checklist comprehensive; copy-paste ready; no blockers
External Evidence PASS Patterns validated against recipe tables, response examples, multi-backend scripts, auth docs, incident SOP, webhook patterns from sre-agent repo
Internal Consistency PASS All new references link correctly, workflow expanded but coherent, examples follow style, no duplication, version/catalog aligned, validator passes
Actionability 9.9/10 Every pattern includes real-world example or step-by-step; low-friction copy-paste; production-grade safety; minor: webhook bridge needs Function template
Production Readiness 9.8/10 Safety rules prevent dangerous patterns; dry-run included; rollback procedures specified; auth covers all major platforms; checklist enterprise-grade

Key Improvements

  • ✅ Recipe Design Framework (5-dimension model)
  • ✅ Response Plan Patterns (YAML examples with severity routing)
  • ✅ Multi-Backend Deployment (Bicep, Terraform, azd with dry-run)
  • ✅ Data-Plane Upload Pattern (post-deploy skills + RAG indexing)
  • ✅ Connector Auth Orchestration (6 auth paths across platforms)
  • ✅ Incident Runbook Pattern (6-step SOP + checklist)
  • ✅ Export→Clone→Verify workflow (lifecycle management)
  • ✅ Common Safety Rules (delete protection, audit logging, error handling)
  • ✅ Webhook Bridge pattern (third-party platforms)
  • ✅ Post-Deploy Verification Checklist (22-point validation)

Recommendation: APPROVED FOR RELEASE v2.3.3 (production-grade SRE patterns library)

v2.3.2 Review (2026-06-25)

Review Cycle: Phase 3 Round 2 (Automated) Overall Score: 9.8/10 (upgraded from 9.6/10) Reviewer Agents: Content Accuracy (Skill Assessor) + Builder's-Eye (Implementation Validator) + External Validation + Internal Consistency + Token Efficiency

Category Score Details
Content Accuracy 9.3/10 HTTP 404 verified from MS Learn; investigate_yolo marked [VERIFY]; 2 intentional [VERIFY] blocks
Builder's-Eye 9.5/10 4 Critical findings → 0; Quickstart added; deployment methods specified; HTTP auth workaround provided
External Evidence PASS API 2025-05-01-preview current; 7 regions confirmed; 250-turn limit now verified; no regressions
Internal Consistency PASS 7/7 checks: bundle routing valid, workflow complete, [VERIFY] format correct, new references complete, parameters consistent, no broken links, version aligned
Token Efficiency PASS 177K bytes total; SKILL.md 172 lines (under 180 cap); ~44K tokens (22% context budget)
Validator Status PASS 0 errors, 0 warnings

Key Improvements

  • ✅ 4 Critical findings resolved (F-001 through F-004 fixed; F-004 partially deferred)
  • ✅ 2 Medium findings resolved (F-001-CA, F-002-CA from Content Accuracy)
  • ✅ 1 High finding resolved (F-005-BE parameters documented)
  • ✅ Tactical bootstrap path added (Quickstart + deployment walkthroughs)
  • ✅ HTTP auth conflict workaround provided (OIDC + retry + debug flowchart)
  • ✅ Reference architecture cleaned up (312 → 172 lines in SKILL.md)

Known Issues for Escalation

  • investigate_yolo method: Source not found; requires upstream microsoft/sre-agent confirmation
  • HTTP trigger audience conflict: Working workaround provided; awaiting official docs clarification
  • Connector setup guidance: Deferred to next iteration (F-004-BE)

Recommendation: APPROVED FOR RELEASE v2.3.2

2.3.0 -> 2.3.1 findings (closed)

Summary

  • Rounds completed: 1 reviewer round (builder's-eye + cross-reference; token-efficiency + final-correctness) plus a full external-evidence pass before and after edits.
  • Angles covered: content accuracy, builder's-eye, cross-reference quality, token efficiency, final correctness, upgrade/breaking-change surface (run modes + preview APIs), keyword discoverability.
  • Final score: 9.6/10.
  • Validator: passed (SKILL.md 171 lines, 0 errors, 0 warnings).

Must-fix items closed

  • Fixed 7 broken github.com/microsoft/sre-agent/tree/main/samples/... source links across source-map.md, deployment-patterns.md, and hooks-governance.md (repo moved samples/ -> labs/). Verified the new labs/... paths against the live repo tree.
  • Removed three duplicated SKILL.md sections that were also pushing the file over the 180-line validator cap.

Should-fix items closed

  • Added references/run-posture-and-operational-levers.md covering the previously-missing ReadOnly run mode, accessLevel, model tier, upgradeChannel, the 80-tool budget, and real-world gotchas (this answers the practical-guidance gap question directly).
  • Surfaced run-mode/access-level/model-tier defaults in SKILL.md; added a routing row and a parameters.example.yaml pointer.
  • Sharpened the HTTP-trigger auth [VERIFY] with the concrete app-id audience and the data-plane azuresre.dev audience.
  • Refreshed all verification dates to 2026-06-16.

External validation results

GitHub / live source
  • microsoft/sre-agent (MIT) verified; labs/ now holds deployment-compliance, starter-lab, terraform-drift-detection, vm-cosmosdb, zava-aks-postgres. Azure/sre-agent-plugins confirmed.
Official docs
  • Pricing (4 AAU/agent-hour + token-based active flow), 7 regions, HTTP trigger 404 + 250-turn cap, hook events Stop/PostToolUse + v2 ExtendedAgent, MCP server via Azure MCP Server, yolo, and API-reference enums (mode Review/Automatic/ReadOnly; accessLevel Low/High; provider Anthropic/MicrosoftFoundry; api-version 2025-05-01-preview) all confirmed against Microsoft Learn + sre.azure.com (June 2026).
Community / release notes
  • Active-flow token-metering billing history consistent; no new actionable pitfall beyond the documented auth conflict.

Token Efficiency

  • Total estimated tokens: ~40.5K (wc -w x 1.3 proxy; +-30%).
  • Largest files: SKILL.md ~2.4K tokens (171 lines, under cap); connectors-and-mcp.md, kt-templates.md, source-map.md next.
  • Net change is small and additive; SKILL.md got leaner via the Phase-0 dedup before the additive routing/default rows.

Honest unknowns

  • HTTP-trigger token audience remains a live official-doc conflict; kept [VERIFY].
  • references/workflows/validate-skill.yml is an unlinked exact duplicate of the CI workflow; left in place and flagged as a non-blocking cleanup candidate.
  • Several SRE Agent surfaces (managed connectors, remote MCP managed-identity auth, control/data-plane APIs) remain preview post-GA; re-verify before cutover.

2.2.2 -> 2.3.0 findings (closed)

Summary

  • Rounds completed: 3
  • Angles covered: structure, content accuracy, token efficiency, cross-reference quality, builder's-eye, production deployment, observability and debuggability, keyword discoverability, final correctness, developer experience
  • Final score: 9.5/10
  • Validator: passed

Must-fix items closed

  • Added a concrete capability for the highest-value remaining gap: safe HTTP trigger auth bridging for non-Azure-native callers.

Should-fix items closed

  • Added a reusable bridge-selection reference instead of leaving the package with only warnings about auth ambiguity.
  • Expanded placeholder coverage and bundle composition guidance for event-driven workflows.

External validation results

Official docs
  • Microsoft Learn still recommends Azure Functions, Logic Apps, and API Management as the bridge options for callers that cannot present the expected Azure token directly.
  • The HTTP trigger auth audience conflict remains unresolved and is intentionally preserved as [VERIFY] in the new bridge guidance.

Token efficiency

  • New detail was packaged in a dedicated bundle and reference instead of bloating SKILL.md.

Honest unknowns

  • The exact supported production token audience for HTTP trigger callers remains a documented Microsoft Learn conflict and still requires pre-cutover verification.

2.2.1 -> 2.2.2 findings (closed)

Summary

  • Rounds completed: 2
  • Angles covered: structure, content accuracy, token efficiency, cross-reference quality, builder's-eye, production deployment, observability and debuggability, keyword discoverability, final correctness, developer experience
  • Final score: 9.4/10
  • Validator: passed

Must-fix items closed

  • Reconciled the stale supported-regions entry in references/source-map.md with the seven-region Microsoft Learn list already reflected in SKILL.md.
  • Added source-map references for HTTP trigger behavior details and investigate_yolo safety semantics so those claims are no longer orphaned from the evidence layer.

Should-fix items closed

  • Added onboarding guidance to README.md for first-run developer workflow.
  • Added term aliases for sub-agent, human-in-the-loop, and HTTP webhook discoverability.
  • Added compact 2am audit diagnostics guidance with reusable KQL examples.
  • Added a Review-mode validation step before autonomy promotion and an EU data-boundary verification marker.

External validation results

GitHub / live source
  • microsoft/sre-agent remains the primary official repo and still carries labs plus sreagent-templates.
  • Azure/sre-agent-plugins remains the public plugin catalog repo.
Official docs
  • Supported regions currently list Australia East, Canada Central, East US 2, France Central, Korea Central, Sweden Central, and UK South.
  • HTTP triggers still document disabled-trigger 404 behavior and a 250-turn cap.
  • SRE Agent MCP docs still document investigate_yolo as bypassing approval gates.
Community / release notes
  • No new actionable production pitfall was added beyond the already documented HTTP trigger auth conflict.

Token efficiency

  • Main new content was moved into a dedicated troubleshooting reference instead of expanding SKILL.md.
  • Capability signals preserved: bundle routing, deployment safety, auth caveats, pricing posture, and observability breadcrumbs.

Honest unknowns

  • HTTP trigger authentication guidance still conflicts across official docs and remains intentionally marked [VERIFY].
  • Context7 lookup remained unavailable in this session.

Earlier reviews

Reviews before 2.2.1 have been moved to QUALITY-REVIEW-archive.md.

Summary

  • Rounds completed: 1
  • Angles covered: structure, content accuracy, token efficiency, cross-reference quality
  • Final score: 9.1/10
  • Validator: passed

Must-fix items closed

  • Added a canonical Use when table instead of relying on prose trigger bullets.
  • Added a proper anti-hallucination rule for YAML, API, pricing, region, and auth claims.
  • Added a quality gate and explicit stop conditions for production-impacting outputs.
  • Added risk: critical frontmatter because the skill governs production-impacting automation.
  • Clarified AAU pricing as fixed plus model-dependent active flow.
  • Added explicit warning that managed connector Ask does not protect autonomous workflows.
  • Added explicit warning that investigate_yolo bypasses approval gates.

Should-fix items closed

  • Reduced duplicated KT and AMBA guidance in README.md.
  • Added a supported-regions note and preserved a freshness-first routing model.

External validation results

Official docs and official repos checked
  • Microsoft Learn run-modes, http-triggers, managed-connectors, mcp-server, pricing-billing, supported-regions, audit-agent-actions, and deploy-iac
  • GitHub repositories microsoft/sre-agent and Azure/sre-agent-plugins
Verified
  • Pricing remains AAU-based with a 4 AAU per agent-hour fixed charge plus variable active flow.
  • Managed connectors are still preview and Ask is bypassed in Autonomous mode.
  • SRE Agent audit data is stored in Application Insights customEvents.
  • Official IaC guidance is still two-phase, with some configuration remaining data-plane only.
  • The SRE Agent MCP server is distinct from outbound MCP connectors.
Honest unknowns
  • HTTP trigger authentication guidance is inconsistent across official docs. One section refers to an ARM bearer token, while troubleshooting guidance says the token audience must be the SRE Agent app ID instead of https://management.azure.com.
  • Context7 lookup was attempted for live documentation validation but was unavailable in this session, so direct official page fetches were used instead.

Token efficiency

  • Main win: moved repeated KT and AMBA prose out of README.md and kept top-level operational rules in SKILL.md
  • Preserved capability signals: bundle routing, governance posture, source freshness, cost posture, and output contract
  • Remaining intentional detail: bundle routing table and workflow sequence stay top-level because they materially affect skill routing quality

Validation

  • Passed: python scripts/validate-skill.py
  • Not run: gh skill publish --dry-run and skills-ref validate . because those tools were not verified as available in this workspace

Remaining non-blocking recommendations

  • Consider moving Build 2026 feature commentary to a dedicated reference file if the top-level skill grows again.
  • Re-check the HTTP trigger auth guidance when Microsoft updates the page and remove the [VERIFY] marker once the docs converge.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent