All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referenceshooks-governance.md

≈2.8k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Hooks and Governance Controls

Use this guide to enforce quality and safety gates in production.

Hook Types

Two hook events are supported:

  1. Stop: intercept final response before completion.
  2. PostToolUse: inspect tool outcomes after execution. Note: PostToolUse fires only after a tool succeeds; failed tool calls do not trigger it, so do not rely on PostToolUse to catch tool errors.

Two configuration levels:

  1. Agent-level hooks (apply broadly). Configure in Builder > Hooks.
  2. Custom-agent-level hooks (apply only to that custom agent). Configure via Agent Canvas > Manage Hooks or the v2 API.

When an agent-level and a custom-agent-level hook match the same event, both run and the agent-level hook fires first. Hooks complement run modes, they do not replace them: run modes control what the agent can do; hooks control how well it does it and what happens with results. Keep approval and run-mode controls in place even when hooks are present.

Prompt vs Command Hooks

  1. Prompt hooks: nuanced LLM evaluation and policy checks.
  2. Command hooks: deterministic checks (regex/policy/audit logic).

Approval-Gate Pattern for Sensitive Actions

Use a Stop prompt hook for write-action gates.

Template:

api_version: azuresre.ai/v2
kind: ExtendedAgent
metadata:
  name: @@CUSTOM_AGENT_NAME@@
spec:
  hooks:
    Stop:
      - type: prompt
        prompt: |
          Review the response for infrastructure-modifying actions.
          If any write/change action is present, reject and request explicit user approval.
          If read-only/reporting only, approve.
          $ARGUMENTS
          Return JSON:
          {"ok": true, "reason": "Read-only response"}
          or
          {"ok": false, "reason": "Approval required before remediation."}
        timeout: 30
        failMode: block
        maxRejections: 3

Audit Pattern for Tool Use

Use a PostToolUse command hook for audit trails.

Template:

hooks:
  PostToolUse:
    - type: command
      matcher: "*"
      timeout: 30
      failMode: allow
      script: |
        #!/usr/bin/env python3
        import json, sys
        context = json.load(sys.stdin)
        tool = context.get("tool_name", "unknown")
        print(json.dumps({
          "decision": "allow",
          "hookSpecificOutput": {
            "additionalContext": f"[AUDIT] Tool executed: {tool}"
          }
        }))

Built-in execution safety (verified 2026-08-10)

Independent of hooks, the agent enforces platform-level safety on Azure CLI actions (see https://learn.microsoft.com/azure/sre-agent/execute-mitigations):

  1. Delete/remove operations are refused. The agent never runs delete or remove commands and returns an error directing the user to the Azure portal.
  2. All az keyvault commands are blocked to prevent credential exposure.
  3. Management locks are respected. Resources with ReadOnly locks cannot be modified.
  4. Subscription IDs are validated (GUID format) before execution.
  5. Every az command is logged as an AgentAzCliExecution custom event in the agent's Application Insights.

Hooks add guardrails on top of these; do not assume a hook is the only control on a write path.

Tool Access Policies (documented 2026-08-25)

Policies are a first-class control alongside hooks and run modes (https://learn.microsoft.com/azure/sre-agent/tool-access-policies). Three scopes:

Scope Who sets it Rules allowed
Global Admin (Settings > Permissions) Allow, Ask, Deny
Custom agent Admin or author Allow only (can widen inside global bounds, never weaken a global deny)
Thread Any user Allow only

Evaluation order, highest priority first: Hooks → Tool access policies → Connector Ask / RequiresApproval → Run modes. A matched Allow runs immediately even in Review mode; a matched Ask pauses in Review and auto-approves in Autonomous.

Safety-critical interaction: a user-defined hook returning allow overrides policy rules, including a global deny (system hooks cannot trigger this override; overrides are audit-logged; only Administrators create hooks). Consequence: a permissive self-judging hook silently defeats your deny posture; the same failure class as the self-approving gate documented above, one layer up. Read every hook's prompt AND know which of your hooks return allow.

Patterns are tool-name globs; command tools also match arguments with toolGlob(argGlob) syntax; e.g. bash(az * delete *), RunKubectlWriteCommand(kubectl delete *), RunAzCliReadCommands. Argument-level matching expresses boundaries RBAC cannot (distinguish "scale the app" from "scale the load generator"). Configure via API:

# Global (Settings > Permissions equivalent)
curl -X PUT "$ENDPOINT/api/v2/agent/settings/global" \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{"permissions": {"allow": [...], "ask": [...], "deny": [...]}}'
# Custom-agent allow-only / thread allow-only
PUT $ENDPOINT/api/v2/extendedAgent/{name}/permissions
PUT $ENDPOINT/api/v2/threads/{threadId}/permissions

Limits: max 1,000 patterns per scope. Paths above are documentation-sourced; mark [VERIFY] against your agent before scripting.

Safe Defaults

  1. Always include reason when rejecting, especially for Stop.
  2. Keep timeout short (30s default), increase only when needed.
  3. Use failMode: allow for observability hooks and failMode: block for strict policy.
  4. Set maxRejections to prevent endless loops (default 3, valid 1-25).
  5. Log diagnostics to stderr for command hooks.
  6. Do not rely on managed connector Ask permissions as a human approval control in Autonomous workflows. Use Review mode, scoped connectors, or hooks that block unsafe output before promotion.

Hook Response and Exit-Code Semantics (verified 2026-08-10; re-verified against docs 2026-08-25)

  1. Hooks must output JSON. Simple format (recommended for prompt hooks): {"ok": true} or {"ok": false, "reason": "..."}. Expanded format (command hooks): {"decision": "allow"} / {"decision": "block", "reason": "..."}, optionally with hookSpecificOutput.additionalContext (injected as a user message).
  2. Command hooks can also use exit codes: 0 with no output = allow; 0 with JSON = parse decision; 2 = always block (stderr becomes the reason); any other code uses failMode (allow or block).
  3. failMode default is allow; use block for strict policy hooks.
  4. maxRejections valid range 1–25 (default 3), applies to prompt-type Stop hooks only; command-type Stop hooks have no implicit limit. When multiple prompt hooks set different values, the maximum is used.
  5. timeout default 30 s, range 1–300 s; values above 300 are flagged during CLI validation.
  6. Prompt hooks receive context via the $ARGUMENTS placeholder; if absent it is appended automatically. When a transcript is available, prompt hooks also get ReadFile and GrepSearch tools to reason over full conversation history. Command hooks receive context as JSON on stdin (tool_name, tool_input, tool_result, final_output, stop_rejection_count, ...); execution_summary points at a transcript file path.
  7. Caution — Stop-hook rejection without a reason is treated as approval, and the agent stops normally. Always provide a reason when rejecting.

Matcher and script constraints (documented 2026-08-25)

  • Matcher regexes are evaluated against tool names anchored as ^(pattern)$ and matched case-sensitively; empty/null match nothing; * matches all tools. This officially corroborates the live finding that snake_case matchers silently select zero PascalCase tools.
  • Command-hook scripts run in a sandboxed code interpreter: max size 64 KB, shebangs limited to #!/bin/bash and #!/usr/bin/env python3.
  • Prompt hooks carry a model option selecting the evaluating model tier, ReasoningFast; tiers are Reasoning / Fast Reasoning (default) / General Purpose / Fast / Long Context. Default suits per-call latency; reserve Reasoning for complex policy hooks where accuracy outweighs cost.

Audit Event Baseline

Use Application Insights customEvents as the baseline audit trail for agent behavior. Azure SRE Agent emits events for responses, model generation, tool execution, agent execution, meta-agent decisions, handoffs, incident activity snapshots, Azure CLI execution, and approval decisions.

Track these fields in operational reviews:

  1. TraceId, SpanId, ThreadId, and CorrelationId for investigation reconstruction.
  2. AgentToolExecution and AgentAzCliExecution for high-impact tool usage.
  3. ApprovalDecision for approval quality and denial patterns.
  4. IncidentActivitySnapshot for assisted versus mitigated incident outcomes.

API and UX Notes

  1. Use v2 extended-agent APIs for full hook configuration.
  2. Agent Canvas YAML view can omit hook details for v1 display.
  3. Manage operationally in Builder > Hooks or v2 API workflows.

AKS/Container Apps Governance Examples

Use approval gates for:

  1. az aks upgrade, node pool scale, disruptive maintenance actions
  2. Container App revision activation/deactivation or traffic shifts
  3. NSG rule changes and subnet routing changes

Use audit hooks for:

  1. all write-capable tool calls
  2. high-risk command classes
  3. privileged connector invocations

KT Compliance Stop Hook (P1/P2 or Write Actions)

Use this pattern to enforce KT section completeness when required:

api_version: azuresre.ai/v2
kind: ExtendedAgent
metadata:
  name: @@CUSTOM_AGENT_NAME@@
spec:
  hooks:
    Stop:
      - type: prompt
        prompt: |
          Review the response below.

          If incident severity is P1/P2 OR the response proposes a production write action,
          verify it includes these sections:
          1. Situation Appraisal
          2. Problem Analysis
          3. Decision Analysis
          4. Potential Problem Analysis

          If all required sections are present and meaningful, approve.
          Otherwise reject and explain which section is missing.

          $ARGUMENTS

          Return JSON:
          {"ok": true, "reason": "KT sections complete"}
          or
          {"ok": false, "reason": "Missing KT section(s): @@MISSING_SECTION_LIST@@."}
        timeout: 30
        failMode: block
        maxRejections: 3

Pair with kt-methodology.md and kt-templates.md. Canonical hook file: ../bundles/governance-kt/hooks/kt-completeness-gate.yaml

Sources

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent