All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referenceslive-verified-operations.md

≈9.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Live-Verified Operations

Production facts about Azure SRE Agent behaviour that were verified against a live deployed agent, not inferred from config files or docs. Core facts were verified on a live deployment 2026-08-10/11; the run-mode vocabulary, region list, path table, and property surface were re-verified against official sources 2026-08-25. Where a claim contradicts a reference or guide written earlier, the live-verified claim in this file is authoritative and the older file must be updated (see the contradiction register in QUALITY-REVIEW.md).

Load this file when: promoting a workflow to autonomous, writing or auditing guardrail hooks, debugging why an agent did not engage, verifying incident routing, or working with the data-plane API.

An agent that has never remediated has not earned Review

A trap seen on a real build 2026-08-10, worth checking on every engagement.

It is easy to reach a state where the agent reliably receives incidents, matches the right response plan, and writes accurate investigations , and to read that as "the loop works". It does not. Investigation and remediation are different capabilities with different risk, and only the first was demonstrated.

The condition hides because nothing looks wrong: in Review mode with a fault that a human or a script resolves, the agent is never required to propose an action, so the absence of remediation produces no error, no gap in the incident record, and no missing telemetry. IncidentMitigatedByAgent: False is also the correct value for Review mode, so it does not distinguish "proposed and awaiting approval" from "never proposed anything".

Check explicitly rather than inferring:

  • Has the agent ever produced a proposed action, not just a summary?
  • Has a human ever approved one, and did it resolve the incident?
  • Does the fault used in testing actually require remediation, or does it self-heal or get reverted by a script?

Until all three are yes, the defensible claim is "signal to alert to agent investigation", not "closed loop". Say that plainly in the write-up; the demo is still worth showing, and overstating it is what loses trust when someone asks.

Where run mode actually lives: ARM is agent-wide, the data plane is per workflow

Verified 2026-08-10. Getting this wrong turns "promote one workflow" into "promote everything".

Surface Field Scope
ARM, Microsoft.App/agents properties.actionConfiguration.mode Agent-wide fallback
Data plane, response-plans.json agentMode on each plan Per workflow

Setting the ARM field to Autonomous promotes every workflow at once, which is precisely what autonomy guidance says never to do. The per-plan field is the one you want:

{
  "id": "slo-burn-rate",
  "targetResourceType": "microsoft.resources/subscriptions/resourcegroups",
  "targetResource": "<resource-group-id>",
  "priorities": ["Sev0", "Sev1"],
  "titleContainsAny": ["SloAvailabilityFastBurn"],
  "handlingAgent": "alert-investigator",
  "agentMode": "review",                    // "review" | "autonomous"
  "maxAutomatedInvestigationAttempts": 2    // the rate limit
}

Response plans are applied through the data-plane phase (apply-extras.sh in the microsoft/sre-agent templates). The ARM API also documents a Microsoft.App/agents/incidentFilters sub-resource (PUT/GET/DELETE, base64 envelope) ; verify it against your pinned api-version before switching plan management to ARM; the data-plane semantics below remain the live-verified authority. maxAutomatedInvestigationAttempts is where the required rate limit is expressed.

Per-plan promotion works directly — an earlier "ceiling" reading was wrong. While properties.actionConfiguration.mode is Review, attempts to set agentMode: "automatic" on a response plan were rejected:

PUT  (create) with automatic  -> 400 Cannot create filter with automatic
POST (update) with automatic  -> 400 Cannot update handler with automatic

That rejection was initially read as "ARM mode is a ceiling; raise ARM first, then set plans". Re-testing proved the real cause: automatic is simply an invalid value everywhere (Autonomous/autonomous is valid). Set a plan's agentMode directly; raise ARM mode only when you intend to change the agent-wide default, remembering it governs incidents matching no plan.

Both surfaces use Autonomous. Automatic is not valid anywhere. Verified by rejection from each:

ARM PATCH  mode: "Automatic"       -> InvalidActionMode: must be null or one of
                                      'Review', 'Autonomous', 'ReadOnly'
Data plane agentMode: "automatic"  -> 400 Cannot update handler with automatic
Data plane agentMode: "autonomous" -> 200 OK   (also accepts "Autonomous")
Surface Field Valid values
ARM Microsoft.App/agents properties.actionConfiguration.mode ReadOnly, Review, Autonomous
Data plane response plan agentMode review, autonomous (case-insensitive)

The data-plane error message is misleading: "Cannot update handler with automatic" reads like a problem with the handler, and it is not; the handler is irrelevant (empty and every subagent name fail identically). It is an invalid-enum-value message. Do not go looking at handlingAgent.

Every plan you want to stay supervised must be explicitly review, because the agent-level mode is the fallback that governs any incident matching no plan.

That sequencing is genuinely awkward to get right, and it is the opposite of what "promote one workflow at a time" implies is possible. Plan for it: enumerate every response plan and set its mode explicitly before changing the agent-level default, so there is no window in which an unmatched incident runs autonomously.

Route pipeline-health alerts to their own plan. An alert saying "the SLO telemetry stopped arriving" needs the opposite response from "users are failing". Sharing a handler between them invites the agent to remediate a workload that is not broken.

Incident titles are the rule GROUP name, so titleContainsAny routing misfires

Verified live 2026-08-10, and it silently breaks per-alert routing.

For a Prometheus alert, the incident title the agent matches on is the rule group resource name (prg-<prefix>-sli-burnrate) not the individual alert rule name (SloAvailabilityFastBurn). So a plan filtering on specific alert names never matches on those names:

"titleContainsAny": ["SloAvailabilityFastBurn", "SloLatencyFastBurn"]   // never matches

What actually matched was the incidental substrings: burnrate in the group name, and the resource group name. The result on one outage: six incidents raised across all three response plans, including a pipeline-health plan whose filters (SloInputMetricAbsent, etc.) share no substring with the title at all. Routing was effectively arbitrary.

Consequences worth designing around:

  • One rule group per routing intent. If availability fast-burn must route differently from latency slow-burn, put them in separate prometheusRuleGroups, because the group name is the only part of the identity the response plan can see.
  • Filtering on alert-rule names inside a shared group does not work. Neither does relying on severity alone, since priorities is matched in addition to, not instead of, the title.
  • Verify routing by reading ResponsePlanId off the IncidentActivitySnapshot events rather than assuming the filter did what it says.

Overlapping response plans silently suppress investigation

The single most surprising behaviour found while proving an autonomous loop, and it presents as the agent simply doing nothing.

Because incident titles are the rule group name (above), several plans can match the same incident. When they do, the agent creates an incident per plan, marks each IncidentHandledOn within about a minute, and runs no tools at all. No error, no rejection, no diagnostic; the incident record looks handled.

Measured on the same fault, same agent, same autonomy setting, changing only which plans were enabled:

Plans enabled Tool executions Outcome
3 plans matching the same alert 0 No investigation, no action, 22 min
Only the autonomous plan 31+ Investigated, approved, remediated

So an agent that appears inert may simply be losing every incident to plan contention. Check for it by grouping IncidentActivitySnapshot events by ResponsePlanId; several distinct IncidentId values for one underlying alert is the signature.

Design rule: response plans must be mutually exclusive. With title matching this coarse, that usually means one Prometheus rule group per plan. Overlapping titleContainsAny sets are not merely redundant; they disable the agent.

Confirm autonomy by tool execution, never by handled timestamps

IncidentHandledOn is populated even when the agent does nothing, and AgentAutonomyLevel: Autonomous only reports configuration. Neither is evidence of work. The presence of AgentToolExecution events is.

A verified autonomous remediation, end to end

Recorded 2026-08-10 so the shape of a working loop is on file. Fault: a Deployment scaled to zero, which does not self-heal.

  1. SloAvailabilityFastBurn fired about 5 minutes after injection.
  2. The autonomous plan matched; the agent investigated with 31+ tool calls, reading its skills and querying Log Analytics.
  3. It called SearchMemory, found its own earlier incident recording an RBAC blocker, and *chose to re-verify rather than trust it: "the memory might be stale". It was: the missing role had since been granted. Worth noting as a positive behaviour, and as a reason to keep agent memory accurate.
  4. It handed to the remediation-advisor subagent, which approved the restore.
  5. It executed the scale through RunInTerminal, not RunKubectlWriteCommand. Guardrail hooks evaluated all 11 mutating calls (hooksRun=4, FinalDecision=pass, RejectionCount=0).
  6. Service recovered to 2/2 ready; the external probe returned HTTP 200. Injection to recovery, roughly 20 minutes.

The remediation arrived via a general shell tool. Any guardrail matcher that lists only RunKubectlWriteCommand and RunAzCliWriteCommands would have missed it entirely. Include Terminal and RunInTerminal in every matcher, and audit by watching which tool actually performs the change rather than assuming the specific one.

Plan-level review does NOT restrain an agent-level Autonomous

Controlled test reproduced twice (2026-08-10): agent at ARM actionConfiguration.mode: Autonomous, every response plan explicitly agentMode: review, then a fault that cannot self-heal (a Deployment scaled to zero). Result: 8 mutating RunInTerminal calls, the Deployment scaled back to 2/2 with no approval, handled by a deliberately inert control plan, incident snapshot reporting AgentAutonomyLevel: review throughout.

Design consequences:

  • ARM actionConfiguration.mode is the governing control. Autonomous there is a global write permission; plan-level agentMode did not restrain it downward in either reproduction.
  • Supervised operation means ARM mode: Review, whatever the plans say.
  • Bound autonomy with controls that were observed to hold: namespace-scoped Kubernetes RBAC on the agent identity, and hooks whose matchers actually bind to the mutating tool names.
  • A per-tier design ("fast burn autonomous, slow burn supervised") is not achievable through plan mode alone ; use separate agents, RBAC/hook differentiation, or accept everything is autonomous once the ARM ceiling is raised.

The two-agent pattern for graduated autonomy

When a design genuinely needs "urgent incidents may act, everything else must ask", deploy two agents with two identities (Microsoft.App/agents, stable api-version 2026-01-01, ordinary Bicep):

supervised autonomous
ARM mode Review Autonomous
Handles everything except the urgent class the urgent class only
Monitoring Contributor at subscription, if investigation needs it Reader, at resource-group scope
Kubernetes AKS RBAC Writer at namespace custom role, only the verbs its one workflow needs
Owns its own config may hold SRE Agent Administrator on itself never

Four rules that matter more than the table:

  1. Give the autonomous agent its own managed identity. Sharing one makes role assignments additive, and the narrow grants are silently subsumed by the broad ones.

  2. Azure Kubernetes Service RBAC Writer is not a narrow role. Namespace scoping makes it sound bounded, but its dataActions include configmaps/*, pods/*, persistentvolumeclaims/*, services/* and ingresses/*, all with delete. For an agent that acts without asking, write a custom role listing the exact dataActions its workflow needs. Avoid wildcards such as apps/deployments/*; the wildcard silently includes delete.

    Two deploy-blockers when writing that custom role, neither caught at compile time. Bicep validates the schema, not the contents, so both surface only as a failed deployment:

    • AssignableScopeMismatch — "at least one of the scopes that are available for assignment must be within the request scope". A Microsoft.Authorization/roleDefinitions created inside a resource-group-scoped module must list an assignable scope within that resource group. subscription().id is broader, not within, and fails. Use resourceGroup().id — which is also the better posture, since the role then cannot be assigned anywhere else in the subscription.
    • InvalidDataActionOrNotDataAction — every dataAction string must exist verbatim. Plausible-looking actions do not all exist: Microsoft.ContainerService/managedClusters/pods/log/read is not a data action even though pods/log is a real Kubernetes subresource. Azure does not enumerate it, so log access rides on pods/read.

    Validate the list before deploying rather than after:

    az provider operation show --namespace Microsoft.ContainerService -o json

    Diff your dataActions against that output in CI. One invented string fails the entire deployment, and the error names only the first offender.

  3. Never grant an autonomous agent SRE Agent Administrator on itself, and never Monitoring Contributor over the alerts that judge it. An agent that can rewrite its own response plans can widen its own remit; one that can silence its own alerts is not supervisable. Apply configuration with an operator identity instead.

  4. RBAC cannot express every boundary that matters, so pair it with an allowlist hook. RBAC cannot distinguish "scale the application" from "scale the load generator" when both are Deployments in one namespace , and scaling a load generator changes the SLI denominator and fakes a recovery. An autonomous agent's PreToolUse hook should therefore deny by default and approve only named resources with bounded parameters. A supervised agent's hook can afford to be denylist-shaped; an autonomous one cannot, because no human sees the call first.

Both agents receive every alert for a resource group they both watch, and each applies its own plans independently. Splitting workloads across two agents is therefore a change to both: remove the urgent class from the supervised agent's plans as well as adding it to the autonomous one. Otherwise the incident is handled twice, and the approval request the operator receives is for an action that has already been taken.

Also verify which plan actually handled each incident (ResponsePlanId on IncidentActivitySnapshot); with coarse title matching, do not assume the intended plan did.

An agent without incidentManagementConfiguration fails late and cryptically

A newly deployed agent without incidentManagementConfiguration provisions Succeeded, reports runningState: Running, accepts skills and hooks, then rejects every response-plan create with a bare 405 and an empty body. The neighbouring error actively misleads: POST to the same path answers 404 Incident filter not found. Use PUT to create a new filter.; PUT and POST each blame the other, neither mentions the real cause. An incident filter has nothing to filter until the agent has an incident source. Set it alongside the rest of the agent properties:

incidentManagementConfiguration: {
  type: 'AzMonitor'
  connectionName: 'azmonitor'
}

Diagnose this class of problem by diffing the broken agent against a working one (az resource show on both) rather than by reading the API errors. Note you cannot repair it with az resource update: that does GET-then-PUT, the GET returns logConfiguration.applicationInsightsConfiguration.connectionString as null because it is redacted on read, and the round-trip fails its own validation with InvalidApplicationInsightsConfiguration. Redeploy the template instead.

RBAC scope is not incident-ingestion scope

Verified against a live deployment: an agent scoped (RBAC-wise) to two resource groups ingested incidents only from the Sev1 Prometheus alert rule on the Azure Monitor workspace in one of them. Alert rules that fired in the second group (Health Model health-state alerts carrying the failing entity) were never ingested, even though the agent's identity had Monitoring Reader over that group. The AzMonitor incident source watches the alert rules the source is wired to, not every alert the identity can read. Health-state alerts are a human-triage feed (they carry the failing entity and auto-resolve); if a remediation workflow genuinely needs them, the alert rule must be reachable by a configured incident source, or route the alert into the source that is.

Auditing autonomous actions: Activity Log will not show them

A namespace-scoped custom-RBAC write (apps/deployments via the AKS data plane) is a Kubernetes API call, not an ARM control-plane call, so it writes nothing to Azure Activity Log. kubectl get events sees the write but is ephemeral (~1h TTL) and unqueryable. IncidentActivitySnapshot is not a reliable substitute: across two independent live scaled-to-zero remediations, IncidentMitigatedByAgent/IncidentAssistedByAgent read False on both even minutes after the Kubernetes event log showed the scale-up had happened; the snapshot proves an agent investigated, not what it did or which identity did it.

The durable answer is the AKS control plane audit log shipped via a kube-audit-admin diagnostic setting:

AzureDiagnostics
| where Category == "kube-audit-admin"
// objectRef / verb / user fields name the acting identity

This is the only signal here that names the acting identity, not just the acting agent resource. On one build it was the only thing that ever caught a Review-mode agent silently writing: two "independent live remediations" attributed to the autonomous agent had actually been performed by the supervised agent's identity via its broad AKS RBAC Writer role. Do not present mode: Review as a proven safety boundary without this audit-log check in the loop.

When narrowing RBAC afterward, verify the fix rather than trusting a clean deployment: re-list role assignments at the namespace scope and confirm the identity named in the kube-audit-admin record changed. And mind the parameter trap: the module's agentPrincipalId must be the user-assigned identity's principal ID (actionConfiguration.identity), not the agent resource's system-assigned identity.principalId; the wrong value deploys cleanly while granting the narrow role to an identity that never performs writes.

Check hooks for a self-approving gate before trusting their name

The supervised-agent escape above was root-caused to a PreToolUse hook literally named for human approval whose actual prompt was a self-judged four-condition gate: "DENY unless ALL of the following hold… Approve those when 1–4 hold." Nothing in that logic checks for a human decision; the agent judges its own justification and proceeds, and its own investigation satisfies its own conditions trivially. This project had already learned once that a hook's matcher can match zero tools and read as active; the same lesson recurs one layer deeper; a hook can match correctly, fire every time, and still not do what its name says, because its prompt logic grants itself an escape hatch. Read every hook's actual prompt before trusting what its name or description claims it does, and prefer unconditional-deny prompts for anything guarding a supervised agent.

Suspect the aggregation function when a burn-rate query flat-lines

During one total, 100%-failing outage, availability burn rate held flat at 3.46–3.57 against a >14 threshold for over five minutes because the query used sum_over_time: the SLI's cumulative :total had reached 18.3M requests, so minutes of fresh failures barely moved the lookback-averaged ratio. That was a symptom of the wrong PromQL function, not an inherent property of SLI-derived burn rate; rate()/increase() work fine once an SLI series has real history (an earlier misdiagnosis came from testing against a too-fresh series). Switching to increase() took identical fault-injection detection from 49 minutes to 10m06s. If a burn-rate query flat-lines well under threshold during an obvious outage, suspect the aggregation function before the SLI or the alert wiring.

Related demo trap: lowering a live metric alert's threshold via az rest PUT can fail with FailedToAddRoleAssignmentOn403 when the alert's managed identity already exists (the resubmission tries to recreate its Monitoring Data Reader assignment) despite the caller holding User Access Administrator. Prefer a real fault injection sized appropriately over patching a live alert's threshold.

Prerequisites before promoting any workflow to autonomous

Promote one workflow, never an agent. "The agent is reliable" is unmeasurable; "this remediation succeeded 47 of 48 times" is.

Record these in writing first; the act of writing them is what surfaces an unbounded action:

Prerequisite Why
Blast radius Exactly what the action can touch. kubectl rollout restart on one deployment in one namespace is bounded; a resource-limit patch that triggers rescheduling across a node pool is not obviously bounded.
Rollback path What restores the prior state, and who runs it. An action with no rollback cannot be autonomous at any success rate.
Rate limit Maximum actions per hour before the agent stops and escalates. An agent in a retry loop is a larger incident than the one it was fixing.
Evidence Demonstrated success against real incidents, not confidence. A workflow that has never executed cannot be promoted.
Demotion triggers Agreed in advance: two false positives in 30 days, any action that widened an incident, any action outside the stated blast radius.

Demotion is the mechanism working, not a failure. Automation that can only be promoted is automation nobody will trust.

Verifying that the agent received an incident

Verified live 2026-08-10. Use this before concluding an agent is not picking up alerts, the failure is usually in the verification method, not the agent.

The agent emits an IncidentActivitySnapshot custom event to its own Application Insights on every incident transition. This is the reliable check, and it is discoverable from the agent resource itself:

AGENT=$(az resource list --resource-type Microsoft.App/agents --query '[0].id' -o tsv)
az rest --method GET --url "https://management.azure.com${AGENT}?api-version=2026-01-01" \
  --query "properties.logConfiguration.applicationInsightsConfiguration.appId" -o tsv
az monitor app-insights query --app "<agent-app-insights-resource-id>" --analytics-query \
  "customEvents | where name == 'IncidentActivitySnapshot' | extend d=customDimensions | project timestamp, tostring(d.IncidentSeverity), tostring(d.ResponsePlanId), tostring(d.IncidentStatus), tostring(d.IncidentHandledOn), tostring(d.IncidentSummary) | order by timestamp desc" -o json

Useful dimensions: ResponsePlanId (confirms a custom plan matched rather than a built-in, cross-check ResponsePlanCustom), IncidentPlatform, AgentAutonomyLevel, IncidentStatus, IncidentCreatedOn / IncidentHandledOn, IncidentMitigatedByAgent, IncidentAssistedByAgent, and IncidentSummary (the written investigation). Measured pickup latency from alert fired to IncidentHandledOn: 1m26s and 1m12s.

In Review mode expect IncidentMitigatedByAgent: False and IncidentAssistedByAgent: False. Those are the posture working as configured, not a failure to engage, the presence of an IncidentSummary is what proves engagement.

Hook matchers match TOOL NAMES, and the names are PascalCase

Check this on every engagement before trusting any guardrail. Verified live 2026-08-10 on a deployed agent whose two safety hooks were completely inert.

Event-name caveat: this skill observed eventType: "PreToolUse" on a live agent's hooks (telemetry showed HooksRun: 2), but current official docs document only Stop and PostToolUse events. Treat the docs as authoritative for the supported surface and re-verify the event name on your agent before relying on it; the matcher mechanics below apply regardless.

A PreToolUse hook carries a matcher, which is a regex evaluated against the tool name. The agent's tools are PascalCase (RunKubectlWriteCommand, RunAzCliWriteCommands, Terminal) but hooks are commonly authored with snake_case, verb-style patterns copied from other agent frameworks:

"matcher": "^(restart_|scale_).*"     // matches NOTHING in this agent
"matcher": "^(delete_|remove_).*"     // matches NOTHING in this agent

Both were present, both activationMode: always, both matching zero tools. The agent had RunKubectlWriteCommand enabled, so had it been promoted to autonomous it could have scaled or deleted workloads with no guardrail evaluating at all. Nothing reports this: the hooks exist, they are enabled, and the telemetry shows HooksConfigured: 2, HooksRun: 2, FinalDecision: pass, which reads exactly like a guardrail working.

Verify by intersecting matchers with the real tool inventory:

TOKEN=$(az account get-access-token --resource "https://azuresre.dev" --query accessToken -o tsv)
curl -s "$ENDPOINT/api/v2/agent/tools" -H "Authorization: Bearer $TOKEN"      # {"data":[{"name":...,"enabled":...}]}
curl -s "$ENDPOINT/api/v2/extendedAgent/hooks" -H "Authorization: Bearer $TOKEN"

Then confirm each matcher actually selects the state-changing tools. A matcher that selects nothing is worse than no hook, because it looks like coverage.

The state-changing tools to enumerate in a matcher (from a live agent, 76 tools total, 45 enabled):

RunKubectlWriteCommand    RunAzCliWriteCommands    Terminal        RunInTerminal
CreateFile                ReplaceStringInFile      CreateDirectory MultiReplaceStringInFile
SaveFileToBlob

Hook body shape:

{ "name": "...", "type": "GlobalHook", "tags": [],
  "properties": {
    "eventType": "PreToolUse", "activationMode": "always", "description": "...",
    "hook": { "type": "prompt", "prompt": "...", "command": null, "script": null,
              "matcher": "^(RunKubectlWriteCommand|RunAzCliWriteCommands|...)$",
              "timeout": 30, "model": null, "failMode": null, "maxRejections": null, "sources": null } } }

type: "prompt" hooks are LLM-evaluated, so the prompt carries the actual policy; put the deny conditions in it explicitly (namespace bounds, reversibility, required evidence, forbidden targets) rather than relying on the matcher alone to express intent.

Global tool enablement and subagent tool lists are independent

/api/v2/agent/tools reports an enabled flag per tool, and it is easy to read that as the definitive capability list. It is not. A subagent's own tools array grants tools regardless of that flag: observed live, QueryLogAnalyticsByWorkspaceId reported enabled: false globally while the alert-investigator subagent listed it and ran it 54 times in one investigation.

So auditing capability means reading both: the global inventory and every subagent's tools array. Reviewing only the global list will understate what the agent can do.

The data-plane API

Verified working 2026-08-10 against a deployed agent. The token audience is https://azuresre.dev - not the agent's own azuresre.ai hostname, and not ARM:

TOKEN=$(az account get-access-token --resource "https://azuresre.dev" --query accessToken -o tsv)
ENDPOINT=$(az rest --method GET --url "https://management.azure.com<agent-id>?api-version=2026-01-01" --query properties.agentEndpoint -o tsv)
Path Method Resource
/api/v1/incidentPlayground/filters GET only (collection listing) Response plans — list
/api/v1/incidentPlayground/filters/{id} PUT=create, POST=update, DELETE Response plans — write
/api/v2/extendedAgent/skills/{name} GET / PUT Skills
/api/v2/extendedAgent/agents/{name} GET / PUT Subagents
/api/v2/extendedAgent/connectors/{name} PUT Connectors
/api/v1/scheduledtasks GET / POST / DELETE Scheduled tasks (also surfaced as /api/v2/extendedAgent/scheduledtasks/{name} in the API reference — both paths live-verified/docs-sourced, mark canonical [VERIFY])
/api/v2/agent/tools GET Tool inventory with per-tool enabled flag, under {"data":[...]}
/api/v2/agent/tools/configure POST Built-in tool enablement
/api/v2/extendedAgent/hooks/{name} GET / PUT Guardrail hooks
/api/v2/extendedAgent/commonprompts/{name} PUT / GET / PATCH / DELETE Common prompts (API reference)
/api/v2/extendedAgent/plugins/{name} PUT / GET / PATCH / DELETE Plugins (API reference)
/api/v1/httptriggers GET List HTTP triggers
/api/v1/httptriggers/create POST Create an HTTP trigger
/api/v1/httptriggers/{id} GET / PUT / DELETE HTTP trigger details and lifecycle
/api/v1/httptriggers/{id}/enable · /disable POST Toggle a trigger without deleting it
/api/v1/httptriggers/{id}/execute POST Run a trigger manually
/api/v1/httptriggers/{id}/executions GET Trigger execution history
/api/v1/httptriggers/trigger/{id} POST External webhook endpoint (API reference marks it "no auth required" while the HTTP-triggers doc requires a bearer token — docs conflict, mark [VERIFY])
/api/v1/agentmemory/upload POST Upload memory documents (multipart, max 100 MB total, 16 MB per file)
/api/v1/agentmemory/status GET Memory status
/api/v1/agentmemory/indexer-status GET Memory indexer progress
/api/v1/agentmemory/document/{fileName} · /api/v1/agentmemory/documents DELETE Delete one / all uploaded memory documents
/api/v2/repos/{repoName} PUT / GET / DELETE Code repositories connected to the agent
/api/v2/repos/{repoName}/test POST Test repo connectivity
/api/v1/threads · /api/v1/threads/{id} GET Conversation threads
/api/v1/threads/{id}/messages GET / POST Thread messages (POST starts a conversation)
/api/v1/approvals/{threadId} GET Pending approvals
/api/v1/approvals/{threadId}/{approvalId}/decision POST Approve or reject a proposed action
/api/v2/agent/settings/global PUT Global permissions / tool-access policies (allow/ask/deny)
/api/v2/extendedAgent/{name}/permissions · /api/v2/threads/{id}/permissions PUT Per-custom-agent / per-thread allow policies
/agentHub SignalR hub Real-time chat streaming (same bearer token)

(The threads, approvals, repos, and policies rows above are documentation-sourced — API reference and tool-access-policies docs, verified 2026-08-25 — rather than live-probed from this skill's deployments; treat paths as [VERIFY] against your agent before scripting against them.)

Two-phase deployment split, per official deploy-iac docs (2026-08-25): ARM Phase 1 covers the agent plus connectors, skills, subagents, and tools as ARM sub-resources. Phase 2 data-plane remains required for code repositories (Git auth), hooks ("not yet exposed as ARM sub-resources at deploy time"), HTTP triggers (server-generated URLs), knowledge files (binary upload), and plugin configuration.

Endpoints that do not exist and return the SPA's HTML shell with HTTP 200: /api/v1/incidents, /api/v2/extendedAgent/knowledge, /api/v1/memory. Probing these is how the earlier path-guessing went wrong. Note the docs surface memory under /api/v1/agentmemory/* (not /api/v1/memory).

Response plans being called incidentPlayground/filters is the least guessable thing here, and it is why path-guessing failed: earlier probes of /api/v1/incidents returned HTTP 200 with the SPA's HTML shell. A 200 from a SPA host means the host serves a web app, not that the endpoint exists - always check the content type before treating a 200 as a working API.

The response-plan id must be in the PATH, not just the body. Verified 2026-08-11 by probing all four combinations:

Request Result
PUT /api/v1/incidentPlayground/filters 405
POST /api/v1/incidentPlayground/filters 405
PUT /api/v1/incidentPlayground/filters/{id} 409 if it exists, otherwise creates
POST /api/v1/incidentPlayground/filters/{id} 200 — updates

The collection path is GET-only. Writing to it returns 405 with an empty body, which is easy to misread as a transient failure or an auth problem rather than the wrong path; there is no message telling you the id is missing. An id inside the JSON body is not enough.

PUT creates, POST updates. PUT against an existing id returns 409 Incident filter with the same ID already exists. Use POST to update. This matters because the obvious workaround is DELETE-then-PUT, and that is destructive: if the PUT is then rejected for any reason, the response plan is gone and the alerts it handled silently route nowhere. Observed live -- a plan was missing for about a minute because a DELETE succeeded and the following PUT was rejected by a validation rule. Use POST to update; never DELETE a live plan as a step toward changing it.

GET and PUT do not use the same serialization. GET returns array fields as stringified Python-style lists ("tools": "['RunAzCliReadCommands']"), while PUT requires real JSON arrays. Copying a GET response back into a PUT produces 400 The JSON value could not be converted to System.Collections.Generic.List. Build PUT bodies from the documented shape, not from a round-tripped GET.

Skill body:

{ "name": "...", "type": "Skill",
  "properties": { "description": "...", "tools": [], "skillContent": "...", "additionalFiles": [] } }

Subagent body:

{ "name": "...", "type": "ExtendedAgent", "tags": [], "owner": "",
  "properties": { "instructions": "...", "handoffDescription": "...", "handoffs": [],
                  "tools": [], "mcpTools": [], "allowParallelToolCalls": true, "enableSkills": true } }

Reference implementation worth reading before writing your own: scripts/setup-sre-agent.ps1 in lukemurraynz/drasi-aks-sre-agent, which wraps all of the above with retry.

Check whether the skills are actually populated. A deployed agent can carry a skill whose skillContent is an empty string while its subagent instructions say "use the X skill to gather data". That combination produces confident-looking investigations that query the wrong tables, and nothing reports an error. Observed live: an investigation ran five Log Analytics queries against Application Insights tables the workload never populated, returning ZERO_ROWS_RETURNED each time, because the skill it was told to use was empty.

Alerts do not reach the agent through an action group. With incidentManagementConfiguration.type: AzMonitor the agent subscribes to Azure Monitor directly, scoped by knowledgeGraphConfiguration.managedResources. An action group with only an email receiver is not a wiring gap. What matters is that the alert fires within a resource group the agent manages, adding a webhook to "connect the agent" is cargo cult.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent