All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencesaudit-diagnostics-2am.md

≈908 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Use this reference

Use this file when you need a fast path through Azure SRE Agent audit telemetry in Application Insights customEvents.

Event types verified from Microsoft Learn on 2026-06-05:

  • AgentToolExecution
  • ModelGeneration
  • ApprovalDecision
  • IncidentActivitySnapshot

Also use shared correlation fields such as TraceId, SpanId, ParentSpanId, ThreadId, and CorrelationId.

Start here at 2am

  1. Open the agent's Application Insights resource from Monitor > Logs.
  2. Switch the default query from traces to customEvents.
  3. Start with the thread timeline query below.
  4. Pivot to tool usage, approval history, incident outcomes, or token usage as needed.

Thread timeline

customEvents
| where timestamp > ago(7d)
| where tostring(customDimensions.ThreadId) == "<THREAD_ID>"
| project timestamp,
    Event = name,
    EventType = tostring(customDimensions.EventType),
    Tool = tostring(customDimensions.ToolName),
    Agent = tostring(customDimensions.SubAgentName),
    TraceId = tostring(customDimensions.TraceId)
| sort by timestamp asc

Tool execution history

customEvents
| where name == "AgentToolExecution"
| where timestamp > ago(24h)
| project timestamp,
    Tool = tostring(customDimensions.ToolName),
    EventType = tostring(customDimensions.EventType),
    Input = tostring(customDimensions.ToolInput),
    Output = tostring(customDimensions.ToolOutput),
    CallId = tostring(customDimensions.CallId),
    Agent = tostring(customDimensions.SubAgentName)
| sort by timestamp desc

Approval decisions

customEvents
| where name == "ApprovalDecision"
| where timestamp > ago(30d)
| project timestamp, customDimensions
| sort by timestamp desc

Incident outcomes

customEvents
| where name == "IncidentActivitySnapshot"
| where timestamp > ago(30d)
| project timestamp,
    IncidentId = tostring(customDimensions.IncidentId),
    Title = tostring(customDimensions.IncidentTitle),
    Platform = tostring(customDimensions.IncidentPlatform),
    MitigatedByAgent = tostring(customDimensions.IncidentMitigatedByAgent),
    AssistedByAgent = tostring(customDimensions.IncidentAssistedByAgent),
    Autonomy = tostring(customDimensions.AgentAutonomyLevel),
    ResponsePlan = tostring(customDimensions.ResponsePlanId)
| sort by timestamp desc

Token and AAU proxy trend

Use ModelGeneration events as the local signal for token cost. Azure billing remains the source of truth for actual AAU charges.

customEvents
| where name == "ModelGeneration"
| where customDimensions.EventType == "ModelGenerationEnd"
| where timestamp > ago(30d)
| extend Agent = tostring(customDimensions.AgentName),
    Model = tostring(customDimensions.ModelId),
    InputTokens = toint(customDimensions.InputTokens),
    OutputTokens = toint(customDimensions.OutputTokens)
| summarize TotalInput = sum(InputTokens), TotalOutput = sum(OutputTokens), Calls = count() by Agent, Model
| sort by TotalInput desc

What to check before autonomy promotion

  • repeated approval decisions for the same action pattern
  • tool failures or repeated retries in AgentToolExecution
  • high token use with low incident mitigation value
  • false positives or noisy response plans in IncidentActivitySnapshot

Source

Primary source: references/source-map.md → Microsoft Learn audit-agent-actions

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent