All skills
lukemurraynz avatar

/azure-sre-agent

@2cc2455

Design, configure, review, and operate production-grade Azure SRE Agent capabilities: response plans, scheduled tasks, HTTP triggers, custom agents, autonomous and review workflows, approval guardrails, AMBA observability, source RCA, connectors, MCP, governance hooks, WAF reviews, AI Foundry posture, Digital Native governance, postmortem generation, and KT discipline.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/azure-sre-agent

This session only. Nothing lands on disk.

referencesknowledge-files-pattern.md

≈1.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Knowledge Files Pattern

Source inspiration: MSFTLabs/msftlabs-sre-agent-demo

Azure SRE Agent supports uploading Knowledge Files: markdown documents that the agent reads before investigating any incident. This pattern gives the agent the same situational awareness a senior SRE builds over months of on-call duty for a specific system.

What belongs in a knowledge file

A knowledge file is a structured operational document, not general documentation. Include:

  • Resource inventory — names, SKUs, resource groups, subscriptions, and key configuration settings for all Azure resources the agent will work with.
  • Normal-state baselines — what healthy metrics look like, health probe endpoints, expected error rates.
  • Key telemetry signals — which KQL queries surface what problems (e.g., "spike in exceptions table = unhandled exception in web app").
  • Dependency map — how resources connect (App Gateway → App Service → SQL → Key Vault).
  • Common failure modes — the three or four most likely causes for each service, with triage shortcuts.
  • Managed identity and RBAC layout — which identity the agent uses and what permissions it has.
  • Environment-specific connection strings or app settings format (no secrets ; structure only).

What NOT to include

  • Secrets, connection strings with passwords, or storage keys.
  • Generic Microsoft documentation — the agent already knows this.
  • Rarely-changing content that is already in ARM metadata.

Recommended knowledge files per workload type

App Gateway + App Service + SQL workload (e.g., MSFTLabs demo)

File Contents
application-architecture.md Full resource inventory, DB schema, controller endpoints, chaos triggers, managed identity grants, normal-state baselines.
appgw-health-probe-troubleshooting.md Probe config, triage decision tree, ready-to-use KQL queries (probe timeline, SQL error correlation, UnhealthyHostCount trend), symptom → cause → fix table, recovery verification steps.

AKS workload

File Contents
aks-cluster-overview.md Node pool config, namespaces, critical deployments, PDB settings, HPA min/max, KEDA scaler sources.
aks-triage-runbook.md Common failure modes (OOMKilled, CrashLoopBackOff, node pressure), standard kubectl commands, escalation path.

Azure OpenAI / AI Foundry workload

File Contents
ai-foundry-overview.md Deployed models, PTU vs. pay-as-you-go split, quota limits, APIM policy summary, RAI policy names.
ai-cost-baselines.md Normal token consumption rates per model, expected PTU utilisation band, cost anomaly thresholds.

How to upload knowledge files

  1. Go to sre.azure.com and open your SRE Agent.
  2. Navigate to Knowledge → Add knowledge file.
  3. Upload the .md file. The agent indexes it and uses it in all subsequent investigations.
  4. Update knowledge files when the architecture changes; stale knowledge is worse than no knowledge.

Triage investigation prompt pattern (from MSFTLabs demo)

When configuring an SRE Agent for a specific workload, use a structured investigation prompt that references these steps:

1. Confirm the outage: identify which backend services are down from the alert metadata.
2. Check resource health: inspect gateway/probe health and recent error rates.
3. Check Activity Log: query for write/delete operations on impacted resources (last 24h).
4. Check Application Insights: look for connectivity exceptions and failed dependencies.
5. Correlate changes to impact: determine if a recent change aligns with the outage start time.
6. Propose targeted rollback: confirm with user before executing any rollback.
7. Acknowledge and close the alert after remediation is verified.

Constraints:
- If no recent changes found, state clearly and suggest escalation — do not guess.
- Limit rollback to the specific change correlated with the outage.
- Do not repeat the same diagnostic steps.

Value proposition

The knowledge files pattern complements skill-based proactive reviews:

Custom Skills Knowledge Files
Best for Proactive governance, WAF, FinOps, posture reviews Reactive incident investigation of a specific system
Agent context Generic Azure knowledge + skill instructions Deep environment-specific context
Maintenance Update when skill logic changes Update when infrastructure changes
Format YAML front matter + structured markdown Plain markdown, any structure

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The Azure SRE Agent skill provides a production-grade framework for managing Azure infrastructure using AI agents. It incorporates extensive safety documentation, approval-based hooks, and least-privilege role templates. The 'low' verdict is assigned due to the inherent risk of indirect prompt injection when the agent processes external incident data and source code, a necessary function for its SRE capabilities.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
compatibility
Azure SRE Agent; GitHub Copilot agent skills; new projects only
Other metadata
metadata
{
  "last_verified": "2026-08-25",
  "version": "2.23.3",
  "risk": "critical",
  "last_updated": "2026-08-25"
}

README badge

README badge for lukemurraynz/hve-agent-skills/azure-sre-agent