Knowledge Files Pattern
Source inspiration: MSFTLabs/msftlabs-sre-agent-demo
Azure SRE Agent supports uploading Knowledge Files: markdown documents that the agent reads before investigating any incident. This pattern gives the agent the same situational awareness a senior SRE builds over months of on-call duty for a specific system.
What belongs in a knowledge file
A knowledge file is a structured operational document, not general documentation. Include:
- Resource inventory — names, SKUs, resource groups, subscriptions, and key configuration settings for all Azure resources the agent will work with.
- Normal-state baselines — what healthy metrics look like, health probe endpoints, expected error rates.
- Key telemetry signals — which KQL queries surface what problems (e.g., "spike in
exceptionstable = unhandled exception in web app"). - Dependency map — how resources connect (App Gateway → App Service → SQL → Key Vault).
- Common failure modes — the three or four most likely causes for each service, with triage shortcuts.
- Managed identity and RBAC layout — which identity the agent uses and what permissions it has.
- Environment-specific connection strings or app settings format (no secrets ; structure only).
What NOT to include
- Secrets, connection strings with passwords, or storage keys.
- Generic Microsoft documentation — the agent already knows this.
- Rarely-changing content that is already in ARM metadata.
Recommended knowledge files per workload type
App Gateway + App Service + SQL workload (e.g., MSFTLabs demo)
| File | Contents |
|---|---|
application-architecture.md |
Full resource inventory, DB schema, controller endpoints, chaos triggers, managed identity grants, normal-state baselines. |
appgw-health-probe-troubleshooting.md |
Probe config, triage decision tree, ready-to-use KQL queries (probe timeline, SQL error correlation, UnhealthyHostCount trend), symptom → cause → fix table, recovery verification steps. |
AKS workload
| File | Contents |
|---|---|
aks-cluster-overview.md |
Node pool config, namespaces, critical deployments, PDB settings, HPA min/max, KEDA scaler sources. |
aks-triage-runbook.md |
Common failure modes (OOMKilled, CrashLoopBackOff, node pressure), standard kubectl commands, escalation path. |
Azure OpenAI / AI Foundry workload
| File | Contents |
|---|---|
ai-foundry-overview.md |
Deployed models, PTU vs. pay-as-you-go split, quota limits, APIM policy summary, RAI policy names. |
ai-cost-baselines.md |
Normal token consumption rates per model, expected PTU utilisation band, cost anomaly thresholds. |
How to upload knowledge files
- Go to sre.azure.com and open your SRE Agent.
- Navigate to Knowledge → Add knowledge file.
- Upload the
.mdfile. The agent indexes it and uses it in all subsequent investigations. - Update knowledge files when the architecture changes; stale knowledge is worse than no knowledge.
Triage investigation prompt pattern (from MSFTLabs demo)
When configuring an SRE Agent for a specific workload, use a structured investigation prompt that references these steps:
1. Confirm the outage: identify which backend services are down from the alert metadata.
2. Check resource health: inspect gateway/probe health and recent error rates.
3. Check Activity Log: query for write/delete operations on impacted resources (last 24h).
4. Check Application Insights: look for connectivity exceptions and failed dependencies.
5. Correlate changes to impact: determine if a recent change aligns with the outage start time.
6. Propose targeted rollback: confirm with user before executing any rollback.
7. Acknowledge and close the alert after remediation is verified.
Constraints:
- If no recent changes found, state clearly and suggest escalation — do not guess.
- Limit rollback to the specific change correlated with the outage.
- Do not repeat the same diagnostic steps.Value proposition
The knowledge files pattern complements skill-based proactive reviews:
| Custom Skills | Knowledge Files | |
|---|---|---|
| Best for | Proactive governance, WAF, FinOps, posture reviews | Reactive incident investigation of a specific system |
| Agent context | Generic Azure knowledge + skill instructions | Deep environment-specific context |
| Maintenance | Update when skill logic changes | Update when infrastructure changes |
| Format | YAML front matter + structured markdown | Plain markdown, any structure |