All skills
microsoft avatar

/azure-diagnostics

@ae5e585
by microsoftmicrosoft/skills3.1k stars
351

Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, analyze logs, KQL, insights, image pull failures, cold start issues, health probe failures, resource health, root cause of errors, troubleshoot event hubs, troubleshoot service bus, messaging SDK error, AMQP connection failure, message lock lost, service bus dead letter.

Use this Skill: https://skilld.dev/gh/microsoft/skills/azure-diagnostics

This session only. Nothing lands on disk.

troubleshootingaksaks-troubleshooting.md

≈1.6k tokens on demand. Your agent reads this file only when SKILL.md points to it.

AKS Troubleshooting Guide

Primary AKS troubleshooting guide for incidents routed from ../../SKILL.md.

When to Use This Guide

  • lifecycle, access, node, kube-system, workload, ingress, DNS, or scaling issues
  • kubectl cannot connect, nodes are NotReady, or pods are unhealthy

Scenario Playbooks

Scenario Reference
broad cluster investigation general-diagnostics.md
workload, crash, image pull, readiness, or pending pod issues pod-failures.md
node health, scaling, pressure, upgrade, or zone issues node-issues.md
service, ingress, DNS, or network policy issues networking.md

Tool Selection For Diagnostics

When gathering AKS diagnostic evidence, prefer mcp_azure_mcp_aks, then the smallest discovered AKS-MCP tool that fits the read, then supporting Azure tools such as mcp_azure_mcp_applens, mcp_azure_mcp_monitor, or mcp_azure_mcp_resourcehealth. Use raw az aks and kubectl only when the AKS-MCP surface cannot perform the needed check.

When standard diagnostics do not reveal root cause, use Inspektor Gadget for real-time, low-level node and pod observability (DNS traces, TCP traces, process snapshots, file access traces). See references/inspektor-gadget.md for the gadget catalog, the run-ig script, and symptom-to-gadget mapping.

See references/aks-mcp.md, references/structured-input-modes.md, references/command-flows.md

Required Inputs

  • subscription or active Azure context
  • resource group and cluster name
  • symptom summary
  • first observed time or recent change window
  • impacted namespace, workload, service, or ingress when known

If cluster identity is missing, stop and ask for it.

Scope Buckets

  • Lifecycle: create, update, start, stop, upgrade, or provisioning failures
  • API access: kubeconfig, auth, private endpoint, DNS, or reachability problems
  • Nodes: missing nodes, NotReady, pressure, CNI, kubelet, certificate, or VMSS drift
  • kube-system: CoreDNS, metrics-server, konnectivity, ingress, CNI, CSI, or add-on failures
  • Workloads: Pending, CrashLoopBackOff, OOMKilled, PVC, quota, secret, readiness, or dependency issues
  • Connectivity and DNS: pod -> service -> endpoints -> ingress/load balancer -> DNS -> network controls
  • Scaling: node pool sizing, pending pods, autoscaler config, metrics, quota, or subnet constraints

Evidence Order

  1. Azure-side state first: cluster state, resource health, recent operations, node pool state, detector or monitoring output.
  2. Kubernetes-side state second: cluster reachability, nodes, kube-system, events, affected namespace, pod detail, logs.
  3. Use detector, warning-event, or metrics modes when the incoming data already matches them.
  4. Deep diagnostics; when steps 1–3 do not reveal root cause, use inspektor-gadget.md for real-time tracing and snapshots on the affected node.

Workflow

  1. Get cluster context.
  2. Classify the problem by scope bucket.
  3. Prefer Azure-side evidence before Kubernetes-side evidence.
  4. Use the matching AKS-MCP path first, then the documented CLI fallback if MCP cannot perform that read.
  5. Return evidence, failure domain, confidence, next checks, remediation, and escalation.

Error Patterns

  • No cluster context: ask for subscription, resource group, and cluster name.
  • MCP unavailable: fall back to safe az aks and kubectl reads.
  • kubectl blocked: separate auth problems from network reachability.
  • Logs or metrics missing: use events, node state, and resource descriptions.
  • Detector noise: ignore emergingIssues, prefer critical findings, rank the most actionable signal first.

Safe Fallback Checks

When AKS-MCP cannot perform the baseline read, run the aks-baseline script. It executes the read-only cluster + Kubernetes baseline sweep (provisioning state, node pools, activity log, node readiness, unhealthy pods, kube-system health, warning events) and returns a single labeled digest:

# bash
./scripts/aks-baseline.sh -g <resource-group> -n <cluster-name> [--namespace <namespace>]
# PowerShell
.\scripts\aks-baseline.ps1 -ResourceGroup <resource-group> -Cluster <cluster-name> [-Namespace <namespace>]

Then deep-dive on a specific pod as the digest indicates:

az aks show -g <resource-group> -n <cluster-name>
az aks nodepool list -g <resource-group> --cluster-name <cluster-name>
kubectl cluster-info
kubectl get nodes -o wide
kubectl get pods -n kube-system
kubectl get events -A --sort-by=.lastTimestamp
kubectl describe pod <pod-name> -n <namespace>
kubectl logs <pod-name> -n <namespace> --previous

For unhealthy pods, gather the full read-only evidence bundle (describe, current + previous logs, resources vs usage) with the pod-evidence script instead of running the commands one by one — ../../scripts/pod-evidence.sh / ../../scripts/pod-evidence.ps1:

../../scripts/pod-evidence.sh <pod-name> -n <namespace>
../../scripts/pod-evidence.sh --all-failing
../../scripts/pod-evidence.ps1 <pod-name> -Namespace <namespace>
../../scripts/pod-evidence.ps1 -AllFailing

See pod-failures.md for how to interpret the digest.

Keep these read-only unless the user explicitly asks for remediation.

Guardrails

  • default to read-only diagnostics
  • do not restart, delete, cordon, drain, scale, upgrade, or reconfigure resources unless the user explicitly asks for remediation
  • do not conclude root cause without quoting the evidence that supports it

Output Checklist

Return scope and impact, evidence, failure domain, root cause, confidence, next checks, remediation, and escalation.

Source: SKILL.md on GitHub

1 warning15d4 checks · Risk SAFE
  • Gen Agent Trust Hub15d

    This skill is designed for Azure diagnostics and troubleshooting, providing a comprehensive set of guides and scripts that utilize standard tools like the Azure CLI and kubectl. It includes some security considerations, such as the ingestion of logs which provides a surface for indirect prompt injection, and the use of privileged debug pods for advanced diagnostics. These are used within the skill's intended functionality and are accompanied by appropriate guidance for user approval.

  • Socket15d

    No alerts

  • Snyk15d

    Risk: LOW · No issues

  • Runlayer7mo

    4/4 files flagged

Signed by skilld at ae5e585. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
metadata
{
  "author": "Microsoft",
  "version": "1.2.6"
}

README badge

README badge for microsoft/skills/azure-diagnostics