All skills
microsoft avatar

/azure-diagnostics

@ae5e585
by microsoftmicrosoft/skills3.1k stars
351

Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, analyze logs, KQL, insights, image pull failures, cold start issues, health probe failures, resource health, root cause of errors, troubleshoot event hubs, troubleshoot service bus, messaging SDK error, AMQP connection failure, message lock lost, service bus dead letter.

Use this Skill: https://skilld.dev/gh/microsoft/skills/azure-diagnostics

This session only. Nothing lands on disk.

troubleshootingaksreferencescommand-flows.md

≈1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

AKS Command Flows

Cluster Baseline Flow

Resolve subscription -> resolve resource group -> resolve cluster -> inspect cluster state -> inspect node pools -> inspect resource health -> inspect recent operations

CLI fallback when AKS-MCP cannot perform the cluster baseline read — run the aks-baseline script, which gathers cluster state, node pools, and recent operations as one read-only digest:

# bash
./scripts/aks-baseline.sh -g <resource-group> -n <cluster-name>
# PowerShell
.\scripts\aks-baseline.ps1 -ResourceGroup <resource-group> -Cluster <cluster-name>

Kubernetes Baseline Flow

Check API reachability -> inspect nodes -> inspect kube-system -> inspect events -> inspect affected namespace -> inspect pod details and logs

CLI fallback when AKS-MCP cannot perform the Kubernetes baseline read — the same aks-baseline script also covers node readiness, unhealthy pods, kube-system health, and recent warning events. Pass --namespace to include an affected namespace, then deep-dive on a specific pod:

kubectl cluster-info
kubectl get nodes -o wide
kubectl get pods -n kube-system
kubectl get events -A --sort-by=.lastTimestamp
kubectl get pods -n <namespace>

For pod detail and logs, gather the read-only evidence bundle (describe, current + previous logs, resources vs usage) with the pod-evidence script — ../../../scripts/pod-evidence.sh / ../../../scripts/pod-evidence.ps1:

../../../scripts/pod-evidence.sh <pod-name> -n <namespace>
../../../scripts/pod-evidence.sh --all-failing
kubectl describe pod <pod-name> -n <namespace>
kubectl logs <pod-name> -n <namespace> --previous
../../../scripts/pod-evidence.ps1 <pod-name> -Namespace <namespace>
../../../scripts/pod-evidence.ps1 -AllFailing
# PowerShell
.\scripts\aks-baseline.ps1 -ResourceGroup <resource-group> -Cluster <cluster-name> -Namespace <namespace>

Connectivity Flow

pod -> service -> endpoints -> ingress or load balancer -> DNS -> network controls

CLI fallback when AKS-MCP cannot perform the connectivity read:

kubectl get pods -n <namespace> -o wide
kubectl get svc -n <namespace>
kubectl get endpoints -n <namespace>
kubectl get ingress -n <namespace>
kubectl describe ingress <ingress-name> -n <namespace>

Detector Flow

resolve cluster resource ID -> list detectors or choose category -> select a focused time window -> run the detector or category -> rank critical findings above warnings -> ignore emerging issues when choosing the primary root cause

Monitoring Flow

check resource health -> inspect metrics -> verify diagnostics settings -> inspect control plane logs if available -> correlate with Application Insights or namespace symptoms

Scheduling Flow

pod events -> node capacity -> taints and tolerations -> affinity rules -> PVC state -> quotas

CLI fallback when AKS-MCP cannot perform the scheduling read:

kubectl describe pod <pod-name> -n <namespace>
kubectl get nodes -o wide
kubectl describe node <node-name>
kubectl get pvc -n <namespace>
kubectl describe quota -n <namespace>

Deep Diagnostics Flow (Inspektor Gadget)

Standard diagnostics inconclusive -> select gadget from symptom-to-gadget map -> run `scripts/run-ig.sh` (or `run-ig.ps1`; resolves node, applies timeout) -> interpret output -> correlate with prior evidence

Use when steps 1–3 of the evidence order (Azure-side, Kubernetes-side, and detector evidence) do not reveal root cause. See inspektor-gadget.md for the full gadget catalog and command patterns.

Safety Boundary

Treat the following as change operations and avoid them unless the user explicitly asks for remediation:

  • deleting or restarting pods
  • cordon and drain operations
  • scaling workloads or node pools
  • cluster upgrade operations
  • DNS, route, NSG, or firewall changes

Source: SKILL.md on GitHub

1 warning15d4 checks · Risk SAFE
  • Gen Agent Trust Hub15d

    This skill is designed for Azure diagnostics and troubleshooting, providing a comprehensive set of guides and scripts that utilize standard tools like the Azure CLI and kubectl. It includes some security considerations, such as the ingestion of logs which provides a surface for indirect prompt injection, and the use of privileged debug pods for advanced diagnostics. These are used within the skill's intended functionality and are accompanied by appropriate guidance for user approval.

  • Socket15d

    No alerts

  • Snyk15d

    Risk: LOW · No issues

  • Runlayer7mo

    4/4 files flagged

Signed by skilld at ae5e585. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
metadata
{
  "author": "Microsoft",
  "version": "1.2.6"
}

README badge

README badge for microsoft/skills/azure-diagnostics