All skills
microsoft avatar

/azure-diagnostics

@ae5e585
by microsoftmicrosoft/skills3.1k stars
351

Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, analyze logs, KQL, insights, image pull failures, cold start issues, health probe failures, resource health, root cause of errors, troubleshoot event hubs, troubleshoot service bus, messaging SDK error, AMQP connection failure, message lock lost, service bus dead letter.

Use this Skill: https://skilld.dev/gh/microsoft/skills/azure-diagnostics

This session only. Nothing lands on disk.

troubleshootingaksnetworking.md

≈1.4k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Networking Troubleshooting

For CNI-specific issues, check CNI pod health and review AKS networking concepts.

Service Unreachable / Connection Refused

Diagnostics - always start here:

# 1. Verify service exists and has endpoints (read-only)
kubectl get svc <service-name> -n <ns>
kubectl get endpoints <service-name> -n <ns>

# 2. Optional connectivity test from inside the namespace
# This creates a temporary pod. Prefer read-only checks first.
# Only use it after the user explicitly approves a mutating test.
kubectl run netdebug --image=curlimages/curl -it --rm -n <ns> -- \
  curl -sv http://<service>.<ns>.svc.cluster.local:<port>/healthz

Decision tree:

Observation Cause Fix
Endpoints shows <none> Label selector mismatch Align selector with pod labels; check for typos
Endpoints has IPs but unreachable Port mismatch or app not listening Confirm targetPort = actual container port
Works from some pods, fails from others Network policy blocking See Network Policy section
Works inside cluster, fails externally Load balancer issue See Load Balancer section
ECONNREFUSED immediately App not listening on that port Check listening ports in the pod

Pods that are running but not Ready are removed from Endpoints. Check kubectl get pod <pod> -n <ns>.

Deep diagnostics with Inspektor Gadget (when the above checks are inconclusive):

Use scripts/run-ig.sh (or run-ig.ps1) with --pod <pod-name> --ns <ns> and these gadgets:

  • snapshot_socket — check what ports the pod is listening on
  • trace_tcp — trace connect/accept/close events
  • trace_tcpretrans — packet retransmissions

See references/inspektor-gadget.md.


DNS Resolution Failures

Diagnostics:

The live DNS test creates a temporary pod. Prefer get, describe, logs, or exec into an existing pod first. Only use it after the user explicitly approves creating the test pod.

# Confirm CoreDNS is running and healthy (read-only)
kubectl get pods -n kube-system -l k8s-app=kube-dns -o wide
kubectl top pod -n kube-system -l k8s-app=kube-dns

# Optional live DNS test from the same namespace as the failing pod
kubectl run dnstest --image=busybox:1.28 -it --rm -n <ns> -- \
  nslookup <service-name>.<ns>.svc.cluster.local

# CoreDNS logs - errors show here first
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=100

DNS failure patterns:

Symptom Cause Fix
NXDOMAIN for svc.cluster.local CoreDNS down or pod network broken After confirming the diagnostics above, coordinate with the cluster operator to restart or redeploy CoreDNS and verify CNI
Internal resolves, external NXDOMAIN Custom DNS not forwarding to 168.63.129.16 Fix upstream forwarder
Intermittent SERVFAIL under load CoreDNS CPU throttled Remove CPU limits or add replicas
Private cluster - external names fail Custom DNS missing privatelink forwarder Add conditional forwarder to Azure DNS
i/o timeout not NXDOMAIN Port 53 blocked by NetworkPolicy or NSG Allow UDP/TCP 53 from pods to kube-dns ClusterIP

⚠️ Warning: The fixes in this table can change cluster state. Use them only after performing the read-only diagnostics above, and only with explicit confirmation from the cluster owner or operator.

kubectl get svc kube-dns -n kube-system -o jsonpath='{.spec.clusterIP}'

Custom VNet DNS must forward .cluster.local to the CoreDNS ClusterIP and other lookups to 168.63.129.16.

Deep diagnostics with Inspektor Gadget (when the above checks are inconclusive):

Use scripts/run-ig.sh (or run-ig.ps1) with --pod <pod-name> --ns <ns> and trace_dns. Key signals: rcode=3 (NXDOMAIN), rcode=2 (SERVFAIL), high latency values, queries going to unexpected destinations.

See references/inspektor-gadget.md.


Detailed Networking Guides

Source: SKILL.md on GitHub

1 warning15d4 checks · Risk SAFE
  • Gen Agent Trust Hub15d

    This skill is designed for Azure diagnostics and troubleshooting, providing a comprehensive set of guides and scripts that utilize standard tools like the Azure CLI and kubectl. It includes some security considerations, such as the ingestion of logs which provides a surface for indirect prompt injection, and the use of privileged debug pods for advanced diagnostics. These are used within the skill's intended functionality and are accompanied by appropriate guidance for user approval.

  • Socket15d

    No alerts

  • Snyk15d

    Risk: LOW · No issues

  • Runlayer7mo

    4/4 files flagged

Signed by skilld at ae5e585. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
metadata
{
  "author": "Microsoft",
  "version": "1.2.6"
}

README badge

README badge for microsoft/skills/azure-diagnostics