All skills
microsoft avatar

/azure-diagnostics

@ae5e585
by microsoftmicrosoft/skills3.1k stars
351

Debug Azure production issues on Azure using AppLens, Azure Monitor, resource health, and safe triage. WHEN: debug production issues, troubleshoot app service, app service high CPU, app service deployment failure, troubleshoot container apps, troubleshoot functions, troubleshoot AKS, VM RDP, Linux SSH, VM black screen, can't connect to VM, reset VM password, NSG or firewall blocking, kubectl cannot connect, kube-system/CoreDNS failures, pod pending, crashloop, node not ready, upgrade failures, analyze logs, KQL, insights, image pull failures, cold start issues, health probe failures, resource health, root cause of errors, troubleshoot event hubs, troubleshoot service bus, messaging SDK error, AMQP connection failure, message lock lost, service bus dead letter.

Use this Skill: https://skilld.dev/gh/microsoft/skills/azure-diagnostics

This session only. Nothing lands on disk.

troubleshootingaksgeneral-diagnostics.md

≈514 tokens on demand. Your agent reads this file only when SKILL.md points to it.

General AKS Investigation & Diagnostics

"What happened in my cluster?"

When a user asks a broad question like "what happened in my AKS cluster?" or "check my AKS status", follow this systematic flow:

  1. Cluster health
  2. Recent events
  3. Node status
  4. Unhealthy pods
  5. All pods overview
  6. System pods health
  7. Activity log

Run the aks-baseline script instead of issuing these commands one by one. It performs the entire read-only sweep above and prints a single labeled digest (provisioning state, node pool summary, recent activity log, node readiness, unhealthy pods, kube-system health, and recent warning events), so you get one summarized result instead of seven raw dumps.

# bash
./scripts/aks-baseline.sh -g <rg> -n <cluster> [--namespace <ns>]
# PowerShell
.\scripts\aks-baseline.ps1 -ResourceGroup <rg> -Cluster <cluster> [-Namespace <ns>]

After reviewing the digest, deep-dive into a specific pod with kubectl describe / kubectl logs.


AKS CLI Tools

# Get cluster credentials (required before kubectl commands)
az aks get-credentials -g <rg> -n <cluster>

# View node pools
az aks nodepool list -g <rg> --cluster-name <cluster> -o table

AppLens (MCP) for AKS

For AI-powered diagnostics:

mcp_azure_mcp_applens
  intent: "diagnose AKS cluster issues"
  command: "diagnose"
  parameters:
    resourceId: "/subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.ContainerService/managedClusters/<cluster>"

💡 Tip: AppLens automatically detects common issues and provides remediation recommendations using the cluster resource ID.


Best Practices

  1. Start with kubectl get/describe - Always check basic status first
  2. Check events - kubectl get events -A reveals recent issues
  3. Use systematic isolation - Pod -> Node -> Cluster -> Network
  4. Document changes - Note what you tried and what worked
  5. Escalate when needed - For control plane issues, contact Azure support

Source: SKILL.md on GitHub

1 warning15d4 checks · Risk SAFE
  • Gen Agent Trust Hub15d

    This skill is designed for Azure diagnostics and troubleshooting, providing a comprehensive set of guides and scripts that utilize standard tools like the Azure CLI and kubectl. It includes some security considerations, such as the ingestion of logs which provides a surface for indirect prompt injection, and the use of privileged debug pods for advanced diagnostics. These are used within the skill's intended functionality and are accompanied by appropriate guidance for user approval.

  • Socket15d

    No alerts

  • Snyk15d

    Risk: LOW · No issues

  • Runlayer7mo

    4/4 files flagged

Signed by skilld at ae5e585. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
metadata
{
  "author": "Microsoft",
  "version": "1.2.6"
}

README badge

README badge for microsoft/skills/azure-diagnostics