General AKS Investigation & Diagnostics
"What happened in my cluster?"
When a user asks a broad question like "what happened in my AKS cluster?" or "check my AKS status", follow this systematic flow:
- Cluster health
- Recent events
- Node status
- Unhealthy pods
- All pods overview
- System pods health
- Activity log
Run the aks-baseline script instead of issuing these commands one by one. It performs the entire read-only sweep above and prints a single labeled digest (provisioning state, node pool summary, recent activity log, node readiness, unhealthy pods, kube-system health, and recent warning events), so you get one summarized result instead of seven raw dumps.
# bash
./scripts/aks-baseline.sh -g <rg> -n <cluster> [--namespace <ns>]# PowerShell
.\scripts\aks-baseline.ps1 -ResourceGroup <rg> -Cluster <cluster> [-Namespace <ns>]After reviewing the digest, deep-dive into a specific pod with kubectl describe / kubectl logs.
AKS CLI Tools
# Get cluster credentials (required before kubectl commands)
az aks get-credentials -g <rg> -n <cluster>
# View node pools
az aks nodepool list -g <rg> --cluster-name <cluster> -o tableAppLens (MCP) for AKS
For AI-powered diagnostics:
mcp_azure_mcp_applens
intent: "diagnose AKS cluster issues"
command: "diagnose"
parameters:
resourceId: "/subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.ContainerService/managedClusters/<cluster>"💡 Tip: AppLens automatically detects common issues and provides remediation recommendations using the cluster resource ID.
Best Practices
- Start with kubectl get/describe - Always check basic status first
- Check events -
kubectl get events -Areveals recent issues - Use systematic isolation - Pod -> Node -> Cluster -> Network
- Document changes - Note what you tried and what worked
- Escalate when needed - For control plane issues, contact Azure support