Cluster Autoscaler: Debugging & Performance
Live Visibility Logs
- Asset:
assets/log-autoscaler-events.sh <cluster-name>(Live tail).
messageId Cheat Sheet
| ID | Meaning | Fix |
|---|---|---|
scale.up.error.out.of.resources |
GCE Stockout | Add zone/family fallback in ComputeClass. |
scale.up.error.quota.exceeded |
Project quota cap | Raise regional quota. |
scale.up.error.ip.space.exhausted |
Subnet full | Expand pod IP ranges. |
scale.up.no.scale.up |
No priority match | Check Pod requests vs ComputeClass bounds. |
Pending Pod Checklist
kubectl describe pod: Check events for "insufficient cpu" or "taints".- Hit
--max-nodes? Check pool limits. - Selector Conflict? Pod Pins
gke-spot=truewhile ComputeClass is On-Demand. - node pool auto-creation Enabled? Check
nodePoolAutoCreation.enabled: true. - Visibility Logs: Read
noDecisionStatus.noScaleUpfor exact rejection reason. - EKS to GKE Selector Translation: If migrating from EKS/Karpenter, ensure the user translates AWS-style or generic selectors (
machine-family) to GKE-native ones (cloud.google.com/machine-family). A common cause ofscale.up.no.scale.upis a Pod asking formachine-family: c3while GKE only recognizescloud.google.com/machine-family: c3. - Machine Series Support: If node pool auto-creation fails to provision nodes for a specific
machineFamilyorinstance-type(e.g., N4, C3A), verify the GKE version supports that series for node pool auto-creation / Autopilot. Old GKE versions will ignore unsupported series. Check GKE release notes or node pool auto-creation docs for version requirements. - Brand-new reservation? A reservation created in the last ~30 min may not be in Cluster Autoscaler's cache yet. Targeting it before the cache catches up makes Cluster Autoscaler back off that reservation and stall. Wait ≥30 min after creating the reservation before driving scale-up against it (see
ca-optimization.md).
Finding Scale-down Blockers
- Asset:
./assets/find-scale-down-blockers.sh(Scan cluster for blockers).
Common Causes
- Bare Pods: No controller (Deployment/Job); autoscaler won't evict.
- Local Storage:
emptyDiron local SSD orhostPath. - Annotation:
cluster-autoscaler.kubernetes.io/safe-to-evict: "false". - PDBs: Currently allowing zero disruptions.
- Floor:
min-nodesortotal-min-nodes> 0.
Performance & Sluggishness
- Required Anti-affinity: Explodes scheduler cost at scale. Use
preferredortopologySpreadConstraints. - Pool Count: Beyond ~200 pools, autoscaling slows down. Consolidate near-duplicate ComputeClasses.
- Spot Grace Period: Preemption notice is ~30s and is not extensible via ComputeClass fields. Keep
terminationGracePeriodSecondsand SIGTERM handling within it; rely on replicas and PDBs to absorb churn.
Segregating System Pods (Expert Pattern)
Symptom: kube-system pods (metrics-server, coredns) land on expensive nodes and pin them.
Fix: Segregate via namespace default ComputeClass.
- Apply a "cheap"
system-poolComputeClass. - Label
kube-systemnamespace:kubectl label ns kube-system cloud.google.com/default-compute-class-non-daemonset=system-pool