All skills
google avatar

/gke-upgrades

@becc4b8
by googlegoogle/skills21k stars
1,698

Plans, executes, and validates Google Kubernetes Engine (GKE) cluster upgrades and maintenance operations for both Standard and Autopilot clusters. Produces upgrade plans, pre/post-upgrade checklists, maintenance runbooks with gcloud commands, release channel strategy, and troubleshooting guides. Handles node pool upgrade strategies (surge, blue-green), version compatibility, PDB management, and workload-specific concerns (stateful, GPU, operators). Use this skill whenever the user mentions GKE upgrades, Kubernetes version bumps, node pool maintenance, GKE patching, cluster version management, release channel selection, maintenance windows, surge upgrades, stuck upgrades, or any GKE lifecycle management task — even casual mentions like "we need to upgrade our clusters" or "plan our next GKE maintenance" or "our upgrade is stuck." Don't use for GKE cluster creation, application onboarding, general networking/routing setup, or security policy configurations (use gke-basics or relevant GKE skills instead).

Use this Skill: https://skilld.dev/gh/google/skills/gke-upgrades

This session only. Nothing lands on disk.

referenceschecklists.md

≈853 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Checklist Templates

Adapt these to the user's environment. Fill in cluster names, versions, and remove items that don't apply.

Pre-Upgrade Checklist

Pre-Upgrade Checklist
- [ ] Cluster: ___ | Mode: Standard / Autopilot | Channel: ___
- [ ] Current version: ___ | Target version: ___

Compatibility
- [ ] Target version available in release channel (`gcloud container get-server-config --zone ZONE --format="yaml(channels)"`)
- [ ] No deprecated API usage (check GKE deprecation insights dashboard or check metrics: `kubectl get --raw /metrics | grep apiserver_request_total | grep deprecated`)
- [ ] GKE release notes reviewed for breaking changes between current → target
- [ ] Node version skew within 2 minor versions of control plane
- [ ] Rollout Sequencing configured and verified (if upgrading across environments)
- [ ] Third-party operators/controllers compatible with target version
- [ ] Admission webhooks tested against target version

Workload Readiness
- [ ] PDBs configured for critical workloads (not overly restrictive)
- [ ] No bare pods — all managed by controllers
- [ ] terminationGracePeriodSeconds adequate for graceful shutdown
- [ ] StatefulSet PV backups completed, reclaim policies verified
- [ ] Resource requests/limits set on all containers (mandatory for Autopilot)
- [ ] GPU driver compatibility confirmed with target node image (if applicable)
- [ ] Postgres/database operator compatibility verified (if applicable)

Infrastructure (Standard only)
- [ ] Node pool upgrade strategy chosen (surge / blue-green / autoscaled blue-green)
- [ ] Surge settings configured per pool: maxSurge=___ maxUnavailable=___
- [ ] Sufficient compute quota for surge nodes
- [ ] Maintenance window configured (off-peak hours)
- [ ] Maintenance exclusions set for freeze periods (if applicable)

Ops Readiness
- [ ] Monitoring and alerting active (Cloud Monitoring / Prometheus)
- [ ] Baseline metrics captured (error rates, latency, throughput)
- [ ] Upgrade window communicated to stakeholders
- [ ] Rollback plan documented
- [ ] On-call team aware and available

Post-Upgrade Checklist

Post-Upgrade Checklist

Cluster Health
- [ ] Control plane at target version: `gcloud container clusters describe CLUSTER --zone ZONE --format="value(currentMasterVersion)"`
- [ ] All node pools at target version: `gcloud container node-pools list --cluster CLUSTER --zone ZONE`
- [ ] All nodes Ready: `kubectl get nodes`
- [ ] System pods healthy: `kubectl get pods -n kube-system`
- [ ] No stuck PDBs: `kubectl get pdb --all-namespaces`

Workload Health
- [ ] All deployments at desired replica count: `kubectl get deployments -A`
- [ ] No CrashLoopBackOff or Pending pods: `kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded`
- [ ] StatefulSets fully ready: `kubectl get statefulsets -A`
- [ ] Ingress/load balancers responding
- [ ] Application health checks and smoke tests passing

Observability
- [ ] Metrics pipeline active, no collection gaps
- [ ] Logs flowing to aggregation
- [ ] Error rates within pre-upgrade baseline
- [ ] Latency (p50/p95/p99) within pre-upgrade baseline

Cleanup
- [ ] Old node pools removed (if blue-green)
- [ ] Surge quota released (automatic for surge upgrades)
- [ ] Upgrade documented in changelog
- [ ] Lessons learned captured

Source: SKILL.md on GitHub

No alerts1d3 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    This skill provides comprehensive guidance for GKE upgrades and maintenance, generating specific gcloud and kubectl commands tailored to a user's environment. Security considerations include the generation of infrastructure management commands and the processing of user-provided cluster details, which are appropriate for its intended administrative function.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: LOW · No issues

Signed by skilld at becc4b8. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 weeks ago
metadata
{
  "version": "1.0.0",
  "category": "Containers"
}

README badge

README badge for google/skills/gke-upgrades