All skills
lukemurraynz avatar

/aks-cluster-architecture

@2cc2455

AKS cluster architecture decisions for new Azure Kubernetes Service projects: AKS Automatic vs Standard, networking topology, dual-stack (IPv4/IPv6), Kubernetes version and OS currency, node pool strategy, identity, production NetworkPolicy, namespaces, autoscaling, ingress and Gateway API, observability, operations, resilience, GPU and AI workloads, GPU partitioning (MIG, time-slicing, MPS), batch scheduling (Kueue), AKS on bare metal, AI Runway and KAITO model serving, AKS MCP server access, kars (Agent Reference Stack for Kubernetes) for agent isolation, Kata MicroVM pod sandboxing, Azure Kubernetes Fleet Manager, multi-cluster governance, update orchestration, resource placement, cross-cluster networking, and cost. WHEN: designing new AKS clusters, reviewing production readiness, choosing CNI or outbound connectivity, planning node pools, defining namespace, network and security controls, evaluating Fleet Manager, deploying AI agent runtimes on AKS, or making hard-to-reverse infrastructure decisions.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/aks-cluster-architecture

This session only. Nothing lands on disk.

bundlesoperations-resilienceguide.md

≈31k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Operations Resilience Bundle

Cluster-scope operations: observability, upgrade strategy, GitOps repo layout, IaC pipeline, backup/DR, traffic management controller selection, CVE response, incident response.

Load this bundle for ingress/Gateway API, App Routing, TLS ownership, observability, upgrades, GitOps, backup/restore, multi-region strategy, cost optimisation, known pitfalls, and Drasi-on-AKS considerations. For production namespace, NetworkPolicy, replica, PDB, and autoscaling manifests, also load production-workload-controls. For multiple clusters, fleet governance, staged multi-cluster updates, workload placement, managed fleet namespaces, or cross-cluster networking, also load fleet-management.

<!-- toc --> <!-- /toc -->

Assume new AKS projects. Prefer supported Azure-native operations patterns and avoid new long-term dependencies on deprecated or near-EOL ingress/logging approaches.

Traffic Management

Need Preferred direction Notes
Simple supported ingress now AKS application routing add-on - managed NGINX is a supported transitional bridge with a two-stage EOL cliff (OSS maintenance ended March 2026; AKS managed-NGINX security-patch-only through November 2026). AGC (Application Gateway for Containers) or App Routing with Gateway API is the strategic target — name it explicitly as the target architecture and document any NGINX bridge step with its migration path and EOL deadline. Transitional only; do not start new long-term ingress designs on NGINX.
Long-term Gateway API direction Application routing Gateway API (GA) or Application Gateway for Containers Verify TLS/DNS limitations, controller capabilities, and coexistence with service mesh add-on.
Enterprise L7 ingress with WAF/private frontend Application Gateway for Containers or Application Gateway pattern Confirm private frontend and AKS Automatic support before choosing.
Service mesh AKS Istio add-on only when mesh capabilities are required Do not adopt mesh just for ingress.
Internal-only workloads Internal ingress/Gateway, private DNS, no public load balancer Validate client network paths and DNS.

Guidance:

  • Retirement clifftops driving migration urgency (verify against current AKS release notes before acting): upstream ingress-nginx reaches end-of-life in March 2026; the AKS managed-NGINX add-on (legacy nginxIngress configuration) reaches end of support in November 2026. Treat App Routing on Gateway API or Application Gateway for Containers (AGC) as the migration target - both are GA. Do not start new long-term ingress designs on the legacy NGINX path.
  • Do not recommend upstream/self-managed ingress-nginx as the default for new long-term production designs.
  • When designing for a clean-slate cluster, name AGC explicitly as the target ingress and treat managed NGINX as a transitional choice with a documented migration path.
  • Treat Gateway API implementation choice as difficult; CRDs, controllers, TLS ownership, DNS, and annotations differ.
  • Check whether the selected controller owns TLS, DNS, certificate rotation, WAF, and private frontend behaviour.
  • For app routing Gateway API, verify current feature flag requirements and limitations before production use. Gateway API-based ingress is GA (April 2026), but TLS/DNS limitations, controller coexistence, and migration compatibility should be checked per deployment.
  • Do not recommend application routing Gateway API for production unless its support state, TLS/DNS limitations, coexistence with the Istio service mesh add-on, migration path, rollback plan, and controller ownership are documented.
  • Do not assume Gateway API, Application Gateway for Containers, app-routing managed NGINX, and service mesh are interchangeable. Their CRDs, annotations, TLS ownership, WAF/private frontend capability, and operational responsibilities differ.

AGC WAF (WebApplicationFirewallPolicy, OWASP DRS)

Application Gateway for Containers exposes the Azure Web Application Firewall through the WebApplicationFirewallPolicy CRD (alb.networking.azure.io/v1). A WebApplicationFirewallPolicy does not define WAF rules inline. It binds a pre-provisioned Azure WAF policy (created in ARM via CLI/Bicep/Portal, where you select the managed rule set and mode) to a Gateway, a specific listener (sectionNames), or an HTTPRoute, so that AGC inspects north-south ingress traffic for payload-level attacks (SQLi, XSS, protocol anomalies) before it reaches cluster pods.

The ARM WAF policy is a one-to-one mapping to an AGC security policy resource (currently the only security-policy type is waf). Create the WAF policy in Azure first, then reference it by full resource ID from Kubernetes.

# Kubernetes manifest — binds an existing Azure WAF policy to an HTTPRoute.
# Verify the CRD kind and field names, the ARM resource ID, and the DRS version
# configured on the referenced Azure WAF policy against your deployment. [VERIFY]
apiVersion: alb.networking.azure.io/v1
kind: WebApplicationFirewallPolicy
metadata:
  name: payments-waf
  namespace: payments-prod
spec:
  targetRef:
    group: gateway.networking.k8s.io
    kind: HTTPRoute          # Gateway scopes to all listeners; HTTPRoute narrows blast radius
    name: payments-route
    namespace: payments-prod
    # sectionNames: ["https"]  # optional: target a specific listener on a Gateway
  webApplicationFirewall:
    id: /subscriptions/00000000-0000-0000-0000-000000000000/resourceGroups/payments-rg/providers/Microsoft.Network/applicationGatewayWebApplicationFirewallPolicies/payments-waf-azure

The referenced Azure WAF policy (created separately, e.g. az network application-gateway waf-policy create) holds the managed rule set. As of 2026 the recommended set is DRS 2.2 (based on OWASP CRS 3.3.4); DRS 2.1 is the previous version. Only the latest three DRS releases are supported for new policies. [VERIFY DRS version against your region/SKU before pinning.]

Guidance:

  • Start the referenced Azure WAF policy in Detection mode, observe Hubble + AGC diagnostic logs for false positives, then promote to Prevention with an exclusion list if needed.
  • Prefer the narrowest scope that covers the exposure: HTTPRoute per route over Gateway (all listeners).
  • The WAF is complementary to, not a replacement for, pod-level Cilium L7 policy and default-deny: the WAF sees request payloads on permitted routes; Cilium L7 sees method/path and east-west traffic the WAF cannot. See production-workload-controls - Layered AKS ingress and east-west security.
  • Do not attempt to define WAF rules inline in the Kubernetes manifest. The rule set, mode, exclusions, and custom rules live on the Azure WAF policy; the K8s CRD only binds it by resource ID.

Ingress / CNI Migration Rollback

Ingress controller and CNI changes are Difficult-to-Permanent. Plan a rollback gate before you start the cutover.

Ingress controller swap (e.g., managed NGINX → AGC, NGINX → Istio gateway):

  • Run the new controller in parallel with the existing one, on a separate GatewayClass / IngressClass so old traffic continues to flow.
  • Cut over per-service by changing the IngressClass annotation or the Gateway parent; never flip all services at once.
  • Hold the old controller in place for a defined soak window (recommended ≥ 4 hours of business traffic) with rollback criteria: 4xx/5xx rate, p95/p99 latency, TLS handshake errors, DNS resolution failures.
  • Keep the old DNS record and TLS certificate binding until the soak window passes; failover means flipping DNS or the Gateway back to the old controller.
  • Decommission only after the rollback criteria are clean.

CNI / dataplane change (e.g., kubenet → Azure CNI Overlay, Azure CNI → Azure CNI Powered by Cilium):

  • Treat as Permanent in practice for an existing cluster - most production teams cut over by building a parallel cluster with the new CNI, draining workloads via DNS/traffic shift, and decommissioning the old cluster.
  • In-place CNI swap is supported in some narrow scenarios; verify support against current Microsoft Learn guidance and accept that NetworkPolicy semantics, pod IP behaviour, and observability tooling may change.
  • Document rollback as "fail back to the old cluster" rather than "revert CNI on the same cluster."

Validation gates (apply to both):

  • Run synthetic and real-user health checks during and after the cutover.
  • Watch NetworkPolicy enforcement: a CNI change can silently change which connections were previously allowed.
  • Confirm Hubble / ACNS / Container Insights observability is wired before cutover, not after - diagnosing a broken cutover without flow visibility is significantly harder.

Gateway API / App Routing Production Gate

Before selecting a Gateway or ingress implementation, answer these checks explicitly:

  • Is the feature GA or preview in the target region and cluster mode?
  • Does the controller support the required TLS termination, certificate rotation, DNS automation, private frontend, WAF, and internal/external exposure model?
  • Can it coexist with any required Istio service mesh add-on, or must one be disabled first?
  • What CRDs, annotations, GatewayClasses, and route types will become migration dependencies?
  • How will rollback work if the controller, GatewayClass, or TLS ownership model changes?
  • Who owns certificate renewal, DNS records, gateway upgrades, and incident response?

If these answers are unknown, recommend a conservative supported ingress pattern and record Gateway API as an ADR follow-up rather than treating it as a default.

Cert-Manager and Let's Encrypt

This section covers production-grade cert-manager setup on AKS with Let's Encrypt for automated TLS certificate issuance and renewal. For the broader TLS certificate planning and CA strategy selection, see Secrets, etcd Encryption, and Data-at-Rest in the cluster-foundations bundle.

Installation

cert-manager can be installed via Helm or kubectl. Helm is the recommended approach for production because it provides fine-grained control over installation options, namespace scoping, and upgrade management.

# Add the Jetstack Helm repository
helm repo add jetstack https://charts.jetstack.io
helm repo update

# Install cert-manager CRDs
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.16.3/cert-manager.crds.yaml

# Create namespace and install cert-manager
kubectl create namespace cert-manager

helm upgrade --install cert-manager jetstack/cert-manager \
  --namespace cert-manager \
  --version v1.16.3 \
  --set installCRDs=false \
  --set global.leaderElection.namespace=cert-manager \
  --set webhook.timeoutSeconds=30

[VERIFY] cert-manager version: confirm current stable release and Kubernetes version compatibility against the cert-manager supported versions matrix before choosing a version. The v1.16.x series is used as an illustrative example.

Key Helm values for AKS:

Value Recommendation Reason
installCRDs false Manage CRDs separately so GitOps controls CRD versions explicitly.
webhook.timeoutSeconds 30 Avoid webhook timeouts in slower or private clusters.
global.leaderElection.namespace cert-manager Clean leader-election lock namespace.

Azure Arc extension (public preview). As of April/May 2026, Microsoft offers cert-manager as an Azure Arc Kubernetes extension for Arc-connected clusters. This provides a Microsoft-supported packaging of cert-manager and trust-manager. The extension is primarily documented for Arc-enabled clusters (on-premises, edge, other clouds); for AKS-native clusters, Helm remains the standard installation path. [VERIFY] current Arc extension GA/preview status and AKS support before choosing this path.

Core Resources

cert-manager defines three primary CRDs:

Resource Scope Purpose
Issuer Namespace Represents a certificate authority for a single namespace. Use for tenant isolation where each namespace configures its own CA.
ClusterIssuer Cluster Represents a certificate authority for the entire cluster. Use for shared infrastructure issuers (e.g., Let's Encrypt).
Certificate Namespace Requests a certificate from an Issuer or ClusterIssuer and stores the resulting private key + certificate in a Kubernetes Secret.

Rule of thumb: Use a ClusterIssuer for shared Let's Encrypt issuers used by all namespaces. Use namespace-scoped Issuer only when different teams need different CAs or ACME accounts.

Let's Encrypt Issuer Configuration

Let's Encrypt offers two ACME environments:

Environment Endpoint Use
Staging https://acme-staging-v02.api.letsencrypt.org/directory Test and development. Has relaxed rate limits.
Production https://acme-v02.api.letsencrypt.org/directory Production certificates. Enforces rate limits (see Known Pitfalls).

Always test against staging first before switching to production. A misconfigured production issuer can exhaust Let's Encrypt rate limits and block certificate issuance for a week.

Cert-manager proves domain ownership via an ACME challenge. Two challenge types are supported:

HTTP-01 Challenge

The HTTP-01 challenge proves domain ownership by serving a token on port 80 at http://<domain>/.well-known/acme-challenge/<token>. cert-manager creates an Ingress or HTTPRoute resource with the required path.

Prerequisites:

  • The domain must resolve to a public IP that routes to the cluster.
  • An ingress controller must be installed and serving HTTP traffic on port 80.
  • The ingress controller must be reachable from the internet on port 80 for the challenge domain.

Limitations:

  • Does not support wildcard certificates (*.example.com).
  • Requires the domain to be publicly resolvable, not suitable for internal-only services.
  • The ingress controller must be fully operational before certificates can be issued (chicken-and-egg: a new cluster needs the ingress controller before it can prove domain ownership).

ClusterIssuer example (HTTP-01):

apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-staging
spec:
  acme:
    server: https://acme-staging-v02.api.letsencrypt.org/directory
    email: admin@example.com
    privateKeySecretRef:
      name: letsencrypt-staging-account-key
    solvers:
    - http01:
        ingress:
          class: webapprouting.kubernetes.azure.com  # AKS App Routing ingress class, or "nginx" for NGINX
apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-prod
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: admin@example.com
    privateKeySecretRef:
      name: letsencrypt-prod-account-key
    solvers:
    - http01:
        ingress:
          class: webapprouting.kubernetes.azure.com

The ingress.class field must match the ingress class of the installed ingress controller. For AKS App Routing managed NGINX this is webapprouting.kubernetes.azure.com; for self-managed NGINX it is typically nginx.

Application Gateway for Containers (AGC) HTTP-01 example:

When using AGC with the Gateway API, cert-manager creates an HTTPRoute for the challenge. The ClusterIssuer uses an http01 solver with the gatewayHTTPRoute option instead of the ingress option. See the Microsoft Learn guide for the current configuration shape: Cert-manager and Let's Encrypt with Application Gateway for Containers. [VERIFY] current AGC Gateway API annotation requirements and cert-manager compatibility.

DNS-01 Challenge

The DNS-01 challenge proves domain ownership by creating a TXT record in the DNS zone. cert-manager creates the record, the CA queries DNS to verify, and cert-manager removes the record after validation.

Advantages:

  • Supports wildcard certificates (*.example.com).
  • Does not require the ingress controller to be operational first, certificates can be issued before any HTTP routing is configured.
  • Works for internal-only domains where the certificate is used for internal ingress or mTLS.

Prerequisites:

  • The DNS zone must be managed by Azure DNS (public zone).
  • The DNS zone must be publicly resolvable. Azure Private DNS Zones cannot be used for ACME challenges because Let's Encrypt must query the public DNS.
  • The AKS cluster must have OIDC issuer and Workload Identity enabled.
  • A managed identity with DNS Zone Contributor on the Azure DNS zone.
  • A federated identity credential linking the cert-manager service account to the managed identity.

Azure DNS / Workload Identity setup:

# 1. Enable OIDC issuer and Workload Identity on the AKS cluster (if not already enabled)
az aks update \
  --name <cluster-name> \
  --resource-group <cluster-rg> \
  --enable-oidc-issuer \
  --enable-workload-identity

# 2. Get the OIDC issuer URL
az aks show \
  --name <cluster-name> \
  --resource-group <cluster-rg> \
  --query "oidcIssuerProfile.issuerUrl" \
  --output tsv

# 3. Create a managed identity for cert-manager
az identity create \
  --name <cert-manager-identity-name> \
  --resource-group <identity-rg>

# 4. Assign DNS Zone Contributor role on the Azure DNS zone
az role assignment create \
  --assignee-object-id $(az identity show --name <cert-manager-identity-name> --resource-group <identity-rg> --query principalId --output tsv) \
  --role "DNS Zone Contributor" \
  --scope /subscriptions/<subscription-id>/resourceGroups/<dns-rg>/providers/Microsoft.Network/dnsZones/<zone-name>

# 5. Create a federated identity credential linking the cert-manager ServiceAccount to the managed identity
az identity federated-credential create \
  --name cert-manager-federated-credential \
  --identity-name <cert-manager-identity-name> \
  --resource-group <identity-rg> \
  --issuer <oidc-issuer-url> \
  --subject system:serviceaccount:cert-manager:cert-manager \
  --audience api://AzureADTokenExchange

[VERIFY] Update the --subject value to match the actual ServiceAccount name if cert-manager is installed with a custom service account. The default Helm deployment uses cert-manager as the ServiceAccount name in the cert-manager namespace.

ClusterIssuer example (DNS-01 with Azure DNS):

apiVersion: cert-manager.io/v1
kind: ClusterIssuer
metadata:
  name: letsencrypt-dns
spec:
  acme:
    server: https://acme-v02.api.letsencrypt.org/directory
    email: admin@example.com
    privateKeySecretRef:
      name: letsencrypt-dns-account-key
    solvers:
    - dns01:
        azureDNS:
          subscriptionID: <subscription-id>
          resourceGroupName: <dns-rg>
          hostedZoneName: <zone-name>
          # Use managed identity for authentication — no clientSecret
          managedIdentity:
            clientID: <managed-identity-client-id>

Important: Do not use clientSecret authentication for production cert-manager Azure DNS integration. Always use Workload Identity with a managed identity. Storing Azure service principal secrets in Kubernetes Secrets defeats the purpose of automated TLS management and creates a secret-rotation problem.

Choosing between HTTP-01 and DNS-01:

Factor HTTP-01 DNS-01
Wildcard certificates No Yes
Requires public ingress Yes No
Cluster bootstrap order Ingress before certs Certs before ingress
DNS provider dependency No Yes (Azure DNS RBAC)
Setup complexity Low Medium
Private/internal domains No (must be public) No (must be public DNS)

Certificate Resource

Once a ClusterIssuer is configured, request certificates by creating Certificate resources.

Standard certificate example (used with HTTP-01):

apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: app-tls
  namespace: myapp
spec:
  secretName: app-tls-cert
  issuerRef:
    name: letsencrypt-prod
    kind: ClusterIssuer
  commonName: app.example.com
  dnsNames:
  - app.example.com

Wildcard certificate example (requires DNS-01 issuer):

apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: wildcard-tls
  namespace: istio-system  # Namespace of the ingress gateway
spec:
  secretName: wildcard-tls-cert
  issuerRef:
    name: letsencrypt-dns
    kind: ClusterIssuer
  commonName: "*.example.com"
  dnsNames:
  - "*.example.com"

Ingress annotation-based certificate (ingress-shim):

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: myapp-ingress
  annotations:
    cert-manager.io/cluster-issuer: letsencrypt-prod
spec:
  ingressClassName: webapprouting.kubernetes.azure.com
  tls:
  - hosts:
    - app.example.com
    secretName: app-tls-cert  # cert-manager creates this secret
  rules:
  - host: app.example.com
    http:
      paths:
      - path: /
        pathType: Prefix
        backend:
          service:
            name: myapp-service
            port:
              number: 80

The cert-manager.io/cluster-issuer annotation tells cert-manager to automatically create a Certificate resource for the TLS section of the Ingress. This is the most common pattern and avoids managing Certificate resources separately.

Ingress Controller TLS Integration

Controller TLS integration approach Notes
AKS App Routing NGINX Ingress annotation cert-manager.io/cluster-issuer cert-manager creates the secret; App Routing auto-detects TLS from the Ingress spec.
Self-managed NGINX Ingress annotation cert-manager.io/cluster-issuer Same pattern; requires manual NGINX management.
Application Gateway for Containers Gateway API Certificate resource referencing cert-manager-managed secret AGC supports Key Vault certs via secretProviderClass as an alternative to cert-manager; [VERIFY] current AGC + cert-manager integration guidance.
Istio ingress gateway Certificate resource in the istio-system namespace; gateway references the TLS secret Wildcard certs stored as kubernetes.io/tls secrets referenced by Istio Gateway servers.tls.credentialName.

Renewal and Monitoring

cert-manager automatically renews certificates before expiry. By default it attempts renewal when 2/3 of the certificate lifetime has elapsed (60 days into a 90-day Let's Encrypt certificate).

Monitor certificate status:

# Check certificate ready status
kubectl get certificate -A

# Inspect certificate conditions (Ready, Issuing, etc.)
kubectl describe certificate <name> -n <namespace>

# Check ACME order and challenge status
kubectl get orders -A
kubectl get challenges -A
kubectl describe order <name> -n <namespace>

Prometheus metrics. cert-manager exposes metrics on port 9402 by default. Key metrics for alerting:

Metric Type What it signals
certmanager_certificate_expiry_seconds Gauge Time until certificate expiry. Alert when approaching renewal deadline.
certmanager_certificate_ready_status Gauge 1 = Ready, 0 = Not Ready. Alert on 0 for any production certificate.
certmanager_clock_time_seconds Gauge cert-manager's current wall-clock time. Useful for detecting clock skew issues.
certmanager_acme_client_request_count Counter Rate of ACME API requests. Spikes can indicate misconfiguration.

Recommended Prometheus alerts:

# Alert when a certificate will expire within 14 days
- alert: CertificateExpiresSoon
  expr: certmanager_certificate_expiry_seconds < 1209600
  for: 1h
  labels:
    severity: warning

# Alert when a certificate is not ready
- alert: CertificateNotReady
  expr: certmanager_certificate_ready_status{ready=="false"} == 1
  for: 5m
  labels:
    severity: critical

Troubleshooting

Symptom Likely cause Diagnosis Mitigation
ACME order stuck or failing (DNS-01) Azure DNS integration not working kubectl describe order <name>; check challenge status. Verify managed identity has DNS Zone Contributor on the zone. Re-run federated identity credential setup. Verify the clientID in the ClusterIssuer matches the managed identity.
ACME order stuck or failing (HTTP-01) Ingress controller not reachable on port 80 Check the cluster's public ingress is functioning. Try curl http://<domain>/.well-known/acme-challenge/ — should return a 404. Confirm ingress controller is healthy and public-facing. Check NSGs and firewall rules.
Webhook errors creating resources cert-manager webhook not reachable (common in private clusters) kubectl get pods -n cert-manager — verify webhook pod is running. Ensure cert-manager webhook service is reachable from the API server subnet. Check VNet peering, NSGs, and egress firewall rules.
Certificate stuck in IssuancePending Rate limited, DNS propagation delay, or CA unreachable kubectl describe challenge; check state and reason fields. Wait for rate-limit window to reset. Use staging issuer first to confirm configuration.
Let's Encrypt rate limit hit 50 certificates/week/domain. Each SAN counts separately. Check logs: kubectl logs -n cert-manager -l app.kubernetes.io/name=cert-manager Use staging issuer for testing. Consolidate SANs into fewer certificates.
Certificate secret missing after renewal cert-manager renewed but old secret not updated, or ingress controller caching Verify kubectl get certificate shows Ready=True. Check secret tls.crt timestamp. Restart ingress controller pods if caching is suspected.

Egress Requirements for Private Clusters

cert-manager requires outbound HTTPS access to the following endpoints:

Endpoint Port Purpose
acme-v02.api.letsencrypt.org 443 Production ACME API
acme-staging-v02.api.letsencrypt.org 443 Staging ACME API

For clusters using forced-tunnel egress through Azure Firewall or a NVA, ensure these FQDNs are allowed in the firewall policy or the AKS egress allowedFqdnRules.

Bootstrap Ordering

cert-manager follows the standard add-on bootstrap sequence defined in Bootstrap Ordering:

  1. Layer 2 (CRDs): cert-manager CRDs (certificaterequests.cert-manager.io, certificates.cert-manager.io, challenges.acme.cert-manager.io, clusterissuers.cert-manager.io, issuers.cert-manager.io, orders.acme.cert-manager.io) must be installed first.
  2. Layer 4 (Platform add-ons): cert-manager Helm release, followed by ClusterIssuer resources, then Certificate resources.
  3. Layer 6 (Workload dependencies): Ingress/Gateway resources that reference TLS secrets created by cert-manager.

Chicken-and-egg: HTTP-01 with first-time cluster bootstrap. If using HTTP-01 challenges, the ingress controller must be operational before cert-manager can issue the first certificate. Mitigations:

Approach How it works Trade-off
DNS-01 first Issue the first certificate via DNS-01 (doesn't need ingress) Requires Azure DNS RBAC from day one
Self-signed bootstrap Use a self-signed or internal CA certificate for initial ingress setup, then replace with Let's Encrypt Transient insecure cert; extra step to swap
Staging issuer Use Let's Encrypt staging (same challenge flow, no rate-limit pressure) to validate the flow first Staging certs are untrusted

Do not use HTTP-01 on a fresh cluster that has no functional ingress controller. Prefer DNS-01 for initial certificate issuance, or provision the ingress controller HTTP endpoint with a self-signed cert as a bootstrap step.

Once the ingress controller exists, two more first-time-bootstrap traps commonly stack with each other and produce the identical symptom (self-check context deadline exceeded, zero bytes); diagnose both, do not stop at the first plausible cause: App Routing's default externalTrafficPolicy: Local breaking in-cluster hairpin, and a default-deny-ingress NetworkPolicy with no rule matching cert-manager's dynamically-named solver pod. See the Known Pitfalls table rows for both, verified live together on a single first-time AKS App Routing + cert-manager bootstrap (2026-08-23).

Observability Baseline

Enable observability from day one:

  • Azure Monitor Container Insights with ContainerLogV2.
  • Managed Prometheus for Kubernetes and workload metrics.
  • Control plane metrics via Managed Prometheus (GA August 2026). API server/etcd/scheduler/controller-manager/autoscaler telemetry collected by a component outside the ama-metrics add-on: enable with --enable-control-plane-metrics together with --enable-azure-monitor-metrics - enabling the Prometheus add-on alone does not collect control-plane metrics. Requires managed-identity authentication; Private Link scenarios are not supported; self-hosted Prometheus cannot scrape control-plane metrics (managed service only). Default scrape targets are apiserver and etcd; kube-scheduler, kube-controller-manager, cluster-autoscaler, and node-auto-provisioning stay off until enabled in the ama-metrics-settings-configmap (controlplane-metrics: section, schema v2). Verify collection with az aks show --query azureMonitorProfile.metrics plus actual series in the workspace, not add-on state alone.
  • Azure Managed Grafana or an approved equivalent dashboarding path.
  • Platform alerts for node readiness, pod restarts, pending pods, OOMKilled, high throttling, SNAT/egress errors, and API server errors.
  • Application SLOs with RED or four-golden-signals metrics.
  • Correlation IDs and structured logs for app workloads.
  • GPU metrics and health monitoring when GPU pools exist.
  • Network visibility with ACNS/Cilium/Hubble-equivalent capabilities where supported and justified.
  • NetworkPolicy validation evidence for production namespaces, including default deny and explicit allow paths.
  • Cluster Health Monitor (Preview) — AKS-managed add-on (--enable-continuous-control-plane-and-addon-monitor) that runs in-cluster health checks against DNS resolution (CoreDNS, LocalDNS), API server connectivity (synthetic create/get/delete), and metrics-server availability. Exposes Prometheus metrics on port 9800 with no Managed Prometheus dependency. Provides CoreDNS auto-remediation (deletes unhealthy pod after 5 min with guardrails). Lightweight complement to Managed Prometheus and Container Insights — not a replacement. Requires aks-preview extension and Azure CLI 2.73.0+.
  • Node Problem Detector (NPD). An open-source Kubernetes add-on (DaemonSet) that runs on every node and surfaces kernel-level node faults (OOM kills, NFS mount failures, disk pressure, network unreachable) as NodeCondition objects and Kubernetes Event entries before they surface as pod failures. NPD is not included in AKS by default but is available via the Helm chart (node-problem-detector/node-problem-detector). For production clusters where silent node degradation is a concern, NPD provides earlier detection than relying on pod restart loops to expose the symptom. Pair with an alert on custom NodeConditions in Managed Prometheus.
  • Inspektor Gadget (AKS extension, Preview). eBPF-based kernel-level observability (DNS queries, TCP connections, process exec, file opens) enriched with pod/namespace metadata: fills the gap between logs/metrics and what actually happened at the kernel. Now installable as a managed AKS cluster extension (az k8s-extension create --extension-type microsoft.inspektorgadget --release-train preview; images from MCR, minor-version auto-upgrades) rather than a self-managed Helm chart; run gadgets via the krew plugin (kubectl krew install gadget). Gotchas: extension config changes are not picked up live; restart the gadget DaemonSet (kubectl rollout restart daemonset/gadget -n gadget) after any az k8s-extension update, and the update command is additive (failed settings are not auto-reverted). Best on non-production clusters initially; useful complement to ACNS/Hubble for deep plumbing debugging. [VERIFY] the gadget catalog and extension GA status before standardising on it.
  • Falco (open-source runtime security). Falco is a CNCF graduated project that uses eBPF (or kernel module) to detect unexpected syscalls and network activity at runtime. Defender for Containers provides a managed equivalent, but Falco is a common complementary layer for teams that need vendor-neutral, customisable detection rules or that operate in environments where Defender is not available. If adopting: install via the falcosecurity/falco Helm chart, route alerts to the SOC via Falcosidekick → Azure Event Hub → Sentinel, and treat the Falco DaemonSet as a system workload (system node pool toleration, PriorityClass). Falco is not AKS-supported software ; plan community support and upgrade ownership separately.

Do not mark a deployment complete on successful IaC alone. Require runtime evidence.

Runtime validation examples:

kubectl get nodes -o wide
kubectl get pods -A --field-selector=status.phase!=Running
kubectl top nodes
kubectl top pods -A
kubectl get events -A --sort-by=.metadata.creationTimestamp | tail -50
kubectl get networkpolicy -A
kubectl get hpa,pdb -A
az aks show --resource-group <rg> --name <cluster> --query "addonProfiles" --output yaml

Audit policy and SIEM forwarding

Kubernetes audit logs are the primary detective control for cluster compromise. Tune audit policy levels for signal vs. ingest cost, and forward to a SIEM with immutable retention.

Audit policy levels (per-rule, set via the cluster audit policy):

  • RequestResponse (full body) - Secret reads/writes, pods/exec, pods/attach, RBAC mutations (Role, RoleBinding, ClusterRole, ClusterRoleBinding), and privileged-pod create.
  • Metadata - everywhere else by default.
  • None - high-volume noisy verbs (get/list/watch on events, leases, endpointslices, tokenreviews, subjectaccessreviews) to keep ingest cost bounded.

Forwarding path:

  • AKS diagnostic settings → Log Analytics workspace (kube-audit, kube-audit-admin categories).
  • Log Analytics → Microsoft Sentinel via the Sentinel Kubernetes / AKS data connector.
  • Immutable retention for compliance windows: Azure Storage immutable blob (WORM) export with a time-based retention policy - typically 1–7 years depending on framework (PCI-DSS ~1y, HIPAA 6y, SOX 7y - verify against the controlling framework).

Starter detection rules (Sentinel analytics):

  • Anonymous (system:anonymous) or unauthenticated get/list on secrets.
  • pods/exec or pods/attach against any production namespace.
  • create/update on ClusterRoleBinding outside the GitOps service account.
  • Pod create with securityContext.privileged: true or hostPID/hostNetwork/hostIPC: true.
  • ServiceAccount token mount (automountServiceAccountToken: true) in a namespace where workloads should not mount tokens.

Upgrade and Patch Strategy

Pre-Upgrade Checklist and Runbook

Use this section as the quick-reference runbook for AKS cluster and node-pool upgrades. Each time-bucketed checklist consolidates validation steps from across this bundle into an practical sequence. For detailed rationale, CLI examples, and edge cases, follow the cross-references to the subsections below.

For staged multi-cluster upgrades, run this checklist per fleet stage or member cluster and also see fleet-management - Update Orchestration and Auto-Upgrade Profiles and Emergency CVE Patching.

[!IMPORTANT] AKS control-plane minor upgrades are irreversible. There is no in-place control-plane downgrade. Node-pool rollback is a separate (Preview) recovery path; it does not restore control-plane state. Plan the primary recovery path (parallel cluster, restore-from-backup, or node-pool snapshot) before starting any production upgrade.

Upgrade type Preferred recovery path Cross-reference
Control-plane minor regression Parallel cluster or restore-from-backup Upgrade Irreversibility and Recovery
Node-pool version/image regression Node pool rollback (GA) if eligible (7-day window); otherwise snapshot or replace Node pool rollback (GA)
Emergency security patch blocked by PDB Force upgrade (break-glass only) CVE Response Decision Matrix
T-1 week
  • Confirm scope: control plane, all node pools, specific node pool, or node-image-only. Determine target Kubernetes version and whether the upgrade is manual, auto-upgrade channel, or a Fleet update run.
  • Confirm maintenance window (aksManagedAutoUpgradeSchedule for Kubernetes, aksManagedNodeOSUpgradeSchedule for node OS), approvers, stop conditions, and primary recovery path.
  • Check version skew across control plane, node pools, and client tooling. Ensure no node pool is more than 2 minors behind the control plane (Version Skew Policy).
  • Check surge headroom against the VMSS instance cap: AKS upgrade validation rejects a node-pool upgrade when current pool size plus effective surge would exceed the 1,000-instance VMSS limit (AKS release 2026-08-07) - lower max-surge or stage large pools rather than discovering the failure mid-upgrade.
  • Run API deprecation inventory for the target minor, check removed and deprecated APIs in live cluster manifests, Helm releases, and the GitOps repo. Use pluto, kube-no-trouble/kubent, or kubeconform (Pre-Upgrade API Deprecation Inventory).
  • Run CRD storage-version migration checks for conversion-sensitive CRDs (cert-manager, Istio, Knative, etc.). Run kube-storage-version-migrator for affected CRDs before the upgrade window (Pre-Upgrade API Deprecation Inventory).
  • Inventory admission and conversion webhooks. Every failurePolicy: Fail webhook must exclude kube-system and the GitOps namespace. Verify webhook HA and apiserver reachability (Admission and Conversion Webhook Pre-Upgrade Checklist).
  • Verify add-on and operator compatibility for the target Kubernetes minor. Check each component's compatibility matrix (Add-on / operator compatibility gate).
  • Run the upgrade end-to-end in a clone of production with production-like policies, CRDs, and replayed traffic. Use node-pool snapshots or Velero restore to match production data-plane state (Pre-upgrade validation in a clone of prod).
  • Verify AKS pre-upgrade validations will pass:
    • Quota: sufficient VM core quota for current nodes + surge nodes (az vm list-usage --location <region>).
    • Subnet IPs: enough addresses for all nodes, surge nodes, and pods (Total IPs = (nodes + maxSurge) * (1 + maxPods)).
    • Certificates / service principals: no expired credentials on the cluster or node pools.
    • Managed resource locks: no locks on the MC_ node resource group that block operations.
    • PDB configuration: no PDB has maxUnavailable=0 or minAvailable equal to current replica count (blocks all evictions).
  • Confirm node-pool upgrade settings on every pool that will be upgraded: maxSurge, maxUnavailable, drain timeout, and node soak time. For multi-zone pools, set maxSurge to a multiple of 3 to keep zone balance during surge.
  • If using auto-upgrade channels, configure AKS Communication Manager for in-advance and in-window upgrade notifications.
T-1 day
  • Freeze nonessential cluster, GitOps, and add-on changes until the maintenance window completes.
  • Confirm backup and restore readiness: Azure Backup for AKS or Velero, persistent volume snapshots, and GitOps repo mirror. Verify the recovery path is tested and the runbook is accessible.
  • Recheck quota headroom, subnet IP availability, target version release notes, and the AKS release tracker for region rollout status.
  • Review PodDisruptionBudget settings across production namespaces. Check ALLOWED DISRUPTIONS per PDB. Scale up replicas where the PDB math blocks eviction of even a single pod.
  • Decide whether undrainable node behavior should remain at the default (Schedule) or use Cordon during this upgrade (--undrainable-node-behavior Cordon). Cordon quarantines undrainable nodes and lets the upgrade proceed, you remediate them after.
  • Confirm autoscaler/NAP pause impact. If the upgrade window exceeds expected scaling demand, over-provision a small warm "pause pool" sized to absorb it (Cluster autoscaler / NAP behaviour during upgrades).
T-1 hour
  • Verify cluster health: all nodes Ready, no unexpected Pending or CrashLoopBackOff pods, no unresolved warning events.
  • Confirm ingress, DNS, certificate expiry, and external dependency health. Run a smoke test against the application endpoint.
  • Start upgrade monitoring in a dedicated terminal:
    kubectl get events --field-selector source=upgrader --watch
    kubectl get nodes -w
  • Open dashboards for cluster health, node health, and application SLO metrics.
  • Reconfirm rollback decision owner and communication channel for go/no-go decisions.
During the upgrade
  • Watch upgrader events (Surge, Drain, Update, Delete), node readiness transitions, and application error-rate or latency signals.
  • Expect cluster autoscaler and NAP to pause scaling for the upgrade duration. Pods that need new capacity remain Pending until the upgrade completes.
  • If drains block on PDBs (FailedDrain events), remediate by scaling replicas or temporarily widening the PDB. Commit PDB changes to Git, do not edit live (PDB-induced drain stalls).
  • If nodes are undrainable after drain timeout expires, use quarantined-node workflow:
    1. Identify quarantined nodes (kubectl get nodes --show-labels | grep Quarantined).
    2. After the upgrade, remove the offending PDB or resolve the pod termination issue.
    3. Remove the kubernetes.azure.com/upgrade-status=Quarantined label.
    4. Reconcile cluster state: az aks update then scale the node pool to restore original size.
  • If using a node soak time, validate application health before each node batch proceeds.
  • The MaxUnavailable fallback (Preview) allows surge and in-place upgrade to coexist, if surge nodes cannot be provisioned, AKS falls back to surging 1 node (K8s >= 1.35), then to in-place upgrade using maxUnavailable (MaxUnavailable fallback (Preview)).
Rollback decision flow
  1. If platform and application health remain within agreed thresholds throughout the upgrade, continue.
  2. If drain failures are the only issue, fix PDB or replica configuration, or use undrainable-node Cordon behavior before aborting.
  3. If a node-pool regression occurs and rollback is eligible (within 7 days, one step, auto-upgrade disabled), use node pool rollback (GA).
  4. If an urgent security patch is blocked by PDBs and break-glass approval exists, use force upgrade (--enable-force-upgrade --upgrade-override-until) as the last resort. This bypasses PDB protections and all other drain configurations. [VERIFY] CLI/API version requirements against current Microsoft Learn.
  5. If the control-plane upgrade fails or broader platform state is degraded, fail over to parallel cluster or restore-from-backup.

[!WARNING] Force upgrade bypasses Pod Disruption Budgets and can cause complete service unavailability during the upgrade window. Use only for urgent CVE response scenarios where the risk of not patching outweighs the disruption. Requires Azure CLI 2.79.0+ or AKS API version 2025-09-01+. When force upgrade is active, undrainable node behavior settings are not applied.

Post-upgrade
  • Validate cluster state: az aks show confirms provisioningState: Succeeded and the target Kubernetes version. Control plane and all node pools are on the intended version.
  • Validate node state: all nodes Ready, correct node image version. No quarantined nodes left unresolved.
  • Validate platform state: add-ons, webhooks, CoreDNS, ingress/Gateway API, and NetworkPolicy behaviour are working. Check kubectl get events -A --sort-by=.metadata.creationTimestamp for post-upgrade anomalies.
  • Validate workload state: Deployments, StatefulSets, and DaemonSets are available. No unexpected Pending, CrashLoopBackOff, or ImagePullBackOff pods.
  • Validate application state: smoke tests pass, SLO/error-rate and latency metrics are within baseline, key transaction paths succeed.
  • If any validation step fails, use Post-Upgrade Pod Failure Triage before deciding recovery.

For further troubleshooting and emergency response, see CVE Response Decision Matrix, Known Pitfalls, and Stop Conditions.

Upgrade Irreversibility and Recovery

AKS does not support rollback of a control-plane Kubernetes minor upgrade. There is no az aks downgrade. Once the control plane moves from 1.N to 1.N+1, it cannot move back. However, node pools have a first-class recovery path (see below).

Recovery paths if an upgrade goes wrong:

  • Node pool rollback (GA). az aks nodepool rollback reverts a node pool to its pre-upgrade Kubernetes version and node image. Available for 7 days after upgrade, one step only (no chaining). Requires disabling cluster auto-upgrade first or it re-upgrades immediately. Not available for control-plane rollback. See Node pool rollback below.
  • Parallel cluster at the previous minor. Build a new cluster at the prior minor from IaC, restore workloads via GitOps, shift DNS/traffic to it. Fastest recovery when IaC and GitOps are reproducible.
  • Restore from backup into a new cluster. Azure Backup for AKS or Velero restore into a freshly built cluster at the prior minor. Use when GitOps state alone is insufficient (stateful workloads).
  • Node-pool snapshot restore for node-image-only regressions. az aks nodepool add --snapshot-id <prior-snapshot> to replace nodes with a known-good OS+disk image. Only addresses node-image issues, not control-plane regressions.

Design implications:

  • Keep IaC reproducible end-to-end; a half-coded cluster is unrecoverable under pressure.
  • Keep DNS and traffic management outside the cluster (Front Door, Traffic Manager, Application Gateway in a separate RG) so a clean-slate build does not require DNS re-issuance.
  • Automate parallel cluster build to < 2 hours from IaC apply to workloads serving traffic - the fastest path when IaC and GitOps are reproducible.
  • Plan restore-from-backup (Velero / Azure Backup for AKS) for < 4 hours for stateful workloads - slower because PV restore is data-volume bound; rehearse against representative data sizes.
  • Plan node-pool snapshot rollback for < 30 minutes for node-image-only regressions - fastest of the three, but only covers OS/disk-level regressions, not control-plane or API regressions.
  • Run upgrade rehearsals in a clone of prod before every minor.

Node pool rollback (GA)

Node Pool Rollback is generally available as of AKS release 2026-08-07; no preview flag or extension is required. Reference: Roll back node pool versions.

az aks nodepool rollback reverts a node pool to the Kubernetes version and node image (VHD) it ran before the last upgrade. This covers both full Kubernetes version upgrades and node-image-only updates. Both version and image are rolled back together, no mixed state. If only a node-image update happened within the window (no version change), rollback restores the previous VHD while keeping the Kubernetes version.

Key constraints:

  • 7-day window from the upgrade completion. After that, the previous version is no longer available.
  • One step only. You cannot chain rollbacks to skip multiple versions. Only the immediate prior version is available.
  • No concurrent cluster operations during rollback - abort any in-flight operation first.
  • Must disable cluster auto-upgrade first (az aks update --disable-cluster-autoupgrade), or AKS will re-upgrade the pool immediately after rollback completes. Same for Fleet auto-upgrade profiles: remove the pool from its update group first.
  • Cannot roll back to a version outside AKS support.
  • OS-SKU changes are not revertible by rollback (e.g., Ubuntu → Azure Linux) - the pre-change image belongs to a different OS SKU and is rejected; revert an OS-SKU change with az aks nodepool update --os-sku instead.
  • Security regression risk: rolling back removes the newer version's security patches - plan the re-upgrade, don't camp on the rolled-back state.

Check available rollback versions first:

az aks nodepool get-rollback-versions \
  --name <pool> \
  --cluster-name <cluster> \
  --resource-group <rg>

If the node pool has never been upgraded, this returns an empty/error response, there is nothing to roll back to.

Perform the rollback:

az aks nodepool rollback \
  --name <pool> \
  --cluster-name <cluster> \
  --resource-group <rg>

The rollback is manual to trigger but fully automatic once started. AKS rolls all nodes back to the previous version state. The operation is all-or-nothing, if any node fails, the entire operation fails, leaving the cluster in a defined state. Monitor via the Azure Portal Activity Log on the cluster, or the Operation Status API for real-time progress; the rollback appears as a standard AKS operation.

When to use it: Production incidents where an upgrade breaks application compatibility, introduces performance regressions, or causes unexpected infrastructure behaviour that cannot be fixed quickly. Treat rollback as temporary recovery, re-upgrade once the root cause is resolved. Re-upgrade within days for critical security issues, within weeks for app compatibility problems, no more than 30 days in any case.

Window vs triage reality: the 7-day clock runs whether or not humans have finished triaging. If your aksManagedAutoUpgradeSchedule fires upgrades outside business hours, the window can close before anyone diagnoses the failure - pair rollback-capable pools with Node Disruption Policy gating and a maintenance-window design that leaves triage time, and keep blue/green pool capacity as the durable recovery strategy beyond day 7.

What it is not: This is a version rollback, not a full state restore. Workload changes, configuration changes, and anything outside the node pool version are not affected. It also does not replace the control-plane irreversibility above.

Version Skew Policy

Kubernetes version skew limits are narrow. AKS enforces them on upgrade; falling behind on a node pool will block its upgrade.

  • Control plane vs kubelet: kubelet may be up to N-3 minors behind the control plane on current Kubernetes (verify against the upstream version-skew policy and the AKS support matrix for the target minor).
  • client-go and kubectl: typically within ±1 minor of the control plane. Older kubectl may fail on newer API verbs; newer kubectl emits deprecation warnings.
  • Order of operations: upgrade the control plane first, then node pools, then client tooling.
  • Idle node pools: node pools left without upgrades across multiple minors hit skew and fail to upgrade or to rejoin.

Guardrail: do not let any production node pool fall more than 2 minors behind the control plane. Catch this in preflight by comparing kubernetesVersion across az aks nodepool list.

Admission and Conversion Webhook Pre-Upgrade Checklist

A webhook with failurePolicy: Fail on the apiserver path can wedge an upgrade - kube-system pods cannot reach a down webhook and the upgrade stalls.

Before every minor upgrade, inventory and verify:

kubectl get mutatingwebhookconfigurations -o json \
  | jq '.items[] | {name:.metadata.name, fp:[.webhooks[].failurePolicy], ns:[.webhooks[].namespaceSelector]}'
kubectl get validatingwebhookconfigurations -o json \
  | jq '.items[] | {name:.metadata.name, fp:[.webhooks[].failurePolicy], ns:[.webhooks[].namespaceSelector]}'

Required:

  • Every failurePolicy: Fail webhook must exclude kube-system and the GitOps namespace via namespaceSelector (e.g., matchExpressions excluding kubernetes.io/metadata.name in (kube-system, flux-system, argocd)).
  • Webhook pod has a PodDisruptionBudget and runs HA across zones.
  • Webhook service is reachable from the apiserver subnet (private-cluster gotcha).
# Minimum production PDB for an admission webhook — adapt the selector to the controller.
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: cert-manager-webhook
  namespace: cert-manager
spec:
  minAvailable: 1
  selector:
    matchLabels:
      app.kubernetes.io/name: webhook
      app.kubernetes.io/instance: cert-manager

Component-specific:

  • Ratify and Kyverno image-verification webhooks - consider temporarily setting failurePolicy: Ignore during the maintenance window, then restoring after. Image admission down for the upgrade window is lower risk than a wedged apiserver.
  • CRD conversion webhooks (cert-manager, Istio, Knative, others) - verify the webhook pod is healthy and that the CRD's spec.versions[*].served=true set still includes a version supported by the target Kubernetes minor.

Pre-Upgrade API Deprecation Inventory

Before approving a Kubernetes minor upgrade (cluster or fleet update run), scan in-cluster manifests for deprecated or removed APIs. AKS does not block an upgrade just because removed APIs are in use - workload Deployments will fail to apply after the upgrade.

Tooling options (any one is acceptable; pick what fits the pipeline):

  • kubectl deprecations (kubectl plugin, krew install) - quick interactive scan.
  • pluto (Fairwinds) - scans live cluster manifests and static files; CI-friendly.
  • kube-no-trouble / kubent - scans live cluster Helm releases and manifests for deprecated APIs relative to a target version - stale upstream since August 2024; verify it still covers your target minor before relying on it. Prefer pluto for active maintenance.
  • kubeconform / kubeval - schema validation against a target Kubernetes minor; catches removed group/version pairs.
  • kubectl-convert - official Kubernetes tool (kubectl-convert binary, separate from kubectl) that converts manifests from an old apiVersion to a target version in place. Use it alongside pluto/kubent: once a deprecated resource is identified, kubectl-convert updates the manifest file rather than requiring a manual edit for each field. Install from kubernetes/kubernetes releases or the kubernetes.io tools page. Run on GitOps manifests as part of the pre-upgrade remediation step.

Example pipeline step:

# Scan live cluster and Helm releases for APIs deprecated or removed in the target minor.
kubent --target-version 1.30 --output json

# Scan static manifests in the GitOps repo against the target minor.
pluto detect-files -d ./manifests --target-versions k8s=v1.30.0

Acceptance criteria before the upgrade is approved:

  • Zero removed findings for the target minor (these will fail at apply time).
  • All deprecated findings have a tracked remediation owner and target date.
  • Helm chart vendors have been checked for upcoming chart releases that drop the deprecated APIs.

This check is in addition to AKS upgrade preflight checks; it covers application manifests, which AKS does not inspect.

AKS does hard-block the control-plane upgrade on live deprecated-API usage , and force-upgrade is a footgun. Distinct from the static-manifest scan above: AKS upgrade preflight inspects the API server's actual deprecated-API usage over roughly the last 12 hours and blocks the control-plane upgrade with an error of the form "Control Plane upgrade is blocked due to recent usage of a Kubernetes API deprecated in the specified version," naming the group/version (a common blocker upgrading to 1.32 is flowcontrol.apiserver.k8s.io/v1beta3). The caller is frequently an add-on or controller that watches the deprecated API (Kueue, policy/monitoring controllers, service-mesh components), not your own manifests, so a clean pluto/kubent scan of the GitOps repo does not guarantee a clean gate. Fix the caller (upgrade the add-on to a build using the stable API); do not clear it with the force-upgrade override, which passes the gate but leaves those API calls failing against the new API server. [VERIFY] the exact deprecated group/version for your target minor via the AKS "deprecated APIs" diagnostic before scheduling.

CRD storage version migration. A Kubernetes minor upgrade may remove a previously-served CRD storage version (e.g., cert-manager.io/v1alpha2 → v1). Existing CR objects stored at the old version must be re-written before the version is removed, or the apiserver loses the ability to decode them. Use the kube-storage-version-migrator controller to drive a no-op update across all instances at the new storage version. Run this as a routine pre-upgrade step on every minor, not only when a known migration is published.

# Run a no-op update across every CR of the named resource so etcd stores it at the current served version.
apiVersion: storagemigration.k8s.io/v1alpha1
kind: StorageVersionMigration
metadata:
  name: certificates-cert-manager-io
spec:
  resource:
    group: cert-manager.io
    resource: certificates
    version: v1
kubectl apply -f storage-migration-certificates.yaml
kubectl get storageversionmigration certificates-cert-manager-io -o yaml \
  | yq '.status.conditions'
# Expect: type=Succeeded, status=True before removing the old served version from the CRD.
Area Default Notes
Kubernetes upgrades Stable/patch channel aligned to org risk appetite Confirm current AKS channel semantics.
Maintenance windows Two required for production: aksManagedAutoUpgradeSchedule for Kubernetes and aksManagedNodeOSUpgradeSchedule for node OS Configure both; do not let one window absorb both upgrade classes.
Node pool upgrades Surge/blue-green/canary where appropriate Validate PDBs, topology spread, HPA/KEDA min replicas, and NetworkPolicy before drain tests.
Preview features Non-prod first Include rollback and support impact.

IaC surface anchors. Maintenance windows shown here are the AKS-managed configuration names (aksManagedAutoUpgradeSchedule, aksManagedNodeOSUpgradeSchedule). They are exposed in Bicep/ARM under Microsoft.ContainerService/managedClusters/maintenanceConfigurations and in Terraform under nested maintenance_window_auto_upgrade / maintenance_window_node_os blocks on azurerm_kubernetes_cluster. Verify current azurerm provider support before relying on either nested block.

Add-on / operator compatibility gate

Every cluster-installed controller has its own supported Kubernetes range. The table below is illustrative - verify the target minor against each project's compatibility matrix before the upgrade.

Component Target K8s minor compatibility source Owner
Flux (AKS extension) Microsoft.KubernetesConfiguration extension version + Flux release notes Platform
Argo CD Argo CD release notes / supported-versions doc Platform
KEDA (managed add-on or self-hosted) AKS KEDA add-on docs / KEDA project compatibility matrix Platform
Cilium / Azure CNI Powered by Cilium / ACNS AKS networking docs for the target minor Platform
NGINX ingress (app routing managed) AKS app routing add-on docs Platform
Application Gateway for Containers (AGC) AGC release notes and AKS supported-versions Platform
cert-manager cert-manager supported Kubernetes versions Platform
Istio add-on AKS Istio add-on supported revisions per K8s minor Platform
Kyverno / Gatekeeper / Ratify Project release notes; verify webhook CRD versions Security

Cluster autoscaler / NAP behaviour during upgrades

  • Cluster autoscaler pauses scaling during cluster upgrades. Pods that need new capacity remain Pending until the upgrade completes.
  • Node Auto-Provisioning (NAP) pauses provisioning for the duration of the upgrade.
  • Mitigation: over-provision a "pause pool" (a small static node pool with a few warm nodes) ahead of high-risk upgrades, sized to absorb expected scaling demand for the upgrade window.
  • --max-surge on AKS node pools: default 10%. Recommend 33% for stateless pools (faster upgrades) and 1 node for stateful/quorum-sensitive pools. Surge nodes count against subscription vCPU quota - confirm headroom before the maintenance window. The MaxUnavailable fallback (Preview, June 2026) extends surge upgrades with a multi-strategy fallback when quota or capacity blocks the preferred surge value - see MaxUnavailable fallback (Preview) below.

MaxUnavailable fallback (Preview)

[!IMPORTANT] This feature is in preview. It requires the aks-preview Azure CLI extension and Azure CLI 2.34.1 or later. Preview features are excluded from the service-level agreement and limited warranty. Not for production use.

When both maxSurge and maxUnavailable are greater than 0 on a node pool, AKS follows a three-strategy fallback during upgrades:

  1. Attempt full surge — AKS tries the configured maxSurge value first.
  2. Fall back to surge of 1 — If full surge isn't possible (quota, capacity, subnet IPs), AKS attempts a surge of 1 node instead. This step only applies to agent pools running Kubernetes 1.35 or later.
  3. Fall back to in-place upgrade — If even a single surge node can't be provisioned, AKS falls back to an in-place upgrade using maxUnavailable, cordoning and draining existing nodes without adding new ones.

Configure it on an existing node pool:

az aks nodepool update \
  --resource-group <rg> \
  --cluster-name <cluster> \
  --name <pool> \
  --max-surge 33% \
  --max-unavailable 1

Verify the settings applied:

az aks nodepool show \
  --resource-group <rg> \
  --cluster-name <cluster> \
  --name <pool> \
  --query upgradeSettings

You should see both maxSurge and maxUnavailable in the output. If not, check the CLI version and aks-preview extension.

Behavioural differences from pure surge upgrades:

  • maxUnavailable does not provision new nodes. It cordons existing nodes and evicts pods into a pool that is already under capacity pressure.
  • Pod Disruption Budgets are more likely to block the drain during the fallback path because there are fewer available nodes to reschedule pods onto. Check kubectl get pdb -A, any PDB with ALLOWED DISRUPTIONS at zero will stall the drain.
  • maxUnavailable cannot be set on system node pools. This only applies to user node pools.

Before testing, check whether the cluster has headroom to absorb the disruption:

kubectl get pdb -A
kubectl top nodes
kubectl get nodes

Recommendation: Start with --max-surge 33% --max-unavailable 1 to keep surge as the preferred path and limit the fallback to one unavailable node at a time. Only increase maxUnavailable after testing in a non-production cluster. This is a preview feature, test before relying on it in production.

[VERIFY] MaxUnavailable fallback Preview status, regional availability, and Kubernetes version requirements against current Microsoft Learn before production consideration.

Pre-upgrade validation in a clone of prod

"Production-like" is not enough. Upgrade dress rehearsals need the data-plane state and admission configuration of production to surface webhook, CRD, and policy regressions.

Options:

  • AKS node-pool snapshots - az aks snapshot create then az aks nodepool add --snapshot-id <id> to clone node OS+disk state into a staging cluster.
  • Velero / Azure Backup for AKS - restore Kubernetes resources and PVs into the staging cluster so admission policies, CRDs, and stateful workloads match production.
  • Shadow / replayed traffic - mirror production traffic at the ingress layer (AGC, NGINX mirror annotation, service mesh) for the duration of the rehearsal.

Run the full upgrade (control plane + every node pool) in the clone with replayed traffic before scheduling the production maintenance window.

Node OS auto-upgrade channel

Pick exactly one node OS channel per cluster. This is separate from the cluster-level Kubernetes auto-upgrade channel and is the primary control for CVE response.

Channel Updates owner Cadence OS-specific behaviour When to use
None You Never Nodes get no security updates Do not use for production.
Unmanaged OS distro Ubuntu/Azure Linux apply unattended upgrades ~daily around 06:00 UTC; Windows behaves like None New nodes are unpatched at allocation; patches eventually applied in-place Legacy/migration only - newly allocated nodes are unpatched until the distro patcher runs.
SecurityPatch AKS Weekly Updates the node VHD with security-only patches honouring maintenance windows and surge; live-patches nodes in place when possible (drain/reimage only when a patch requires it, e.g. certain kernel packages); disables Linux unattended upgrades; not supported on Windows pools or Azure Linux GPU VMs Use when nodes must stay on the current VHD image and only receive security patches at a faster cadence than the weekly image; acceptable production alternative to NodeImage when image-version drift across the fleet must be minimised.
NodeImage AKS Weekly Replaces VHD with newly patched build; disruptive but follows maintenance window/surge; no extra VHD storage cost; disables Linux unattended upgrades Recommended default for production.

Set it explicitly on cluster create or update:

# On cluster create
az aks create --resource-group <rg> --name <cluster> \
  --node-os-upgrade-channel NodeImage \
  <other-flags>

# On an existing cluster
az aks update --resource-group <rg> --name <cluster> \
  --node-os-upgrade-channel NodeImage

IaC surface anchors. The channel name shown here is the CLI form (--node-os-upgrade-channel). In Bicep/ARM, the property is properties.autoUpgradeProfile.nodeOSUpgradeChannel. In Terraform's azurerm_kubernetes_cluster, the argument is node_os_upgrade_channel. The same value strings (None, Unmanaged, SecurityPatch, NodeImage) apply across surfaces - verify against the current azurerm provider and ARM API version.

Changes to the node OS channel take up to 24 hours to take effect. Node image versions are valid for 90 days from publish - design maintenance windows and staged rollouts so the slowest member still patches inside that window.

Rules:

  • Never leave production clusters without both aksManagedAutoUpgradeSchedule and aksManagedNodeOSUpgradeSchedule maintenance windows.
  • Never leave production on --node-os-upgrade-channel None; default to NodeImage.
  • Do not wait until N-3/N-4 pressure forces upgrades.
  • Test upgrades in lower environments with production-like add-ons and policies.
  • Use canary node pools or staged workload migration for disruptive node image, OS SKU, GPU, or kernel changes.
  • Monitor the AKS Security Bulletins feed and nodeImageVersion per pool. When an advisory names a patched VHD build, drive the response through the Fleet Manager NodeImage auto-upgrade profile (see fleet-management) or, for single clusters, an out-of-band az aks nodepool upgrade --node-image-only run.

Shared Maintenance Windows (Public Preview)

Planned-maintenance schedules have historically been inline per-cluster properties, which lets windows drift apart across a fleet until an upgrade fires at the wrong time. Shared Maintenance Windows turn the schedule into a standalone ARM resource you define once and link into as many clusters as needed.

# One-time, subscription-scoped, auto-approved feature registration
az feature register --namespace Microsoft.ContainerService --name AKSSharedMaintenanceWindowPreview
az provider register --namespace Microsoft.ContainerService

# Create the standalone window resource (resource-group scope)
az aks maintenancewindow create \
  --resource-group <platform-rg> \
  --name fleet-standard-window \
  --schedule-type Weekly \          # Daily | Weekly | AbsoluteMonthly | RelativeMonthly
  --day-of-week Sunday \
  --interval-weeks 1 \
  --duration 4 \                    # 4-24 hours
  --utc-offset +00:00 \
  --start-time 01:00

# Link it into per-cluster maintenance configurations (either config name)
az aks maintenanceconfiguration add \
  --resource-group <rg> --cluster-name <cluster> \
  --name aksManagedAutoUpgradeSchedule \    # or aksManagedNodeOSUpgradeSchedule
  --maintenance-window-id "/subscriptions/<sub>/resourceGroups/<platform-rg>/providers/Microsoft.ContainerService/maintenanceWindows/fleet-standard-window"

Design considerations:

  • Fleet-scale drift control. One window resource governs every linked cluster's aksManagedAutoUpgradeSchedule and/or aksManagedNodeOSUpgradeSchedule; changing the window updates all linked clusters. This is the "when", Fleet Manager update runs remain the "in what order"; the two are complementary, not alternatives (see fleet-management).
  • Mutual exclusivity is enforced. --maintenance-window-id cannot be combined with inline schedule flags (--schedule-type, --day-of-week, …) in the same command, and cannot be combined with --config-file; set maintenanceWindowId inside the JSON instead. To unlink a cluster, re-issue the configuration with an inline schedule and omit the window ID; other clusters referencing the window are unaffected. [VERIFY] whether deletion of a window is blocked while configurations still reference it (reported behaviour) via az aks maintenancewindow CLI reference.
  • Management surface. az aks maintenancewindow (aks-preview extension) and ARM/Bicep (Microsoft.ContainerService/maintenanceWindows, preview API 2026-04-02-preview per the Bicep template reference); no first-class Azure portal or azurerm Terraform support yet; Terraform consumers need azapi. Blackout date ranges (notAllowedDates) are only settable via the JSON config file.
  • Preview gate. Self-service, opt-in, excluded from SLA ; do not migrate production upgrade windows to shared windows yet; validate on non-production clusters that upgrades fire inside the linked window before fleet rollout. [VERIFY] API version and flag surface against the maintenance configuration CLI reference at implementation time.

Node Disruption Policy (Public Preview)

AKS Node Disruption Policy (cluster property nodeDisruptionProfile) lets platform teams control when nodes can be reimaged during routine cluster configuration changes (distinct from Kubernetes version upgrades). Without it, any az aks update or add-on change that triggers a reimage may do so immediately, regardless of business hours or active workloads. A reimage is a full node recreation: node-local ephemeral data (EmptyDir, local paths) is lost; persistent volumes are unaffected.

Policy value Behaviour Use case
Allow (default) Reimage-triggering operations proceed immediately Dev/test where speed beats stability
AllowDuringMaintenanceWindow Held until the aksManagedNodeOSUpgradeSchedule window is active Production default — changes land without manual action, inside a pre-approved window
Block API call fails with an error; nothing is applied Short-term freeze during incidents/high-traffic events — a hard gate, not a deferral queue; toggle to Allow to apply the change
# Register the preview feature
az feature register --namespace Microsoft.ContainerService --name NodeDisruptionProfile
az provider register --namespace Microsoft.ContainerService

# Set at create or update time (aks-preview extension)
az aks create ... --node-disruption-policy AllowDuringMaintenanceWindow
az aks update --resource-group <rg> --name <cluster> --node-disruption-policy Block

Design considerations:

  • The silent-fallback trap. AllowDuringMaintenanceWindow is satisfied only by an aksManagedNodeOSUpgradeSchedule window; the default window and aksManagedAutoUpgradeSchedule do not qualify. With no qualifying window configured, all disruptive operations are allowed (the policy has nothing to gate against), so the policy silently becomes a no-op. This is a second independent reason the "never leave production without an aksManagedNodeOSUpgradeSchedule window" rule above is non-negotiable.
  • Coverage is a matrix, not a blanket. Covered (reimage gated): network policy/CNI-Overlay changes, node OS channel changes to/from Unmanaged, IPv6 dual-stack enablement, Cilium data plane on/off, HTTP proxy config, custom CA certs, kubelet identity, private DNS zone, API Server VNet Integration enablement, eBPF host routing; pool-level: LocalDNS profile, Trusted Launch toggles, artifact streaming, Windows GMSA, Capacity Reservation Group attachment. Version-dependent surface (K8s 1.37+): SSH node-access config changes, IMDS restriction, network-isolated bootstrap profile (artifactSource/containerRegistryId), and cluster outbound-type changes now trigger an immediate node reimage as of AKS release 2026-08-07 - gate them with this policy instead of the pre-1.37 pattern of applying az aks nodepool upgrade --node-image-only manually afterwards; cluster-wide service-principal → managed-identity conversion also reimages every pool. Never controlled: node image and Kubernetes version upgrades, node pool rollback, admin restore, node identity credential rotation.
  • Mixed changes escape the gate. If a single configuration change bundles a covered operation together with an upgrade, the covered part is not controlled by the policy. Split bundled changes when the reimage portion must be gated.
  • Azure override. Azure reserves the right to perform urgent/critical maintenance regardless of the policy setting; Block is not an absolute freeze.
  • Not a substitute for upgrade windows. Kubernetes minor-version upgrades, node OS auto-upgrade channel runs, and CVE-response --node-image-only runs follow the aksManagedNodeOSUpgradeSchedule window independently. Node Disruption Policy covers the additional reimage surface from cluster config mutations.
  • Preview gate. [VERIFY] the feature-flag name (NodeDisruptionProfile per the CLI extension PRs; the ARM property is nodeDisruptionProfile) and CLI flag surface against the Node Disruption Policy docs and Azure/AKS #5017 before committing to production designs.

Post-Upgrade Pod Failure Triage

After a Kubernetes minor or node-image upgrade, certain failure modes recur. Map symptom → likely cause before reaching for rollback (pool-scoped rollback is available within the 7-day window when eligible - see Node pool rollback (GA); control-plane regressions have no undo).

  • ImagePullBackOff - admission policy regression around the image registry allow-list; Workload Identity / imagePullSecret regression from a token format change.
  • CrashLoopBackOff - liveness probe regression against a new kubelet version (probe semantics tighter); readiness probe timing now stricter; Pod Security Admission bumped from baseline to restricted on a namespace via admission policy drift.
  • Pending - cluster autoscaler / NAP paused during the upgrade window; insufficient subscription quota for --max-surge; PDB-stuck drain blocking a node from coming back into scheduling.
  • Networking regressions - CoreDNS ConfigMap reset to defaults (custom Corefile lost); newly-applied default-deny NetworkPolicy; Cilium / CiliumNetworkPolicy CRD version skew between operator and dataplane.

Triage commands:

kubectl get events -A --sort-by=.metadata.creationTimestamp | tail -100
kubectl describe pod <pod> -n <ns>
kubectl get pdb -A
kubectl logs <pod> -n <ns> -p   # previous container, post-crash
kubectl get pods -A -o wide --field-selector=status.phase!=Running

CVE Response Decision Matrix

When an AKS Security Bulletin, upstream Kubernetes CVE, or third-party component CVE applies to a production cluster, the response path depends on which layer is affected. Use this matrix to choose the right action surface.

Affected component Response surface Single-cluster command Fleet command
Node OS / kernel / container runtime (e.g., the AKS-2026-0003 "Copy Fail" / Dirty Frag kernel LPE class referenced elsewhere in this skill - verify against the current AKS Security Bulletin before acting) New node image VHD az aks nodepool upgrade --node-image-only --resource-group <rg> --cluster-name <cluster> --nodepool-name <pool> az fleet autoupgradeprofile generate-update-run --resource-group <rg> --fleet-name <fleet> --auto-upgrade-profile-name <node-image-profile>
Kubernetes control plane / kubelet (in-tree CVE) Patched K8s minor or patch version az aks upgrade --kubernetes-version <patched-version> Fleet Kubernetes auto-upgrade profile run or scheduled update run
Managed add-on (App Routing NGINX, App Gateway for Containers, KEDA, Managed Prometheus, etc.) Add-on patched release az aks addon update / az aks update for the relevant flag Fleet update run scoped to add-on, where supported
Cilium / ACNS data plane Patched AKS release for the CNI engine az aks upgrade (control plane carries the engine version) Coordinated fleet update run; treat as Kubernetes-track risk
Self-managed component (BYO ingress, BYO service mesh, BYO Helm chart) Application repo / Helm chart GitOps PR + reconcile GitOps PR + reconcile per cluster
Application dependency in your own image (OpenSSL, log4j, etc.) New image build, not an AKS control CI/CD image rebuild + Flux/Argo redeploy Same, then optional staged rollout via fleet

Urgency tiers (illustrative - align to your incident process):

  • Critical RCE or privilege escalation: patch within 24–48 hours; consider out-of-band node-image runs and an emergency change ticket. Pause routine update strategies until the emergency run completes.
  • High severity: patch within 7 days inside the next maintenance window.
  • Medium severity: patch in the standing maintenance cadence.

Always confirm against the current AKS Security Bulletin: the affected nodeImageVersion or Kubernetes version, the patched VHD build identifier, and any interim mitigation (kernel module disable, NetworkPolicy block, etc.) before deciding speed and surface.

See also: fleet-management - Emergency CVE response runbook for the multi-cluster reaction path. The single-cluster fallback is az aks nodepool upgrade --node-image-only.

See also: fleet-management - Update Orchestration for cross-cluster staged rollouts.

Cluster certificate rotation

AKS cluster certificates (control-plane and node certificates) have a finite validity. Letting them expire degrades or breaks the cluster, and recovery from a fully expired state is painful, treat rotation as a scheduled maintenance task, not a surprise.

  • Rotate proactively with az aks rotate-certs (a disruptive operation, node pools are reimaged/restarted; run it in a maintenance window). Clusters with auto-rotation enabled still benefit from a documented manual path.
  • Track certificate expiry as a monitored signal; alert well before expiry rather than discovering it through an outage.
  • Service-principal-based clusters (legacy) also need credential rotation (az aks update-credentials); new clusters should be on managed identity (see cluster-foundations) to avoid this class of work.
  • [VERIFY] current rotation behaviour, auto-rotation defaults, and downtime expectations against Microsoft Learn before scheduling.

Incident Response (Runtime Compromise)

When runtime indicators (audit alerts, EDR, anomalous egress) suggest a workload or node is compromised, follow this sequence. Preserve evidence - destruction-first responses lose forensic data.

Pod quarantine. Isolate before remediating.

  1. Apply a deny-all NetworkPolicy that selects the suspect pod by label; deny both ingress and egress.
  2. Cordon the node (kubectl cordon <node>).
  3. Drain other workloads off the node (kubectl drain <node> --ignore-daemonsets --delete-emptydir-data --skip-wait-for-delete-timeout=30).
  4. Do not delete the suspect pod. Leave it Running on the cordoned node so memory and process state are preserved for forensics.
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: quarantine-suspect
  namespace: <ns>
spec:
  podSelector:
    matchLabels:
      incident: quarantine
  policyTypes: [Ingress, Egress]
  ingress: []
  egress: []

Credential and federated-credential revocation. Cut the workload's identity before rotating the data it touched.

  1. Revoke the workload-identity federated credential first (az identity federated-credential delete) - severs the pod's path to Entra tokens immediately.
  2. Rotate the Key Vault key version(s) the workload referenced; update CSI/secret references on healthy workloads.
  3. Revoke active Entra tokens via Conditional Access (sign-in risk policy + revoke session) and PIM where applicable. Force re-authentication for the affected principal.

Node disk snapshotting for forensics.

  1. Identify the node VM / scale-set instance (kubectl get node <node> -o jsonpath='{.spec.providerID}').
  2. Take an Azure managed-disk snapshot of the OS disk before re-imaging or deleting the node (az snapshot create --source <disk-id>).
  3. Preserve /var/log/, container runtime state (/var/lib/containerd/), and journald output in the snapshot. Hand to the forensics team before re-imaging.

Regulator notification clocks.

  • GDPR - personal-data breach notification within 72 hours of awareness to the supervisory authority (Art. 33).
  • HIPAA (US) - individual notification within 60 days of discovery; HHS notification within 60 days for breaches ≥ 500 individuals, otherwise annually.
  • Jurisdiction-specific - CCPA, PIPEDA, APPI, NIS2, and sector regulators have their own clocks. Confirm with legal/privacy before the clock starts.

Status page communication template.

[Status: Investigating | Identified | Monitoring | Resolved]
Summary: <one-line impact>
Affected Kubernetes minor: <e.g., 1.30.x>
Affected node-image build: <AKSUbuntu-2204gen2containerd-YYYY.MM.DD>
Scope: <namespaces / clusters / regions impacted>
Current action: <quarantine applied | rotation in progress | restore in progress>
Next update / ETA: <UTC timestamp>

See also: cluster-foundations - Threat Model for the STRIDE/ATT&CK mapping that feeds detection rules.

GitOps and Release Safety

  • Prefer pull-based GitOps such as Flux where it fits the organisation.
  • Use Argo CD only if it is the organisational standard and current extension/support status has been verified.
  • Keep emergency break-glass documented and reconciled back into Git after incident resolution.
  • Use progressive delivery for risky workload changes where supported by the app platform.
  • For IaC changes, use plan/review/apply gates and stop on destructive network, identity, or cluster replacement changes.
  • Use private or controlled ACR integration, image vulnerability scanning, immutable image tags/digests for production where supported, and admission controls for unsigned or non-compliant images.
  • Separate platform add-on changes from application releases when blast radius or rollback ownership differs.

GitOps Repository Structure

Pick a layout before onboarding workloads - retrofitting structure later breaks promotion flows and admission policy versioning.

Recommended defaults:

  • Monorepo with env-per-folder (environments/{dev,test,stage,prod}/<cluster>/<namespace>/) - simplest for promotion via PR + Kustomize overlay; default for small/medium estates.
  • App-of-apps / fleet-of-fleets - Flux Kustomization or Argo CD ApplicationSet referencing per-team repos; better for large estates with strong RBAC.
  • Helm vs Kustomize - prefer HelmRelease for third-party charts where upstream releases drive upgrades, and Kustomization for first-party manifests where the team owns the deltas. Mixing is fine; pick a per-app standard.
  • Branch model - trunk-based with environment promotion via PR is the default. Per-environment branches are acceptable only when a team can keep them in sync; long-lived branches drift.
  • Bootstrap separation - keep the GitOps controller bootstrap (Flux/Argo install) in a separate repo or top-level folder so platform changes are reviewable independently of app changes.

Operational guardrails:

  • PR approval gates: required reviewer from the namespace owner team plus a platform-team reviewer for cluster-scoped resources.
  • Per-environment branch protection: require status checks (manifest validation, policy scan, image signature verification) before merge.
  • Tag/release the GitOps repo at every production rollout so a rollback target exists.
  • Mirror the admission policy bundle (Kyverno/Gatekeeper) into the GitOps repo so CI validation and cluster admission stay in sync.

Progressive Delivery

  • Argo Rollouts and Flagger are the canary / blue-green controllers in scope. Both require a metrics provider (Managed Prometheus is the most common on AKS) and a traffic shifter (Istio add-on, AGC, Gateway API HTTPRoute weights, or NGINX annotations).
  • Define canary success metrics aligned to the SLO: request error rate, p95/p99 latency, plus the workload's primary business metric. Default thresholds drift; pick numbers that match the SLO.
  • Treat canary metric definitions as code in the same GitOps repo as the application. A canary that lives outside Git is a canary nobody reviews.

Automated Dependency Updates

  • Run Renovate or Dependabot against the GitOps repo, scoped to image digests, Helm chart versions, and Terraform module versions.
  • Configure digest-pinning (Renovate pinDigests: true). Tag-only references hide image content drift.
  • Separate PRs for major bumps. Auto-merge only patch versions with green CI; require manual approval for major/minor bumps.
  • Apply the same automation to the GitOps platform itself (Flux / Argo controller image versions, Microsoft.KubernetesConfiguration extension version) - platform drift is just as risky as workload drift.

Ephemeral Preview Environments

  • Per-PR namespace with a TTL via Argo CD ApplicationSet PR generator, or Flux ImageUpdateAutomation plus a cleanup CronJob.
  • Each preview namespace gets a ResourceQuota and a LimitRange to bound cost; cleanup CronJob deletes namespaces older than N days (default 7).
  • High-use for catching admission-policy regressions and CRD compatibility issues before merge - the preview applies the same policies as production.

GitOps Bootstrap

Chicken-and-egg: the cluster needs Flux/Argo to apply manifests, but Flux/Argo must be installed somehow.

Pattern:

  • Install Flux via the Microsoft.KubernetesConfiguration/extensions Flux extension at cluster create - declared in Bicep / Terraform / Azure Policy alongside the cluster.
  • The bootstrap pipeline runs once per cluster. From then on the GitOps controller reconciles everything else, including its own GitRepository, Kustomization, and HelmRelease resources.
  • Keep the bootstrap repo separate from the workload repo so the controller can be re-installed without depending on workload state. A workload-state-coupled bootstrap is impossible to recover when the controller is the thing that's broken.

IaC Pipeline and Drift Detection

AKS architecture decisions land via Bicep, Terraform, or ARM/CLI. The pipeline that ships them needs guardrails of its own.

Recommended gates on every IaC change:

  1. Plan / what-if - terraform plan or az deployment ... what-if; require the diff in the PR description.
  2. Static security scan - tfsec, checkov, terrascan, or Bicep psrule-rules-azure for Terraform/Bicep; fail the build on high-severity findings.
  3. Azure Policy compliance preview - for ALZ/landing-zone environments, validate the change against Azure Policy assignments in a non-prod subscription before merge.
  4. Manual approval for destructive change - node pool delete, CIDR change, identity model change, cluster delete, Fleet Manager hub access change must require a second approver.
  5. Apply via OIDC federated identity - pipeline runner uses a federated workload identity scoped to the target subscription/resource group; no long-lived service principal secrets.

Drift detection:

  • Run terraform plan -refresh-only (or equivalent Bicep/ARM tooling) on a schedule against production; treat non-zero diff as an alert.
  • Track Flux reconciliation status: alert on Kustomization or HelmRelease Ready=False, on Suspended=True (someone paused reconciliation), and on reconciliation latency above a threshold.
  • Audit kubectl writes against production clusters via Kubernetes audit logs forwarded to ContainerLogV2; reconcile out-of-band changes back to Git or remove them, never silently accept drift.
  • Periodically diff cluster state against the GitOps repo using flux diff kustomization or argocd app diff.

CI runner network path to private API server. Private API server clusters block GitHub Actions hosted runners and Azure DevOps Microsoft-hosted agents - neither can reach the API server. Treat this as a Difficult decision: it affects every IaC and GitOps pipeline targeting the cluster.

Options, with trade-offs:

  • Self-hosted runners in a peered spoke VNet - direct path, full audit trail. Operational cost (scaling, patching the runner pool) is the trade-off.
  • JIT bastion-tunnel pattern - lowest infrastructure cost, slower per-job (tunnel setup), harder to audit cleanly.
  • Ephemeral public-IP allowlist on the API server authorizedIPRanges - simplest to set up, weakest posture. Avoid for regulated workloads.

Choose once, per environment, before standing up the GitOps controller - switching later requires re-running cluster IaC and potentially redeploying runners.

Backup, Restore, and DR

Backup design must cover both Kubernetes objects and persistent data.

Layer Requirement
Manifests/IaC Source-controlled and reproducible.
Cluster state Backup extension or approved equivalent for Kubernetes resources.
Persistent volumes Snapshot/backup policy aligned to RPO/RTO.
External dependencies Database, Key Vault, ACR, DNS, identities, and firewall rules included in recovery plan.
Restore testing Scheduled restore tests; do not trust untested backups.
GitOps repository Mirror to a second Git provider (e.g., Azure DevOps Repos ↔ GitHub); branch protection enforced; alert on force-push to main and on branch-protection bypass; periodic restore drill: re-clone the mirror into a fresh cluster bootstrap.

Multi-Region Strategy

Pattern Use when Notes
Single region, zone-redundant Most production workloads with regional SLA acceptance Simpler and cheaper.
Warm standby Higher availability with controlled cost Requires tested promotion and data replication.
Active-active Global or mission-critical workloads Highest complexity; design data consistency first.
Fleet management Multi-cluster governance, update orchestration, resource placement, managed namespaces, or cross-cluster networking Load fleet-management; verify hubless vs hub, preview gates, support limits, and operational ownership.

Choose the data architecture before the cluster topology. Multi-region AKS cannot compensate for a single-region database dependency.

High Availability Framework

Before applying the specific controls in this bundle, establish a structured HA methodology. Use the four pillars below as the organising model, and the SPOF identification process as the starting point for every HA design.

Identify single points of failure

  1. Map the critical path. Trace every component between the client request and the response. DNS, ingress/Gateway, application pods, middleware, databases, caches, and external dependencies.
  2. Classify each component against the four HA pillars below. Any component on the critical path that lacks redundancy, monitoring, recovery, or checkpointing (where applicable) is a single point of failure.
  3. Eliminate systematically. Apply the corresponding pillar to each SPOF. Even a replicated component is a SPOF if it lacks monitoring, failure goes silently undetected.

The four HA pillars

Pillar What it does Kubernetes constructs
Redundancy Run multiple identical instances so a single failure does not take down the service. Deployment/StatefulSet with multiple replicas, podAntiAffinity or topologySpreadConstraints across zones/nodes, HPA/KEDA for dynamic replica sizing, multiple node pools.
Monitoring Detect failures and signal when a component is unhealthy or degraded. Liveness, readiness, and startup probes; ContainerLogV2 and Managed Prometheus metrics; kube-state-metrics and node-exporter dashboards; Azure Monitor alerts on pod restarts, node readiness, and error rates.
Recovery Isolate the faulty instance, redirect traffic, and restore health — automatically or through orchestrated action. Service abstraction (endpoint removal on probe failure), leader election for singletons, restartPolicy, preStop hooks for graceful drain, PDBs for controlled disruption during node drains and upgrades, deployment rollback strategies, and the upgrade/rollback runbooks in this guide.
Checkpointing Persist state so a new instance can resume work without starting from zero. PVCs (Azure Disk/File/Elastic SAN), application-level replication to a database (Azure Cosmos DB, Azure PostgreSQL, Azure SQL), write-ahead logs, ZRS disks for zone resilience.

Pillars are applied in order: redundancy first (run more than one copy), monitoring second (verify the copies are healthy), recovery third (act when they are not), checkpointing last (for stateful components that cannot reconstruct from source).

HA vs DR

High availability (HA) tolerates component, node, and zone failures within a region through multi-zone redundancy, probes, PDBs, and automatic recovery. Disaster recovery (DR) tolerates a regional outage through multi-region deployment, data replication, and traffic failover. The same four pillars apply, but DR adds cross-region networking, data consistency, and failover orchestration. See Multi-Region Strategy below for DR patterns.

For HA within a single cluster, the controls in this bundle and in production-workload-controls cover every pillar. The checklist below and the Production Readiness Checklist at the end of this guide consolidate them into practical validation steps.

Resilience Validation

Replicas, PDBs, topology spread, and multi-zone design are hypotheses until something fails. Rehearsing upgrades and testing restores (covered above) proves the planned paths; fault injection proves the unplanned ones. Build a resilience-validation cadence so the HA design is verified, not assumed.

HA and cost optimisation directly conflict, redundancy costs money, consolidation saves it. Validate that your replica counts, zone spread, and PDB settings are tied to a documented SLO, not fear. See Cost Controls for the trade-off analysis.

  • Failure drills to run. Zone outage (cordon/drain a zone's nodes and confirm workloads reschedule and stay within SLO), single-node loss, pod kill (confirm PDB + readiness behaviour), DNS/CoreDNS degradation, dependency latency/outage (database, Key Vault, downstream API), and disk/PV failure for stateful workloads.
  • Tooling. Azure Chaos Studio injects AKS faults (pod failures, network latency/loss, stress) with a controlled blast radius and an experiment-as-code definition; community tools (LitmusChaos, kube-monkey) cover pod-level chaos. Start in non-prod, define abort conditions, and graduate to controlled game-days in production.
  • Tie each drill to a hypothesis and a signal. "Losing zone 2 keeps p95 < X and triggers no SLO burn" is testable; "we think it's HA" is not. Feed failures back into PDB sizing, replica counts, topology spread, and the runbooks.
  • Validate the validators. Confirm alerts fire, runbooks are reachable, and on-call can act, a drill that nobody is paged for only tests the platform, not the response.

KEDA ScaledJob for Batch and Queue Workloads

KEDA supports two primary patterns. ScaledObject (covered in production-workload-controls) scales running Deployments or StatefulSets. ScaledJob is a separate pattern that creates and terminates Kubernetes Job objects dynamically as queue/event items arrive: the Job (and its pod) exists only for the duration of one unit of work.

apiVersion: keda.sh/v1alpha1
kind: ScaledJob
metadata:
  name: queue-processor
  namespace: processing
spec:
  jobTargetRef:
    template:
      spec:
        containers:
          - name: processor
            image: myregistry/queue-processor@sha256:<digest>
            resources:
              requests:
                cpu: "500m"
                memory: "512Mi"
        restartPolicy: Never
  pollingInterval: 15
  maxReplicaCount: 50
  scalingStrategy:
    strategy: "accurate"   # or "default" (each item → one job)
  triggers:
    - type: azure-servicebus
      metadata:
        queueName: work-items
        namespace: my-servicebus
        messageCount: "1"
      authenticationRef:
        name: keda-workload-identity

Key design rules for ScaledJob:

  • Each queue message (or configurable batch) maps to one Job. The scalingStrategy controls whether KEDA creates one Job per message (accurate) or scales to a target replica count (default).
  • Set maxReplicaCount to prevent burst-consuming more node capacity than the cluster can provision. Pair with node autoscaler or NAP max-node limits.
  • Jobs must complete (exit 0) or be cleaned up. Configure spec.ttlSecondsAfterFinished on the job template to avoid accumulating completed Job objects.
  • Pair KEDA ScaledJob with Kueue for capacity-aware admission when jobs compete for GPU or constrained resources: KEDA creates the Job, Kueue decides when it starts.
  • Authentication: use KEDA TriggerAuthentication with Workload Identity (not connection-string secrets) for Azure Service Bus, Event Hubs, and Storage Queue triggers.

VPA Update Modes

VPA (Vertical Pod Autoscaler) supports four updateMode values, not just Auto:

Mode Behaviour When to use
Off Recommendations computed but never applied Monitoring-only; use kubectl get vpa -o yaml to read recommendations periodically
Initial Requests set at pod creation time; existing pods not evicted Safer than Auto for production; new pods get right-sized requests, live pods are undisturbed
Recreate Recommendations applied by evicting and recreating pods Acceptable only when pod restart is tolerable and HPA is not on the same target
Auto Same as Recreate for now; may live-migrate in future Fully automated; never combine with HPA on CPU/memory without a tested design

Production recommendation: Use Off first for 2–4 weeks to gather recommendations. Promote to Initial for greenfield namespaces with accurate baseline data. Reserve Auto/Recreate for batch or dev workloads.

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: api-server-vpa
  namespace: production
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: api-server
  updatePolicy:
    updateMode: "Initial"   # safe production default
  resourcePolicy:
    containerPolicies:
      - containerName: api
        minAllowed:
          cpu: "100m"
          memory: "128Mi"
        maxAllowed:
          cpu: "4"
          memory: "4Gi"

Cost Controls

  • Right-size monthly with VPA recommendations and Azure Monitor usage data.
  • Track two efficiency ratios for every production namespace and node pool:
    • Request efficiency = actual usage / requested resources
    • Allocation efficiency = requested resources / allocatable node resources Low request efficiency indicates over-requesting; low allocation efficiency indicates poor bin-packing or pool/SKU mismatch.
  • Use autoscaling/NAP for variable workloads; set sensible min/max values.
  • Use spot pools only for interruptible workloads.
  • Use reserved capacity/savings plans only after baseline demand is proven.
  • Track NAT Gateway, public IP, Application Gateway/AGC, Log Analytics, Managed Prometheus, and GPU costs explicitly.
  • Alert on idle GPU pools, over-requested CPU/memory, and log ingestion spikes.

Cost visibility and showback tooling

Right-sizing requires attribution, you cannot optimise what you cannot see per team. The Azure bill shows VM/disk/network totals but not which namespace or team consumed them.

  • AKS Cost Analysis add-on (--enable-cost-analysis), the built-in, OpenCost-based add-on that surfaces cost allocation by cluster, namespace, and Azure Compute/Network/Storage resource directly in Microsoft Cost Management. Requires the Standard or Premium tier (not Free) and a managed-identity cluster; Kubernetes cost views are limited to Enterprise Agreement / Microsoft Customer Agreement offers. The lowest-friction first step for namespace-level showback.
  • Kubecost / OpenCost (self-hosted) — install via Helm when you need deeper allocation (controller/label/pod granularity), chargeback reports, or cost-anomaly alerts beyond what the add-on surfaces; integrate Azure billing export for actual (not estimated) rates.
  • Resource tags for allocation. Propagate team / cost-center / environment tags to the node resource group and node pools (az aks nodepool update --tags ...) so Cost Management can group spend by owner, pair with the namespace cost-center labels from production-workload-controls.
  • Budgets and alerts. Set Azure Cost Management budgets with threshold alerts on the cluster's node resource group so spend spikes surface before the invoice.
  • Cost anomaly correlation playbook. For each cost spike, correlate spend deltas with the same-window workload and platform events: deployment rollouts, replica jumps (HPA/KEDA), node-pool scale-out/NAP activity, and telemetry config changes (log or metric ingestion). Keep this as an explicit incident runbook step so cost regressions become engineering actions, not monthly surprises.

Known Pitfalls

Pitfall Symptom Mitigation
Load balancer SNAT exhaustion Intermittent outbound timeouts or cannot assign requested address NAT Gateway or firewall egress with adequate SNAT capacity.
HPA without requests HPA cannot calculate utilisation Require requests on all scaled containers.
Default deny without allow policies DNS, ingress, or dependencies fail Roll out with test pods and explicit DNS/ingress/east-west/egress rules.
One replica production app Outage during node drain or rollout Use at least two replicas, usually three across zones, or document singleton exception.
PDB too strict Node drains/upgrades hang Use maxUnavailable: 1 or tested percentages for scalable deployments.
CoreDNS mistakes Cluster-wide name resolution failures Minimise customisation and test first.
CoreDNS split-DNS trap (NAT Gateway egress + Private Endpoints) Pods cannot resolve Azure private-zone names; Azure DNS 168.63.129.16 is unreachable from pods because NAT Gateway egress does not route the link-local Azure DNS address Patch CoreDNS for split-DNS post-provision: forward Azure private zones to 168.63.129.16 and public zones to an external resolver (e.g. 8.8.8.8). Keep the Corefile in Git and test in non-prod. See cluster-foundations - CoreDNS at scale.
CoreDNS overload / ndots amplification DNS timeouts and elevated latency cluster-wide under high QPS; every external lookup tries cluster search suffixes first Lower ndots (or use FQDNs), enable NodeLocal DNSCache / LocalDNS, and confirm CoreDNS replica scaling. See cluster-foundations - CoreDNS at scale.
Manual GitOps drift Changes revert unexpectedly Change Git source or use documented break-glass.
Untested backups Restore fails during incident Test restore quarterly or per release-critical cadence.
Ingress controller lock-in Migration blocked by annotations/TLS/DNS coupling Prefer Gateway API-compatible patterns and document ownership.
Preview feature surprise Production unsupported behaviour Confirm GA/preview and support scope before adoption.
GPU cold start Slow model availability after scale-out Pre-warm, set min replicas, or design queue/SLO accordingly.
Stale node image / missed CVE patch nodeImageVersion older than 30 days or older than the latest AKS Security Bulletin patched build; the AKS-2026-0003 'Copy Fail'/Dirty Frag class is an illustrative worked example; verify the actual bulletin, CVE IDs, and patched VHD build against the current AKS Security Bulletins feed before acting Configure --node-os-upgrade-channel NodeImage and a weekly aksManagedNodeOSUpgradeSchedule; for multi-cluster fleets, run a Fleet NodeImage auto-upgrade profile and use az fleet autoupgradeprofile generate-update-run for emergency response; monitor AKS Security Bulletins and nodeImageVersion drift.
Let's Encrypt rate limit exhausted Certificate orders fail with urn:ietf:params:acme:error:rateLimited. Production limit is 50 certificates/week/registered domain. Each SAN counts separately. Test all issuer configuration against Let's Encrypt staging first. Consolidate SANs into fewer certificates when practical. Use a single ClusterIssuer per domain.
HTTP-01 chicken-and-egg on new clusters No ingress controller means no HTTP-01 challenge means no certificate — circular dependency for first-time TLS setup Use DNS-01 for initial certificate issuance (does not require a functioning ingress controller). See Bootstrap Ordering.
cert-manager webhook unreachable in private cluster Error creating: Internal error occurred: failed calling webhook "..." on any cert-manager resource operation Ensure the cert-manager webhook service is reachable from the API server subnet. Verify VNet peering, NSG rules, and AKS egress firewall rules permit required FQDN traffic.
ACME DNS-01 challenge fails with Azure Private DNS Zone Let's Encrypt tries to query the TXT record but the zone is not publicly resolvable DNS-01 requires a public Azure DNS zone. Private DNS Zones cannot be used for ACME challenges. Target the public zone explicitly in the ClusterIssuer if both exist.
Certificate secret missing after renewal cert-manager renewed the certificate but the old secret was not updated, or the ingress controller is caching the old TLS secret Verify kubectl get certificate shows Ready=True. Check the secret's tls.crt timestamp. Restart the ingress controller pods if caching is suspected.
Workload Identity credential subject mismatch Federated identity credential subject does not match the cert-manager ServiceAccount name and namespace The default subject is system:serviceaccount:cert-manager:cert-manager. Verify against kubectl get sa -n cert-manager <sa-name> -o jsonpath='{.metadata.name}'.
App Routing add-on externalTrafficPolicy: Local breaks in-cluster hairpin cert-manager HTTP-01 self-check (or any in-cluster probe hitting the cluster's own public ingress hostname/IP) hangs with context deadline exceeded, zero bytes — looks like a DNS or firewall problem, is neither. Local only forwards LB traffic to a node that has a matching backend pod; a pod calling back out to the cluster's own external IP does not land on a qualifying node. Verified live 2026-08-23. kubectl patch svc nginx -n app-routing-system -p '{"spec":{"externalTrafficPolicy":"Cluster"}}'. The NginxIngressController CRD (approuting.kubernetes.azure.com/v1alpha1) does not expose this field (only loadBalancerAnnotations, loadBalancerSourceRanges, etc. as of 2026-08-23) — patch the add-on-managed Service directly and re-apply on every provision in case a reconcile reverts it. Trade-off: Cluster loses real client-IP preservation at the ingress unless something downstream reads a forwarded-for header instead.
NetworkPolicy default-deny blocks cert-manager's dynamic solver pods Same context deadline exceeded symptom as the hairpin trap above and easy to conflate with it — check both. A default-deny-ingress policy (podSelector: {}) with only named-app allow-rules has no rule matching the solver pod cert-manager creates per challenge (labelled acme.cert-manager.io/http01-solver: "true", unpredictable name), so its ingress traffic is silently dropped. Verified live 2026-08-23. Add an explicit allow rule scoped to the label, not a static pod/service name: podSelector: {matchLabels: {acme.cert-manager.io/http01-solver: "true"}}, ingress: [{}]. The solver only ever serves a static .well-known/acme-challenge/<token> response, so open ingress on it is not a real exposure.

PDB-induced drain stalls

When a PDB is too strict for the replica count, node drain stalls indefinitely with events like FailedDrain and Cannot evict pod ... violates PodDisruptionBudget. Diagnose with kubectl get events --field-selector reason=FailedDrain -A and kubectl get pdb -A (inspect ALLOWED DISRUPTIONS). Mitigation, in order:

  1. Scale the workload replicas up so the PDB's minAvailable / maxUnavailable math allows at least one eviction.
  2. Temporarily widen the PDB (maxUnavailable: 1, or lower minAvailable) - commit the change to Git, do not edit live.
  3. Last resort: kubectl drain <node> --disable-eviction --force and accept the brief unavailability.

Never delete the PDB without a restore plan committed in Git - losing the PDB removes the protection the next maintenance window will rely on.

OOMKilled Triage Runbook

OOMKilled is the most common runtime container failure on AKS. The container exceeded its memory limit and the Linux kernel OOM killer terminated it. Step-by-step triage:

  1. Identify the killed container.

    kubectl get pods -A --field-selector=status.phase!=Running | grep OOMKilled
    kubectl describe pod <pod-name> -n <namespace> | grep -A5 "Last State"
    # Look for: Reason: OOMKilled and Exit Code: 137
  2. Check configured limits vs actual usage.

    kubectl top pod <pod-name> -n <namespace> --containers
    kubectl get pod <pod-name> -n <namespace> -o jsonpath='{.spec.containers[*].resources}'
  3. Check node-level OOM events.

    # On the node (via kubectl debug or kubectl exec into node-debug pod)
    dmesg | grep -i oom | tail -20
    journalctl -u kubelet | grep -i oom | tail -20
  4. Review Managed Prometheus / Container Insights metrics. Check container_memory_working_set_bytes vs container_spec_memory_limit_bytes over the past 24–48 hours. A rising trend approaching the limit with sudden drops is the classic OOMKill signature.

  5. Mitigate. Choose one:

    • Increase the container resources.limits.memory if the workload legitimately needs more memory (base this on measured P99 usage, not guesses).
    • Use VPA recommendations (kubectl get vpa -n <namespace> -o yaml | grep target) to get a data-driven limit suggestion.
    • Fix a memory leak in application code (watch container_memory_working_set_bytes over time; a monotonically rising line that never plateaus is a leak, not a limit problem).
  6. Prevent recurrence. Set resources.requests.memory and resources.limits.memory explicitly on every container. Use VPA in Off mode to generate recommendations. Add a Managed Prometheus alert: container_memory_working_set_bytes / container_spec_memory_limit_bytes > 0.9 for 5 minutes.

kubectl Tooling for Live Cluster Inspection

kubectl neat strips the noisy managed-fields, resourceVersion, and kubectl.kubernetes.io/last-applied-configuration from kubectl get -o yaml output, producing a clean manifest that is easy to review and diff:

# Install via krew
kubectl krew install neat

# Use
kubectl get pod <pod-name> -n <namespace> -o yaml | kubectl neat
kubectl get deployment <name> -n <namespace> -o yaml | kubectl neat > deployment-clean.yaml

Other high-value kubectl plugins for AKS production operations:

Plugin Install Purpose
krew krew.sigs.k8s.io Plugin manager; install first
neat kubectl krew install neat Strip managed fields from YAML output
tree kubectl krew install tree Show owner references as a tree (e.g., Deployment → ReplicaSet → Pod)
view-secret kubectl krew install view-secret Base64-decode and display Kubernetes Secrets cleanly
kubelogin Azure/kubelogin Required for Azure AD (Entra ID) kubeconfig auth; convert kubeconfig with kubelogin convert-kubeconfig
ctx + ns kubectl krew install ctx ns Fast context/namespace switching (kubectx/kubens)

kubelogin is not a krew plugin but is required for any cluster using Azure RBAC + Entra ID login; include it in developer onboarding documentation.

Stop Conditions

Stop and confirm before approving operations-resilience controls, or before promoting a cluster to production, if any of the following is true:

  • Ingress / Gateway controller has been chosen without confirmed support status (GA/preview), TLS/DNS ownership, and a migration path off any bridge controller (managed NGINX, self-managed NGINX).
  • Managed Prometheus, ContainerLogV2, and a baseline alert set are not configured, or audit logs are not flowing to Microsoft Sentinel (or equivalent SIEM) with immutable retention.
  • An AKS auto-upgrade channel is set, but no aksManagedAutoUpgradeSchedule maintenance window exists, or windows are routinely overridden so upgrades skip more than one cycle.
  • A --node-os-upgrade-channel other than NodeImage is configured for production without a documented reason (SecurityPatch requires documented support, maintenance-window ownership, and explicit acceptance of its faster patch cadence; Unmanaged/None is not acceptable for production).
  • A Kubernetes minor upgrade is being planned without running the Pre-Upgrade API Deprecation Inventory (pluto / kube-no-trouble), the Admission/Conversion Webhook checklist, and the CRD storage-version migration check.
  • Add-ons or operators (Flux, Argo CD, KEDA, Cilium/ACNS, NGINX, AGC, cert-manager, Istio, Kyverno, Gatekeeper, Ratify) have not been verified against the target Kubernetes minor's compatibility matrix.
  • No recovery design exists for the case where a Kubernetes minor upgrade fails: parallel cluster build, restore-from-backup, or node-pool snapshot path is not documented and rehearsed.
  • Backups (etcd/cluster resources, persistent volumes, ACR/Key Vault/DNS dependencies, GitOps repo) exist but restore has never been tested end-to-end.
  • GitOps reconciliation is configured, but break-glass access is not time-bound, audited, and reconciled back into Git after use.
  • IaC pipeline lacks pre-merge gates (tfsec/checkov, Azure Policy what-if, drift detection) or has destructive-change protection disabled for production state.
  • CVE response runbook for AKS Security Bulletins is not in place, or az fleet autoupgradeprofile generate-update-run (multi-cluster) / az aks nodepool upgrade --node-image-only (single cluster) is not pre-approved as the emergency node-image path.
  • Multi-region or DR pattern is claimed but data plane (database primary location, traffic failover, image/secret availability in the failover region) has not been validated.

Resolve each condition (or capture an explicit ADR exception with expiry) before treating the operations posture as production-ready.

Drasi on AKS Considerations

When running Drasi on AKS:

  • Enable OIDC issuer and Workload Identity where Azure resource access is required.
  • Do not assume az aks get-credentials exec-auth kubeconfig works inside Drasi source containers; static service-account-token based kubeconfig may be required for Kubernetes source integration.
  • Store kubeconfig material in Kubernetes Secrets and scope RBAC tightly.
  • Plan for Dapr sidecar and actor behaviour during source crashes; document recovery steps.
  • Use AKS observability to monitor source, query, reaction, and Dapr component health.

Production Readiness Checklist

  • Ingress/Gateway choice has GA/preview status and limitations documented.
  • TLS and DNS ownership are explicit.
  • cert-manager installation method and version are documented; ClusterIssuer is configured against the correct ACME endpoint (staging vs production).
  • Certificate renewal monitoring is in place (Prometheus CertificateExpiresSoon and CertificateNotReady alerts or equivalent).
  • Azure DNS Workload Identity integration is tested (if using DNS-01); DNS Zone Contributor RBAC and federated identity credential are verified.
  • cert-manager compatibility with the current AKS Kubernetes minor is confirmed against the cert-manager supported versions matrix.
  • Runtime validation commands exist for every deployment.
  • Production namespaces have ResourceQuota, LimitRange, NetworkPolicy, Pod Security labels, and owner/cost/data labels.
  • Managed Prometheus, ContainerLogV2, dashboards, and alerts are configured.
  • Upgrade and node OS maintenance windows are configured.
  • Backup and restore are tested.
  • Multi-region design aligns with data dependencies and RTO/RPO.
  • Fleet Manager decision is documented when more than one cluster is in scope, including hubless vs hub, update stages, placement, and preview gates.
  • Cost controls cover egress, logs, gateways, autoscaling, namespace quotas, fleet hub/cross-cluster networking, and GPU pools.
  • Cost attribution is in place (AKS Cost Analysis add-on or Kubecost) with namespace/team showback and budget alerts.
  • HA design is validated by fault injection (zone loss, node loss, dependency outage), not assumed; findings fed back into PDB/replica/topology settings.
  • Cluster certificate rotation path is documented and expiry is monitored.
  • Break-glass process exists and reconciles through GitOps/IaC.

Use With

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The skill is a comprehensive architecture and configuration guide for Azure Kubernetes Service (AKS). it emphasizes security best practices, including RBAC, NetworkPolicy, workload identity, and kernel-level isolation for AI agents. No malicious patterns or security risks were detected.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-26"
}
Other metadata
argument-hint
workload=<type>; region=<azure-region>; availability=<SLO>; network=<hub-spoke|standalone>; scope=<new-cluster|production-review|fleet>

README badge

README badge for lukemurraynz/hve-agent-skills/aks-cluster-architecture