All skills
lukemurraynz avatar

/aks-cluster-architecture

@2cc2455

AKS cluster architecture decisions for new Azure Kubernetes Service projects: AKS Automatic vs Standard, networking topology, dual-stack (IPv4/IPv6), Kubernetes version and OS currency, node pool strategy, identity, production NetworkPolicy, namespaces, autoscaling, ingress and Gateway API, observability, operations, resilience, GPU and AI workloads, GPU partitioning (MIG, time-slicing, MPS), batch scheduling (Kueue), AKS on bare metal, AI Runway and KAITO model serving, AKS MCP server access, kars (Agent Reference Stack for Kubernetes) for agent isolation, Kata MicroVM pod sandboxing, Azure Kubernetes Fleet Manager, multi-cluster governance, update orchestration, resource placement, cross-cluster networking, and cost. WHEN: designing new AKS clusters, reviewing production readiness, choosing CNI or outbound connectivity, planning node pools, defining namespace, network and security controls, evaluating Fleet Manager, deploying AI agent runtimes on AKS, or making hard-to-reverse infrastructure decisions.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/aks-cluster-architecture

This session only. Nothing lands on disk.

bundlesproduction-workload-controlsguide.md

≈17k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Production Workload Controls Bundle

Workload-scope guardrails: namespace contract, NetworkPolicy YAML, ResourceQuota/LimitRange, image supply chain, replicas/PDB/topology spread, HPA/KEDA/VPA examples. For cluster-level node pools, identity, and GPU pool design, see workload-platform.

Load this bundle for production namespace design, Kubernetes NetworkPolicy, Cilium/ACNS policy considerations, replica strategy, PodDisruptionBudgets, topology spread, HPA/KEDA/VPA/NAP interactions, ResourceQuota, LimitRange, and workload guardrails.

<!-- toc --> <!-- /toc -->

Assume new AKS projects. Prefer least-privilege, label-driven, GitOps/IaC-managed controls. Treat the YAML below as implementation patterns to adapt, not blind copy/paste manifests.

Production Workload Contract

Every production workload namespace should have an explicit contract:

Area Minimum production expectation
Namespace boundary One bounded app/service/domain per namespace unless a platform standard says otherwise.
Labels app.kubernetes.io/*, environment, owner, cost-center, data-classification, and network-policy=enabled where policy is enforced.
Access Namespace-scoped RBAC groups; no broad developer access to kube-system, gatekeeper-system, flux-system, ingress, or monitoring namespaces.
Security Pod Security Admission labels set to restricted where possible, or baseline with a documented exception.
Resources ResourceQuota and LimitRange in every team/app namespace.
Network Default deny ingress and egress, followed by explicit allow policies.
Availability Production stateless services run at least two replicas, usually three across zones when the app can support it.
Disruption PDBs, topology spread, readiness probes, and graceful termination before node upgrades or autoscaling are trusted.
Scaling HPA/KEDA min/max, cluster autoscaler/NAP limits, and VPA mode are designed together.

Namespace Model

Use namespaces as deployment, ownership, policy, and cost boundaries. Do not use namespaces as the only security boundary for hostile multi-tenancy.

Recommended labels:

apiVersion: v1
kind: Namespace
metadata:
  name: payments-prod
  labels:
    app.kubernetes.io/part-of: payments
    environment: prod
    owner: payments-team
    cost-center: cc-1234
    data-classification: confidential
    network-policy: enabled
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/audit: restricted
    pod-security.kubernetes.io/warn: restricted

Use separate namespaces for:

  • application workloads by domain/team/environment
  • platform components such as ingress, monitoring, GitOps, policy, secrets, and service mesh
  • high-risk or high-noise workloads that need separate quotas or RBAC
  • regulated data workloads where audit, egress, and network policy differ

Avoid:

  • one giant shared apps namespace for all production systems
  • namespace names that hide ownership or environment
  • allowing developers to create arbitrary namespaces without policy/bootstrap controls
  • relying on namespace isolation without RBAC, network policy, quotas, and admission policy

Namespace Resource Controls

Add ResourceQuota and LimitRange early. These are reversible, but retrofitting them after teams deploy workloads can break releases.

Example namespace quota:

apiVersion: v1
kind: ResourceQuota
metadata:
  name: namespace-quota
  namespace: payments-prod
spec:
  hard:
    requests.cpu: "20"
    requests.memory: 80Gi
    limits.cpu: "40"
    limits.memory: 160Gi
    pods: "80"
    services.loadbalancers: "0"
    persistentvolumeclaims: "20"

Example default requests and bounds:

apiVersion: v1
kind: LimitRange
metadata:
  name: default-container-resources
  namespace: payments-prod
spec:
  limits:
    - type: Container
      defaultRequest:
        cpu: 100m
        memory: 128Mi
      default:
        cpu: 500m
        memory: 512Mi
      min:
        cpu: 50m
        memory: 64Mi
      max:
        cpu: "2"
        memory: 4Gi

Recommendations:

  • Quotas should reflect SLO and cost budgets, not arbitrary round numbers.
  • Use separate quotas for GPU, batch, and stateful namespaces.
  • Do not set tiny default requests that hide under-provisioning and create noisy-neighbour behaviour.
  • Use VPA recommender data and production telemetry to tune quotas monthly.

Node allocatable vs capacity

A node's allocatable resources are not its full VM size. AKS reserves CPU and memory for the kubelet, container runtime, and OS, plus a hard-eviction memory threshold, so a node advertised as 4 vCPU / 16 GiB schedules noticeably less. This is the usual answer to "the node looks empty but my pod is Pending," and it skews every quota, bin-packing, and HPA-headroom calculation.

  • Memory reservation is regressive (a larger percentage on small nodes, tapering on large ones) and includes a fixed hard-eviction threshold; CPU reservation is also regressive. The exact formula changes across AKS releases, [VERIFY] the current reservation table in Microsoft Learn rather than hard-coding percentages.
  • Size ResourceQuota and node-pool capacity against allocatable, not VM spec. Check it directly: kubectl get node <node> -o jsonpath='{.status.allocatable}' vs {.status.capacity}.
  • Fewer large nodes waste proportionally less to reservation than many small nodes, but concentrate blast radius, weigh against zone spread and PDB math.
  • Set pod requests from real usage (VPA recommender), because requests, not limits, drive scheduling against allocatable and the cluster-autoscaler/NAP provisioning signal. Over-requesting strands allocatable capacity; under-requesting causes noisy-neighbour eviction.
Cost and capacity accounting model

Track capacity in four layers so rightsizing and autoscaling decisions are based on scheduler reality, not VM marketing numbers:

Layer Meaning What to use it for
Provisioned Raw VM SKU capacity you pay for Infrastructure spend baseline
Allocatable Capacity left after AKS/system reservations Scheduler and node-pool headroom
Requested Sum of pod requests Bin-packing and autoscaler signal quality
Consumed Real workload usage (P50/P95) Rightsizing and quota tuning

If these four numbers are not reviewed together, teams usually optimize the wrong bottleneck.

Image Supply Chain and Admission

Production workloads should treat the container image and the admission layer as a single security boundary. The defaults below are illustrative - confirm current Microsoft support status and feature availability against Microsoft Learn and your registry/admission tooling before implementation.

Build pipeline and image hardening

The image that lands in ACR should already be minimal, reproducible, and accompanied by build-time evidence. Treat the build pipeline as the first admission gate.

  • Multi-stage Dockerfiles. Use a builder stage for compilers, package managers, and test tooling; copy only the final binary, static assets, and runtime dependencies into the runtime stage. Build tooling (gcc, npm, pip, curl, package caches) must not appear in the published layer.
  • Distroless or minimal base images. Prefer Mariner-distroless, Microsoft chiseled Ubuntu, or Google distroless for the runtime stage to remove shells, package managers, and unused libraries. Confirm CVE feed coverage, organisational support, and debugging tooling (ephemeral debug containers, kubectl debug) before committing to a specific distro - distroless images intentionally exclude sh and standard utilities.
  • Reproducible builds. Use BuildKit with --output type=oci (or docker buildx build --provenance=true) so the build emits an OCI image plus provenance metadata. Pin base images by digest in FROM statements (FROM mcr.microsoft.com/...@sha256:...) so a moving tag cannot silently change the base layer between builds.
  • SBOM generation per image. Generate an SBOM for every image with Syft or the Microsoft SBOM tool, publish it in SPDX or CycloneDX format, and push it as an OCI artefact alongside the image in ACR. The admission layer can then verify the SBOM is present before allowing deploy.
  • in-toto attestations / SLSA. State a target SLSA level for the build process - SLSA Build L2 is a reasonable starting target for new projects (version-controlled build, hosted build platform, signed provenance); SLSA Build L3 for regulated workloads (hardened build platform, non-falsifiable provenance). SLSA is a build-side control that documents how an image was produced; it complements the Notation/Cosign signature verification described below, which is the admission-side control that proves which image is allowed to run. Do not treat one as a substitute for the other.

Image provenance and pull

  • Pin images by immutable digest (@sha256:...) in production manifests; tags can be mutated and are not a supply-chain control on their own.
  • Use Azure Container Registry (ACR) with Workload Identity for image pull (AcrPull role on the kubelet identity or a workload-scoped identity). Do not use admin credentials or shared imagePullSecrets for new clusters.
  • Place ACR behind a Private Endpoint with private DNS, and disable public network access for regulated workloads. Allow pull only from the AKS subnets that need it.
  • Enable image vulnerability scanning via Microsoft Defender for Containers and configure Defender CSPM/agentless scanning on the registry and the runtime; alert on critical/high CVEs with a documented remediation window.
  • Set imagePullPolicy: IfNotPresent as the production default (not Always). Always forces a registry round-trip on every pod start, adding latency proportional to registry response time and increasing the blast radius of a registry outage. IfNotPresent uses the node-cached image when a matching digest is present, which is safe when images are pinned by digest (the digest is the cache key). Use Always only in development or when you are explicitly testing image-refresh behaviour.

Image pull performance and registry resilience

ImagePullBackOff and slow rollouts at scale are usually registry-side, not cluster-side. Design the pull path, not just the image:

  • ACR throttling. Azure Container Registry enforces per-registry read/write operation limits. A large rollout, a node-image refresh, or a thundering-herd scale-out can hit them, surfacing as intermittent ImagePullBackOff across many pods at once. Use the Premium ACR tier for higher limits and throughput, stagger large rollouts, and alert on registry throttling metrics. [VERIFY] current ACR throttling limits per tier.
  • ACR geo-replication. For multi-region clusters, geo-replicate the Premium registry so each region pulls from a local replica, this cuts pull latency and cross-region egress and removes a single-region registry as a failure domain for every cluster. Pair with regional Private Endpoints.
  • Artifact Streaming (Preview). For large images (common with GPU/AI and model-serving workloads), Artifact Streaming pulls only the layers needed for startup, cutting time-to-pod-ready substantially for images up to ~30 GB. For a file larger than ~30 GB that the pod needs at start, mount it as a volume instead of baking it into a layer. [VERIFY] Preview status and feature flag. This is the registry-layer answer to the GPU cold-start pitfall in operations-resilience - Known Pitfalls.
  • Pull mechanics. Pin by digest (above), keep runtime images small (multi-stage + distroless, above), and pull with Workload Identity / AcrPull rather than long-lived pull secrets so credential rotation never stalls a rollout.

Admission-time signature verification

  • For regulated workloads, sign release images with Notation (CNCF/Notary v2) or Cosign at build time. Store signatures in ACR alongside the image.
  • Enforce signature verification at admission with Ratify (or an equivalent admission webhook) so unsigned images cannot be deployed. Treat the webhook as a critical-path component: configure HA, monitor health, define break-glass. Ratify recently moved under the Notary Project (notaryproject/ratify, CNCF Sandbox); the example below targets the current config.ratify.deislabs.io/v1beta1 API, with v2alpha1 in development - pin to v1beta1 in production today and re-check the API group, support state, and project velocity against the Notary Project repository before adoption.
  • When Ratify or Notation is not yet acceptable, use Azure Policy / Deployment Safeguards to constrain which registries are allowed and to deny :latest or untagged images.

Tie SBOM generation to admission by emitting the SBOM as a Cosign attestation in CI and requiring Ratify to verify it alongside the image signature. This closes the loop between the SBOM artefact pushed to ACR and the runtime admission gate.

# CI emits an SBOM attestation alongside the signed image.
syft <registry>/payments-api:<tag> -o spdx-json > sbom.spdx.json
cosign attest --key azurekms://<vault>/<key> --type spdx \
  --predicate sbom.spdx.json <registry>/payments-api@sha256:<digest>
# Ratify verifier policy — require both a Notation signature AND an SPDX SBOM attestation.
apiVersion: config.ratify.deislabs.io/v1beta1
kind: Verifier
metadata:
  name: notation-and-sbom
spec:
  policy:
    artifactTypes:
      - application/vnd.cncf.notary.signature
      - application/spdx+json
    requiredVerifiers:
      - notation
      - sbom

Admission then rejects images that lack both the signature and an attested SBOM, closing the gap between build-side SLSA provenance and runtime enforcement.

Signing-key custody is the load-bearing control here - a compromised signing key undoes every downstream admission check. Make the following non-negotiable for production signing:

  • Signing keys MUST live in Azure Key Vault Premium (HSM-backed). Software-protected keys, exported PFX files, and developer workstation keystores are not acceptable for production signing.
  • Notation: use the notation-azure-kv plugin so signing operations call into Key Vault and the private key never leaves the HSM. Cosign: use --key azurekms://<vault>/<key> so the signature is produced by Key Vault, not by local material.
  • No exported PFX. No GitHub Actions long-lived secrets. No Service Principal client secrets. A leaked SP secret with Key Vault Crypto User is equivalent to a leaked signing key.
  • The pipeline authenticates to Azure via OIDC workload-identity federation for the signing step. The signing identity is a workload-identity-federated pipeline identity scoped to Key Vault Crypto User on the signing key only - not Key Vault Administrator, not subscription-wide, not reusable for non-signing operations.
  • Define a key-rotation cadence (annual at minimum, more often for high-assurance workloads) and rehearse it before the first rotation is forced by an incident. Verify the Ratify trust policy supports multiple active keys (current and previous signing keys both trusted during a rollover window) so rotation does not require a fleet-wide redeploy in a single window.

VEX exceptions for known-not-affected CVEs

Scanners frequently flag CVEs in base layers that do not apply to the running workload - the vulnerable code path is not reached, or a runtime mitigation neutralises the issue. Treating every such finding as a blocker drowns the real signal. Use OpenVEX or CSAF-VEX documents, signed and attested alongside the image (the same Cosign attestation pattern as the SBOM), to suppress those findings both at admission and in the vulnerability dashboard. VEX statements must carry an explicit expiry date (max 90 days) and a named owner; treat an expired VEX as a policy violation, not a silent pass. Trivy, Grype, and Defender for Containers all consume VEX - pick one consumer and document it as the source of truth. For new AKS projects, Trivy with --vex is a reasonable default because it runs equally in CI and as a Kubernetes operator, so the same VEX document gates the build and re-scans running workloads.

Manifest validation in CI

Workload guardrails are only as strong as the pipeline that enforces them. For every production workload repository, run pre-merge validation:

  • Schema: kubeval or kubeconform against the target Kubernetes minor.
  • Policy: conftest with Rego, Kyverno CLI, or Gatekeeper policies covering: required labels (owner, environment, cost-center, data-classification), resource requests present, image pinned by digest, securityContext.runAsNonRoot=true, readOnlyRootFilesystem=true, no privileged containers, PDB defined for production Deployments, NetworkPolicy present in the namespace.
  • Best-practice scan: Polaris or kube-score as an advisory gate (informational, not blocking) to catch regressions.
  • Secret scanning: gitleaks or trufflehog against the manifests, Helm values, and Kustomize overlays - fail the build on any finding; rotate-and-replace before merge.

ValidatingAdmissionPolicy (CEL-based, GA from Kubernetes 1.30)

ValidatingAdmissionPolicy (VAP) is a native Kubernetes admission control mechanism that evaluates Common Expression Language (CEL) rules inline in the API server, without requiring a separate webhook server. For simple structural rules, VAP replaces or supplements OPA/Gatekeeper ConstraintTemplates.

apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
  name: disallow-latest-tag
spec:
  failurePolicy: Fail
  matchConstraints:
    resourceRules:
      - apiGroups: ["apps"]
        apiVersions: ["v1"]
        operations: ["CREATE","UPDATE"]
        resources: ["deployments","statefulsets","daemonsets"]
  validations:
    - expression: >
        object.spec.template.spec.containers.all(c,
          !c.image.contains(':latest') && c.image.contains('@sha256:'))
      message: "Container images must be pinned by digest (@sha256:...) and must not use ':latest'."
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
  name: disallow-latest-tag-binding
spec:
  policyName: disallow-latest-tag
  validationActions: [Deny]
  • Good candidates for VAP: disallow-latest-tag, require-resource-limits, require-labels, deny-privileged-containers, require-runasnonroot.
  • Use Gatekeeper/Kyverno for rules that need complex data lookups (external data, Rego functions), multi-resource correlation, or mutation. VAP is evaluate-only (no mutation support in standard VAP).
  • MutatingAdmissionPolicy is in alpha as of 1.32; do not use for production.
  • VAP validations expressions that fail produce events visible via kubectl get events; add these to your admission audit runbook.
  • For mutating webhook ordering (reinvocation): set reinvocationPolicy: IfNeeded on mutating webhooks whose output could be changed by a subsequent webhook; this is particularly important when multiple operators (Istio, Dapr, secrets-store) inject into the same pod spec.
  • Treat the same Kyverno/Gatekeeper policy bundle as the source of truth for cluster admission too - keep CI and runtime policies in sync via GitOps so a manifest that passes CI also passes admission.

The matrix below summarises which tools belong at which stage and what they typically fail on. Use it to avoid duplicating an admission control as advisory in CI, or leaving an advisory check as the only line of defence.

Tool Stage Gate type Typical failures
kubeval / kubeconform CI Blocking Schema mismatch against target Kubernetes minor, deprecated API versions.
conftest (Rego) CI Blocking Required labels missing, mutable image tag, missing resource requests.
Kyverno CI + Admission Blocking Same policy bundle in both stages - single source of truth via GitOps.
Gatekeeper (OPA) Admission Blocking ConstraintTemplate violations enforced at admission.
Ratify Admission Blocking Unsigned image, missing SBOM attestation, untrusted signing key.
Azure Policy / Deployment Safeguards Admission Blocking or Audit Managed control library - registry allow-list, restricted PSA, image-tag controls.
Polaris CI Advisory Best-practice deviations, securityContext misses, anti-pattern surfaces.
kube-score CI Advisory Manifest-quality regressions, probe/limit/requests hygiene.
Trivy / Defender for Containers CI + Runtime Blocking on critical CVEs in CI, alerting at runtime Critical/high CVEs in image layers; runtime detections feed Sentinel.
Syft (SBOM) CI Artefact (not a gate) Generates the SBOM that downstream gates (Ratify, attack-path analysis) consume.
gitleaks / trufflehog CI Blocking Hard-coded credentials, connection strings, private keys committed to manifests or Helm values.

Runtime security baseline

The Deployment skeleton later in this guide already sets runAsNonRoot, readOnlyRootFilesystem, allowPrivilegeEscalation: false, capabilities.drop: [ALL], and seccompProfile.type: RuntimeDefault. Reinforce these via:

  • Pod Security Admission namespace labels - prefer restricted for app namespaces, baseline only with a documented exception.
  • Defender for Containers runtime threat detection where available - review the alert taxonomy and route to the SOC channel.
  • Azure Policy (or Gatekeeper/Kyverno) to enforce runAsNonRoot, readOnlyRootFilesystem, restricted capabilities, and seccomp RuntimeDefault cluster-wide except for explicitly exempted namespaces.

seccompProfile: RuntimeDefault — non-negotiable for all production containers. The default Kubernetes seccompProfile before 1.27 was Unconfined (all syscalls allowed). On modern AKS, RuntimeDefault is a GA feature that restricts the pod's syscall surface to what containerd considers a safe default profile, with zero application code changes required. Always set it explicitly; do not rely on defaults that vary by Kubernetes minor:

securityContext:
  seccompProfile:
    type: RuntimeDefault
  runAsNonRoot: true
  runAsUser: 1000
  allowPrivilegeEscalation: false
  readOnlyRootFilesystem: true
  capabilities:
    drop: ["ALL"]

allowPrivilegeEscalation: false — always set for application containers. Even when the container runs as non-root, omitting allowPrivilegeEscalation: false leaves the door open for setuid/setcap binaries inside the image to escalate. This field defaults to true (for backward compatibility) unless set explicitly. It is a required field in PSA restricted mode, but set it explicitly in your pod spec so the intent is auditable regardless of the enforced PSA level.

Defender for Containers vs Defender CSPM. These are two distinct Defender plans covering different lifecycle stages - enable both for regulated production workloads.

Plan Where it runs What it covers
Defender for Containers In-cluster runtime sensor (eBPF on Linux nodes) Alerts on suspicious process, network, and file events at runtime; route alerts to Sentinel for triage.
Defender CSPM Agentless, control-plane / cloud-resource scanning Vulnerability scanning of ACR images, attack-path analysis across cloud resources (image -> cluster -> identity -> data); no in-cluster agent required.

Container escape and sandboxing. The default container runtime (containerd) is shared-kernel: every container on a node shares the host kernel. Container escape via a kernel CVE or a misconfigured capability is the dominant runtime-isolation risk on a stock AKS node pool. Use isolation tiers appropriate to each workload's threat model:

Tier Mechanism Isolation boundary Use when
Baseline PSA restricted, seccomp RuntimeDefault, drop all caps, AppArmor where supported Linux security controls on shared kernel All production workloads - required, not optional
Pod Sandboxing runtimeClassName: kata-vm-isolation (AKS Kata Containers, GA) Lightweight VM per pod - kernel escape → hypervisor escape Untrusted-code execution, hostile multi-tenancy, or shared-node workloads needing VM-level isolation
Confidential VM node pools (AMD SEV-SNP) Whole-node confidential VM pool on CVM-capable SKUs Hardware-rooted encrypted memory at the node; operator and hypervisor cannot read node memory Regulated workloads (HIPAA, PCI, financial data), AI inference on sensitive data, workloads with contractual confidentiality requirements. Per-pod AKS Confidential Containers (Kata + SEV-SNP) was a preview that has been retired (node images removed March 31, 2026) — use Confidential VM node pools for new work; see note below.

Pod Sandboxing (Kata Containers) - declare on the pod spec:

spec:
  runtimeClassName: kata-vm-isolation

Verify supported VM instance families (D-series with nested virtualisation), networking mode (Azure CNI Overlay is required; Azure CNI with pod subnet has constraints), Windows pool exclusion, and feature GA status against Microsoft Learn before standardising. Kata containers have longer start times than standard containers - factor into startup probes and HPA cooldowns.

⚠ Deprecation notice (verified 2026-06-16 against Microsoft Learn). Per-pod AKS Confidential Containers (preview) (Kata + AMD SEV-SNP) has been retired; node images were removed on March 31, 2026, after which affected node pools can no longer scale. Do not adopt per-pod Confidential Containers for new AKS projects. For "data in use" protection on new clusters, use Confidential VM node pools (AMD SEV-SNP), the whole-node pattern, which remains supported. Existing per-pod confidential-container workloads need a migration plan. See cluster-foundations - Confidential Compute for the cluster-level decision and the deprecation banner.

Confidential VM node pools (AMD SEV-SNP) - the supported "data in use" pattern for new clusters; a dedicated node pool using CVM-capable VM SKUs (DCasv5, ECasv5 families) where the whole node runs as a confidential VM. (The retiring per-pod Confidential Containers preview used a confidentialComputingAddon and a Kata confidential RuntimeClass; do not build new designs on it.) Key design decisions before adoption:

  • Attest the Trusted Execution Environment (TEE) using the Microsoft Azure Attestation service or an approved equivalent; build attestation verification into the admission or startup flow.
  • Encrypted memory incurs a performance overhead (~10–20% throughput reduction is common); benchmark under realistic load before production sizing.
  • GPU passthrough into a confidential pod has additional constraints - validate NVIDIA CVM support and AKS preview status before committing.
  • Volume encryption and key release must be designed explicitly; use Azure Key Vault with attestation-gated key release policies for data at rest within the pod.
  • Debugging tooling (kubectl exec, ephemeral debug containers) is significantly restricted in confidential pods by design - plan your observability and troubleshooting runbook before deployment.

Treat exceptions to baseline controls (privileged sidecars, hostPath mounts, host network) as ADR-worthy with a named owner and expiry date. Node pool selection for Pod Sandboxing or Confidential VM node pools is a cluster-architecture decision; capture it as an ADR before workload deployment.

See also: operations-resilience - CVE Response Decision Matrix for the patch-and-redeploy reaction when a signed image's base layer turns up a CVE.

Network Policy Strategy

Kubernetes starts with flat pod networking. For production, use least-privilege policy:

[!IMPORTANT] Azure Network Policy Manager (NPM) is being deprecated. Azure NPM is the legacy network policy engine for AKS. Microsoft has announced its intent to deprecate NPM and recommends all customers transition to Cilium Network Policy. For new AKS clusters on Linux, Cilium is the only supported network policy engine (requires Azure CNI Powered by Cilium). For existing clusters using Azure NPM, plan migration to Cilium. For Windows node pools, AKS does not natively support Kubernetes NetworkPolicy; use Calico (--network-policy calico) as the supported option for Windows workloads.

  1. Add default deny ingress and egress per workload namespace.
  2. Allow DNS egress explicitly.
  3. Allow ingress only from the ingress/Gateway namespace or approved client namespaces.
  4. Allow east-west pod traffic by namespace and pod labels, not pod IPs.
  5. Allow egress to Azure PaaS through Private Link/private IPs or Cilium/ACNS FQDN policy where supported.
  6. Test policy in non-prod using runtime connectivity checks before enforcing broadly.

For new Linux AKS clusters, prefer Azure CNI Powered by Cilium where supported. Cilium is the preferred AKS direction for richer policy, performance, and capabilities, but verify engine support for the target cluster mode, OS pools, AKS Automatic/Standard choice, and network policy settings before implementation.

Important cautions:

  • Kubernetes NetworkPolicy is namespaced and additive. A pod is restricted for ingress or egress only when selected by at least one policy for that direction.
  • Standard Kubernetes NetworkPolicy is L3/L4. Use Cilium policy / ACNS policy for FQDN, HTTP, or L7 patterns when supported and approved.
  • --enable-acns alone enables only FQDN filtering. L7 policy (HTTP/gRPC method, path, header rules) requires the additional --acns-advanced-networkpolicies L7 flag on az aks create/az aks update; the L7 value is a superset that also covers FQDN policy. A CiliumNetworkPolicy containing rules.http/rules.grpc is silently ignored when L7 is not enabled, there is no error, the rule simply does not enforce.
  • L7 rules are not supported in CiliumClusterwideNetworkPolicy (CCNP) — only in namespaced CiliumNetworkPolicy. Placing an L7 rule in a CCNP yields silent no-enforcement. Keep L7 policy namespaced.
  • Avoid ipBlock rules for pod or node IPs on Cilium-based clusters; use selectors for in-cluster traffic.
  • LoadBalancer source/destination rewrites can affect how external traffic appears to NetworkPolicy. Validate the actual path.
  • Do not apply default deny to platform namespaces such as kube-system without a platform-specific policy design.

See also: cluster-foundations - Azure CNI Powered by Cilium for engine selection and CIDR planning that precedes NetworkPolicy authoring.

Network Policy Examples

Every YAML block below is illustrative - adapt labels, namespaces, ports, and selectors to your workload before applying. Validate with runtime connectivity tests in non-prod first.

1. Default deny all ingress and egress in a workload namespace

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: default-deny-all
  namespace: payments-prod
spec:
  podSelector: {}
  policyTypes:
    - Ingress
    - Egress

2. Allow DNS egress to CoreDNS

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-dns-egress
  namespace: payments-prod
spec:
  podSelector: {}
  policyTypes:
    - Egress
  egress:
    - to:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: kube-system
          podSelector:
            matchLabels:
              k8s-app: kube-dns
      ports:
        - protocol: UDP
          port: 53
        - protocol: TCP
          port: 53

3. Allow ingress controller or Gateway namespace to reach the public API pods

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-ingress-to-api
  namespace: payments-prod
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: payments-api
  policyTypes:
    - Ingress
  ingress:
    - from:
        - namespaceSelector:
            matchLabels:
              kubernetes.io/metadata.name: ingress-system
      ports:
        - protocol: TCP
          port: 8080

4. Allow frontend pods to call API pods

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-frontend-to-api
  namespace: payments-prod
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: payments-api
  policyTypes:
    - Ingress
  ingress:
    - from:
        - podSelector:
            matchLabels:
              app.kubernetes.io/name: payments-web
      ports:
        - protocol: TCP
          port: 8080

5. Allow API pods to reach a database service in the same namespace

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-api-to-db
  namespace: payments-prod
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: payments-db-proxy
  policyTypes:
    - Ingress
  ingress:
    - from:
        - podSelector:
            matchLabels:
              app.kubernetes.io/name: payments-api
      ports:
        - protocol: TCP
          port: 5432

6. Allow egress to an approved private endpoint IP

Use this only for external/private endpoint IPs, not pod or node IPs on Cilium-based clusters.

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: allow-egress-to-private-sql
  namespace: payments-prod
spec:
  podSelector:
    matchLabels:
      app.kubernetes.io/name: payments-api
  policyTypes:
    - Egress
  egress:
    - to:
        - ipBlock:
            cidr: 10.42.8.15/32
      ports:
        - protocol: TCP
          port: 1433

7. Cilium/ACNS FQDN egress pattern

Use CiliumNetworkPolicy or ACNS network policy only after confirming all of the following: the cluster is running Azure CNI Powered by Cilium (not legacy CNI + Calico); ACNS is enabled with the correct tier for the policy type, --enable-acns alone enables FQDN filtering, while L7 policy (and FQDN policy) requires --enable-acns --acns-advanced-networkpolicies L7 on both az aks create and az aks update; the cilium.io/v2 API is exposed (kubectl api-resources --api-group=cilium.io); and your platform team owns Cilium policy lifecycle and ingestion of Hubble/ACNS observability. Mixed-OS clusters with Windows pools have additional constraints - validate Windows policy enforcement separately. Prefer Private Link to Azure PaaS over FQDN egress where the service supports it. Known failure mode (AKS issue #4525): an FQDN egress policy combined with a default-deny NetworkPolicy that does not also allow CoreDNS egress causes Cilium DNS lookups to crash-loop pods that depend on the policy - always pair FQDN egress rules with an explicit DNS allow rule, and validate the path in non-prod first.

# Illustrative — verify cilium.io API version and ACNS configuration in your cluster before applying.
apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: allow-selected-fqdns
  namespace: payments-prod
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: payments-api
  egress:
    - toFQDNs:
        - matchName: login.microsoftonline.com
        - matchPattern: "*.vault.azure.net"
      toPorts:
        - ports:
            - port: "443"
              protocol: TCP

8. Cilium L7 HTTP policy (requires ACNS)

Use Cilium L7 policies when you need to control traffic based on HTTP methods, paths, headers, or gRPC/Kafka application protocols rather than just IP and port. L7 policies require ACNS (Advanced Container Networking Services) enabled with L7 enforcement on: az aks update --enable-acns --acns-advanced-networkpolicies L7 (also applies to az aks create). --enable-acns alone enables only FQDN filtering. L7 rules (and FQDN policy) require the additional --acns-advanced-networkpolicies L7 flag; the L7 value is a superset that also covers FQDN. Without it, the rules.http/rules.grpc blocks below are silently ignored. These policies work only with namespaced CiliumNetworkPolicy (not standard Kubernetes NetworkPolicy, and not CiliumClusterwideNetworkPolicy. L7 rules are not supported in cluster-wide policies). Verify cilium.io/v2 API availability and ACNS L7 support before adopting.

apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: allow-frontend-get-to-api
  namespace: payments-prod
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: payments-api
  ingress:
    - fromEndpoints:
        - matchLabels:
            app.kubernetes.io/name: payments-web
      toPorts:
        - ports:
            - port: "8080"
              protocol: TCP
          rules:
            http:
              - method: "GET"
                path: "/api/.*"

9. Cilium L7 gRPC policy (requires ACNS)

apiVersion: cilium.io/v2
kind: CiliumNetworkPolicy
metadata:
  name: allow-grpc-list-orders
  namespace: payments-prod
spec:
  endpointSelector:
    matchLabels:
      app.kubernetes.io/name: payments-api
  ingress:
    - fromEndpoints:
        - matchLabels:
            app.kubernetes.io/name: payments-web
      toPorts:
        - ports:
            - port: "50051"
              protocol: TCP
          rules:
            grpc:
              - service: "orders.OrderService"
                method: "ListOrders"

Validating NetworkPolicy with Cilium Hubble

Before enforcing a default-deny policy in production, observe current traffic flows to identify which connections need explicit allow rules. Hubble provides real-time flow visibility on clusters with Azure CNI Powered by Cilium and ACNS enabled.

# Observe live traffic flows in a namespace (run in non-prod first)
kubectl exec -n kube-system ds/cilium -- hubble observe -f --namespace <namespace>

# Show only dropped/denied traffic (reveals policy gaps)
kubectl exec -n kube-system ds/cilium -- hubble observe --namespace <namespace> --verdict DROPPED

# Show flows between specific pods by label
kubectl exec -n kube-system ds/cilium -- hubble observe \
  --namespace <namespace> \
  --label app.kubernetes.io/name=payments-api

# List all Cilium endpoints and their identities
kubectl exec -n kube-system ds/cilium -- cilium endpoint list

Use Hubble flow data to build your allowlist before writing policies, then re-run after enforcement to confirm no legitimate traffic is blocked. Pair this with the runtime connectivity tests in Validation Commands for a complete picture.

RBAC model for NetworkPolicy management

Separate policy ownership between platform and application teams to reduce misconfiguration risk:

Role Scope Can create/modify NetworkPolicy for
Platform security team All namespaces Baseline default-deny policies, platform namespace policies, cluster-wide egress rules
Application team Own namespace only App-specific allow rules within their namespace boundary
Cluster administrators All namespaces CiliumClusterwideNetworkPolicy (cluster-scoped, L3/L4 only — L7 rules are not supported in CCNP, see caution above), troubleshooting override

Example RBAC binding granting app-team policy management in their namespace:

apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
  name: network-policy-manager
  namespace: payments-prod
rules:
  - apiGroups: ["networking.k8s.io"]
    resources: ["networkpolicies"]
    verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
  - apiGroups: ["cilium.io"]
    resources: ["ciliumnetworkpolicies"]
    verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: RoleBinding
metadata:
  name: payments-team-policy-manager
  namespace: payments-prod
subjects:
  - kind: Group
    name: payments-team
    apiGroup: rbac.authorization.k8s.io
roleRef:
  kind: Role
  name: network-policy-manager
  apiGroup: rbac.authorization.k8s.io

This split prevents app-team changes from affecting other namespaces while still allowing teams to self-serve within their own boundary. Platform-owned baseline policies (default-deny, DNS allow, ingress-controller allow) should be applied by a GitOps pipeline or cluster-wide admission control, not delegated to namespace scopes.

Layered AKS ingress and east-west security (AGC + WAF + Cilium L7)

For internet-facing workloads, the most resilient posture layers three independent controls so each catches what the others miss. No single layer is sufficient; the layers are complementary, not redundant.

Layer Control Catches Misses (why you need the next layer)
1. North-south front door Application Gateway for Containers (AGC) via Gateway API Undesired hostnames/paths never reaching the cluster; TLS termination; routing Payload attacks on permitted requests
2. WAF on the gateway Azure WAF bound to AGC through a WebApplicationFirewallPolicy CRD referencing an ARM WAF policy (OWASP DRS 2.2 recommended) [VERIFY DRS version] Payload-level attacks on permitted routes — SQLi, XSS, protocol anomalies — before traffic reaches pods Cannot see east-west traffic; cannot enforce app-specific method/path contracts
3. Pod-level L7 + east-west/egress ACNS Cilium L7 CiliumNetworkPolicy + default-deny + DNS carve-out Method/path/Header violations, lateral movement, unauthorized egress — inside the cluster Cannot inspect deep payloads (that is the WAF's job)

The complementarity is the point: the WAF (layer 2) inspects the body of an HTTP request whose method and path the Cilium L7 policy (layer 3) already permitted, catching SQLi/XSS that a method/path allowlist cannot see. Conversely, the Cilium L7 policy catches method/path violations and east-west lateral movement between pods that the WAF, operating only on north-south ingress at the gateway, never observes. A pod compromised through an allowed path still cannot fan out east-west because the default-deny and L7 ingress policies block it.

Building the layered posture from the existing examples in this guide:

  • Default-deny baseline + DNS carve-out: examples 1 and 2.
  • Allow ingress only from the AGC/ingress namespace: example 3.
  • Per-pod HTTP method/path enforcement (AGC has routed the request, now constrain what the pod accepts): example 8.
  • gRPC method enforcement for service-to-service calls: example 9.
  • FQDN egress allowlist so compromised pods cannot exfiltrate to arbitrary endpoints: example 7.
  • WAF on the gateway via WebApplicationFirewallPolicy referencing an ARM WAF policy (OWASP DRS 2.2): see operations-resilience - AGC WAF.

Prerequisites for the L7 layers: Azure CNI Powered by Cilium, ACNS enabled with --acns-advanced-networkpolicies L7, and Hubble wired into your observability pipeline to confirm each layer is actually enforcing (dropped flows should appear at the layer that denies them).

Replica and Availability Rules

Use replicas to express availability, not just load handling.

Workload type Starting replica guidance Notes
Dev/test stateless 1-2 Keep cost low, but test production rules elsewhere.
Production stateless API/web Minimum 2, prefer 3 across zones 3 gives better disruption and rollout tolerance.
Critical stateless services 3+ with topology spread and PDB Validate zone failure and node drain.
Singleton workers 1 only when singleton semantics are required Implement leader election for warm standby; see guidance below.
Queue/event workers HPA/KEDA min 0-2 depending on cold-start tolerance Do not scale to zero if SLO requires warm capacity.
Stateful workloads Follow app quorum/storage rules Do not blindly set replicas without storage and failover design.
GPU inference Min 1+ if cold start is expensive Pair with KEDA/HPA only after metrics and quota validation.

Rules:

  • Do not set production stateless workloads to one replica unless the service is explicitly non-critical or protected by another HA layer.
  • Use odd replica counts such as 3 for quorum-like services when the application requires it; otherwise size from SLO and load evidence.
  • Set Deployment rolling update settings intentionally. For high availability APIs, prefer maxUnavailable: 0 with a controlled maxSurge when capacity allows.
  • PDBs protect against voluntary disruption, not crashes, OOM, bad releases, or zone loss.
  • Set restartPolicy: Always (the default) for production workloads. Only use Never or OnFailure with a documented operational reason, each restart policy changes how the kubelet handles container exit codes, which interacts differently with HPA scale-down, PDB eviction math, and node drain behaviours.

Example production Deployment skeleton:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: payments-api
  namespace: payments-prod
  labels:
    app.kubernetes.io/name: payments-api
    app.kubernetes.io/part-of: payments
spec:
  replicas: 3
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0
      maxSurge: 25%
  selector:
    matchLabels:
      app.kubernetes.io/name: payments-api
  template:
    metadata:
      labels:
        app.kubernetes.io/name: payments-api
        app.kubernetes.io/part-of: payments
    spec:
      serviceAccountName: payments-api
      # Set to false unless the pod needs Kubernetes API access or a projected identity token.
      automountServiceAccountToken: false
      terminationGracePeriodSeconds: 30
      securityContext:
        runAsNonRoot: true
        seccompProfile:
          type: RuntimeDefault
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels:
              app.kubernetes.io/name: payments-api
      containers:
        - name: app
          image: <registry>/payments-api:<tag>
          securityContext:
            allowPrivilegeEscalation: false
            readOnlyRootFilesystem: true
            capabilities:
              drop:
                - ALL
          ports:
            - containerPort: 8080
          resources:
            requests:
              cpu: 250m
              memory: 512Mi
            limits:
              memory: 1Gi
          readinessProbe:
            httpGet:
              path: /health/ready
              port: 8080
            periodSeconds: 10
          livenessProbe:
            httpGet:
              path: /health/live
              port: 8080
            periodSeconds: 20

Security notes for the skeleton:

  • Keep automountServiceAccountToken: false unless the pod needs Kubernetes API access or a projected token for Workload Identity; validate Azure Workload Identity webhook injection when the workload authenticates to Azure.
  • Use a writable emptyDir or mounted volume for paths that must be writable when readOnlyRootFilesystem: true is enabled.
  • Add startupProbe for slow-starting services, especially Java, large images, or model-serving workloads, so liveness probes do not kill legitimate startup.
  • Pin images by immutable digest for production where the organisation supports it.

Leader election for singleton workloads

Some components (schedulers, controllers, single-writer adapters) cannot run more than one active instance. A single replica without a standby is a cold standby, on failure the new pod starts from scratch, and availability drops during the start-up window. Leader election converts this to a warm standby: two (or more) pods run, but only the leader serves; the standby is ready and waiting to acquire the lease.

Kubernetes-native leader election uses a Lease object in coordination.k8s.io/v1:

apiVersion: coordination.k8s.io/v1
kind: Lease
metadata:
  name: my-scheduler-lock
  namespace: payments-prod
spec:
  holderIdentity: null
  leaseDurationSeconds: 15
  renewDeadlineSeconds: 10
  acquireDelaySeconds: 5

The application acquires the lease via the client-go leaderelection package (or a framework-native equivalent). Key design rules:

  • Lease duration should be short (15–30 seconds) so failover is fast.
  • Use the same ServiceAccount and RBAC so standby pods can read/write the Lease.
  • Both pods must pass readiness probes. The standby is ready but idles, readiness proves it could serve, not that it is serving.
  • Do not rely on leader election for crash tolerance. PDBs, health probes, and anti-affinity still apply to both pods.
  • Test lease expiry and re-acquisition. A network partition that isolates the leader from the apiserver (but not from clients) can produce a split-brain window, design for this in the application logic.

When leader election is too heavyweight, consider queue partitioning (each worker owns a shard of the work queue) or idempotent workers (any instance can process any work item without coordination).

Pre-stop hooks and graceful shutdown

A preStop lifecycle hook runs before the SIGTERM signal is sent to the container. Its primary HA purpose is to remove the pod from the service endpoint before the process terminates, so in-flight requests complete without being dropped. This is critical during HPA scale-down events, node drains, and rolling updates, especially when the application does not natively handle SIGTERM gracefully.

lifecycle:
  preStop:
    exec:
      command: ["/bin/sh", "-c", "sleep 30"]

Key guidance:

  • Paired with terminationGracePeriodSeconds. The grace period must exceed the preStop hook duration plus expected drain time. If the hook runs for 30 seconds and the application takes up to 10 seconds to shut down, set terminationGracePeriodSeconds: 45 or higher.
  • What the hook does under the hood. Kubernetes marks the pod as Terminating, removes it from the Service's Ready endpoints (so new requests stop arriving), runs the preStop hook, then sends SIGTERM. After terminationGracePeriodSeconds expires, SIGKILL is sent.
  • When to use it. Any production pod that serves traffic and does not implement a SIGTERM handler. Even when the app handles SIGTERM, a short preStop sleep (1–5 seconds) can act as a buffer for DNS and proxy caches to converge.
  • When not to use it. Jobs, batch workers, or pods that own their drain logic via a SIGTERM handler and terminationGracePeriodSeconds alone.

Example PDB:

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: payments-api
  namespace: payments-prod
spec:
  maxUnavailable: 1
  selector:
    matchLabels:
      app.kubernetes.io/name: payments-api

Automatic PDB Management (Preview)

AKS cluster extension microsoft.evictionautoscaler (based on the open-source eviction-autoscaler; installed via az k8s-extension create --extension-type microsoft.evictionautoscaler) has two behaviours: it auto-creates PodDisruptionBudget resources for Deployments that do not yet have one, and it temporarily scales up replicas when an existing PDB blocks eviction on a cordoned node during a drain, then scales back down after a cooldown (default 60s). Deployments only; StatefulSets are not supported and using the extension with them is not recommended.

# Recommended: targeted auto-protection (namespace-scoped)
az k8s-extension create \
  --resource-group <rg> \
  --cluster-name <cluster> \
  --cluster-type managedClusters \
  --extension-type microsoft.evictionautoscaler \
  --name eviction-autoscaler \
  --release-train stable \
  --configuration-settings controllerConfig.pdb.create=true controllerConfig.namespaces.actionedNamespaces="{kube-system,production}" \
  --auto-upgrade-minor-version true

Prerequisites: managed-identity cluster (extensions do not work with service-principal clusters), Azure CLI 2.64.0+, Microsoft.KubernetesConfiguration provider registered. The extension installs its own service account and RBAC to read/write Deployments, PDBs, and Events in the namespaces it manages.

Design considerations:

  • What the auto-created PDB actually is. minAvailable is set to the deployment's current replica count (or, under HPA/KEDA, the autoscaler's minimum replica floor), and it is continuously reconciled as replicas change. Consequence: eviction is blocked until the extension surges replicas; that is by design, not a misconfiguration.
  • How the surge coordinates with autoscalers. Deployment-only: raises replicas directly. With HPA: raises the HPA minReplicas floor (so HPA cannot scale the surge away mid-drain). With KEDA: raises minReplicaCount on the ScaledObject (bypassing the KEDA→HPA→deployment sync lag). Deployment + KEDA + a separate HPA is unsupported; the extension marks its EvictionAutoScaler status Degraded and skips the surge; remove the duplicate HPA to fix.
  • Capacity is a precondition. The surge pods need somewhere to schedule; without headroom, cluster autoscaler or NAP scale-up time joins the upgrade critical path. Plan pool headroom (or CA/NAP) before relying on this for drain reliability.
  • Only voluntary disruptions. Upgrade drains and manual cordon/drain , not node failures or pod crashes.
  • Scope control. actionedNamespaces is fixed at install; changing any configuration requires deleting and recreating the extension ; treat the initial namespace selection as a platform decision. Runtime opt-in/out per namespace: eviction-autoscaler.azure.com/enable: "true"/"false" annotation; per-deployment opt-out of PDB creation: eviction-autoscaler.azure.com/pdb-create: "false". Deployments with a non-zero maxUnavailable rolling-update strategy are skipped automatically. Auto-created PDBs carry ownedBy: EvictionAutoScaler and are garbage-collected with the deployment/namespace/extension; strip the annotation to take manual control. Manually created PDBs are never modified.
  • Not a substitute for deliberate PDB design. The auto-generated budget tracks replica counts; it does not know your SLO. For workloads with strict availability requirements, define explicit PDBs with tuned minAvailable/maxUnavailable values and ownership labels.
  • Preview gate. [VERIFY] current preview/GA status before relying on it for production drain operations. See Automatic PDB management.

Autoscaling Strategy

Design pod autoscaling, node autoscaling, and disruption policy together.

Need Recommendation Cautions
Steady web/API traffic HPA on CPU plus app/custom metrics where useful Requires accurate requests and Metrics API.
Bursty queue/event workloads KEDA with explicit min/max and cooldown Validate scaler auth, poison messages, and cold starts.
Right-sizing CPU/memory VPA recommender Off mode first Avoid automatic mutation with HPA on same CPU/memory target until tested.
Static node pools Cluster autoscaler with explicit min/max Avoid zero-min for critical pools unless startup SLO allows it.
Heterogeneous workloads NAP/Karpenter Requires accurate requests, scheduling constraints, and supported networking/policy combination.
GPU inference KEDA/HPA only with reliable model and GPU metrics Cold start and quota can dominate scaling behaviour.

HPA example:

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
  name: payments-api
  namespace: payments-prod
spec:
  scaleTargetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: payments-api
  minReplicas: 3
  maxReplicas: 12
  metrics:
    - type: Resource
      resource:
        name: cpu
        target:
          type: Utilization
          averageUtilization: 65
  behavior:
    scaleDown:
      stabilizationWindowSeconds: 300
    scaleUp:
      stabilizationWindowSeconds: 60

VPA recommender example:

apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: payments-api
  namespace: payments-prod
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: payments-api
  updatePolicy:
    updateMode: "Off"

KEDA example pattern using Azure Workload Identity rather than connection-string secrets:

apiVersion: v1
kind: ServiceAccount
metadata:
  name: payments-worker
  namespace: payments-prod
  annotations:
    azure.workload.identity/client-id: <managed-identity-client-id>
---
apiVersion: keda.sh/v1alpha1
kind: TriggerAuthentication
metadata:
  name: payments-worker-keda-auth
  namespace: payments-prod
spec:
  podIdentity:
    provider: azure-workload
    # Optional when the ServiceAccount annotation supplies the client ID.
    # identityId: <managed-identity-client-id>
---
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
  name: payments-worker
  namespace: payments-prod
spec:
  scaleTargetRef:
    name: payments-worker
  minReplicaCount: 1
  maxReplicaCount: 20
  pollingInterval: 30
  cooldownPeriod: 300
  triggers:
    - type: azure-servicebus
      metadata:
        queueName: <queue-name>
        namespace: <service-bus-namespace>
        messageCount: "50"
      authenticationRef:
        name: payments-worker-keda-auth

Autoscaling guardrails:

  • Never enable HPA/KEDA without resource requests, metrics availability, scaler authentication, and max replica/cost limits.
  • Prefer Workload Identity for Azure scaler authentication; use Kubernetes secrets only when the scaler lacks identity support and document rotation ownership.
  • With TriggerAuthentication podIdentity.provider: azure-workload, do not set authModes: "bearer" on the trigger - KEDA then expects a static bearer-token secret that does not exist and the ScaledObject stays READY=False ("bearer token is required when bearer auth is enabled"). Leave authModes unset; the pod-identity provider handles token acquisition.
  • Keep HPA/KEDA minimum replicas aligned to SLO and cold-start tolerance.
  • Do not configure HPA max replicas higher than node autoscaling or downstream dependencies can absorb.
  • Watch for scale thrash: rapid scale up/down, long startup probes, slow image pulls, or queue lag.
  • For NAP, resource requests and scheduling constraints are the provisioning signal; inaccurate requests produce bad nodes.
  • Validate drain behaviour with PDBs before relying on cluster autoscaler, NAP consolidation, upgrades, or blue-green node upgrades.

See also: workload-platform - Autoscaling Decision Matrix for the node-pool/NAP side of the same decision.

Debugging Production Workloads with Ephemeral Containers

Distroless and minimal base images intentionally exclude shells, package managers, and debugging utilities. The correct way to debug such containers in production without breaching the image supply chain is kubectl debug with ephemeral containers.

# Attach an ephemeral debug container to a running pod (shares the pod's namespaces)
kubectl debug -it <pod-name> -n <namespace> \
  --image=mcr.microsoft.com/azurelinux/base/core:3.0 \
  --target=<main-container-name>

# Clone a pod with modifications (useful for crash-loop debugging)
kubectl debug -it <pod-name> -n <namespace> \
  --image=mcr.microsoft.com/azurelinux/base/core:3.0 \
  --copy-to=debug-pod --share-processes
  • Ephemeral containers are GA from Kubernetes 1.25. They are injected into a running pod and share the pod's network, PID, and IPC namespaces (with --target), enabling netstat, curl, strace, and process inspection without modifying the running image.
  • Approve a set of debug images in the cluster image allow-list (Ratify/Azure Policy). Using an unapproved image in a debug session is a supply-chain exception - treat it as one.
  • Ephemeral containers are visible in kubectl get pod -o yaml under .status.ephemeralContainerStatuses. They are not deleted on command exit; run kubectl delete pod <debug-pod> (for copies) or wait for the ephemeral container to reach Terminated.
  • Do not use kubectl exec into distroless containers as a substitute; the shell is not present. Ephemeral containers are the sanctioned pattern.
  • For node-level debugging (kernel, kubelet, containerd): use kubectl debug node/<node-name> -it --image=mcr.microsoft.com/azurelinux/base/core:3.0; this runs a privileged pod on the named node with access to the host filesystem at /host.

Validation Commands

# Namespace and policy inventory
kubectl get ns --show-labels
kubectl get resourcequota,limitrange -A
kubectl get networkpolicy -A
kubectl describe networkpolicy -n <namespace>

# Runtime connectivity tests
kubectl run network-debug -n <namespace> --rm -it --image=<approved-network-debug-image> -- /bin/sh
kubectl exec -n <namespace> deploy/<deployment> -- nslookup kubernetes.default.svc.cluster.local
kubectl exec -n <namespace> deploy/<deployment> -- curl -sS http://<service>.<namespace>.svc.cluster.local:<port>/health

# Replica, disruption, and autoscaling
kubectl get deploy,statefulset,hpa,pdb -A
kubectl describe hpa -n <namespace> <hpa-name>
kubectl get events -n <namespace> --sort-by=.metadata.creationTimestamp | tail -50
kubectl top pods -n <namespace>
kubectl top nodes

# Node placement and zone spread
kubectl get pods -n <namespace> -o wide
kubectl get nodes -L topology.kubernetes.io/zone,kubernetes.azure.com/agentpool,node.kubernetes.io/instance-type

Stop Conditions

Stop and confirm before approving production workload controls, or before enforcing them on a live namespace, if any of the following is true:

  • The namespace has no ResourceQuota and no LimitRange, or has placeholder values that do not reflect SLO and budget.
  • pod-security.kubernetes.io/enforce is unset, or restricted is requested for a namespace whose workloads are known to violate it (privileged sidecars, hostPath, host network) without a documented exception.
  • A default-deny NetworkPolicy is about to be enforced without runtime connectivity tests covering DNS, ingress, east-west, identity, telemetry, and approved egress paths in non-prod.
  • Production stateless workloads are deployed at one replica, or with no PodDisruptionBudget and no topology spread, without a documented singleton/availability exception.
  • HPA / KEDA is enabled without resource requests, accurate metrics, scaler authentication via Workload Identity (or a documented secret-rotation path), and reconciled maxReplicas against node and downstream capacity.
  • Image references are mutable tags (:latest, branch tags) rather than immutable digests, OR images are not scanned for vulnerabilities, OR admission does not enforce a registry allow-list.
  • CI pipeline does not validate manifests against the same policy set the cluster enforces (Kyverno/Gatekeeper drift between CI and admission).
  • VPA in Auto mode is being enabled on a workload that also has HPA on CPU or memory without a tested design.
  • A workload writes its own Kubernetes Secret with Azure connection strings or keys instead of using Workload Identity, with no rotation owner.

Resolve each condition (or capture an explicit ADR exception with expiry) before treating the workload as production-ready.

Review Checklist

  • Every production namespace has owner, environment, cost, data, security, and network labels.
  • Every production namespace has ResourceQuota and LimitRange or an approved exception.
  • Default deny ingress and egress exists for workload namespaces.
  • DNS egress and required ingress/Gateway paths are explicitly allowed.
  • App-to-app and app-to-data paths use selectors or approved private endpoints, not broad CIDRs.
  • Production stateless services have at least two replicas, usually three across zones.
  • Deployment strategy, readiness probes, PDBs, and topology spread are consistent.
  • HPA/KEDA/VPA/NAP responsibilities are documented and do not conflict.
  • Autoscaling max values align to node capacity, downstream limits, and budget.
  • NetworkPolicy and autoscaling behaviour is validated in non-prod with runtime tests.
  • Production images are pinned by immutable digest and pulled with Workload Identity from a private endpoint registry.
  • Pre-merge CI runs manifest validation (kubeval/kubeconform + Kyverno/Conftest) covering required labels, security context, and resource requests.
  • Admission-time policy (Gatekeeper, Kyverno, or Azure Policy / Deployment Safeguards) enforces the same rules as CI.

Use With

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The skill is a comprehensive architecture and configuration guide for Azure Kubernetes Service (AKS). it emphasizes security best practices, including RBAC, NetworkPolicy, workload identity, and kernel-level isolation for AI agents. No malicious patterns or security risks were detected.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-26"
}
Other metadata
argument-hint
workload=<type>; region=<azure-region>; availability=<SLO>; network=<hub-spoke|standalone>; scope=<new-cluster|production-review|fleet>

README badge

README badge for lukemurraynz/hve-agent-skills/aks-cluster-architecture