All skills
lukemurraynz avatar

/aks-cluster-architecture

@2cc2455

AKS cluster architecture decisions for new Azure Kubernetes Service projects: AKS Automatic vs Standard, networking topology, dual-stack (IPv4/IPv6), Kubernetes version and OS currency, node pool strategy, identity, production NetworkPolicy, namespaces, autoscaling, ingress and Gateway API, observability, operations, resilience, GPU and AI workloads, GPU partitioning (MIG, time-slicing, MPS), batch scheduling (Kueue), AKS on bare metal, AI Runway and KAITO model serving, AKS MCP server access, kars (Agent Reference Stack for Kubernetes) for agent isolation, Kata MicroVM pod sandboxing, Azure Kubernetes Fleet Manager, multi-cluster governance, update orchestration, resource placement, cross-cluster networking, and cost. WHEN: designing new AKS clusters, reviewing production readiness, choosing CNI or outbound connectivity, planning node pools, defining namespace, network and security controls, evaluating Fleet Manager, deploying AI agent runtimes on AKS, or making hard-to-reverse infrastructure decisions.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/aks-cluster-architecture

This session only. Nothing lands on disk.

bundlesworkload-platformguide.md

≈11k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Workload Platform Bundle

Cluster-level workload platform: node pools, VM/OS selection, identity binding, GPU pools, scheduling. For per-namespace YAML guardrails (NetworkPolicy, ResourceQuota, replicas, PDB, HPA examples), see production-workload-controls.

Load this bundle for node pool strategy, VM family decisions, system/user/spot/GPU pools, Workload Identity, Entra/RBAC, scheduling, bin packing, and workload placement trade-offs. For detailed production NetworkPolicy, namespace, replica, PDB, and autoscaling YAML examples, also load production-workload-controls.

<!-- toc --> <!-- /toc -->

Assume new AKS projects and prefer modern AKS-native capabilities. Do not preserve legacy node pool or identity patterns unless the task is explicitly a migration.

Node Pool Baseline

Pool type Default guidance Difficulty Notes
System pool Dedicated Linux system pool; no user workloads Difficult Use taints/labels to keep platform components isolated.
General user pool Separate pool for stateless app workloads Reversible to difficult Enable autoscaling or NAP.
Stateful pool Dedicated pool when storage, IO, or maintenance windows differ Difficult Align zones, disk type, PDBs, and upgrade sequencing.
Spot pool Optional for interruptible/batch workloads only Reversible Never place critical or stateful workloads here.
Windows pool Only when workload requires Windows containers Difficult Lifecycle dates differ from Linux.
GPU/AI pool Separate tainted accelerator pool Difficult Verify quota, supported VM SKUs, driver model, and monitoring.

Minimum expectations:

  • Set resource requests for every container before enabling HPA, KEDA, NAP/Karpenter, or bin-packing decisions.
  • Use topology spread constraints or anti-affinity for production replicas.
  • Use PodDisruptionBudgets for services that must survive node drains.
  • Define namespace ResourceQuota, LimitRange, NetworkPolicy, and Pod Security Admission before treating a workload as production-ready.
  • Keep system workloads and application workloads isolated.
  • System node pool minimum 2 nodes. A single system node is a single point of failure for kube-proxy, CoreDNS, and CNI. Scale to at least 2 nodes (--node-count 2 or autoscaler --min-count 2).
  • Treat node pool OS SKU, disk type, VM family, zones, and mode as disruptive decisions.
  • In-place node pool VM size change (Preview). az aks nodepool update --node-vm-size <new-sku> lets you resize the VM SKU of an existing VMSS node pool without recreating it. The operation uses the rolling upgrade engine (surge + cordon/drain + delete). Requires AKS API 2026-01-02-preview and the aks-preview CLI extension; blocked when --max-surge 0. Treat as a Difficult, not Permanent, decision for VM family ; plan for the drain window and PDB impact. [VERIFY] preview status and supported SKU transition matrix before committing. See Resize node pools.
  • Use v5 or newer VM SKU families for production node pools; avoid B-series burstable SKUs.
  • Prefer ephemeral OS disks where the VM SKU supports the OS disk size; validate per SKU family.
  • Enable Azure Key Vault CSI driver (addonProfiles.azureKeyvaultSecretsProvider.enabled: true) from day one, do not store secrets as Kubernetes Secrets in etcd.

Account for node allocatable vs capacity when sizing pools and bin-packing. AKS reserves CPU/memory/eviction headroom, so schedulable capacity is below VM spec. See production-workload-controls - Node allocatable vs capacity.

Windows node pools

Only add a Windows node pool when the workload needs Windows containers, the operational surface differs materially from Linux and is easy to underestimate:

  • No Linux system pool substitute. Windows pools are always user pools; AKS still requires a Linux system pool. You run a mixed-OS cluster, not a Windows-only one.
  • OS version + lifecycle. Pick the Windows Server version deliberately (e.g., Server 2022/2025); versions have independent end-of-support dates from Linux, and in-place OS-version moves mean node-pool replacement. Patch via the node OS upgrade channel just like Linux, but expect larger, slower node-image updates.
  • Image size and pull time. Windows container images are large; cold starts and rollouts are slower. Apply the registry-resilience guidance (ACR Premium/throttling, geo-replication) from production-workload-controls, it bites Windows harder.
  • NetworkPolicy. AKS does not support Cilium/Azure NetworkPolicy enforcement on Windows; use Calico (--network-policy calico) for Windows policy, and validate Windows policy separately from Linux.
  • Pod Security Admission gap. PSA restricted semantics are Linux-centric; many controls (seccomp, capabilities, readOnlyRootFilesystem) do not map cleanly to Windows. Use HostProcess containers only with explicit justification, and gMSA for Active Directory-integrated identity instead of secrets.
  • Scheduling. Use the kubernetes.io/os: windows nodeSelector and taints/tolerations so Linux DaemonSets and workloads never land on Windows nodes (and vice versa). Verify GPU/confidential/Pod-Sandboxing features, most are Linux-only.

[VERIFY] current Windows Server version support, NetworkPolicy options, and feature parity against Microsoft Learn before committing a Windows design.

VM and OS Selection

Workload shape Common starting point Notes
General web/API D-series or equivalent balanced VM Validate with load test and VPA recommendations.
Memory-heavy E-series or memory-optimised VM Watch request-to-usage ratio.
CPU-heavy F-series or compute-optimised VM Validate CPU throttling and HPA behaviour.
IO-heavy/stateful Premium-capable VM, planned disk/cache profile Confirm storage class, zone, and backup behaviour.
GPU inference/training Dedicated NVIDIA GPU pool, or managed GPU node pool where supported Verify region quota, driver support, model startup time, and scale policy.

Do not hard-code a VM SKU as a universal default. Recommend a starting family and require sizing evidence from performance tests, VPA, Azure Monitor, or production telemetry.

VM SKU selection rules:

  • Prefer v5 or newer VM SKU families for production node pools. Older generations lack performance, security (AMD SEV-SNP for confidential, vTPM), and feature parity. Verify SKU availability per region before committing.
  • Do not use B-series (burstable) VM SKUs for production node pools. B-series VMs accumulate CPU credits during idle periods and consume them under load. Sustained production traffic depletes credits, causing CPU throttling and unpredictable application latency. Use D-series or equivalent balanced VMs as the minimum starting point for any production workload.
  • Match OS disk type to workload I/O requirements. Use ephemeral OS disks (osDiskType: Ephemeral) for node pools where the VM SKU supports it (the VM's temp/cache disk size must be ≥ the OS disk size). Ephemeral OS disks provide lower write latency, reduced OS disk billing, and faster node image operations. For node pools where the VM SKU cannot accommodate ephemeral OS disks, use managed OS disks with Premium SSD if the workload is I/O-sensitive.
  • Set osDiskSizeGB explicitly when using managed (non-ephemeral) OS disks. The AKS default OS disk size varies by VM SKU and may be too small for workloads with large container images or aggressive image-pull concurrency. Size based on observed node-level disk usage rather than the default.

Prepared Image Specification (Preview, July 2026)

AKS Prepared Image Specification (public preview, 2026-07-29) lets you build node images preloaded with container images and initialisation tasks so nodes come up with workloads already on disk, aimed at large-scale, AI/GPU, and Windows workloads where node bootstrap/startup latency dominates rollout time. Source: Azure Updates (https://azure.microsoft.com/updates).

  • Preview gate. Confirm current preview/GA status, regional availability, and support boundary before committing; treat as non-production by default.
  • Trigger only on measured latency. Use when node startup/cold-start is a demonstrated bottleneck (large image sets, GPU models, Windows image size). For most workloads, ACR geo-replication + imagePullPolicy/registry-resilience tuning is sufficient, see production-workload-controls.
  • Keep it in IaC. A prepared image is still a node image; model its build pipeline and lifecycle like any node image artifact.

Identity and Access

For new clusters:

  • Enable OIDC issuer and Workload Identity at cluster creation.
  • Use user-assigned managed identities for workload access to Azure services when identity lifecycle must be controlled outside the pod.
  • Avoid AAD Pod Identity and service principal secrets.
  • Use Azure RBAC for Kubernetes authorization unless the organisation has a documented Kubernetes RBAC-only model.
  • Use namespace-scoped roles and managed namespaces or policy controls where useful.
  • Keep workload namespace ownership, quota, network policy, and Pod Security labels under Git/IaC control.
  • Use Key Vault CSI Driver or workload identity-based secret retrieval; do not bake secrets into manifests or images.
  • For large multi-cluster estates, consider AKS identity bindings only after checking current preview/GA status, tenant restrictions, support scope, and whether the simpler federated credential model is sufficient.

Validation:

az aks show --resource-group <rg> --name <cluster> --query "{oidc:oidcIssuerProfile, workloadIdentity:securityProfile.workloadIdentity}" --output yaml
kubectl auth can-i <verb> <resource> --namespace <namespace> --as <principal-or-group>

See also: cluster-foundations - Secrets, etcd Encryption, and Data-at-Rest for secrets storage and rotation, and cluster-foundations - API Server Hardening for anonymous-auth, PIM, Conditional Access, break-glass, and audit policy.

Service Account and Workload Identity Pattern

For workloads that access Azure services:

  • Use a dedicated Kubernetes ServiceAccount per workload or bounded application component.
  • Annotate the ServiceAccount with the managed identity client ID and label workload pods for Azure Workload Identity injection.
  • Grant the managed identity only the Azure RBAC roles needed at the narrowest practical scope.
  • Avoid sharing one broad managed identity across many namespaces or environments.
  • Set automountServiceAccountToken: false for pods that do not need Kubernetes API access or projected identity tokens.
  • For KEDA, distinguish the identity used by the scaler from the identity used by the application container when the least-privilege scopes differ.

Secretless Entra app authentication (app-registration federation)

Workload Identity also federates Entra app registrations (not only managed identities) to Kubernetes service accounts. Use it when an OIDC/OAuth middleware component must redeem tokens on its own behalf - e.g. oauth2-proxy fronting an app that cannot do modern authentication itself, with Istio external authorization (envoyExtAuthzHttp extension provider) delegating every protected request to the proxy. The result is a conventional OIDC sign-in for users and no long-lived Entra credential to store or rotate.

  • Create the app registration (az ad app create + az ad sp create) with the redirect URI known up-front, then federate it directly: az ad app federated-credential create --id <client-id> with the AKS OIDC issuer, subject system:serviceaccount:<ns>:<sa>, audience api://AzureADTokenExchange. Issuer, subject, and audience must match exactly - mismatches fail at token exchange while the pod looks healthy.
  • Do not run az ad app credential reset - it creates exactly the client secret this pattern removes. oauth2-proxy equivalent: --provider=entra-id --entra-id-federated-token-auth=true; the projected SA token is sent as the client assertion at the Entra token endpoint (authorization code flow with PKCE).
  • Restart deployments after changing SA annotations/labels or the client ID - the Workload Identity webhook configures pods only at creation; retrofitting requires kubectl rollout restart.
  • "Secretless" means no long-lived Entra credential, not zero sensitive state: the cookie-encryption secret remains sensitive runtime state with its own rotation owner.
  • Keep the auth boundary single-pathed: convert upstream services from LoadBalancer to ClusterIP so no second public path bypasses the gateway's external authorization policy, and keep policy path exclusions narrow (/oauth2/*, ACME challenge only). Authentication on one ingress path does not protect a second public path.
  • Istio mesh extension providers live in istio-shared-configmap-asm-<revision> in aks-istio-system: merge into the existing data.mesh, never apply a sample wholesale (it replaces the entire map). Extension tools are outside the Istio add-on support boundary.

Source: Azure Global Black Belt - Secretless Entra ID Authentication for AKS with Istio and oauth2-proxy (2026-07-29). Flags verified against official oauth2-proxy 7.15.x Microsoft Entra ID provider docs (2026-08-23): provider=entra-id and entra_id_federated_token_auth are current documented options (boolean, default false) and client_secret may be omitted when federated token auth is enabled.

Two further production traps from the same provider docs:

  • Group overage. Groups claims cover up to 200 memberships; beyond that oauth2-proxy calls Graph transitiveMemberOf, which needs delegated User.Read. Without scope=openid User.Read (and user/admin consent), users with 200+ groups silently authenticate with zero groups - breaking allowed_groups or any group-based authorization downstream.
  • Multi-tenant apps require oidc_issuer_url=https://login.microsoftonline.com/common/v2.0 plus insecure_oidc_skip_issuer_verification=true; restrict tenants explicitly with --entra-id-allowed-tenant (the provider still validates each token's issuer against the per-tenant template). Prefer single-tenant app registrations unless multi-tenancy is an explicit requirement.

Autoscaling Decision Matrix

Need Use Notes
Scale replicas from CPU/memory/custom metrics HPA Requires accurate requests, Metrics API, min/max replicas, and reliable metrics.
Scale from queue/event/external metrics KEDA Good for async, bursty workloads; validate scaler auth, cooldown, poison-message handling, and cold start.
Right-size requests/limits VPA recommender first Avoid automatic mutation for latency-sensitive production until tested.
Add/remove nodes for existing pools Cluster autoscaler or AKS-managed equivalent Keep min/max explicit and test PDB/drain behaviour.
Provision new node shapes dynamically NAP/Karpenter Best for heterogeneous workloads; verify CNI compatibility (Azure CNI Overlay/Cilium required - Calico and Azure CNI Pod Subnet are unsupported combinations), NetworkPolicy support, resource requests present, and az aks show --query nodeProvisioningProfile state before enabling.
GPU event-driven scaling KEDA + GPU metrics Validate DCGM/Managed Prometheus metrics, min replicas, cold start, and quota.

Rules:

  • Do not enable HPA on workloads without CPU/memory requests.
  • Do not run HPA and VPA automatic resource mutation on the same CPU/memory target without a documented design.
  • Use VPA recommendations as monthly rightsizing evidence even if you do not auto-apply them.
  • Use NAP when workload diversity makes static node pool design brittle, but still define scheduling constraints, namespace quotas, disruption budgets, and max-cost guardrails.
  • Artifact Streaming on NAP-managed pools is configured per AKSNodeClass (spec.artifactStreaming), not with the per-pool az aks nodepool update --enable-artifact-streaming path; it still requires a Premium-tier ACR attached to the cluster like any Artifact Streaming enablement. Verify current AKSNodeClass field surface before wiring it into provisioning templates.
  • For production stateless services, start with at least two replicas and usually three across zones unless an explicit singleton or cost exception exists.
  • For autoscaling YAML examples and guardrails, use production-workload-controls.

Autoscaler control-loop failure modes

Failure mode Typical trigger Practical mitigation
Early node scale-out with low real usage Inflated pod requests Fix requests first (VPA recommender + telemetry), then retune HPA/KEDA thresholds
Replica and node churn Aggressive HPA/KEDA scale-down with short stabilization Increase scale-down stabilization windows and cooldowns; verify PDB/drain behavior
Conflicting autoscaler decisions HPA and VPA auto-mutation on the same CPU/memory signal Keep VPA in Off/recommendation mode unless a tested design proves compatibility
Pods stay pending during burst or Spot loss Strict scheduling constraints and no fallback pool capacity Use explicit fallback scheduling, keep baseline reliable capacity, and validate max limits across HPA/KEDA + autoscaler/NAP

See also: production-workload-controls - Autoscaling Strategy for HPA/KEDA/VPA YAML examples and guardrails.

Scheduling and Placement

Use labels, taints, tolerations, affinity, and topology spread to express intent. Avoid relying on accidental placement.

Requirement Scheduling pattern
Keep platform components isolated System pool taint and no app toleration.
Keep GPU workloads on GPU nodes only GPU taint, workload toleration, nvidia.com/gpu request.
Keep critical services zone-resilient topologySpreadConstraints across topology.kubernetes.io/zone.
Isolate noisy tenants Dedicated pool labels and namespace policy.
Prefer but not require a pool preferredDuringSchedulingIgnoredDuringExecution.
Require a pool requiredDuringSchedulingIgnoredDuringExecution plus taints/tolerations.

For bin packing, make request accuracy the first control. Scheduler profile and autoscaler tuning are secondary.

topologySpreadConstraints and minDomains

topologySpreadConstraints with minDomains (GA from Kubernetes 1.28) prevents the scheduler from treating a spread as satisfied when fewer zones are available than the design requires. Without minDomains, a two-zone cluster that loses one zone silently considers a single-zone deployment fully spread.

topologySpreadConstraints:
  - maxSkew: 1
    topologyKey: topology.kubernetes.io/zone
    whenUnsatisfiable: DoNotSchedule
    minDomains: 3        # require all three zones to be present
    labelSelector:
      matchLabels:
        app: api-server
  • Set minDomains to the number of zones the cluster spans (typically 3 for production AKS in supported regions).
  • Pair with whenUnsatisfiable: DoNotSchedule (hard) for critical workloads; ScheduleAnyway for non-critical ones that accept degraded spread.
  • minDomains causes pods to stay Pending if fewer than minDomains zones are available; this is the intended behaviour for critical services, not a bug.

readinessGates for Traffic Readiness

Standard Kubernetes readiness probes signal whether a container is ready, but some controllers (Application Gateway for Containers, Application Load Balancer, external load balancers) need an additional external signal before routing traffic. readinessGates extend the pod readiness check with a named PodCondition managed by an external controller.

spec:
  readinessGates:
    - conditionType: "target-health.alb.networking.azure.io/2f5a15ef..."
  • AGC and the ALB controller inject readinessGates automatically on pods matching a BackendLBPolicy or HTTPRoute. If a pod is Running and the container is ready but shows 0/1 ready in kubectl get pods, check kubectl get pod <name> -o yaml | grep -A5 conditions; a failed readiness gate is the most common cause.
  • Do not remove readiness gates from pod specs manually; they are injected by the controller at runtime and removing them from the manifest will cause them to be re-injected on the next rollout.
  • Plan for readiness gates when designing rolling-deploy strategies: a deployment will not complete if readiness gates prevent pods from becoming Ready, even if the container is healthy.

GPU / AI Workloads

Use this section when designing inference, model serving, batch training, or accelerator-heavy workloads.

Default guidance:

  • Use dedicated GPU pools with taints and labels.
  • Validate subscription GPU quota and regional SKU availability before committing architecture.
  • Use AKS-managed GPU node pools where supported and appropriate; verify NVIDIA-only support and migration constraints.
  • Use automatic GPU driver installation unless you have a documented reason to manage drivers yourself.
  • GPU driver compatibility gate (verified 2026-08-11, AKS 2026-07-17): NVadsA10v5 and NCadsA10v4 node pools must run node image 202606.08.1 or later to stay compatible with NVIDIA v18.x host drivers; older node images move into an unsupported host/guest driver configuration (Azure/AKS #5875). Pin GPU pools to a current node image and include this check in the validation checklist.
  • NVIDIA RTX PRO 6000 Blackwell Server Edition GPU supported (AKS 2026-06-19). AKS supports the RTX PRO 6000 Blackwell Server Edition on Ubuntu node pools with the NVIDIA GRID driver. Relevant for workloads requiring professional-grade GPU rendering, virtualisation, or workstation compute. [VERIFY] VM size naming, regional quota, and driver version before designing around this SKU. Source: AKS 2026-06-19 release.
  • Secure Boot is now supported with GPUs on Azure Linux (AKS 2026-07-17): Trusted Launch (vTPM + Secure Boot) is no longer restricted to non-GPU pools on Azure Linux, and can now be toggled on existing Linux node pools.
  • Add GPU health monitoring, DCGM metrics, and workload-level SLOs.
  • Scale LLM serving fleets on user-facing latency signals, not just GPU utilisation. A GPU can be 60% utilised while TTFT p99 is unacceptable under long-context bursts. KEDA can query Azure Managed Prometheus (e.g., a histogram_quantile TTFT p99 query) with TriggerAuthentication podIdentity.provider: azure-workload, a UAMI granted Monitoring Data Reader on the Azure Monitor Workspace, and the trigger's authModes left unset; pair worker scaling with cluster autoscaler/NAP on the GPU pool so pods pending on GPUs trigger new nodes. Leave authModes: "bearer" unset - see production-workload-controls autoscaling guardrails. Source: GBB - TTFT-Driven Autoscaling for Disaggregated LLM Inference with NVIDIA Dynamo on AKS (2026-05-11).
  • Consider KAITO / the AKS AI toolchain operator for supported self-hosted open-source model serving, but verify current OS SKU, region, model, GPU, and AKS Automatic compatibility before production design.
  • Do not mix GPU and non-GPU workloads on the same pool unless there is a deliberate cost/utilisation design.
NUMA-Aware Scheduling for GPU and HPC Workloads

NUMA (Non-Uniform Memory Access) topology matters for latency-sensitive GPU and HPC workloads that span multiple NUMA nodes on a host, causing cross-NUMA memory traffic and higher latency.

  • topologyManagerPolicy on the kubelet controls how CPU and device resources are aligned to NUMA nodes. Valid values: none (default, no alignment), best-effort (try to align, but schedule even if alignment fails), restricted (require alignment; pods that cannot be aligned are rejected), and single-numa-node (all resources must fit in a single NUMA node).
  • For GPU workloads on large VMs (e.g., Standard_NC96ads_A100_v4, Standard_ND96asr_A100_v4) where PCIe topology matters, prefer restricted or single-numa-node to avoid cross-NUMA GPU-to-CPU transfers.
  • AKS node pools do not expose topologyManagerPolicy directly in the cluster API. Use a NodeConfig custom kubelet configuration (--kubelet-config) or AKS node pool kubeletConfig.topologyManagerPolicy property. [VERIFY] current AKS surface for kubeletConfig.topologyManagerPolicy.
  • Pair with CPUManagerPolicy: static if CPU pinning is also required (common for real-time or HPC workloads); this requires exclusive CPU requests (integer, no HPA on CPU).
Serverless model serving via Microsoft Foundry / Azure OpenAI (no GPU pool required)

Most LLM inference on AKS does not need a GPU pool: a Foundry (AIServices) or Azure OpenAI deployment is called over HTTPS from CPU pods. The pattern is zero-secret by default, the pod exchanges its projected service-account token for a Microsoft Entra token and calls the model endpoint; no API key is stored in Kubernetes.

  • Use a dedicated ServiceAccount per workload, annotated with azure.workload.identity/client-id (+ azure.workload.identity/tenant-id) and labeled azure.workload.identity/use: "true"; label the deployment pod template and selector with azure.workload.identity/use: "true".
  • Create the federated identity credential with subject system:serviceaccount:<namespace>:<service-account> and the cluster OIDC issuer (az aks show --query oidcIssuerProfile.issuerUrl -o tsv). Create it before rollout; retrofitting the annotation to running pods requires kubectl rollout restart.
  • AKS Automatic enables OIDC issuer + Workload Identity by default — no explicit --enable-oidc-issuer/--enable-workload-identity flags are needed on the Automatic SKU.
  • For Foundry model serving, call https://<resource>.services.ai.azure.com/models with azure-ai-inference + DefaultAzureCredential, and assign the managed identity the Cognitive Services User role at the AIServices account scope (not Cognitive Services OpenAI User, not the Foundry project roles). See microsoft-foundry. Serverless Model Inference Endpoint and Workload Identity.
  • Enforce Entra-only authentication on the Foundry resource: create the AIServices account with disableLocalAuth: true so API keys are rejected even if leaked - workload identity becomes the only auth path, not merely the preferred one.
  • New user-assigned managed identities replicate slowly: retry az role assignment create with a short sleep loop and verify before rollout.
  • When no local Docker is available (CI runners, thin clients), use ACR cloud build (az acr build) to compile and push the image without a local daemon.
  • Do not hardcode the model endpoint or deployment name in the image; pass via env (MODEL_ENDPOINT, MODEL_DEPLOYMENT) so promotion across environments only changes the manifest.

Worked example: chengliangli0918/aks-automatic-foundry. AKS Automatic + Foundry gpt-5.4-mini + workload identity, Streamlit chatbot, ACR cloud build.

Inference gateway for self-hosted LLM serving (AGC, Public Preview)

When self-hosted model servers (vLLM, etc.) are fronted by Application Gateway for Containers, the AGC inference gateway (Public Preview, June 2026) adds model-aware, load-aware routing on top of the standard Gateway API surface. It implements the Kubernetes Gateway API Inference Extension:

  • InferencePool — backend resource grouping model-server pods with an Endpoint Picker (EPP).
  • Endpoint Picker (EPP) — customer-provided extension that scores pods (queue depth, KV-cache utilization, prefix-cache affinity) and picks the target pod per request.
  • Body-Based Router (BBR) — AGC-managed request processor that reads the model field of OpenAI-compatible bodies and exposes X-Gateway-Model-Name for model-aware routing; no separate proxy tier to run.
  • InferenceObjective — request-serving objectives (e.g. priority) for requests sharing an InferencePool.

Deployment constraints (verified 2026-08-11 against AGC inference gateway docs):

  • Not supported via the AKS ALB add-on — install ALB Controller via Helm with --set albController.aiGateway=true (installs Inference Extension CRDs v1.3.1). BYO deployment model.
  • The EPP is on the request path; size and monitor it as a critical component; it needs fresh model-server metrics for good routing.
  • Pairs with AGC WAF for AI-traffic protection before requests reach model servers.
  • Scale model servers on inference signals (queue depth, KV-cache utilization) via HPA/KEDA; the EPP routes around saturated replicas.
  • Validate streaming (SSE) latency/timeout behavior and version-match Inference Extension CRDs to the extension release.
  • Route AI-ingress architecture decisions here; KAITO/AI Runway remains the model deployment path, this is the model traffic path.

See Configure the inference gateway for the end-to-end vLLM walkthrough.

GPU Partitioning (sharing a single physical GPU)

When multiple workloads must share a physical GPU, choose a partitioning strategy. The one-to-one (one pod = one GPU) default wastes capacity for workloads that do not fully consume a GPU.

Strategy Management Isolation Use for
Multi-Instance GPU (MIG) AKS-managed (or user-managed via GPU Operator) Hardware partitioning (dedicated SMs, memory, cache) Production multi-tenant GPU sharing on A100/H100/H200
Time-slicing User-managed via NVIDIA GPU Operator Software scheduling (no hard isolation) Experimentation with variable GPU loads
Multi-Process Service (MPS) User-managed via NVIDIA GPU Operator CUDA-level process multiplexing Low-latency, high-throughput inference

MIG is the only production-grade sharing strategy because it provides hardware-level isolation. Constraints:

  • MIG partitioning is static at the node-pool level - changing the profile requires node reprovisioning. AKS now validates the requested MIG instance-profile slice width against the VM SKU's capacity at request time (AKS release 2026-08-07), so unsupported profiles on lower-capacity GPU SKUs fail fast instead of being accepted and failing later.
  • Supported on A100, H100, and H200 series only; choose the MIG profile (1g.5gb, 2g.10gb, 3g.20gb, etc.) at pool creation.
  • AKS fully managed GPU node pools (preview) expose MIG via gpuProfile.nvidia.managementMode (Managed/Unmanaged) and gpuProfile.nvidia.migStrategy (None/Single/Mixed). [VERIFY] managed GPU node pool GA status before production design.
  • Time-slicing and MPS are user-managed through the NVIDIA GPU Operator and are not AKS-supported in the same way as MIG; treat them as experimentation-only unless the team owns the GPU Operator lifecycle.

See GPU node partitioning strategies and Create a managed MIG node pool (preview).

Batch and AI Training Scheduling (Kueue)

For batch training, fine-tuning, data processing, or any job that can tolerate a delayed start but cannot tolerate partial resource allocation, the default Kubernetes scheduler is insufficient. Use Kueue for admission control and fair-share queuing.

Kueue is a Kubernetes-native job queueing system with first-class AKS documentation. It uses a two-level model:

  • ClusterQueue (cluster-scoped) defines a shared resource pool (CPU, memory, GPU quota) and fair-sharing rules across cohorts.
  • LocalQueue (namespace-scoped) is the tenant-facing submission point; jobs target a LocalQueue, which routes to a ClusterQueue that decides admission.

When to introduce Kueue:

  • Multiple teams or tenants compete for finite GPU quota and need fair-share or priority-based admission.
  • Distributed training jobs (Ray, Kubeflow MPIJob, PyTorchJob) that need all resources available before starting (gang scheduling).
  • Batch inference or data pipelines that should queue behind higher-priority online serving workloads.
  • Integration with Ray/KubeRay - Microsoft Learn documents a Ray + Kueue pattern for distributed AI workloads on AKS, where Kueue handles admission control and Ray handles the compute runtime.

Kueue constraints to document:

  • Kueue is open-source software deployed on AKS - it is excluded from AKS SLA and Azure support. Plan community-support and upgrade ownership.
  • Kueue manages admission (whether a job starts); it does not replace the kube-scheduler for pod placement. Pair with node autoscaler or NAP/Karpenter so admitted jobs have nodes to land on.
  • Capacity-on-admission is now a first-class pattern: Kueue's ProvisioningRequest AdmissionCheck drives the AKS cluster autoscaler, so a suspended batch Job gets its nodes provisioned before admission instead of racing the autoscaler after unsuspension - the autoscaler satisfies the request atomically (best-effort-atomic-scale-up class), then Kueue admits. Wire a ProvisioningRequestConfig + AdmissionCheck into the ClusterQueue's admission strategy. See Provision capacity for Kueue batch jobs with the cluster autoscaler (docs added 2026-08-21).
  • For GPU workloads, model GPU quota in the ClusterQueue resourceGroups so Kueue accounts for accelerator capacity, not just CPU/memory.

See Kueue on AKS overview and Run Ray AI workloads with Kueue on AKS (the Ray+Kueue doc family: overview → infrastructure → queue configuration → workload examples).

Validation examples:

kubectl get nodes -L kubernetes.azure.com/agentpool,node.kubernetes.io/instance-type
kubectl describe node <gpu-node> | grep -i nvidia -A5
kubectl get pods -A -o wide --field-selector spec.nodeName=<node>
kubectl top pods -A
kubectl top nodes

See also: production-workload-controls - Image Supply Chain and Admission for signing GPU/AI runtime images and operations-resilience - Upgrade and Patch Strategy for GPU driver/node-image upgrades.

Stop Conditions

Stop and confirm before approving the workload platform design, or before promoting it to production, if any of the following is true:

  • Node pool VM family, SKU, OS, or disk type was chosen without performance evidence (load test, VPA recommender output, or production telemetry).
  • HPA, KEDA, or VPA is being enabled on a workload whose containers do not have CPU and memory requests, or whose Metrics API and scaler source have not been validated.
  • HPA and VPA in Auto or Recreate mode are configured against the same CPU/memory target on the same workload without a documented design and test result.
  • Workload Identity is not enabled at cluster creation, OR a workload that needs Azure access is still using AAD Pod Identity, service principal secrets, or a shared "platform" managed identity.
  • NAP / Karpenter is being introduced without confirmed networking compatibility, unsupported-combination check (e.g., Calico, Dynamic IP Allocation), and resource-request accuracy across the candidate workloads.
  • Production stateless services are running at one replica without an explicit singleton design (leader election, queue partitioning) and ADR-documented availability acceptance.
  • GPU pools are being provisioned without verified subscription GPU quota in the target region, supported VM family list, driver/install model, and a workload that actually requires GPU.
  • GPU partitioning (MIG, time-slicing, or MPS) is being introduced without a documented strategy choice, node-pool-level reprovisioning plan (MIG), and ownership of the NVIDIA GPU Operator lifecycle for time-slicing/MPS.
  • Kueue (or any batch admission controller) is being introduced without acknowledging that it is open-source community-supported software excluded from AKS SLA, and without a node-autoscaler/NAP plan so admitted jobs actually get capacity.
  • Spot, Windows, or stateful pools are being introduced without explicit scheduling rules (taints, tolerations, node affinity, PDB, topology spread) and an upgrade/maintenance window plan.
  • Scaling maximums (HPA maxReplicas, KEDA maxReplicaCount, cluster autoscaler / NAP node maximums) have not been reconciled with downstream dependency capacity (databases, external APIs, NAT Gateway SNAT, ingress) and budget.
  • Workload-to-Azure access paths bypass Workload Identity (e.g., connection-string secrets in environment variables) without a documented rotation owner.

Resolve each condition (or capture an explicit ADR exception) before treating the workload platform as production-ready.

Native Sidecar Containers (Kubernetes 1.29+)

Kubernetes 1.29 introduced a native sidecar pattern via restartPolicy: Always on initContainers entries. Unlike a regular init container (which completes before the main container starts), a native sidecar starts with the pod, runs alongside the main container, and is terminated last during pod shutdown.

initContainers:
  - name: log-forwarder
    image: myregistry/log-forwarder:v1.2
    restartPolicy: Always   # marks this as a native sidecar
    resources:
      requests:
        cpu: "50m"
        memory: "64Mi"
  • Why this matters. Service mesh sidecars, log forwarders, and monitoring agents injected via mutating webhooks (Istio, Dapr, Fluent Bit, OTel Collector) can now be declared as native sidecars so they start before the main app and outlive it during shutdown. This eliminates the "sidecar outlives app" and "app starts before sidecar is ready" races.
  • Pod startup order. Native sidecars run concurrently with subsequent initContainers and the main containers once their startupProbe (or readinessProbe) passes; declare probes on sidecar containers to control sequencing.
  • Compatibility. [VERIFY] current AKS support for native sidecar injection from Istio (App Routing add-on), Dapr, and other managed sidecars. Some service mesh implementations still use webhook-injected regular sidecars; moving to native sidecars requires explicit adoption per project/operator version.
  • Resource accounting. Native sidecar container resources are counted toward pod total requests/limits and affect QoS class; include them in namespace LimitRange and ResourceQuota capacity planning.

QoS Class Strategy

Kubernetes assigns pods a Quality of Service (QoS) class based on container resource declarations, affecting eviction priority under node memory pressure:

QoS Class Condition Eviction priority
Guaranteed All containers have requests == limits (both CPU and memory) Last to be evicted
Burstable At least one container has requests < limits, or limits are missing Middle priority
BestEffort No container has any requests or limits First to be evicted

Recommended strategy for production AKS workloads:

  • Use Burstable QoS as the default for most production workloads: set CPU/memory requests based on measured usage (from VPA recommender or kubectl top), but allow limits to be higher (e.g., memory.limits = 2× requests, CPU uncapped or generous). This gives the scheduler accurate bin-packing signals while allowing burst headroom.
  • Use Guaranteed QoS only for latency-sensitive workloads (e.g., real-time event processors, gRPC serving with P99 SLOs) where unpredictable burst and memory eviction are unacceptable, and only after measuring that the workload does not need burst capacity.
  • Never use BestEffort for production workloads; they are the first candidates for OOM eviction and will be killed when any node is under memory pressure.
  • Setting requests == limits everywhere (a common "safe" default) is not generally optimal: it wastes schedulable capacity, makes HPA harder to tune, and forces Guaranteed QoS on workloads that don't need it.
  • A LimitRange with defaultRequest set to a sane fraction of the default limit prevents BestEffort pods from being admitted without requests.

Readiness Review

Before approving a workload platform design, confirm:

  • system and user workloads are isolated
  • every workload has requests and sensible limits or a documented reason not to
  • HPA/KEDA/VPA/NAP roles do not conflict and max values align with node/downstream capacity
  • namespaces have quota/limit/network/security policy baselines
  • PDBs and topology spread cover production replicas
  • Workload Identity is the default Azure access path
  • GPU/Windows/spot/stateful pools have explicit scheduling and upgrade rules
  • node pool sizes, min/max, and SKU choices are backed by evidence or clear assumptions

Use With

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The skill is a comprehensive architecture and configuration guide for Azure Kubernetes Service (AKS). it emphasizes security best practices, including RBAC, NetworkPolicy, workload identity, and kernel-level isolation for AI agents. No malicious patterns or security risks were detected.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-26"
}
Other metadata
argument-hint
workload=<type>; region=<azure-region>; availability=<SLO>; network=<hub-spoke|standalone>; scope=<new-cluster|production-review|fleet>

README badge

README badge for lukemurraynz/hve-agent-skills/aks-cluster-architecture