---
name: aks-cluster-architecture
description: >-
  AKS cluster architecture decisions for new Azure Kubernetes Service projects: AKS Automatic vs Standard, networking topology, dual-stack (IPv4/IPv6), Kubernetes version and OS currency, node pool strategy, identity, production NetworkPolicy, namespaces, autoscaling, ingress and Gateway API, observability, operations, resilience, GPU and AI workloads, GPU partitioning (MIG, time-slicing, MPS), batch scheduling (Kueue), AKS on bare metal, AI Runway and KAITO model serving, AKS MCP server access, kars (Agent Reference Stack for Kubernetes) for agent isolation, Kata MicroVM pod sandboxing, Azure Kubernetes Fleet Manager, multi-cluster governance, update orchestration, resource placement, cross-cluster networking, and cost. WHEN: designing new AKS clusters, reviewing production readiness, choosing CNI or outbound connectivity, planning node pools, defining namespace, network and security controls, evaluating Fleet Manager, deploying AI agent runtimes on AKS, or making hard-to-reverse infrastructure decisions.
argument-hint: "workload=<type>; region=<azure-region>; availability=<SLO>; network=<hub-spoke|standalone>; scope=<new-cluster|production-review|fleet>"
metadata:
  last_verified: "2026-08-26"
title: aks-cluster-architecture
canonical_url: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/aks-cluster-architecture
last_updated: 2026-09-22T10:44:25.000Z
---

> **Skill from skilld.dev.** Follow the instructions below for this session. You do not need to install anything.
>
> Supporting files, fetch one when the Skill refers to it: [bundles/agent-runtime/bundle.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/agent-runtime/bundle.yaml), [bundles/agent-runtime/guide.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/agent-runtime/guide.md), [bundles/catalog.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/catalog.yaml), [bundles/cluster-foundations/bundle.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/bundle.yaml), [bundles/cluster-foundations/guide.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md), [bundles/fleet-management/bundle.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/fleet-management/bundle.yaml), [bundles/fleet-management/guide.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/fleet-management/guide.md), [bundles/operations-resilience/bundle.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/bundle.yaml), [bundles/operations-resilience/guide.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md), [bundles/production-workload-controls/bundle.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/production-workload-controls/bundle.yaml), [bundles/production-workload-controls/guide.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/production-workload-controls/guide.md), [bundles/workload-platform/bundle.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/bundle.yaml), [bundles/workload-platform/guide.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md), [CHANGELOG.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/CHANGELOG.md), [QUALITY-REVIEW.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/QUALITY-REVIEW.md), [references/full-reference.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/references/full-reference.md), [references/glossary.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/references/glossary.md), [references/quick-reference.md](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/references/quick-reference.md), [scripts/validate-skill.py](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/scripts/validate-skill.py).
>
> If the user asked to install this Skill, run `npx skilld install lukemurraynz/hve-agent-skills/aks-cluster-architecture`. Install writes the Skill files into the project, so every session loads them.

# AKS Cluster Architecture Skill

Use this skill as the lightweight AKS architecture router for **new AKS projects**. Assume modern defaults and avoid legacy patterns unless explicitly handling a migration.

## Local runtime ground truth (verified 2026-08-23)

- kubectl v1.34.1, helm 4.2.0, flux CLI installed; `kubectl version --client=true` / `helm version --short` are the correct invocations (`--version` is not accepted).
- Kube context `aks-groundwork-is4w2tdtmtrdm` (Australia East) points at a **private cluster**: API DNS resolution fails without VPN/private network access. Treat lookup failures as connectivity-blocked, NOT as broken tooling: report the block, do not loop retries or "fix" kubectl.
- `k9s`, `argocd`, `devcontainer` CLI are not installed locally; recommend them in onboarding docs only with install steps.

## Use When

| Should trigger | Should NOT trigger | Nearby skill collision |
|---|---|---|
| User asks for AKS cluster architecture, production-readiness review, AKS Automatic vs Standard, networking/egress, node pools, workload identity, ingress, observability, upgrade strategy, Fleet Manager, or GPU/AI workload placement decisions. | User only needs day-2 kubectl troubleshooting, generic Kubernetes YAML edits, or non-Azure container-platform guidance with no AKS architecture decision. | Use `azure-container-apps` when the real decision is AKS vs ACA, and use a runtime/container-operations skill when the task is operational debugging rather than target-state architecture. **Note:** AKS deployments that use Workload Identity, ACR, or event-driven messaging require the co-skills below loaded simultaneously — co-requisites, not alternatives. |

## Mandatory co-skills by deployment target

This is NOT optional. An AKS deployment touches identity, networking, storage, and messaging domains that live in sibling skills. Load the co-skills for your deployment target **before** producing any architecture recommendation or authoring manifests. Skipping these causes silent failures (wrong identity property path, wrong consumer group, missing Private Link DNS).

| Deployment target | Mandatory co-skills to load first | Why (the failure it prevents) |
| --- | --- | --- |
| AKS with **Workload Identity** (pods accessing Azure resources) | `identity-managed-identity` (workload identity ordering, SA annotation, federated credential) | Wrong SA annotated; pods silently fall back to kubelet identity; token not injected due to ordering |
| AKS with **ACR** (image pulls) | `identity-managed-identity` (kubelet identity ACR pull) | AcrPull assigned to control-plane identity instead of kubelet — image pulls 401/403 |
| AKS with **Event Hubs / Service Bus** (event-driven workloads) | `event-driven-messaging` (consumer groups, partitioning, idempotent consumers) | `$Default` consumer group shared across consumers; silent message loss or duplicate processing |
| AKS with **Private Endpoints** (PaaS access) | `private-networking` (DNS zones, Private Link) | CoreDNS split-DNS trap; pods cannot resolve private zone |
| AKS with **Container Registry supply chain** | `container-supply-chain`, `container-operations` | Untagged image drift; missing SBOM/signing; no vulnerability gate |
| AKS with **Drasi for Kubernetes** | `drasi` (source/reaction identity, EventHub label mapping) | Per-source SA pattern; identity field placement; EventHub entity = Cypher label |
| AKS with **GPU / AI workloads** | `azure-machine-learning` or AI serving guidance | GPU node pool SKU, driver, KAITO/AI Runway placement errors |
| AKS with **Foundry / Azure OpenAI model serving** (pods calling model endpoints) | `microsoft-foundry` (serverless inference reference) + `identity-managed-identity` | Wrong data-plane role (Cognitive Services OpenAI User vs Cognitive Services User) or missing credential scope — 401/403 at inference time |
| AKS **Fleet Manager** (multi-cluster) | Fleet-specific guidance in this skill's Fleet bundle | Hubless vs hub; placement; cross-cluster networking |

If a co-skill listed above is not present in the repository, note it as a gap, do not silently proceed without the guidance.

## Operating Model

1. Classify every architecture recommendation by change difficulty before giving a default:

- **Permanent** - cluster rebuild or parallel-cluster migration is likely.
- **Difficult** - possible, but disruptive and requires outage planning, validation, or staged rollout.
- **Reversible** - safe to change with normal operational controls.

2. Prefer the smallest relevant bundle. Load the full reference only when the bundle does not contain enough detail.
3. Treat fast-moving AKS guidance as verification-driven. Do not hard-code patch versions or preview status in generated IaC without checking current Microsoft guidance.
4. Capture permanent and disruptive choices in an ADR before production provisioning.
5. **Complementary community tooling.** The KubeDeck ecosystem (github.com/kubedeckio) provides agentless AKS/Kubernetes validation (KubeBuddy, 34 AKS best-practice checks), kubeconfig lifecycle management (KubeTidy), resource snapshot/drift detection (KubeSnapIt), and CRD-based operational memory (KubeMemo). These are open-source tools (CNCF Sandbox applicant) that complement, not replace, first-party Azure tooling and this skill's architecture guidance. Reference KubeBuddy specifically as an external validation option in the Freshness Gate and Validation Commands sections.
6. **kubectl plugin ecosystem.** Recommend `krew` (the official kubectl plugin manager) as the standard way to install and update kubectl plugins. Specific plugins to include in developer and SRE onboarding documentation: `neat` (strip managed-fields from YAML), `tree` (owner-reference tree for debugging), `view-secret` (base64-decode Secrets), `ctx`/`ns` (fast context/namespace switch), `deprecations` (API deprecation scan; `kubectl deprecations`). Additionally, `kubelogin` (Azure/kubelogin, not a krew plugin) is required for any cluster using Azure RBAC + Entra ID auth; document this separately in onboarding. Reference for full plugin list: [krew.sigs.k8s.io/plugins](https://krew.sigs.k8s.io/plugins).
7. **`kube-score` in CI.** `kube-score` is a static analysis tool for Kubernetes manifests that checks for common misconfigurations (missing probes, no resource limits, single-replica deployments, missing NetworkPolicy). Use it as an **advisory CI gate** (non-blocking, informational) alongside `kubeconform` for schema validation and Kyverno/Conftest for policy enforcement. `kube-score score <manifest.yaml>` returns exit code 1 on any non-OK finding; run with `--output-format ci` for CI log-friendly output. Treat `kube-score` as a developer feedback layer, not a security gate; its role is to surface best-practice regressions early, not to replace admission policy enforcement.

## Required Output Contract

When this skill is used for a design, review, or readiness assessment, produce a structured answer with these sections unless the user asks for something narrower:

1. **Context and assumptions** - workload, environment, region, compliance, availability target, operational owner, and unknowns.
2. **AKS Automatic vs Standard recommendation** - decision, rationale, and rejected alternative.
3. **Permanent / Difficult / Reversible decision table** - classify each material choice before giving implementation detail.
4. **Recommended architecture** - cluster, networking, identity, node pools, ingress/Gateway, observability, policy, and operations.
5. **Production workload controls** - namespace model, labels, RBAC, ResourceQuota, LimitRange, Pod Security Admission, NetworkPolicy, replicas, probes, PDBs, topology spread, and autoscaling.
6. **Security and identity model** - managed identity, OIDC issuer, Workload Identity, Azure RBAC/Kubernetes RBAC, Key Vault/secret access, and least-privilege boundaries.
7. **Networking and egress model** - CNI/IPAM/data plane, CIDR plan, private API/DNS, ingress, east-west policy, Private Link, NAT Gateway or UDR/firewall egress, and SNAT risk.
   - **CoreDNS split-DNS trap**: When using NAT Gateway egress + Private Endpoints for PaaS, Azure DNS (`168.63.129.16`) becomes unreachable from pods. CoreDNS must be patched post-provision for split-DNS: public zones → external forwarders (8.8.8.8), Azure private zones → 168.63.129.16. See [operations-resilience - Known Pitfalls](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md#known-pitfalls) for the full symptom and fix.
8. **Observability model** - logs (ContainerLogV2), metrics (Managed Prometheus) including control-plane metrics (`--enable-control-plane-metrics`, GA Aug 2026 - see [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) Observability Baseline), dashboards, alerts, SLOs, and runtime validation evidence.
9. **Upgrade and patch strategy** - both maintenance windows (`aksManagedAutoUpgradeSchedule` for Kubernetes, `aksManagedNodeOSUpgradeSchedule` for node OS), whether multi-cluster environments should converge schedules via the Shared Maintenance Windows preview (one standalone `maintenanceWindows` resource linked with `--maintenance-window-id`), Node Disruption Policy (`nodeDisruptionProfile`) for gating config-change reimages, Kubernetes upgrade cadence, node OS auto-upgrade channel choice, and node-image currency tracking against the AKS release tracker.
10. **CVE and security incident response** - AKS Security Bulletins monitoring, illustrative response runbook (treat any worked example as illustrative until verified against current bulletins), Fleet Manager NodeImage auto-upgrade profiles plus the Security patch channel for OS security patches between image releases, and `az fleet autoupgradeprofile generate-update-run` for emergency node-image-only runs. Single-cluster fallback: `az aks nodepool upgrade --node-image-only`, with node-pool rollback as 7-day post-upgrade triage.
11. **Operations model** - GitOps source of truth, break-glass, backup and restore, and DR validation.
12. **Fleet and multi-cluster model, if applicable** - whether Fleet Manager is needed, hubless vs hub, member cluster taxonomy, staged updates, placement, managed namespaces, cross-cluster networking, and preview gates.
13. **Cost and resilience considerations** - zones, replicas, autoscaling max values, gateway/log/GPU/egress/fleet cost drivers, RPO/RTO, and DR pattern.
14. **Validation commands** - CLI and kubectl checks needed before implementation is accepted.
15. **Stop conditions** - choices that must be confirmed before provisioning or production rollout.
16. **ADR entries required** - short list of decisions that must be captured before build starts.
17. **Open questions** - only unresolved items that materially affect the architecture.

Do not produce generic AKS advice when the user asked for a production architecture. Tie every recommendation to a decision, risk, validation step, or trade-off.

## Do Not Use This Skill For

- Generic Kubernetes YAML linting unless the manifest exposes AKS architecture decisions.
- Application code changes that do not affect AKS architecture, networking, identity, scaling, or production readiness.
- Terraform/Bicep syntax implementation details unless the task is deciding AKS architecture or reviewing IaC-impacting choices.
- Non-Azure Kubernetes platforms unless comparing migration, portability, or Azure target-state trade-offs.
- Existing-cluster migration execution plans where a migration/runbook skill is more appropriate; use this skill only for target-state decisions and architectural risk review.
- Day-2 application debugging or kubectl troubleshooting where the answer is `kubectl describe`/`logs`/`events`, not architecture.
- Choosing between AKS and Azure Container Apps, Azure Container Instances, or Azure Functions for a workload - no dedicated comparison skill exists (verified 2026-08-16); use Microsoft Learn `compare-options`, plus `azure-container-apps` (which carries the CMK/runtime-detection AKS steer) for the ACA side.
- On-premises or edge Kubernetes (AKS on Azure Local, AKS Edge Essentials, Arc-only clusters) where the design surface differs from cloud AKS.

## Glossary

| Acronym     | Meaning (AKS context)                                                                                                                              |
| ----------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| ACR         | Azure Container Registry - managed OCI registry for container images, Helm charts, and signed artifacts consumed by AKS.                           |
| ACNS        | Container Networking Services - AKS add-on layering Hubble-style flow logs, DNS metrics, FQDN filtering, eBPF Host Routing performance mode (`--acns-datapath-acceleration-mode BpfVeth`), and (with `--acns-advanced-networkpolicies L7`) Layer-7 policy on Azure CNI Powered by Cilium clusters. `--enable-acns` alone enables only FQDN *filtering*; each feature above has its own gate - L7 *policy* requires the `L7` flag, Host Routing requires the acceleration-mode flag and excludes Static Egress Gateway/CVM/Kata pools. |
| AGC         | Application Gateway for Containers - Azure's strategic managed Gateway API implementation for AKS ingress.                                         |
| AGIC        | Application Gateway Ingress Controller - legacy AKS ingress controller bridging Kubernetes Ingress to Application Gateway v2 (predecessor to AGC). |
| ALZ         | Azure Landing Zone - Microsoft's reference subscription/management-group topology that AKS clusters land into.                                     |
| CNI         | Container Network Interface - pod networking plugin; for AKS typically Azure CNI Overlay Powered by Cilium.                                        |
| Cold standby | Single-replica deployment; on failure a new pod starts from scratch. Acceptable only for non-critical workloads with explicit availability acceptance. See also [Warm standby]. |
| CRD         | CustomResourceDefinition - Kubernetes extension mechanism used by operators (KAITO, Gateway API, Fleet, etc.).                                     |
| CSI         | Container Storage Interface - driver model for Azure Disk, Azure Files, Blob, and Azure Container Storage volumes in AKS.                          |
| CVM         | Confidential VM - SEV-SNP/TDX node SKU enabling confidential AKS node pools and confidential containers.                                           |
| DCGM        | NVIDIA Data Center GPU Manager - exporter for GPU metrics consumed by Managed Prometheus and KEDA in AKS GPU pools.                                |
| FQDN        | Fully Qualified Domain Name - used in AKS for egress allow-listing (Azure Firewall FQDN tags, Cilium FQDN policy) and private DNS.                 |
| GitOps      | Git as source of truth for cluster state - typically Flux (AKS Flux extension) or Argo CD reconciling manifests into AKS.                          |
| HPA         | HorizontalPodAutoscaler - replica autoscaler driven by CPU/memory/custom metrics; pairs with KEDA for event-driven scaling.                        |
| KAITO       | Kubernetes AI Toolchain Operator - AKS add-on that provisions GPU pools and serves supported open-source inference models via CRDs.                |
| KEDA        | Kubernetes Event-Driven Autoscaler - AKS-managed add-on scaling workloads on external event sources (queues, Prometheus, etc.).                    |
| KMS         | Key Management Service - Azure Key Vault-backed envelope encryption for etcd secrets in AKS.                                                       |
| kars        | Agent Reference Stack for Kubernetes — Microsoft's open-source, Kubernetes-native reference stack ([Azure/kars](https://github.com/Azure/kars), not an officially supported product) for running AI agents on AKS: a no-network agent container, a Rust inference router that holds all credentials and governs egress, and an encrypted inter-agent mesh. |
| LTS         | Long-Term Support - AKS Kubernetes version channel offering extended support beyond community N-2.                                                 |
| MCP         | Model Context Protocol - protocol surface for exposing tools to AI agents; AKS MCP connects agents to Azure and Kubernetes operations.             |
| MI          | Managed Identity - Azure AD identity assigned to the cluster control plane, kubelet, or workloads (system- or user-assigned).                      |
| NAP         | Node Autoprovisioning - AKS Karpenter-based capacity autoscaler that provisions nodes from a VM SKU pool on demand.                                |
| OIDC        | OpenID Connect - issuer endpoint AKS exposes so Workload Identity and external IdPs can federate to Entra.                                         |
| PDB         | PodDisruptionBudget - minAvailable/maxUnavailable contract honoured during node drains, upgrades, and Fleet update runs.                           |
| PSA         | Pod Security Admission - namespace-level `privileged`/`baseline`/`restricted` enforcement replacing PodSecurityPolicy on AKS.                      |
| RBAC        | Role-Based Access Control - Kubernetes RBAC and/or Azure RBAC for Kubernetes authorization on AKS.                                                 |
| SBOM        | Software Bill of Materials - image-supply-chain artifact (e.g., SPDX/CycloneDX) attached to ACR images for AKS workloads.                          |
| SIEM        | Security Information and Event Management - typically Microsoft Sentinel ingesting AKS audit, ContainerLogV2, and Defender signals.                |
| SLO         | Service Level Objective - availability/latency target driving AKS replica counts, zones, PDBs, and Fleet staged rollout pace.                      |
| SLSA        | Supply-chain Levels for Software Artifacts - framework for ACR image provenance/attestation feeding AKS admission policy.                          |
| UDR         | User-Defined Route - route-table entry forcing AKS egress through Azure Firewall or NVA in hub-spoke designs.                                      |
| VHD         | Virtual Hard Disk - the AKS node image; weekly VHD builds are the unit of node OS patching and CVE response.                                       |
| VPA         | VerticalPodAutoscaler - recommends/sets pod CPU and memory requests; runs in recommendation mode alongside HPA on AKS.                             |
| Warm standby | Multi-replica deployment using leader election; standby pods are ready and waiting to acquire a lease but do not serve traffic until the leader fails. Faster recovery than cold standby with lower resource overhead than fully active-active. |

- **Permanent decision** - Cannot be changed without rebuilding the cluster (region, CIDRs, CNI/IPAM, availability zones, etcd KMS enablement, hub access mode, hyperscale control plane scaling profile ; creation-only and irremovable; see [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md#hyperscale-control-plane-scaling-profile-public-preview)). Decide once, before provisioning, and capture an ADR.
- **Difficult decision** - Changeable in-place but disruptive: requires node-pool replacement, downtime, or cross-team coordination (API server exposure, OS SKU, Kubernetes minor, RBAC model). Plan a rollback path before committing. Note: **outbound type is now mutable** post-creation with defined migration paths (loadBalancer → NAT GW → none/block), but migrating *away from* `managedNATGatewayV2` to any other outbound type is not supported on a managed VNet once chosen (it is a one-way destination) and any migration changes the cluster's outbound public IPs, treat as a difficult decision with a maintenance window, not a routine change.
- **Reversible decision** - Changeable through normal GitOps or `az aks update` with no rebuild and minimal blast radius (HPA settings, replica counts, NetworkPolicy rules, namespace quotas, CoreDNS ConfigMap edits). Iterate as evidence comes in.
- **Hubless Fleet Manager** - Azure Kubernetes Fleet Manager without a hub cluster. Supports cross-cluster update orchestration only. Cheaper, simpler, can be upgraded to a hub fleet later.
- **Hub Fleet Manager** - Azure Kubernetes Fleet Manager with a managed hub AKS cluster providing the Kubernetes API for `ClusterResourcePlacement`, Managed Fleet Namespaces, and hub-dependent networking. Required for resource placement; hub access mode (public vs private) is set at creation and cannot be changed.

## Quick Routing - I need to…

| I need to…                                            | Load                                                                                                                                                                               |
| ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Respond to an active CVE / kernel advisory            | [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) + [fleet-management](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/fleet-management/guide.md) (Pre-Upgrade API Inventory + CVE Response Decision Matrix) |
| Size pod/service CIDRs, pick CNI                      | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md)                                                                                                                        |
| Decide dual-stack (IPv4/IPv6), plan IPv6 CIDRs        | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (Dual-Stack Networking)                                                                                               |
| Choose a StorageClass / design stateful storage (zonal disk, RWX) | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (Storage and Stateful Data)                                                                               |
| Give a namespace its own fixed egress IP             | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (Static Egress Gateway)                                                                                                |
| Diagnose DNS failures / CoreDNS at scale             | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (CoreDNS at scale)                                                                                                     |
| Attribute AKS cost per team / namespace              | [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) (Cost visibility and showback tooling)                                                                             |
| Write NetworkPolicy YAML for a production namespace   | [production-workload-controls](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/production-workload-controls/guide.md)                                                                                                      |
| Decide hubless vs hub fleet                           | [fleet-management](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/fleet-management/guide.md)                                                                                                                              |
| Pick ingress controller (AGC vs NGINX vs Gateway API) | [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md)                                                                                                                    |
| Design GPU / KAITO / AI inference workload            | [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md) + [production-workload-controls](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/production-workload-controls/guide.md)                                            |
| Serve LLM inference from AKS via Foundry / Azure OpenAI (no GPU) | [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md) (Serverless model serving) + `microsoft-foundry` inference reference |
| Partition a GPU (MIG / time-slicing / MPS)            | [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md) (GPU Partitioning)                                                                                                         |
| Queue batch / training jobs (Kueue)                   | [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md) (Batch and AI Training Scheduling)                                                                                         |
| Design agentic AI / MCP access to AKS                  | [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md) + [production-workload-controls](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/production-workload-controls/guide.md) + [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) |
| Deploy an AI agent runtime with kernel isolation (kars) | [agent-runtime](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/agent-runtime/guide.md) + [microsoft-agent-framework](../microsoft-agent-framework/SKILL.md)                                                                                                                                                                                                                                                                                              |
| Plan a Kubernetes minor upgrade                       | [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) (Pre-Upgrade API Deprecation Inventory)                                                                            |
| Set up GitOps repository structure                    | [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md)                                                                                                                    |
| Plan etcd KMS / customer-managed key                  | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md)                                                                                                                        |
| Map AKS controls to CIS / NIST / PCI-DSS              | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (Compliance Alignment)                                                                                                 |

## Bundle Catalog

Start with [bundles/catalog.yaml](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/catalog.yaml).

| Need                                                                                                                                                                                                                                                                                                           | Load                                                                          |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------- |
| AKS Automatic vs Standard, version currency, deprecations, CNI, CIDRs, private/public API server, outbound type (incl. Static Egress Gateway), DNS and CoreDNS-at-scale, storage/StorageClass and stateful data, pricing tier/SLA, hard-to-reverse cluster decisions (foundational; load first for any new-cluster question)                                                                                                | [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md)                   |
| Node pools, system/user/spot/GPU pools, Workload Identity, Entra/RBAC, scheduler placement, bin packing, GPU/AI workload placement, and AKS MCP server placement                                                                                                                                               | [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md)                       |
| Workload-scope: Production namespace model, NetworkPolicy examples, ResourceQuota, LimitRange, replicas, PDBs, topology spread, HPA, KEDA, VPA, NAP/Karpenter autoscaling guardrails (namespace manifests, NetworkPolicy YAML, HPA/KEDA/VPA examples, PDB and topology spread, ResourceQuota/LimitRange)       | [production-workload-controls](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/production-workload-controls/guide.md) |
| Cluster-scope: Gateway API, App Routing, ingress migration, TLS ownership, observability, upgrades, GitOps, backup/restore, multi-region, cost, known pitfalls, Drasi-on-AKS notes (cluster-wide observability, upgrade strategy, GitOps and release pipelines, backup/DR, ingress controller selection, cost) | [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md)               |
| Azure Kubernetes Fleet Manager, multi-cluster governance, hubless vs hub fleet, staged update runs, resource placement, managed fleet namespaces, cross-cluster networking, and Cilium Cluster Mesh watchpoints                                                                                                | [fleet-management](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/fleet-management/guide.md)                         |
| Agent runtime isolation on AKS: kars (Agent Reference Stack for Kubernetes), Kata MicroVM pod sandboxing, AKS MCP server patterns, agent identity and security boundaries, isolation primitive selection                                                                                                      | [agent-runtime](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/agent-runtime/guide.md)                               |
| Detailed decision tables, validation commands, ADR template, migration notes, and source links                                                                                                                                                                                                                 | [full reference](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/references/full-reference.md)                                |

**Routing rule when both could apply.** When a question mentions both workload manifests and cluster-wide operations, load both bundles but treat **production-workload-controls** as authoritative for YAML examples and per-namespace guardrails, and treat **operations-resilience** as authoritative for cluster-wide decisions (observability stack, upgrade orchestration, GitOps repo layout, traffic management controller choice, DR pattern). If the answer requires both, cite each bundle for the section it owns rather than restating guidance in your response.

```mermaid
graph TD
    CF[cluster-foundations]
    WP[workload-platform]
    PWC[production-workload-controls]
    OR[operations-resilience]
    FM[fleet-management]
    AR[agent-runtime]

    CF --> WP
    CF --> PWC
    WP --> PWC
    CF --> OR
    WP --> OR
    CF --> FM
    OR --> FM
    WP --> AR
    PWC --> AR
```

## Freshness Gate

Before finalising any AKS version, retirement, GA/preview, regional availability, or feature compatibility recommendation, verify at least the relevant current source or command:

```bash
# Versions, previews, and upgrade paths
az aks get-versions --location <region> --output table
az aks get-upgrades --resource-group <rg> --name <cluster> --output table
az provider show --namespace Microsoft.ContainerService --query "resourceTypes[?resourceType=='managedClusters'].apiVersions" --output table
az feature list --namespace Microsoft.ContainerService --output table

# Current node image state per pool (CVE/patch posture)
az aks nodepool list --resource-group <rg> --cluster-name <cluster> \
  --query "[].{name:name,k8s:orchestratorVersion,nodeImage:nodeImageVersion}" --output table
az aks nodepool get-upgrades --resource-group <rg> --cluster-name <cluster> \
  --nodepool-name <pool> --output table

# Fleet extension and auto-upgrade profiles
az extension add --name fleet
az extension update --name fleet
az fleet autoupgradeprofile list --resource-group <rg> --fleet-name <fleet> --output table
```

Also check the current Microsoft Learn page or AKS release notes for the exact feature being recommended, and check the **AKS release tracker** and **AKS Security Bulletins** pages for active node-image advisories and the corresponding patched VHD build identifiers before publishing a node-image or CVE-response recommendation. If verification is not possible, state the assumption and mark it as requiring confirmation before implementation.

**External validation tools.** Consider running an agentless AKS best-practice scanner alongside the Azure CLI checks for an independent compliance perspective:

```bash
# KubeBuddy — agentless AKS best-practice validation (open-source, kubedeckio ecosystem)
# Native Go CLI (primary since v0.0.28)
kubebuddy run --aks --subscription-id <id> --resource-group <rg> --cluster-name <cluster> --html-report --output-path ./reports

# KubeBuddy also generates an AKS Automatic migration readiness report:
kubebuddy run --aks --subscription-id <id> --resource-group <rg> --cluster-name <cluster> --html-report
# Produces: *-aks-automatic-action-plan.html with blockers and migration sequence

# PowerShell wrapper (backward-compatible, wraps native binary):
Invoke-KubeBuddy -HtmlReport -aks -SubscriptionId <id> -ResourceGroup <rg> -ClusterName <cluster>
```

KubeBuddy evaluates 34 AKS-specific checks across Best Practices, Disaster Recovery, Identity & Access, Monitoring, Networking, Resource Management, and Security categories, covering the same surface as this skill's non-negotiables. It also derives an **AKS Automatic migration readiness view** identifying blockers and warnings for Standard → Automatic migration. Use it as a pre-provisioning or post-deployment validation gate, not as a replacement for the `az aks` verification commands above. The native Go CLI (`kubebuddy`) is the primary distribution; the PowerShell module (`Invoke-KubeBuddy`) is a backward-compatible wrapper. `[VERIFY]` CLI version and check IDs against the current KubeBuddy documentation.

## Post-Deployment Validation Checklist

After provisioning or upgrading an AKS cluster, verify every non-negotiable landed correctly. Run these commands against the live cluster and confirm expected values. Use this checklist as the acceptance gate before handing the cluster to workload teams.

### Cluster Security

```bash
# Private API server (or authorized IP ranges if public)
az aks show -g <rg> -n <cluster> --query "apiServerAccessProfile.enablePrivateCluster" -o tsv
az aks show -g <rg> -n <cluster> --query "apiServerAccessProfile.authorizedIpRanges" -o tsv

# OIDC issuer + Workload Identity enabled
az aks show -g <rg> -n <cluster> --query "oidcIssuerProfile.enabled" -o tsv
az aks show -g <rg> -n <cluster> --query "securityProfile.workloadIdentity.enabled" -o tsv

# Azure RBAC for Kubernetes authorization
az aks show -g <rg> -n <cluster> --query "aadProfile.enableAzureRBAC" -o tsv

# Image Cleaner enabled
az aks show -g <rg> -n <cluster> --query "securityProfile.imageCleaner.enabled" -o tsv

# Key Vault CSI driver enabled
az aks show -g <rg> -n <cluster> --query "addonProfiles.azureKeyvaultSecretsProvider.enabled" -o tsv

# Kubernetes Dashboard disabled
az aks show -g <rg> -n <cluster> --query "addonProfiles.kubeDashboard.enabled" -o tsv
# Expected: null or false
```

### Node Pools

```bash
# Multiple node pools exist (system + at least one user pool)
az aks nodepool list -g <rg> -cluster-name <cluster> --query "length([])" -o tsv
# Expected: >= 2

# System pool has >= 2 nodes
az aks nodepool list -g <rg> -cluster-name <cluster> --query "[?mode=='System'].count.count" -o tsv
# Expected: >= 2

# System pool has CriticalAddonsOnly taint
az aks nodepool show -g <rg> -cluster-name <cluster> -n <system-pool> --query "nodeTaints" -o tsv
# Expected: ["CriticalAddonsOnly=true:NoSchedule"]

# No B-series VMs in any pool
az aks nodepool list -g <rg> -cluster-name <cluster> --query "[].vmSize" -o tsv
# Expected: no Standard_B* entries

# Node pool versions vs control plane
az aks nodepool list -g <rg> -cluster-name <cluster> --query "[].{name:name,version:orchestratorVersion}" -o table
az aks show -g <rg> -n <cluster> --query "kubernetesVersion" -o tsv
# Best practice: all node pool versions equal the control plane version.
# Sanctioned exception: skew policy permits pools up to N-3 behind during staged / control-plane-only upgrades.

# Ephemeral OS disks enabled
az aks nodepool list -g <rg> -cluster-name <cluster> --query "[].osDiskType" -o tsv
# Expected: Ephemeral (where VM SKU supports it)

# Custom MC_ resource group name
az aks show -g <rg> -n <cluster> --query "nodeResourceGroup" -o tsv
# Expected: NOT starting with MC_
```

### Upgrade and OS Currency

```bash
# Auto-upgrade channel configured (not None)
az aks show -g <rg> -n <cluster> --query "autoUpgradeProfile.upgradeChannel" -o tsv
# Expected: patch, stable, rapid, or node-image — NOT none/None

# Node OS upgrade channel configured
az aks show -g <rg> -n <cluster> --query "autoUpgradeProfile.nodeOSUpgradeChannel" -o tsv
# Expected: NodeImage or SecurityPatch — NOT None

# Current node image version per pool (verify currency)
az aks nodepool list -g <rg> -cluster-name <cluster> \
  --query "[].{name:name,nodeImage:nodeImageVersion}" -o table
```

### Networking

```bash
# NetworkPolicy enabled (not none)
az aks show -g <rg> -n <cluster> --query "networkProfile.networkPolicy" -o tsv
# Expected: azure, calico, cilium, or None (if using ACNS/Cilium-only)

# Cilium dataplane for new clusters
az aks show -g <rg> -n <cluster> --query "networkProfile.networkDataplane" -o tsv
# Expected: cilium (for new Azure CNI Overlay clusters)

# NAT Gateway or UDR for production egress (not loadBalancer SNAT)
az aks show -g <rg> -n <cluster> --query "networkProfile.outboundType" -o tsv
# Expected: managedNATGateway, managedNATGatewayV2, userDefinedRouting, or userAssignedNATGateway
# NOT: loadBalancer (for production)
```

### Policy and Safeguards

```bash
# Deployment Safeguards enabled at Enforcement level
az aks show -g <rg> -n <cluster> --query "guardrailsProfile.level" -o tsv
# Expected: Enforcement (production) or Warning (dev/staging minimum)

# Azure Policy add-on enabled
az aks show -g <rg> -n <cluster> --query "addonProfiles.azurepolicy.enabled" -o tsv
# Expected: true
```

### Quick KubeBuddy Cross-Check

For an independent agentless validation pass against the same surface:

```bash
kubebuddy run --aks --subscription-id <id> --resource-group <rg> --cluster-name <cluster> \
  --html-report --output-path ./reports
```

Review the HTML report for any FAIL entries in Best Practices, Security, and Networking categories. KubeBuddy's check IDs (AKSBP*, AKSSEC*, AKSNET*) map to the non-negotiables above, a FAIL from KubeBuddy on a check that has a corresponding non-negotiable is a blocking finding.

## Non-Negotiables for New AKS Designs

- Use managed identities and Microsoft Entra Workload ID. Do not recommend service principal secrets for new clusters.
  - When consuming the AVM `container-service/managed-cluster` module, set `managedIdentities` explicitly (`{ systemAssigned: true }` for a single-resource workload). The parameter is optional in the module, so omitting it compiles and lints clean and then fails AKS preflight with `InvalidParameter: Required parameter servicePrincipalProfile is missing (null)`, a message that points at a credential problem when the actual cause is no identity at all. System-assigned is also what creates the kubelet identity that any `AcrPull` assignment must target. Confirmed live 2026-08-09 on AVM `0.14.0`.
- Prefer Azure CNI Overlay **Powered by Cilium** for new general-purpose clusters when supported by the workload and region.
- Do not rely on default load balancer SNAT for production egress. Use NAT Gateway for AKS-managed VNets or UDR through a central firewall for enterprise hub-spoke networks.
- Enable OIDC issuer and Workload Identity at cluster creation.
- Use Azure RBAC for Kubernetes authorization unless a documented exception exists.
- `disableLocalAccounts: true` **requires** either `enableAzureRBAC: true` or `adminGroupObjectIDs` (or both). Setting `disableLocalAccounts` alone produces an unreachable cluster - no local admin kubeconfig, no Azure RBAC, no group-based access is available. This is unrecoverable without redeploying the cluster.
- For CI/CD and multi-user deployments, plan for **who gets initial AKS cluster admin** before provisioning. Use a pre-provision hook (`preprovision`) to capture the deploying principal and grant `Azure Kubernetes Service RBAC Cluster Admin` via a Bicep `Microsoft.Authorization/roleAssignments` resource in the same deployment. Do not rely on manual role assignment after provisioning - the cluster is locked until someone has RBAC access.
- Use an AKS pricing tier suitable for production or at-scale workloads; do not silently default production designs to a free/dev-only posture.
- Use Azure Policy / Deployment Safeguards, Managed Prometheus, ContainerLogV2, and a defined node OS patch channel from day one. For production clusters, set Deployment Safeguards to **Enforcement** (`guardrailsProfile.level: Enforcement`), `Warning` is insufficient as it allows non-compliant deployments through. See [KubeBuddy AKSBP015](https://github.com/KubeDeckio/KubeBuddy) for the corresponding validation check.
  - Apply the built-in Azure Policy initiatives for AKS as a defense-in-depth layer alongside PSA: at minimum **Allowed Container Images** (`K8sAzureV2ContainerAllowedImages` in `deny` enforcement) and **Do not allow privileged containers** (`K8sAzureV2NoPrivilege` in `deny` enforcement). Use the `Kubernetes cluster pod security restricted standards` initiative as the broader baseline. See [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (Compliance Alignment) for the full policy mapping.
  - **Hyperscale carve-out:** on clusters with a control plane scaling profile (hyperscale), Microsoft recommends **not** enabling the Azure Policy add-on; it increases API server load and degrades control-plane performance at scale ([Learn](https://learn.microsoft.com/azure/aks/hyperscale-configuration-aks)). Document the exception in the ADR. Deployment Safeguards is a different admission mechanism and remains recommended on hyperscale clusters.
- Enable **Image Cleaner** (`securityProfile.imageCleaner.enabled: true`) from day one to automatically remove stale, vulnerable container images from cluster nodes. Set `securityProfile.imageCleaner.intervalHours` to match your vulnerability scan cadence (default 24h is a reasonable starting point for production). Image Cleaner (formerly Eraser) reduces the window during which a known-vulnerable image layer is still present on a node after a CVE is published.
- Enable **Azure Key Vault CSI driver** (`addonProfiles.azureKeyvaultSecretsProvider.enabled: true`) from day one. Without it, applications must store secrets as Kubernetes Secrets (base64-encoded in etcd) or environment variables, both bypass Key Vault's access policies, audit trail, and rotation capabilities. The CSI driver mounts secrets, certificates, and keys from Key Vault as volumes via `SecretProviderClass` resources. See [KubeBuddy AKSSEC005](https://github.com/KubeDeckio/KubeBuddy) for the corresponding validation check.
- Choose the node OS deliberately: **Azure Linux 3.0** (`--os-sku AzureLinux`, GA default for K8s 1.32–1.36) and **Ubuntu 24.04** (GA, default for `--os-sku Ubuntu` on K8s 1.35+) are **both first-class, currently-supported options**. Ubuntu 24.04 is not legacy. Pick per your hardening, FIPS (not supported on Ubuntu 24.04), and image-provenance needs; do not carry forward "Azure Linux good, Ubuntu legacy" framing. **Azure Linux 4.0 is Public Preview and not yet an AKS node osSku, do not select it.** Both default SKUs now ship **containerd 2.x**, which arrives with the node image on the 1.33 (Azure Linux) / 1.35 (Ubuntu) upgrade, treat that hop as a runtime major bump (see [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md#containerd-2x-is-the-default-runtime-a-major-bump-that-arrives-with-the-osk8s-upgrade)).
- Use **ephemeral OS disks** for all node pools where the VM SKU's temp/cache disk size meets the OS disk requirement. Ephemeral OS disks reduce write latency and eliminate OS disk billing. The VM SKU's cache size must be ≥ the OS disk size, verify per SKU before setting `osDiskType: Ephemeral`.
- Do **not** use B-series (burstable) VM SKUs for production node pools. B-series VMs depend on CPU credits which deplete under sustained load, causing throttled performance and unpredictable application latency. Use D-series or equivalent balanced VMs as the minimum starting point.
- **Customize the MC_ resource group name** (`--node-resource-group`) at cluster creation with a descriptive name that identifies the cluster. Default MC_ names create organizational confusion in multi-cluster environments and make Azure RBAC scoping harder.
- Configure an explicit AKS **node OS auto-upgrade channel** (`NodeImage` by default for managed weekly VHD updates; `SecurityPatch` as a supported alternative for faster in-place security fixes between weekly VHDs; `Unmanaged` only with documented reason; never leave production on `None`) and an **`aksManagedNodeOSUpgradeSchedule`** maintenance window that does not block patching for more than one cycle.
- For multi-cluster fleets, define a Fleet Manager **NodeImage auto-upgrade profile** alongside the Kubernetes-version channel, choose **Consistent image** for cross-region fleets, and confirm `az fleet autoupgradeprofile generate-update-run` is approved as the emergency CVE response path.
- Treat AKS node images as having a **90-day validity window** - designs must include a recurring patch cadence and a CVE-response runbook tied to the AKS Security Bulletins feed (the AKS-2026-0003 "Copy Fail"/Dirty Frag kernel LPE class is referenced throughout this skill as an illustrative worked example - verify any specific CVE identifier and patched VHD build against the current AKS Security Bulletins feed before acting) rather than ad-hoc node-pool upgrades.
- Use default-deny NetworkPolicy for production workload namespaces, then add explicit label-based allow rules for DNS, ingress, east-west dependencies, and approved egress.
- Give every production namespace ownership labels, ResourceQuota, LimitRange, RBAC boundary, and Pod Security Admission labels or a documented exception.
- Run production stateless services with at least two replicas, usually three across zones, plus readiness probes, PDBs, topology spread, and autoscaling limits.
- Separate system, user, spot, Windows, GPU, and stateful workloads into appropriate node pools or scheduling domains.
- **System node pool minimum 2 nodes.** A single system node is a single point of failure for kube-proxy, CoreDNS, and CNI. During node failure or maintenance, system pods are evicted and the cluster loses DNS resolution and pod networking. Scale system pools to at least 2 nodes (`--node-count 2` or autoscaler `--min-count 2`). See [KubeBuddy AKSBP011](https://github.com/KubeDeckio/KubeBuddy) for the corresponding validation check.
- **Node pools should match the control plane version; skew is bounded, not forbidden.** Out-of-sync node pools can cause subtle API incompatibilities and prevent node image upgrades, so keep every pool at the control-plane version as the steady-state best practice. The version-skew policy (K8s 1.28+) permits pools up to N-3 behind the control plane - this is the sanctioned window for staged upgrades and the control-plane-only path into LTS (validate first, then upgrade pools) - not a resting state. Verify after any control-plane upgrade: `az aks nodepool list --query "[].{name:name,k8s:orchestratorVersion}"`. See [KubeBuddy AKSBP012](https://github.com/KubeDeckio/KubeBuddy) for the corresponding validation check.
- For any persistent data, choose the StorageClass and zone/replication model explicitly. Azure managed disks are **zonal** and pin a pod to one zone (defeating cross-zone spread) - use `WaitForFirstConsumer` binding plus application-level cross-zone replication or ZRS disks, and size `ResourceQuota` against node **allocatable**, not VM spec. See [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (Storage and Stateful Data).
- Avoid upstream/self-managed ingress-nginx as a new long-term production baseline. **Two-stage NGINX EOL cliff to plan for** (per [AKS Engineering Blog 2025-11-13](https://blog.aks.azure.com/2025/11/13/ingress-nginx-update) and [Microsoft Learn - App Routing](https://learn.microsoft.com/en-us/azure/aks/app-routing)): (a) the upstream community Ingress-NGINX project enters end-of-maintenance in **March 2026** - no further OSS releases after that, and (b) the AKS Application Routing add-on's managed NGINX gets **critical security patches only through November 2026**, with no new features. Any new cluster built on App Routing NGINX today will hit the cliff well within its lifetime. For new clusters using the App Routing add-on, AGC (Application Gateway for Containers) is the strategic Gateway API direction. The successor add-on is **Application Routing with Gateway API** (Istio-control-plane backed; **GA April 2026** per [AKS release 2026-04-28](https://github.com/Azure/AKS/releases/tag/2026-04-28) ; the [Azure Updates 567944](https://azure.microsoft.com/updates?id=567944) post published 2026-07-28 is the formalized Azure Updates entry, not the GA date). **Flag trap:** `--enable-app-routing` still enables the legacy NGINX mode; Gateway API mode requires **`--enable-app-routing-istio`** (default-on for new AKS Automatic clusters on AKS 1.36+, which preconfigure Gateway API instead of NGINX). GatewayClass name is **`approuting-istio`** (not `webapprouting.kubernetes.azure.com`). `[VERIFY]` flag/GA status against current Learn [app-routing-gateway-api](https://learn.microsoft.com/en-us/azure/aks/app-routing-gateway-api) at implementation time. Treat ingress controller choice as a Difficult-to-reverse decision: when starting from a clean slate, name AGC (or App Routing with Gateway API) explicitly as the target architecture and document any NGINX bridge step with its migration path and EOL deadline.
- **AKS add-on → cluster-extension migration is an upgrade-break risk.** Several capabilities are moving from the legacy AKS _add-on_ model to the _cluster extension_ / core model; a cluster upgrade can fail or behave unexpectedly mid-migration. Before a minor upgrade, inventory which add-ons have a cluster-extension successor, check the per-component migration guidance, and stage the migration as its own change rather than coupling it to the version upgrade. `[VERIFY]` current add-on/extension status per component.
- Do not recommend preview features for production without explicitly calling out preview status, limitations, rollback approach, and support impact.
- Do not introduce Azure Kubernetes Fleet Manager for a simple single-cluster design unless a multi-cluster operating model is approved with funded delivery plans.
- For multi-cluster platforms, decide explicitly whether Fleet Manager is needed, whether the fleet should be hubless or hub-based, and whether hub access must be private before provisioning.
- Treat Fleet Manager cross-cluster networking, Managed Fleet Namespaces, Arc-enabled member clusters, namespace-scoped placement, and Cilium Cluster Mesh as verification-driven features with explicit support and preview gates.
- Upgrade rollback is **pool-scoped only**: Node Pool Rollback is GA (`az aks nodepool rollback`, [AKS release 2026-08-07](https://github.com/Azure/AKS/releases/tag/2026-08-07)) and restores a node pool's previous Kubernetes version + node image within **7 days** of the upgrade completing - all-or-nothing, no concurrent cluster operations, requires disabling the auto-upgrade channel first, cannot revert OS-SKU changes, and cannot target unsupported versions. Treat it as temporary triage (re-upgrade within ~30 days), not a standing strategy. The **control plane still cannot roll back** - for control-plane failures or regressions past the 7-day window the durable recovery path remains a parallel cluster build with traffic shift, or restore-from-backup - design upgrades so this option is feasible (snapshots, IaC reproducibility, blue/green node-pool capacity, and DNS/traffic management owned outside the cluster). See [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) (upgrade safety) for the rollback constraint checklist.

### Lessons from production deployments

Hard-won corrections from real AKS production deployments (2026-06-23 and 2026-06-27). Verify against current API versions before committing IaC.

- **`managedNATGateway` simplifies AKS-managed egress, but Standard NAT Gateway is *zonal, not zone-redundant*.** `userAssignedNATGateway` referencing a custom VNet fails AKS preflight validation in a single-template Bicep deployment (the VNet does not exist at validation time), so `managedNATGateway` with `managedOutboundIPProfile: { count: N }` is the pragmatic choice for a single-template AKS-managed VNet. But do **not** assume that gives zone-redundant egress: a Standard `managedNATGateway` is a single-zone resource and is a zone-failure single point for outbound. For zone-redundant egress use `managedNATGatewayV2` (StandardV2 SKU, zone-redundant by default; the AKS outbound *type* is still Preview) or UDR to a zone-redundant firewall. See [cluster-foundations. Outbound Connectivity](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md#outbound-connectivity). `[VERIFY]` managedNATGatewayV2 outbound-type status before production.
- **AcrPull must target the kubelet identity, not the control plane identity.** `aksCluster.identity.principalId` is the control plane identity (wrong target for image pulls). The kubelet identity, the one that pulls images, is `aksCluster.identityProfile.kubeletIdentity.objectId`. Assigning AcrPull to the control plane identity leaves pods stuck in `ImagePullBackOff`/`ErrImagePull` with `401 Unauthorized`. In Bicep the kubelet identity is not directly referenceable as a simple property; use `reference(aksCluster.id, '2024-09-02-preview', 'Full').identityProfile.kubeletIdentity.objectId`, or assign the role post-deploy via `az aks show -g <rg> -n <cluster> --query "identityProfile.kubeletidentity.objectId" -o tsv`.
- **Use the correct API property names for `2024-09-02-preview`.** `azureRBACProfile` does not exist → use `aadProfile: { enableAzureRBAC: true, managed: true, adminGroupObjectIDs: [...] }`. `monitoringProfile` does not exist → use `azureMonitorProfile: { metrics: { enabled: true } }`. `oagentAgentProfile` does not exist → omit it. `azurePortalFQDN` is read-only → omit it. Do not copy property names from older API versions without verifying against the target API.
- **Defender config has mutually exclusive states.** If `securityProfile.defender.securityMonitoring.enabled = false`, setting `logAnalyticsWorkspaceResourceId` is a validation error; if `enabled = true`, omitting it is a validation error. Simplest correct configuration: omit `defender` entirely from `securityProfile` unless Defender is explicitly required.
- **KEDA add-on moved from `addonProfiles` to `workloadAutoScalerProfile` in API `2026-03-01`.** Setting `addonProfiles: { keda: { enabled: true } }` silently fails, the cluster provisions without KEDA CRDs, and ScaledObject/TriggerAuthentication resources are rejected with `no matches for kind`. The correct property on newer AKS API versions is `workloadAutoScalerProfile: { keda: { enabled: true } }`. Verify the target API version: if `addonProfiles.keda` doesn't produce KEDA pods post-deploy, check `az aks show --query workloadAutoScalerProfile` to confirm where KEDA landed. `[VERIFY]` KEDA property location per AKS API version.
- **Companion Terraform-provider bug on the same property: `azurerm_kubernetes_cluster` `workload_autoscaler_profile.keda_enabled` perpetual-diff drift.** Removing the `workload_autoscaler_profile` block from HCL doesn't clear KEDA server-side: Azure reports `keda_enabled = false` back to the provider instead of `null`, so Terraform re-plans a change every run even though nothing changed. Reported and, as of this writing, open: [hashicorp/terraform-provider-azurerm#22360](https://github.com/hashicorp/terraform-provider-azurerm/issues/22360). Workaround: set `keda_enabled` explicitly (`true` or `false`) rather than omitting the block, or add `lifecycle { ignore_changes = [workload_autoscaler_profile] }` if the drift is cosmetic for your workflow. `[VERIFY]` issue status/fix version before relying on either workaround.
- **AKS LoadBalancer services fail when the cluster identity lacks Network Contributor on the workload VNet.** Symptom: `kubectl get svc` shows `EXTERNAL-IP: <pending>` indefinitely, with events `SyncLoadBalancerFailed: RESPONSE 403: LinkedAuthorizationFailed ... does not have permission to perform action(s) 'Microsoft.Network/virtualNetworks/subnets/join/action'`. The AKS cluster's system-assigned managed identity needs `Network Contributor` (or broader) on the VNet or subnet where nodes live. This role is NOT automatically granted by `az aks create` when using a custom VNet. Add it in Bicep: `resource aksClusterSubnetRole 'Microsoft.Authorization/roleAssignments@2022-04-01' = { scope: vnet, ... principalId: aksCluster.identity.principalId, roleDefinitionId: Network Contributor }`. Grant on the resource group if the LB also needs to create public IPs (`Microsoft.Network/publicIPAddresses/write`).
- **NetworkPolicy default-deny blocks Azure LoadBalancer traffic.** A `default-deny-ingress-egress` NetworkPolicy that only allows VNet CIDR (`10.0.0.0/16`) will block external LoadBalancer traffic. The Azure LB sends traffic from its frontend public IP to node pods, and the source IP after SNAT may not be in the VNet range. Add `ipBlock: { cidr: 0.0.0.0/0 }` to the ingress allowlist for any pod fronted by a LoadBalancer service, or use `externalTrafficPolicy: Local` and allow the LB health probe source. This is separate from the NetworkPolicy for AGC/Gateway API ingress (which uses the `app-routing-system` namespace selector).
- **AGC add-on compatibility is now a managed pinning model, with Helm as a supported independent path.** Current AKS guidance documents a supported Kubernetes-version-to-Gateway-API-bundle matrix for Managed Gateway API, and add-on behavior pins supported combinations instead of requiring you to manually curate ALB/Gateway API matrixes for managed installs. Since AKS release 2026-08-07 the ALB add-on is aligned to AKS minor versions - AKS automatically selects the compatible ALB controller image during cluster upgrades, reducing controller and feature-flag incompatibilities on the managed path. For managed lifecycle, use the AKS add-ons and validate the bundle mapping for your cluster version in the official matrix: [Managed Gateway API support matrix](https://learn.microsoft.com/azure/aks/managed-gateway-api#supported-kubernetes-versions-for-gateway-api-bundle-versions). If you need ALB controller and Gateway API versions independent of cluster-version pinning, Helm remains a supported self-managed option; tradeoffs are explicit ownership of controller upgrades and workload identity wiring. Version correction for architecture guidance: ALB Controller `v1.11.x` targets Gateway API `v1.5.1`, and `v1.5.1` CRDs are recommended with `v1.11+` for full feature functionality. Source of truth: [ALB controller release notes](https://learn.microsoft.com/azure/application-gateway/for-containers/alb-controller-release-notes). `[VERIFY]` exact chart/version pair and CRD bundle at implementation time.
- **BYO AGC (Application Gateway for Containers) requires `alb-id` annotation, not `alb-name`/`alb-namespace`.** When the AGC traffic controller is created via Bicep/IaC (not by the ALB controller itself), the Gateway resource must reference it with `alb.networking.azure.io/alb-id: <full-resource-id>`. NOT the `alb.networking.azure.io/alb-name`/`alb-namespace` annotation pair (which is for controller-managed AGC). The `alb-id` value is the full ARM resource ID (e.g., `/subscriptions/.../providers/Microsoft.ServiceNetworking/trafficControllers/<name>`). Using the wrong annotation causes the ALB controller to report `ApplicationLoadBalancer CRD not found for Gateway` indefinitely.
- **BYO AGC ALB controller needs data-plane permission on the AGC resource, not just ARM RBAC.** Even with `Contributor` or `Owner` on the AGC resource, the ALB controller's managed identity may still get `PermissionDenied: ALB Controller with object id ... does not have authorization to perform action on Application Gateway for Containers resource`. The AGC uses a data-plane gRPC permission model separate from ARM RBAC. When the AGC is created by the ALB controller add-on, this delegation is automatic. When created via Bicep (BYO), you must wire the identity at AGC creation time via the AVM module or assign the identity as the AGC's managed identity. `[VERIFY]` the correct delegation mechanism for BYO AGC scenarios.
- **Search Provisioner workload identity needs a dedicated ServiceAccount per identity.** The Azure Workload Identity webhook projects tokens based on the `azure.workload.identity/client-id` annotation on the ServiceAccount, NOT the `AZURE_CLIENT_ID` environment variable in the pod. If two different managed identities (e.g., main app identity and search provisioner identity) are federated to the same SA subject, only the SA-annotated client ID gets a valid token. Create a separate ServiceAccount per identity: `kubectl create sa search-provisioner-sa` with its own `azure.workload.identity/client-id` annotation, and create a matching federated credential for the new SA subject.
- **PodSecurity `restricted` rejects jobs without explicit securityContext.** AKS enforces Pod Security Standards on namespaces with `restricted` labels. Job pods without `securityContext.runAsNonRoot: true`, `securityContext.seccompProfile.type: RuntimeDefault`, `allowPrivilegeEscalation: false`, and `capabilities.drop: ["ALL"]` are silently rejected with `FailedCreate: violates PodSecurity "restricted:latest"`. Every Job manifest, including one-shot provisioner jobs, must include the full security context block.
- **Foundry/AI model deployment names must not be hardcoded in K8s manifests.** When the Bicep deploys `gpt-5-mini-support-bot` as the Foundry model deployment but the K8s Deployment YAML hardcodes `Foundry__DeploymentName=gpt-4o-mini-support-bot` as a plain `value:` string, the application silently falls back to local stub providers or crashes on AI calls. The model deployment name changed when upgrading from `gpt-4o-mini` to `gpt-5-mini`, but the K8s manifest wasn't updated. Fix: parameterize the deployment name via ConfigMap or env var from Bicep outputs, OR add the model name to a post-provision config-regeneration script that regenerates ConfigMaps after provisioning. Never hardcode AI model deployment names in static K8s YAML, they change when models are upgraded.
- **Entra ID app registrations need redirect URIs for browser-based MSAL auth.** The SPA frontend uses `@azure/msal-browser` with PKCE flow. If the app registration has no `web.redirectUris` entries, the browser auth redirect fails silently, the frontend loads but all protected features are inaccessible. After deploying behind a new domain (e.g., Front Door, Application Gateway), add the domain as a redirect URI: `az ad app update --id <client-id> --web-redirect-uris "https://<new-domain>/" "http://localhost:5173/"`. Verify the API scope exists: `az ad app show --id <client-id> --query "api.oauth2PermissionScopes"`.
- **`--enable-acns` alone enables only FQDN *filtering*; L7 *policy* is silently ignored without `--acns-advanced-networkpolicies L7`.** A `CiliumNetworkPolicy` containing `toPorts.rules.http` or `rules.grpc` enforces nothing until the cluster is created or updated with `az aks {create,update} --enable-acns --acns-advanced-networkpolicies L7`. There is no admission error and no controller warning, the L7 rule simply passes through unenforced. Compounding trap: **L7 rules are not supported in `CiliumClusterwideNetworkPolicy` (CCNP) at all**, only in namespaced `CiliumNetworkPolicy`; an L7 rule in a CCNP is likewise silently ignored. If you adopted ACNS for L7 policy and see HTTP/gRPC traffic not being dropped, confirm `--acns-advanced-networkpolicies L7` is set (`az aks show --query networkProfile.advancedNetworking`) and that the policy is a namespaced CNP, before debugging selectors. Evidence: [Microsoft Learn. Use ACNS on your AKS cluster (cilium)](https://learn.microsoft.com/azure/aks/use-advanced-container-networking-services). `[VERIFY]` flag syntax per AKS CLI version.

## AI / GPU Workload Routing

When the context includes AI inference, model serving, GPU node pools, or accelerator scheduling, also load [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md), [production-workload-controls](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/production-workload-controls/guide.md), and [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md).

Guidance for new AKS GPU/AI designs:

- Prefer separate GPU node pools with taints, labels, resource requests, topology spread rules, and quota validation.
- Consider AKS-managed GPU node pools where supported; verify NVIDIA-only support, region availability, migration limitations, and whether the pool must be recreated.
- When multiple workloads share a physical GPU, choose a partitioning strategy (**MIG** for production hardware isolation on A100/H100/H200, **time-slicing** or **MPS** for experimentation) and own the NVIDIA GPU Operator lifecycle for user-managed strategies. See [GPU Partitioning](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md#gpu-partitioning-sharing-a-single-physical-gpu).
- When batch training, fine-tuning, or data-processing jobs compete for finite GPU quota or need all resources before starting (gang scheduling), introduce **Kueue** (ClusterQueue + LocalQueue) for admission control. Kueue is open-source community-supported software excluded from AKS SLA; pair it with node autoscaler/NAP so admitted jobs get capacity. See [Batch and AI Training Scheduling](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md#batch-and-ai-training-scheduling-kueue).
- Consider the AKS AI toolchain operator / KAITO add-on for supported self-hosted open-source model inference workloads (as of Build 2026, KAITO is one provider under **AI Runway** - see Build 2026 platform additions below). Verify current model support, GPU VM quota, OS SKU limitations, region support, and AKS Automatic compatibility before recommending it.
- For **self-hosted LLM ingress** (vLLM etc.), evaluate the **AGC inference gateway (Public Preview)**: Gateway API Inference Extension (`InferencePool`/EPP/BBR) for model-aware routing, Helm-only ALB controller install (`--set albController.aiGateway=true`), WAF pairing. See the [August 11, 2026 research pass](#august-11-2026-research-pass) and [workload-platform](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md#inference-gateway-for-self-hosted-llm-serving-agc-public-preview).
- Use KEDA with GPU metrics only after confirming Managed Prometheus/DCGM metric availability, min-replica/cold-start behaviour, PDB/drain behaviour, and GPU quota for the target model.

## Build 2026 platform additions

Announced at Microsoft Build 2026. Verify GA/Preview status, region, and AKS Automatic/Standard compatibility against current Microsoft Learn and the AKS release tracker before committing:

- **AKS on bare metal (Public Preview, June 2026).** Run AKS on dedicated machines with no hypervisor - direct access to NVLink, RDMA, and high-performance networking, using the same AKS control plane and APIs. Target: large training jobs, latency-sensitive inference, and high-throughput data pipelines where the virtualization layer is a measurable cost. Treat as Preview; `[VERIFY]` region/SKU availability and the operational model vs standard node pools ([AKS Engineering Blog, 2026-06-02](https://blog.aks.azure.com/2026/06/02/aks-baremetal-public-preview)). Further coverage and positioning: [Thomas Maurer, July 2026](https://www.thomasmaurer.ch/2026/07/azure-kubernetes-service-aks-on-bare-metal/), confirms on-prem/edge/sovereign scenarios as first-class use cases.

## July 18, 2026 platform updates

Verified via MRC-MCP-Server primary source on 2026-07-18.

- **Encryption in Transit (EiT) for Azure Files NFS v4.1 volumes. GA** (id 567787, Launched July 2026; was Preview May 2025). The Azure File CSI driver now generally supports EiT for NFS v4.1 volumes via the AZNFS mount helper and Stunnel. Data in transit between AKS workloads and Azure Files NFS shares is encrypted using TLS, helping meet security and compliance requirements for production workloads. Configured via storage class without application changes. Available in all Azure regions that support Azure Files Premium (SSD). **Default recommendation:** enable EiT for new NFS-based Azure Files storage classes in security- or compliance-sensitive clusters. Source: https://azure.microsoft.com/updates?id=567787.
- **AKS Node Disruption Policy. Public Preview** (Tech Community, Pixel Robots, July 2026; confirmed in AKS 2026-07-17 release). Control when nodes get reimaged during routine cluster config changes. Lets platform teams block mid-business-hours reimaging and require drain windows. `[VERIFY]` the official Azure Update ID and feature flag before recommending beyond awareness. Source: https://pixelrobots.co.uk/2026/07/aks-node-disruption-policy-is-now-in-preview-control-when-your-nodes-get-reimaged/.

## July 29, 2026 currency updates

Verified against AKS release 2026-07-17 and 2026-05-29 / Microsoft Learn on 2026-07-29.

- **Artifact Streaming. GA** (was Preview in AKS 3.7.x guidance). AKS now streams container images from ACR, only pulling layers needed for initial pod startup. Reduces time-to-pod-readiness by ~15% for images <30 GB. Enabled via `az aks nodepool update --enable-artifact-streaming` on Linux Ubuntu/Azure Linux node pools. Windows node pools not supported. Updates to the AKS release tracker: [AKS #3928](https://github.com/Azure/AKS/issues/3928). Source: [AKS 2026-07-17 release](https://github.com/Azure/AKS/releases/tag/2026-07-17).
- **Windows Server 2025. GA** (was Preview in prior guidance). Windows Server 2025 node pools are now GA, no feature flag required. Requires K8s 1.32+ and Azure CLI 2.87.0+. Ships containerd 2.0, Gen2 VMs, and FIPS by default. Starting K8s 1.37 (Oct 2026), Windows2025 becomes the default Windows OS SKU. Windows Server 2019 retired (March 2026); Windows Server 2022 extended to June 2028. Source: [AKS 2026-05-29 release](https://github.com/Azure/AKS/releases/tag/2026-05-29), [Upgrade Windows OS](https://learn.microsoft.com/azure/aks/upgrade-windows-os).
- **In-place node pool resize. Public Preview.** Resize the VM SKU of an existing VMSS node pool via `az aks nodepool update --node-vm-size <new-sku>`. Uses the rolling upgrade engine (surge + cordon/drain + delete). Requires AKS API `2026-01-02-preview` and `aks-preview` CLI extension. Blocked on `--max-surge 0`. See [Resize node pools](https://learn.microsoft.com/azure/aks/resize-node-pool).
- **Automatic PDB Management. Public Preview.** AKS cluster extension (`microsoft.evictionautoscaler`) that auto-creates PDBs for unprotected Deployments and temporarily scales up replicas to unblock node drain when a PDB would prevent eviction. Installed via `az k8s-extension create`. Supports namespace-scoped opt-in or cluster-wide mode. See [Automatic PDB management](https://learn.microsoft.com/azure/aks/automatic-pod-disruption-budget-management).
- **Azure Container Linux (ACL). GA.** A lightweight, Microsoft-maintained container-optimized OS, now GA as an AKS node OS option starting K8s 1.34+. Separate from Azure Linux; designed for reduced configuration drift and simplified fleet maintenance. Supports in-place OS SKU migration. See [Azure Container Linux](https://learn.microsoft.com/azure/azure-linux/azure-container-linux-overview).
- **NVIDIA RTX PRO 6000 Blackwell Server Edition GPU, supported.** AKS supports these new GPU VM sizes as managed GPUs on Ubuntu node pools (NVIDIA GRID driver). Source: [AKS 2026-06-19 release](https://github.com/Azure/AKS/releases/tag/2026-06-19).
- **FIPS on Ubuntu 22.04. GA.** FIPS 140-3 compliance is now available on Ubuntu 22.04 node pools. Per Microsoft Learn, FIPS-enabled node pools require Kubernetes 1.19 or greater (no higher version gate) with the `Ubuntu` OS SKU; Ubuntu 24.04 doesn't currently support FIPS, so AKS falls back to Ubuntu 22.04 when FIPS is requested. FIPS can be combined with Trusted Launch on Ubuntu 22.04 Gen2 VM sizes. Sources: [AKS 2026-05-29 release](https://github.com/Azure/AKS/releases/tag/2026-05-29), [Enable FIPS for AKS node pools](https://learn.microsoft.com/azure/aks/enable-fips-nodes).
- **`enableCustomCATrust`, retiring September 14, 2026.** After this date the preview property will no longer enable Custom Certificate Authority on node pools. Update affected clusters and remove the property (`--disable-custom-ca-trust`). See [Azure/AKS #5826](https://github.com/Azure/AKS/issues/5826).
- **Kubernetes version status (July 2026).** K8s 1.30 is now deprecated; K8s 1.33 is now LTS-only (standard support ended July 2026). K8s 1.36 GA/LTS is the current recommended target for new clusters. See [supported K8s versions](https://learn.microsoft.com/azure/aks/supported-kubernetes-versions).


    Factor into the AKS Automatic vs Standard decision in [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md).
- **Azure Kubernetes Fleet Manager for Arc-enabled clusters - GA.** Fleet-level update orchestration, policy, and workload placement now extend beyond Azure to Arc-connected clusters. See [fleet-management](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/fleet-management/guide.md).
- **Anyscale on Azure (managed Ray on AKS) - Public Preview.** Ray coordinates distributed execution while AKS handles scheduling and cluster lifecycle; relevant for distributed training/inference. `[VERIFY]` before production.
- **kars. Agent Reference Stack for Kubernetes (open source, July 2026).** An open-source, Kubernetes-native reference stack from Microsoft's Azure Cloud Native team ([Azure/kars](https://github.com/Azure/kars)) for running AI agents on AKS. It is a reference implementation, not an officially supported Microsoft product (no SLA or support contract). Defense-in-depth isolation: a no-network agent container (UID 1000), a Rust inference router (UID 1001) that holds all credentials and governs egress via an L7 allowlist/blocklist, per-sandbox Microsoft Entra Agent ID or cluster Workload Identity via federated OIDC/IMDS, and an encrypted AgentMesh (Signal Protocol) with an A2A gateway. Designed for hosting untrusted agent code, multi-tenant agent platforms, and MCP-server-backed agents (the router includes an MCP gateway). Microsoft Agent Framework is one of its supported runtimes. Optionally combine with AKS Pod Sandboxing (Kata Containers, `kata-vm-isolation`) for a per-pod guest kernel. Route to [agent-runtime](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/agent-runtime/guide.md). Pair with [microsoft-agent-framework](../microsoft-agent-framework/SKILL.md) for agent development and [mcp-server-design](../mcp-server-design/SKILL.md) for MCP server hosting.

## August 11, 2026 research pass

Verified 2026-08-11 via context7/Microsoft Learn, Azure Updates MCP (primary source), AKS releases 2026-07-17 / 2026-06-19 / 2026-05-29, and a public GitHub codebase survey. **Superseded 2026-08-25:** AKS release **2026-08-07** (published 2026-08-11) now exists - see [August 25, 2026 currency updates](#august-25-2026-currency-updates); items below remain valid as of their verification date. Additions below are items from that release and the Azure Updates feed that prior currency passes had not yet captured.

- **AGC inference gateway. Public Preview (Azure Updates 566516, June 2026).** Application Gateway for Containers now supports the Kubernetes [Gateway API Inference Extension](https://gateway-api-inference-extension.sigs.k8s.io/): `InferencePool`, `InferenceObjective`, a customer-provided **Endpoint Picker (EPP)**, and a managed **Body-Based Router (BBR)** that reads the `model` field of OpenAI-compatible requests for model-aware routing. Purpose-built for self-hosted LLM serving on AKS (vLLM etc.): lowers time-to-first-token, routes around saturated replicas, pairs with AGC WAF. **Not supported through the AKS ALB add-on**: ALB Controller must be installed via Helm with `--set albController.aiGateway=true` (installs Inference Extension CRDs v1.3.1). Route AI-ingress designs here: [AGC inference gateway](https://learn.microsoft.com/azure/application-gateway/for-containers/inference-gateway), [how-to](https://learn.microsoft.com/azure/application-gateway/for-containers/how-to-inference-gateway). `[VERIFY]` CRD/EPP chart versions at implementation time.
- **Automatic zone placement** — enabled **globally as of AKS release 2026-08-07** (was Preview at the 2026-07-17 release; see [August 25, 2026 currency updates](#august-25-2026-currency-updates)). AKS dynamically selects the best availability-zone set for a node pool instead of requiring manual zones per region/SKU; existing VMSS pools can move to `availabilityZones=["auto"]` post-rollout. Verify interaction with topology-spread controls before recommending for zone-pinned designs.
- **Full caching mode for Ephemeral OS disks. Public Preview (AKS 2026-07-17).** Caches the entire OS locally so nodes keep running when remote storage is unavailable; improves node resiliency and OS disk performance beyond the default ephemeral behavior. Evaluate for stateful/GPU pools where node uptime during storage outage matters. See [full-cache ephemeral OS disk](https://learn.microsoft.com/azure/aks/full-cache-ephemeral-os-disk).
- **Secure TLS bootstrapping. Enabled by default in `westcentralus` and `eastasia` (AKS 2026-07-17).** Kubelet TLS bootstrap using short-lived credentials; regional default rollout continues. Track regional enablement via [Azure/AKS #5694](https://github.com/Azure/AKS/issues/5694); no action required today, but node-provisioning tooling that assumes long-lived kubelet certs will need review as regions roll out.
- **Trusted Launch (vTPM + Secure Boot) can now be enabled and disabled on existing Linux node pools** (AKS 2026-07-17), and **Secure Boot is now supported with GPUs on Azure Linux**. Previously a create-time-only choice; now a Difficult (not Permanent) decision for new pools. `[VERIFY]` disable path constraints before planning rollback.
- **NAP behavioral changes (AKS 2026-07-17).** NAP-enabled clusters now set `kubernetes.azure.com/mode: user` on the default NodePool (prevents pending system workloads from triggering user-node scale-up) and use an In-VM spot rebalancing signal for proactive spot replacement. Relevant to NAP node-pool taxonomy and spot workload design.
- **kube-proxy `nftables` mode rejected on K8s < 1.33** (AKS 2026-07-17). Previously accepted-then-silently-fell-back to `iptables`; the request now fails at API time. Only affects clusters explicitly configuring kube-proxy mode below 1.33.
- **CNI Overlay Dual-Stack (IPv4/IPv6) on Windows no longer requires preview feature registration** (AKS 2026-07-17). Simplifies Windows dual-stack adoption; verify version gates.
- **Deployment Safeguards Enforce mode now applies default resource requests to DaemonSets and Jobs** in addition to Deployments/StatefulSets (AKS 2026-06-19). If relying on Safeguards to inject resource requests, re-test DaemonSet/Job admission after upgrade. AKS Automatic clusters with managed system node pools can edit `excludedNamespaces`; `/var/log` and `/hostfs` hostPath read-only allowed under Baseline PSA; blocking restrictions on managed system pools: customer SSH keys, `Service` `spec.externalIPs` (ValidatingAdmissionPolicy, aligned with upstream K8s 1.36 deprecation), `kubectl port-forward` on managed pools, `kube-system` secret reads, and unauthorized mutating admission bindings. See [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (AKS Automatic vs Standard) for the managed-system-node-pool decision surface.
- **App Routing with Gateway API: GA date verified April 2026.** The [AKS release 2026-04-28](https://github.com/Azure/AKS/releases/tag/2026-04-28) is the authoritative GA announcement ("Gateway API-based ingress for the application routing add-on is now generally available"); the Azure Updates entry 567944 published 2026-07-28 is the formalized Azure Updates post, not the GA date. Flag trap confirmed: `--enable-app-routing` = legacy NGINX mode; **`--enable-app-routing-istio`** = Gateway API mode; new AKS Automatic clusters on AKS 1.36+ preconfigure Gateway API by default. GatewayClass `approuting-istio`. **Aug 2026 update:** the App Routing DNS/TLS integrations (Azure DNS records + Key Vault certificate TLS termination) now also reconcile Gateways using `gatewayClassName: istio` from the Istio service mesh add-on, previously unsupported. The two Gateway API control planes (`approuting-istio` vs the mesh add-on's `istio`) are mutually exclusive; pick one, and pair it with the always-required `--enable-app-routing`. AKS release 2026-08-07 also fixed a bug where App Routing on Automatic 1.36+ clusters could incorrectly default to NGINX instead of Istio/Gateway API mode during creation - verify the resulting GatewayClass on newly provisioned Automatic clusters rather than assuming either default. See the corrected [non-negotiables entry](#non-negotiables-for-new-aks-designs).
- **Hyperscale configuration via control plane scaling profile: Public Preview** (official doc live Aug 2026). Preprovisioned, guaranteed-capacity control plane tiers (`controlPlaneScalingProfile.scalingSize`: H2/H4/H8) replacing the default *dynamic* control-plane scaling, for clusters with API-server concurrency, pod-scheduling, or etcd pressure at scale. **Permanent decision**: creation-only and irremovable (revert = delete-and-recreate); resizing between H-tiers after creation is supported (`az aks update --control-plane-scaling-size`). Requires Standard/Premium tier, K8s 1.33+, `aks-preview` ≥ 21.0.0b8, and the `ControlPlaneScalingProfilePreview` subscription flag; one cluster per subscription per region during preview; CLI/REST/ARM only (no Terraform/SDK). Avoid the Azure Policy add-on on large hyperscale clusters. Full tier table, etcd partition model, and when-to-use guidance: [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md#hyperscale-control-plane-scaling-profile-public-preview). `[VERIFY]` regional availability and H-tier pricing (unpublished). Source: [hyperscale-configuration-aks](https://learn.microsoft.com/azure/aks/hyperscale-configuration-aks).
- **AKS Automatic on K8s 1.36+ can disable the default application routing (Gateway API) add-on** to use the Istio-based service mesh add-on instead (AKS 2026-06-19). Reconciles with the App Routing GA correction above and the [operations-resilience ingress guidance](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md).
- **Pod Sandboxing (Kata) gates (AKS 2026-06-19).** FIPS (`--enable-fips-image`) is rejected on Kata runtime node pools (no FIPS compliance in the Kata node image); Kata pools now supported on `Standard_DadsV7` series. Feed into [agent-runtime](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/agent-runtime/guide.md) Kata guidance.
- **Confidential VMs with Azure Linux. GA** (AKS 2026-06-19): the recovery path for the retired per-pod Confidential Containers preview. See [cluster-foundations](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) (CVM node pools).
- **Windows gMSA now validates CoreDNS conflicts** (AKS 2026-06-19): enabling gMSA or changing its root domain is rejected when the `coredns-custom` ConfigMap defines the same domain, preventing a duplicate zone that would crash CoreDNS.
- **Istio asm-1-30 (AKS 2026-07-17):** new default proxy redirection uses Istio CNI (init-container model retired for new installs); ISTIO-SECURITY-2026-005 fixed across `asm-1-27`–`asm-1-30`; restart workload pods to re-inject patched proxies. `asm-1-27` deprecated (2026-05-29).
- **GPU driver currency gate (AKS 2026-07-17):** `NVadsA10v5` / `NCadsA10v4` node pools must run node image **202606.08.1 or later** for NVIDIA v18.x host-driver compatibility (Azure/AKS #5875). Include in GPU-pool validation checklist.
- **Karpenter provider v1.14.0** (AKS 2026-07-17), adding the Balanced consolidation policy for lower node churn (2026-08-07 release); **KubeBuddy v0.0.36** (2026-08-13); **kaito v0.12.0** (2026-08-22); **Azure/kars** active (pushed 2026-08-25). Reference-point currency for the ecosystem tools cited in this skill.

## August 25, 2026 currency updates

Verified against [AKS Release 2026-08-07](https://github.com/Azure/AKS/releases/tag/2026-08-07) (published 2026-08-11 - newest release), Microsoft Learn pages current as of 2026-08-24/25, and azure-aks-docs commits through 2026-08-24.

- **Node Pool Rollback. GA.** `az aks nodepool rollback` restores a pool's previous Kubernetes version **and** node image within **7 days** of upgrade completion (node-image-only rollback also works if only the image changed in-window). Constraints: all-or-nothing per pool; no concurrent operations (abort first); disable the auto-upgrade channel before rolling back; cannot revert OS-SKU changes; cannot roll back to an unsupported version; monitor via Activity Log / Operation Status API. Temporary triage only - re-upgrade within ~30 days. Control-plane rollback remains unsupported. See [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) for the constraint checklist and window/maintenance interaction.
- **Automatic availability zone placement enabled globally** (was Public Preview). New VMSS/VirtualMachines node pools accept `availabilityZones=["auto"]`, and existing VMSS pools can be updated to `["auto"]` once regional rollout completes - AKS picks the zone set instead of manual per-region/SKU selection. Still verify interaction with explicit topology-spread controls for zone-pinned designs.
- **Control-plane-only upgrades into LTS permitted** while satisfying version-skew (pools may trail the control plane up to N-3 on K8s 1.28+): upgrade the control plane into an LTS target, validate workloads, then upgrade node pools. This softens "all pools must move with the control plane" from hard rule to best practice - see the corrected post-deployment checklist note.
- **Kubernetes version status (August 2026):** 1.36 GA/LTS remains the recommended target today; **1.37 previews September 2026 and GAs October 2026** (Windows Server 2025 becomes the default Windows OS SKU at 1.37). Plan the 1.37 reimage-trigger list (below) into change-control before adopting.
- **AKS 1.37+ behavioral changes (reimage triggers):** changes to SSH node access configuration, IMDS restriction, network-isolated bootstrap profile, or cluster outbound type now trigger an immediate node reimage; converting a cluster from service-principal to managed-identity auth reimages all pools. Use Node Disruption Policy to gate these windows - details in [operations-resilience](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md) (Node Disruption Policy).
- **Upgrade preflight:** node-pool upgrades are rejected when current pool size plus effective surge would exceed the VMSS 1,000-instance limit - size surge accordingly on very large pools.
- **EiT × Workload Identity for Azure Files. GA** - Azure File CSI driver now generally supports accessing Azure Files via workload identity alongside the July EiT NFS v4.1 GA.
- Minor stamps: Istio Gateway API deployments now set `automountServiceAccountToken: false`; GPU MIG slice-width is validated against VM SKU capacity at request time; Entra ID SSH is rejected on AzureContainerLinux pools (immutable OS); NAP can now be enabled on clusters with restricted `publicNetworkAccess` (incl. private API server VNet-integrated w/ UDR); Capacity Reservation Groups can be associated with **existing** node pools on preview API `2026-01-02-preview+` (zonal pools roll cordon/drain/reboot; regional non-zero pools must scale to zero first).

- **Reference implementations to cite (public GitHub survey 2026-08-11):** the canonical new-cluster reference repos are now **Azure/AKS-Landing-Zone-Accelerator** (Bicep/Terraform/ARM scenario packages, includes AKS Secure Baseline private-cluster scenario; `aks-baseline` repos 404 (they were absorbed into the LZA program)) and **Azure/Aks-Construction** (Bicep templating helper, web UI + CI/CD artifacts). Terraform consumers: `Azure-Samples/aks-openai-terraform`, `paolosalvatori/private-aks-cluster-terraform-devops`, `XenitAB/terraform-modules` all use `workload_autoscaler_profile { keda_enabled }` / `azure_rbac_enabled` / `oidc_issuer_enabled`, live confirmation of the [KEDA API-version lesson](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md) and identity non-negotiables below. Public HCL confirms the Cilium trap: pairing `network_policy = cilium` with the default Azure data plane fails with `NetworkPolicyCiliumRequiresCiliumDataplane` (e.g. TesslateAI/OpenSail `k8s/terraform/azure/aks.tf`).

## August 26, 2026 currency updates

Verified 2026-08-26 against Microsoft Learn pages fetched live, Azure Updates feed, AKS releases API (newest remains 2026-08-07), and azure-aks-docs commits through 2026-08-25.

- **Control plane metrics via Managed Prometheus. GA** (Azure Updates, Aug 2026). Enable with `az aks {create,update} --enable-control-plane-metrics --enable-azure-monitor-metrics`; requires managed-identity auth; not supported with Private Link; self-hosted Prometheus cannot scrape it. Default targets: apiserver + etcd on; scheduler/controller-manager/cluster-autoscaler/NAP off until enabled in configmap schema v2. Route: [operations-resilience Observability Baseline](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/operations-resilience/guide.md#observability-baseline).
- **ACNS eBPF Host Routing. GA June 2026.** Performance mode via `--acns-datapath-acceleration-mode BpfVeth`; cluster-wide only; bypasses and blocks host iptables rules; **incompatible with Static Egress Gateway, Confidential VMs, Pod Sandboxing, Windows nodes, Istio Ambient self-managed**; K8s ≥1.33, CLI ≥2.71, Azure Linux 3.0/Ubuntu 24.04 only. Route: [cluster-foundations eBPF Host Routing](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/cluster-foundations/guide.md#ebpf-host-routing-acns-container-network-performance-ga).
- **Artifact Streaming on NAP pools** configurable per AKSNodeClass (`spec.artifactStreaming`, docs added Aug 2026); Premium-tier ACR prerequisite unchanged. Route: [workload-platform autoscaling rules](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/workload-platform/guide.md).
- **Fleet auto-upgrade profile CLI surface corrected:** `az fleet autoupgradeprofile create --channel {NodeImage|Rapid|SecurityPatch|Stable|TargetKubernetesVersion}` with `--update-strategy-id` (`--upgrade-type`/`--update-strategy-name` are not current flags); Security patch channel documented at [kubernetes-fleet/update-automation](https://learn.microsoft.com/azure/kubernetes-fleet/update-automation) (the prior citation target 404s).

## Workflow

1. Gather region, environment tier, availability target, networking constraints, identity/RBAC model, workload shape, compliance needs, and operational ownership.
2. Decide whether AKS Automatic or Standard is appropriate before selecting lower-level features.
3. Classify decisions as permanent, difficult, or reversible.
4. Define production workload controls before finalising manifests: namespace model, labels, ResourceQuota, LimitRange, NetworkPolicy, replica/PDB/topology spread, and autoscaling roles.
5. Recommend modern defaults for new work, but note when a legacy migration path is safer.
6. If more than one cluster is in scope, decide whether Fleet Manager is required and classify hubless vs hub, hub network access, member labels, update stages, placement, managed namespaces, and cross-cluster networking.
7. Provide validation commands and stop conditions for implementation.
8. Record permanent or disruptive choices in an ADR.

## Implementation Handoff

This skill covers **AKS architecture decisions** and includes production workload-control examples. When the task moves from architecture to implementation, hand off to the relevant repository or organisation standards instead of inventing local conventions.

Use or request the implementation standard for:

- Terraform/Bicep module layout, state isolation, provider versions, and destructive-change gates.
- Azure naming, tagging, management group, subscription, resource group, and Azure Landing Zone placement.
- Private DNS zone ownership, hub-spoke routing, firewall policy, and Private Link governance.
- Log Analytics, Azure Monitor Workspace, Managed Prometheus, Managed Grafana, Sentinel/SIEM, and retention placement.
- Azure Container Registry, Key Vault, Defender for Cloud, Azure Policy/Deployment Safeguards, and workload identity integration.
- GitOps repository layout, environment promotion, pull-request gates, and break-glass reconciliation.
- Kubernetes manifest style, Helm/Kustomize conventions, image signing, admission controls, and progressive delivery.

**IaC surface translation.** This skill uses Azure CLI flag names and Bicep/ARM property names (e.g., `--node-os-upgrade-channel`, `nodeOsUpgradeChannel`, `aksManagedAutoUpgradeSchedule`, `oidcIssuerProfile`, `securityProfile.workloadIdentity`). Terraform's `azurerm_kubernetes_cluster` and `azurerm_kubernetes_fleet_*` resources expose the same concepts under different argument names - verify current azurerm provider documentation when translating. Fleet Manager operations cited via `az fleet ...` have equivalents in `Microsoft.ContainerService/fleets` (Bicep/ARM) and the `azurerm_kubernetes_fleet*` resource family (Terraform); the architectural decision is the same across surfaces.

Use this skill when manifest or IaC work exposes architecture decisions such as node placement, private networking, NetworkPolicy, namespace boundaries, Gateway API, Workload Identity, autoscaling, replicas, PDBs, observability, DR, cost, or production readiness.

## Anti-Hallucination Rule

Before writing AKS CLI commands, Bicep/ARM properties, Terraform arguments, Kubernetes manifests, Gateway API resources, or Fleet Manager configuration, verify exact identifiers against the current AKS, Kubernetes, and provider documentation.

Forbidden shortcuts:

- Do not guess `az aks` or `az fleet` flags; verify with the installed CLI or Microsoft Learn.
- Do not assume Bicep/ARM property names map directly to Terraform argument names.
- Do not invent Kubernetes API versions, NetworkPolicy fields, Gateway API kinds, or workload identity annotations.

### Hidden API Surface Properties (commonly missed)

The `Microsoft.ContainerService/managedClusters` resource schema exposes 100+ properties. These are present in the Bicep/ARM/CLI surface but easily missed when relying on AKS overview documentation:

| Property | Impact |
|---|---|
| `securityProfile.azureKeyvaultKms` | Key Vault key encryption at cluster level (distinct from etcd encryption) |
| `securityProfile.workloadIdentity.enabled` | Cluster-level OIDC/workload identity toggle (skill covers pod-level SA annotation, not the Bicep toggle) |
| `securityProfile.imageCleaner.enabled` + `intervalHours` | Auto-clean unused container images — Bicep shape for Image Cleaner |
| `securityProfile.customCATrustCertificates` | BYO custom CA trust for private registries/proxies |
| `autoUpgradeProfile.upgradeChannel` + `nodeOSUpgradeChannel` | Bicep enums for cluster and node OS upgrade channels — separate from `aksManagedAutoUpgradeSchedule` maintenance windows |
| `networkProfile.networkDataplane: 'azure' \| 'cilium'` | Dataplane enum that determines Cilium vs Azure CNI — skill discusses CNI conceptually but not the enum |
| `networkProfile.podCIDRs[]` / `serviceCIDRs[]` | IPAM CIDR lists as Bicep properties (not just conceptual CIDR planning) |
| `networkProfile.loadBalancerProfile` | Managed LB SKU, backend pool type, quota allocation |
| `networkProfile.natGatewayProfile` | Managed NAT Gateway profile properties |
| `storageProfile.diskCSIDriver` / `fileCSIDriver` / `snapshotController` / `blobCSIDriver` | Per-driver enable/disable toggles at cluster level |
| `ingressProfile.webAppRouting` | Web App Routing add-on at the cluster level |
| `workloadAutoScalerProfile.keda.enabled` | KEDA at the cluster level (not just per-deployment) |
| `metricsProfile` | Managed Prometheus / cost analysis add-on enablement |
| `azureMonitorProfile` | Container Insights enablement at the cluster level |
| `serviceMeshProfile` | Istio-based service mesh add-on configuration |
| `apiServerAccessProfile` | Authorized IP ranges, private cluster (`enablePrivateCluster`), VNet integration |
| `nodeProvisioningProfile.mode: 'Auto' \| 'Manual'` | AKS Automatic node provisioning mode |
| `aiToolchainOperatorProfile.enabled` | KAITO/AI toolchain operator enablement at cluster level (skill covers KAITO conceptually via workload-platform) |
| `bootstrapProfile` | Network-isolated cluster bootstrap (`artifactSource: Cache` + container registry config) pairing with the `none`/`block` outbound types |
| `httpProxyConfig` | Egress through a corporate HTTP proxy (trusted CA + proxy FQDN/IP list) for hub-spoke designs without NAT/UDR inspection |
| `podIdentityProfile` | Legacy pod identity surface - prefer Workload Identity (`securityProfile.workloadIdentity` + OIDC issuer) |
| `diskEncryptionSetID` | Customer-managed-key disk encryption set for node OS/data disks (pairs with etcd CMK guidance in routing) |
| `powerState` | Cluster start/stop state - dev/test cost lever, not for production designs |
| `addonProfiles` (30+ add-ons) | Full Bicep map of add-on enablement (`httpApplicationRouting`, `azurepolicy`, `azureKeyvaultSecretsProvider`, `ingressApplicationGateway`, `openServiceMesh`, `confCompliance`, etc.) |

Asymmetry traps:

- AKS Automatic and AKS Standard have different exposed choices.
- Azure CNI Overlay, Azure CNI Pod Subnet, kubenet, and Cilium have different defaults and support boundaries.
- Kubernetes API support depends on cluster version and installed CRDs.
- **Monitoring enablement flags do not enable collection.** On the ARM/Bicep path, `omsAgent.enabled` (AVM `omsAgentEnabled`) and `azureMonitorProfile.metrics.enabled` each deploy their agent and nothing more. Collection requires a **data collection rule associated to the cluster**, a Container Insights DCR for logs, and a DCE/DCR/DCRA for Managed Prometheus. Verified 2026-08-10: with the addon enabled, the correct workspace configured, and `ama-logs` pods Running, every Container Insights table was empty for the life of the cluster and nothing reported an error. `az aks enable-addons` creates the rule for you; Bicep does not. Check with `az monitor data-collection rule association list --resource <cluster-id>` and expect one association per collection path. See the `observability-monitoring` skill for the DCR shapes.

Safe degraded output:

- Use `[VERIFY] AKS identifier: confirm CLI flag, provider property, or Kubernetes API version` when unverifiable.

## Resource Directories

Load these directories only when the selected workflow needs deeper examples, prompts, or reference material:

- [references/](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/references)

External reference implementations to cite alongside this skill (surveyed 2026-08-11):

- [Azure/AKS-Landing-Zone-Accelerator](https://github.com/Azure/AKS-Landing-Zone-Accelerator) — canonical new-cluster reference implementation (Bicep/Terraform/ARM scenario packages, incl. AKS Secure Baseline private-cluster scenario). The legacy `Azure/aks-baseline` / `Azure-Samples/aks-baseline` repos 404 — they were absorbed into the LZA program.
- [Azure/Aks-Construction](https://github.com/Azure/Aks-Construction) — Bicep templating helper with a web UI; quick artifact generation for CI/CD pipelines.
- [Azure-Samples/aks-openai-terraform](https://github.com/Azure-Samples/aks-openai-terraform), [paolosalvatori/private-aks-cluster-terraform-devops](https://github.com/paolosalvatori/private-aks-cluster-terraform-devops), [XenitAB/terraform-modules](https://github.com/XenitAB/terraform-modules) — production Terraform AKS modules demonstrating the identity/KEDA patterns this skill mandates (`workload_autoscaler_profile { keda_enabled }`, `azure_rbac_enabled`, `oidc_issuer_enabled`). See the Terraform `keda_enabled` perpetual-diff caveat under Lessons from production deployments before wiring these modules' KEDA block.
- [KubeDeckio/KubeBuddy](https://github.com/KubeDeckio/KubeBuddy) — agentless AKS best-practice validation (v0.0.36 as of 2026-08-13; `[VERIFY]` check IDs/count against current release).
- [Azure/kars](https://github.com/Azure/kars) — Agent Reference Stack for Kubernetes (AI agents on AKS; active, pushed 2026-08-11). See [agent-runtime](https://skilld.dev/api/skills-raw/lukemurraynz/hve-agent-skills/aks-cluster-architecture/bundles/agent-runtime/guide.md).

## Final Output Contract

Every engagement produces:
- A decision-oriented AKS architecture recommendation with permanent, difficult, and reversible choices called out.
- A workload-ready control set covering networking, identity, node pools, ingress, observability, upgrades, and fleet considerations when applicable.
- Validation commands, stop conditions, and ADR-worthy decisions that must be captured before implementation.
- Explicit assumptions, unresolved questions, and `[VERIFY]` markers for any version-sensitive or preview-dependent guidance.

## Quality Gate

Do not mark this engagement complete until:
- [ ] All irreversible AKS decisions are classified and justified.
- [ ] Networking, identity, upgrade, and observability guidance is specific enough to implement or review.
- [ ] Validation commands, stop conditions, and any required ADR entries are present.
