All skills
lukemurraynz avatar

/aks-cluster-architecture

@2cc2455

AKS cluster architecture decisions for new Azure Kubernetes Service projects: AKS Automatic vs Standard, networking topology, dual-stack (IPv4/IPv6), Kubernetes version and OS currency, node pool strategy, identity, production NetworkPolicy, namespaces, autoscaling, ingress and Gateway API, observability, operations, resilience, GPU and AI workloads, GPU partitioning (MIG, time-slicing, MPS), batch scheduling (Kueue), AKS on bare metal, AI Runway and KAITO model serving, AKS MCP server access, kars (Agent Reference Stack for Kubernetes) for agent isolation, Kata MicroVM pod sandboxing, Azure Kubernetes Fleet Manager, multi-cluster governance, update orchestration, resource placement, cross-cluster networking, and cost. WHEN: designing new AKS clusters, reviewing production readiness, choosing CNI or outbound connectivity, planning node pools, defining namespace, network and security controls, evaluating Fleet Manager, deploying AI agent runtimes on AKS, or making hard-to-reverse infrastructure decisions.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/aks-cluster-architecture

This session only. Nothing lands on disk.

referencesfull-reference.md

≈14k tokens on demand. Your agent reads this file only when SKILL.md points to it.

AKS Cluster Architecture Full Reference

This reference supports the lightweight router and bundles. It is intentionally verification-driven because AKS versions, add-ons, previews, and support dates change frequently.

Review date: 2026-05-14

Skill version: see bundles/catalog.yaml and individual bundles/*/bundle.yaml files.

Scope: new AKS projects first. Use migration guidance only when the task explicitly involves an existing cluster.

<!-- toc --> <!-- /toc -->

Reference Use Rule

Use this full reference only when the smaller bundles are insufficient. Final answers should still follow the required output contract in SKILL.md: assumptions, AKS Automatic vs Standard decision, decision difficulty table, recommended architecture, workload controls, security, networking, operations, cost/resilience, validation commands, stop conditions, ADRs, and open questions.

Source Notes to Re-check

Before making final implementation choices, re-check the current source for the exact feature:

  • Microsoft Learn - Supported Kubernetes versions in AKS
  • Microsoft Learn - AKS release tracker
  • Microsoft Learn - AKS Automatic overview
  • Microsoft Learn - Azure CNI Overlay, Azure CNI Powered by Cilium, AKS NetworkPolicy, and ACNS network policies
  • Kubernetes documentation - NetworkPolicy, HPA, PDB, ResourceQuota, LimitRange, and Pod Security Admission
  • Microsoft Learn - Node images and OS upgrade guidance
  • Microsoft Learn - Node Auto Provisioning and NAP networking support
  • Microsoft Learn - Application routing add-on and Gateway API implementation
  • Microsoft Learn - AKS ingress concepts
  • Microsoft Learn - AI toolchain operator / KAITO add-on
  • Microsoft Learn - AKS-managed GPU node pools and GPU best practices
  • Microsoft Learn - Deployment Safeguards and Azure Policy for AKS
  • Microsoft Learn - Azure Kubernetes Fleet Manager overview, choosing fleet type, update orchestration, resource placement, rollout strategy, member cluster types, managed namespaces, and multi-cluster networking
  • Cilium documentation - Cluster Mesh, KVStoreMesh, multi-cluster services, and Hubble observability when evaluating Cilium patterns
  • The New Stack and similar industry articles only as trend input; validate implementation decisions against Microsoft Learn, AKS release notes, and support boundaries
  • Azure/AKS GitHub releases for release-note-level drift

Current Retirement Checks for New Designs

Re-check these before final output, and do not recommend retired or near-retired defaults for new production clusters:

  • Azure Linux 2.0: do not use for new node pools; use a supported OS SKU such as Azure Linux 3 or Ubuntu 24.04 where available.
  • Kubenet: do not use as a new target-state production networking baseline.
  • Azure Network Policy Manager for Linux: avoid for new designs; prefer Azure CNI Powered by Cilium with Cilium Network Policy where supported.
  • Upstream/self-managed ingress-nginx: avoid as a new long-term default unless the customer owns the operational lifecycle and a migration path.
  • Preview ingress/Gateway, GPU, NAP, identity binding, or AI toolchain features: non-prod first unless support status and risk acceptance are explicit.

Use these commands during live design work:

az aks get-versions --location <region> --output table
az aks get-upgrades --resource-group <rg> --name <cluster> --output table
az provider show --namespace Microsoft.ContainerService --query "resourceTypes[?resourceType=='managedClusters'].apiVersions" --output table
az feature list --namespace Microsoft.ContainerService --output table
az extension list --query "[?name=='aks-preview']" --output table
az extension add --name fleet
az extension update --name fleet

Current High-Level Currency Snapshot

Snapshot, not source of truth. Every claim below is correct to the best of the author's knowledge on the review date above. Before treating any item as a production decision, re-verify against Microsoft Learn, the AKS release tracker, az aks get-versions, and az feature list --namespace Microsoft.ContainerService. Items are most likely to drift when they involve preview/GA status, retirement deadlines, or regional availability.

Do not copy this snapshot into long-lived implementation output without re-verifying.

  • AKS supports a rolling window of GA Kubernetes minor versions plus LTS options; the exact number of supported minors and the current LTS list change over time. Use the latest regionally available GA minor for most new clusters unless the organisation has a deliberate LTS policy.
  • As reviewed on 2026-05-14, confirm the supported GA minors, LTS list, and regional availability before design approval. Do not assume preview, GA, or rollout status from memory; use the AKS release tracker and az aks get-versions.
  • Azure Linux 2.0 is retired for AKS node images. Prefer Azure Linux 3 or Ubuntu 24.04 where supported.
  • Ubuntu 22.04 requires lifecycle planning; do not choose it as a new long-term default.
  • Windows node pools have separate lifecycle constraints. Use only when Windows containers are required.
  • Upstream/self-managed ingress-nginx should not be a new long-term default. Use supported Azure ingress/Gateway options and plan Gateway API migration.
  • AKS Automatic is a valid new-project option when Microsoft-managed defaults fit; AKS Standard remains the right fit for deep networking, node, GPU, Windows, or compliance control.
  • KAITO / AI toolchain operator and managed GPU nodes are useful for supported AI scenarios, but compatibility and limitations must be rechecked before production recommendations.

ManagedClusters API Volatility and Pinning

Microsoft.ContainerService/managedClusters changes frequently across preview and GA API versions. New properties can appear in a preview, then be removed in the next GA release. Do not use a preview API by default for production IaC.

Recommended rule set:

  • Use a GA API version for production unless a required feature exists only in preview.
  • Pin API versions in Bicep/ARM and review change-log deltas before upgrades.
  • When preview is required, document the dependency in an ADR, include rollback strategy, and assign a revalidation date.
  • Re-check feature status and property availability against Microsoft Learn change-log before each environment promotion.

Validation examples:

az provider show --namespace Microsoft.ContainerService --query "resourceTypes[?resourceType=='managedClusters'].apiVersions" --output table
az feature list --namespace Microsoft.ContainerService --output table

ManagedClusters Property Coverage Watchlist

Use this watchlist when reviewing AKS architecture guidance against IaC surfaces. It is intentionally selective and focused on high-impact design controls that are easy to miss in generic AKS writeups.

Cluster-level properties to verify

  • properties.networkProfile.advancedNetworking (observability, security, transit encryption, performance acceleration)
  • properties.apiServerAccessProfile (enableVnetIntegration, subnetId, privateDNSZone, authorizedIPRanges)
  • properties.azureMonitorProfile (appMonitoring.autoInstrumentation, metrics.kubeStateMetrics)
  • properties.ingressProfile (gatewayAPI.installation, webAppRouting.gatewayAPIImplementations, default NGINX controller type)
  • properties.securityProfile (azureKeyVaultKms, customCATrustCertificates, defender.securityMonitoring, imageCleaner)
  • properties.hostedSystemProfile and properties.nodeProvisioningProfile
  • properties.nodeResourceGroupProfile.restrictionLevel
  • properties.supportPlan (AKSLongTermSupport vs KubernetesOfficial)
  • properties.upgradeSettings.overrideSettings (forceUpgrade, until)

Agent pool properties to verify

  • agentPoolProfiles[].upgradeSettings (maxSurge, maxUnavailable, undrainableNodeBehavior, soak and drain timing)
  • agentPoolProfiles[].localDNSProfile
  • agentPoolProfiles[].networkProfile (allowedHostPorts, applicationSecurityGroups, nodePublicIPTags)
  • agentPoolProfiles[].securityProfile (enableSecureBoot, enableVTPM, sshAccess)
  • agentPoolProfiles[].podIPAllocationMode
  • agentPoolProfiles[].artifactStreamingProfile
  • agentPoolProfiles[].virtualMachinesProfile

Stop and re-check source docs if any required property is assumed from memory or copied from a different API version.

Decision Classification Master

Decision Difficulty Review requirement
Region/data residency Permanent ADR before provisioning.
Cluster name/resource group/subscription layout Permanent Naming and ownership approval.
AKS Automatic vs Standard Permanent to difficult Choose operating model before lower-level design.
VNet/subnet/service CIDR/pod CIDR/DNS service IP Permanent Network/IP plan and overlap check.
CNI/IPAM/data plane Permanent Confirm Azure CNI Overlay Powered by Cilium fit.
Availability zones Permanent Zone support and availability target.
Private API server/DNS mode Difficult Operator access and private DNS design.
Outbound egress path Difficult SNAT, firewall, FQDN, and routing validation.
Kubernetes minor version Difficult Upgrade strategy and support window.
OS SKU and disk type Difficult Node-pool replacement likely.
Node pool VM family Difficult Capacity, quota, and workload validation.
Workload Identity/OIDC Difficult Enable at creation for new clusters.
Azure RBAC/Kubernetes RBAC model Difficult Admin model and namespace model.
Gateway/ingress controller Difficult TLS/DNS/WAF/controller migration impact.
Observability workspace and log schema Difficult Retention/cost/SIEM integration.
Fleet Manager adoption Reversible to difficult Use only when multi-cluster governance or update orchestration creates value.
Fleet without hub vs with hub Difficult Hubless can be upgraded to hub; hub cannot be downgraded.
Fleet hub public vs private access Permanent Hub access mode cannot be changed after creation.
Fleet member cluster taxonomy Difficult Labels/taints drive update rings, placement, cost, compliance, and blast radius.
Fleet resource placement scope Difficult Placement can add, update, or remove resources across clusters.
Fleet cross-cluster networking Preview-gated/difficult Confirm support, network model, failure modes, policy, and ownership.
Policy enforcement level Reversible to difficult Roll out warning before enforcement where risk exists.
Namespace quota/LimitRange Reversible to difficult Can break existing deployments; set before production onboarding.
NetworkPolicy model Reversible to difficult Default deny can break DNS, ingress, dependency, or egress paths; test before enforcement.
Replica/PDB/topology spread design Reversible Requires SLO, drain, and zone-failure validation.
HPA/KEDA/VPA settings Reversible Requires metrics and load validation.
Spot pool use Reversible Workload criticality review.

Architecture Defaults for New AKS Projects

Area Recommended default Exception path
Operating model AKS Automatic where managed defaults fit; AKS Standard where platform control is required Existing platform standard or unsupported feature.
AKS pricing tier Standard or Premium for production or at-scale workloads Dev/test only when explicitly non-production.
CNI Azure CNI Overlay Powered by Cilium VNet-routable pods, unsupported combination, or migration constraint.
Egress NAT Gateway for simple/managed VNets; UDR to firewall for enterprise hub-spoke Dev/test may use simpler egress with explicit risk.
API server Private for regulated production Public with authorized IP ranges for lower-risk/dev/test.
Identity Managed identity + OIDC + Workload Identity No service principal secrets for new designs.
RBAC Azure RBAC for Kubernetes + namespace scoping Kubernetes RBAC-only only with documented reason.
Node OS Azure Linux 3 or Ubuntu 24.04 Windows only for Windows container workloads.
System pool Dedicated Linux system pool AKS Automatic may abstract this.
Workload namespaces Owner/cost/data labels + RBAC + quota + LimitRange + Pod Security labels Exception only for platform namespaces with separate controls.
NetworkPolicy Default deny per production workload namespace + explicit allow rules Dev/test may roll out in phases, but production needs explicit policy.
Replicas/disruption Minimum 2 production stateless replicas, usually 3 across zones + PDB + topology spread Singleton only with documented application reason.
Autoscaling HPA/KEDA + cluster autoscaling/NAP as applicable Static node counts only for tightly controlled workloads with documented reason.
Observability ContainerLogV2 + Managed Prometheus + dashboards/alerts Approved equivalent observability stack.
GitOps Flux or approved GitOps tool Direct deployment only with clear release controls.
Fleet management No Fleet Manager for a single cluster; Fleet Manager for multi-cluster governance/update orchestration Use hubless for updates only; use hub only for placement, managed namespaces, or hub-dependent networking.
Policy Deployment Safeguards / Azure Policy Warning mode for initial rollout where enforcement risk is high.

AKS Automatic vs AKS Standard

Prefer AKS Automatic when

  • The workload is a new general-purpose container platform.
  • The team wants Microsoft-managed defaults for node management, scaling, security, and policy.
  • The platform does not need deep customisation of node pools, CNI details, or egress topology.
  • The supported feature set satisfies ingress, policy, monitoring, and compliance needs.

Prefer AKS Standard when

  • Enterprise hub-spoke networking, custom UDR, custom DNS, or central firewall routing is mandatory.
  • The workload needs GPU, Windows, confidential compute, special OS, or precise VM family control.
  • The platform team needs explicit IaC ownership over node pools and add-ons.
  • The project needs features not currently supported or not yet production-ready in AKS Automatic.

Review prompts

  • Are we accepting Microsoft-managed defaults, or do we need platform control?
  • Does the organisation need private API server, custom egress, or custom DNS from day one?
  • Are GPU/Windows/confidential workloads in scope?
  • Does the selected ingress/Gateway option support AKS Automatic today?
  • Is preview support acceptable for the environment tier?

Networking Architecture

Azure CNI Overlay Powered by Cilium

Use this as the default for new general-purpose clusters when supported.

Benefits:

  • Preserves VNet IP space by avoiding one VNet IP per pod.
  • Uses a modern AKS-supported Cilium eBPF data plane.
  • Aligns with NAP/Karpenter networking guidance better than legacy policy choices.
  • Supports future network security and visibility options such as ACNS where available.

Cautions:

  • Validate Cilium NetworkPolicy behaviour, especially ipBlock limitations around node IPs.
  • Do not assume every upstream Cilium feature is exposed or supported in the AKS managed implementation.
  • Confirm Windows support and mixed OS pool needs before using Cilium-dependent policy assumptions.

CIDR planning

Required checks:

VNet CIDR != pod CIDR
VNet CIDR != service CIDR
pod CIDR != service CIDR
service DNS IP inside service CIDR
no overlap with peered VNets
no overlap with on-premises networks
no overlap with other clusters that need direct routing

Document CIDRs in an ADR and environment registry.

Egress design

Avoid default load balancer outbound SNAT for production.

Recommended patterns:

  1. Managed NAT Gateway for simple AKS-managed VNet designs.
  2. User-assigned NAT Gateway for existing subnet/NAT standards.
  3. UDR through Azure Firewall/NVA for ALZ-style central inspection.
Zonal egress for multi-zone clusters

Azure NAT Gateway is a zonal resource, it can only be pinned to a single availability zone. For a multi-zone AKS cluster, this creates an architectural tension:

  • NAT Gateway (Standard) provides only single-zone egress. If that zone fails, outbound connectivity from all nodes (including nodes in healthy zones) is lost until the zone recovers.
  • Accepted zone-redundancy pattern for Standard NAT Gateway: span node pools across all availability zones for the compute layer (so pods survive a zone failure), and accept that egress is single-zone. Document this as an accepted risk in the ADR. Suitable for dev/test and low-SLO production.
  • Zone-redundant egress (preferred for high SLO): use managedNATGatewayV2 (the underlying StandardV2 NAT Gateway SKU is GA and zone-redundant by default; the AKS managedNATGatewayV2 outbound type is still Preview, [VERIFY] GA status); or deploy a NAT Gateway per zone via a zonal egress subnet design; or route through a zone-redundant Azure Firewall in the hub.
  • Do NOT assume a single Standard NAT Gateway gives a multi-zone cluster zone-redundant egress, multi-zone compute does not compensate for a single-zone egress appliance.

Validation examples:

az aks show --resource-group <rg> --name <cluster> --query "networkProfile.outboundType" --output tsv
kubectl run egress-test --rm -it --restart=Never --image=mcr.microsoft.com/azure-cli -- bash

Stop if required FQDN allowlists, firewall routes, or SNAT capacity are not known.

API server exposure

Private API server is the safer production default for regulated workloads, but it requires an operator connectivity design:

  • private DNS zone and links
  • VPN/ExpressRoute/bastion/jump access
  • CI/CD runner network path
  • emergency access path
  • RBAC and break-glass process

Public API server with authorized IP ranges can be acceptable for dev/test or a risk-accepted production case. Do not recommend an unrestricted public API server for new production clusters.

Production Network Policy

Use this section when a design needs production network controls. For concise YAML examples, use ../bundles/production-workload-controls/guide.md.

Recommended AKS policy direction

For new Linux production clusters, prefer Azure CNI Powered by Cilium where supported. Microsoft recommends Cilium for Kubernetes-native policies and richer capabilities such as L7 and FQDN filtering when policy support is enabled. Verify the selected engine against AKS Automatic/Standard, Linux/Windows node pools, ACNS configuration, and current support status before implementation.

Policy baseline

Production workload namespaces should include:

  • default deny ingress and egress
  • allow DNS egress to CoreDNS
  • allow ingress only from approved ingress/Gateway namespaces or client namespaces
  • allow east-west service calls by namespace and pod selectors
  • allow egress to Azure PaaS through Private Link/private endpoints or supported FQDN/L7 policy
  • explicit test evidence for every allowed path

Design cautions

  • Kubernetes NetworkPolicy is additive and namespaced.
  • A pod is isolated for ingress or egress only when selected by a policy for that direction.
  • Use selectors for in-cluster traffic. Do not build pod-to-pod rules around ephemeral pod IPs.
  • On Cilium-based AKS clusters, avoid ipBlock for pod or node IPs; use selectors instead.
  • LoadBalancer source/destination rewrites can affect policy behaviour for external traffic.
  • Do not default-deny platform namespaces without a platform-owned policy design.
  • Treat FQDN/L7 policy as an AKS/Cilium/ACNS compatibility decision, not plain Kubernetes NetworkPolicy.
  • --enable-acns alone enables only FQDN filtering. L7 policy (HTTP/gRPC method/path rules) requires --enable-acns --acns-advanced-networkpolicies L7 on az aks create/az aks update; without it, CiliumNetworkPolicy L7 rules are silently ignored. L7 rules are not supported in CiliumClusterwideNetworkPolicy (CCNP) at all, only in namespaced CiliumNetworkPolicy.

Production rollout pattern

  1. Map runtime traffic from logs, service mesh/ACNS/Cilium observability, and app owners.
  2. Apply default deny and allow rules in non-prod.
  3. Test DNS, ingress, intra-app, database/private endpoint, identity, telemetry, and external dependency paths.
  4. Add alerts for policy drops or failed dependencies where supported.
  5. Promote via GitOps/IaC with a rollback plan.

Identity and Security

Cluster and workload identity

Default:

  • cluster managed identity
  • kubelet managed identity
  • OIDC issuer enabled
  • Workload Identity for pod-to-Azure access
  • user-assigned managed identity for workloads where lifecycle separation matters

Avoid:

  • service principal secrets
  • AAD Pod Identity for new designs
  • static cloud credentials in Kubernetes Secrets
  • broad cluster-admin grants

Validation:

az aks show --resource-group <rg> --name <cluster> --query "identity" --output yaml
az aks show --resource-group <rg> --name <cluster> --query "oidcIssuerProfile" --output yaml
kubectl get serviceaccounts -A -o yaml | grep -E "azure.workload.identity|client-id" -n

RBAC

Granting the kubelet identity ACR pull permission (Bicep trap)

AKS nodes pull images using the cluster's kubelet identity, not the cluster (control-plane) managed identity. A common Bicep error assigns AcrPull to the wrong principal:

  • Wrong: target aksCluster.identity.principalId (the control-plane identity), image pulls fail with 401/403.
  • Correct: target aksCluster.properties.identityProfile.kubeletIdentity.objectId, this is the identity the kubelet actually uses for image pulls.
resource aksCluster 'Microsoft.ContainerService/managedClusters@2024-... existing = { name: aksName }

resource kubeletAcrPull 'Microsoft.Authorization/roleAssignments@2022-04-01' = {
  scope: containerRegistry
  name: guid(aksCluster.id, containerRegistry.id, 'acrpull')
  properties: {
    roleDefinitionId: subscriptionResourceId('Microsoft.Authorization/roleDefinitions', '7f951dda-4ed3-4680-a7ca-43fe172d538d') // AcrPull
    principalId: aksCluster.properties.identityProfile.kubeletIdentity.objectId  // NOT aksCluster.identity.principalId
  }
}

Validation:

az aks show -g <rg> -n <cluster> --query "identityProfile.kubeletIdentity.objectId" -o tsv
az role assignment list --assignee <kubelet-object-id> --scope <acr-id> --role AcrPull

RBAC

Recommended model:

  • Azure RBAC for Kubernetes authorization at cluster/namespace scope.
  • Microsoft Entra groups for admin, operator, developer, and reader roles.
  • Namespace-scoped access for app teams.
  • Kubernetes RBAC for fine-grained workload service accounts.
  • Privileged operations gated through break-glass and audit.

Policy and safeguards

Use Azure Policy / Deployment Safeguards to enforce baseline controls.

Suggested rollout:

  1. Start in warning/audit mode in dev if policy may break existing manifests.
  2. Fix violations in templates.
  3. Enable enforcement for production once violations are understood.
  4. Keep exceptions time-bound and documented.

Node Pool Strategy

System pool

  • Keep system and user workloads separate.
  • Use a dedicated Linux system pool for AKS Standard unless AKS Automatic abstracts this.
  • Taint the system pool so user workloads do not schedule there.
  • Use reliable VM families and zones for production.

User pools

Use workload-oriented pools when it improves isolation or operations:

  • general stateless apps
  • stateful/IO-sensitive apps
  • memory-heavy apps
  • CPU-heavy apps
  • spot/batch apps
  • GPU/AI apps
  • Windows workloads

Do not create many node pools without a reason. Each pool adds upgrade, capacity, and policy overhead.

OS SKU and disk type

  • Prefer Azure Linux 3 or Ubuntu 24.04 where supported.
  • Ephemeral OS disks can improve node replacement speed, but require compatible VM SKU/temp disk capacity.
  • OS SKU or disk type changes normally mean node pool replacement; plan as difficult.

Autoscaling and Scheduling

HPA

Use for replica scaling from CPU, memory, or custom metrics. Requires resource requests and reliable metrics.

Checklist:

  • CPU/memory requests exist.
  • Metrics source is available.
  • Readiness probes protect startup.
  • Min replicas align with SLO.
  • Scale-down behaviour does not break latency or connection draining.

KEDA

Use for event-driven scaling such as queues, streams, scheduled jobs, external metrics, and GPU metrics when configured.

Checklist:

  • Authentication to scaler source uses Workload Identity where supported, or a secure secret store with documented rotation where identity is not supported.
  • The KEDA operator/scaler identity and the application runtime identity are separated when their Azure permissions differ.
  • Cooldown and polling intervals match workload behaviour.
  • Scale-to-zero is safe for the app and dependencies.
  • Dead-letter/poison message behaviour is defined.

VPA

Use VPA recommender as a rightsizing tool first. Apply automatic mutation only after testing interactions with HPA, disruption budgets, and latency-sensitive workloads.

NAP / Karpenter

Use for heterogeneous workloads where static node pool planning is inefficient.

Before choosing NAP:

  • Verify current GA/preview and regional support.
  • Confirm networking compatibility.
  • Confirm unsupported combinations such as Calico network policy or Dynamic IP Allocation where applicable.
  • Define workload scheduling intent; do not rely only on autoscaler inference.
  • Confirm disruption and consolidation behaviour.

Scheduling controls

Use:

  • labels for pool identity
  • taints/tolerations for hard isolation
  • node affinity for required/preferred placement
  • topology spread constraints for zone resilience
  • PDBs for drain safety
  • priority classes for critical workloads

Avoid:

  • accidental placement on system pools
  • critical workloads on spot pools
  • single-zone replicas for production services
  • GPU nodes consumed by non-GPU workloads

Production Namespace and Workload Controls

Use ../bundles/production-workload-controls/guide.md when generating examples. This section captures the reference model.

Namespace contract

Every production namespace should define:

  • owner, environment, cost centre, data classification, and application labels
  • namespace-scoped RBAC groups
  • ResourceQuota and LimitRange
  • Pod Security Admission labels, preferably restricted where possible
  • default-deny NetworkPolicy and explicit allow rules
  • cost and observability ownership

Namespaces are useful operational boundaries but not a complete hostile multi-tenancy boundary by themselves.

Replica and disruption model

Production stateless services should normally start with at least two replicas and preferably three across zones. One replica is acceptable only for dev/test, true singleton workloads, cost-accepted noncritical services, or workloads protected by another availability layer.

Use:

  • Deployment rolling strategy with explicit maxUnavailable and maxSurge
  • readiness probes before traffic routing
  • startup probes for slow-starting apps and AI/model workloads
  • PDBs for voluntary disruption
  • topology spread across zones and nodes
  • graceful shutdown and termination grace periods

PDBs do not protect against crashes, failed releases, or involuntary disruption. They only constrain voluntary eviction such as node drains and upgrades.

Autoscaling integration

Design scaling controls together:

  • HPA/KEDA controls replica count.
  • Cluster autoscaler or NAP controls node capacity.
  • VPA should begin in recommender/off mode for production right-sizing.
  • Resource requests drive HPA percentages, scheduler placement, and NAP decisions.
  • PDBs and topology spread must still allow drains and consolidation.
  • Max replicas must respect downstream capacity, budget, and node autoscaler maximums.

Avoid conflicting automation such as HPA and VPA automatic mutation on the same CPU/memory target without a tested design.

GPU and AI Workloads

GPU node pools

Default:

  • dedicated GPU pool
  • taint GPU pool
  • workload requests nvidia.com/gpu
  • validate NVIDIA device plugin/driver status
  • collect GPU metrics
  • set budget alerts and idle-node alerts

Verify:

  • regional GPU quota
  • supported VM families
  • driver installation path
  • managed GPU node pool limitations
  • Windows vs Linux support
  • migration limitations from existing GPU pools

KAITO / AI toolchain operator

Consider for supported self-hosted open-source model inference on AKS.

Verify before recommending:

  • current KAITO/add-on version
  • supported OS SKU and region
  • supported model presets and runtime
  • GPU VM quota
  • AKS Automatic compatibility
  • production support status
  • network and data residency constraints

Do not claim broad GA production readiness without source verification.

Traffic Management

Ingress/Gateway decision tree

  1. Does the workload need public ingress, private ingress, or both?
  2. Is WAF required?
  3. Who owns TLS certificates and DNS?
  4. Is Gateway API required or is Ingress API acceptable as a bridge?
  5. Is the selected controller supported on AKS Automatic if using Automatic?
  6. Is the feature GA or preview?
  7. What is the migration path before support windows close?

Options

Option Fit Cautions
Application routing managed NGINX Supported bridge for simple Ingress API workloads Plan migration before Azure support windows close.
Application routing Gateway API Strategic Gateway API path Re-check preview/GA, feature flag, TLS/DNS limitations.
Application Gateway for Containers Enterprise L7/WAF/Gateway API use cases Verify private frontend, region, and AKS Automatic support.
AKS Istio add-on Service mesh and mesh ingress Do not use mesh unless mesh capabilities are required.
Self-managed ingress-nginx Migration or special case Avoid as new long-term default.

Observability

Platform telemetry

Baseline:

  • ContainerLogV2
  • Managed Prometheus
  • Azure Managed Grafana or approved dashboards
  • Container Insights or equivalent
  • Azure Monitor alerts
  • log retention and cost budget

Monitor at minimum:

  • node readiness and pressure
  • pending pods
  • CrashLoopBackOff
  • OOMKilled
  • CPU throttling
  • HPA/KEDA scaling errors
  • API server errors
  • CoreDNS errors
  • ingress 4xx/5xx and latency
  • SNAT/egress errors
  • PV attach/mount failures
  • GPU utilisation and health where applicable

Runtime validation commands

kubectl get nodes -o wide
kubectl get pods -A --field-selector=status.phase!=Running
kubectl get events -A --sort-by=.metadata.creationTimestamp | tail -100
kubectl top nodes
kubectl top pods -A
kubectl describe hpa -A
kubectl get scaledobjects -A
kubectl get pdb -A

Operations

Upgrade strategy

  • Configure maintenance windows.
  • Use a Kubernetes auto-upgrade channel only when it matches organisational risk appetite.
  • Configure node OS upgrade channel separately; do not confuse node image updates with Kubernetes minor upgrades.
  • Use canary or blue-green node pool replacement for high-risk OS, kernel, GPU, or major dependency changes.
  • Validate PDBs before any node drain or upgrade.

See bundles/operations-resilience/guide.md for the Upgrade Irreversibility / Version Skew Policy / Admission Webhook pre-upgrade checklist / CRD storage version migration / Post-upgrade Triage runbook content.

GitOps

Recommended:

  • Flux or approved GitOps tool as source of truth.
  • Break-glass documented and time-bound.
  • Emergency changes reconciled back into Git.
  • No permanent manual kubectl edit drift.

Backup and recovery

Cover:

  • IaC and GitOps repositories
  • Kubernetes resources
  • persistent volumes
  • databases and external dependencies
  • Key Vault, ACR, DNS, identities, firewall rules
  • restore validation

Backups are not complete until restore has been tested.

Azure Kubernetes Fleet Manager and Multi-Cluster Governance

Use Fleet Manager only when the architecture has a genuine fleet problem: multiple AKS clusters, multiple regions, multiple subscriptions, multiple environments, platform-team governance, staged upgrades, or cross-cluster workload placement. A single-cluster workload should not inherit Fleet Manager complexity unless a credible near-term fleet roadmap exists.

Fleet adoption decision

Requirement Preferred fleet option Notes
Single cluster No Fleet Manager by default Keep the design focused on cluster foundations, workload platform, and operations.
Coordinated Kubernetes/node image upgrades across several AKS clusters Fleet Manager without hub cluster Lowest-complexity fleet option; supports update orchestration without hub placement capabilities.
Resource placement, managed fleet namespaces, or DNS load balancing Fleet Manager with hub cluster Hub provides the Kubernetes API/control plane for KubeFleet-style placement and namespace governance.
Hybrid or multicloud Kubernetes members Arc-enabled member clusters supported (GA, Build 2026); validate capability parity Capability coverage differs from AKS member clusters.
Cross-cluster service discovery or traffic Preview-gated pattern Validate DNS/L4/east-west model, failure modes, support scope, and NetworkPolicy behaviour.

Hubless vs hub fleet

  • Hubless fleet: use for grouping and orchestrating updates across member clusters. It avoids hub cost and can be upgraded to hub mode later.
  • Hub fleet: use when resource placement, Managed Fleet Namespaces, or hub-dependent networking capabilities are required.
  • Irreversible access choice: decide public vs private hub access before provisioning. Treat private hub as the production default when operator network access can support it.
  • Managed hub boundary: do not mutate the managed hub AKS cluster or its underlying resources directly.

Member cluster taxonomy

Define labels and taints before production onboarding. Minimum recommended labels:

  • environment=dev|test|stage|prod
  • region=<azure-region>
  • criticality=low|medium|high|mission-critical
  • owner=<team>
  • update-group=canary|early|standard|late|critical
  • data-class=public|internal|confidential|restricted
  • tier=platform|shared|app|data|edge

Labels used for update and placement decisions are platform controls and should be changed only through reviewed IaC or approved platform workflows.

Update orchestration

Fleet update runs should be designed like safe deployment rings:

  1. Lab/platform validation.
  2. Non-production workloads.
  3. Canary production.
  4. Standard production.
  5. Late/critical production with explicit approval.

Required controls:

  • update strategy with stages, groups, wait times, and maximum concurrency where supported
  • member cluster maintenance windows
  • pause/stop criteria based on cluster health, pending pods, PDBs, ingress, NetworkPolicy, and app SLOs
  • separate risk treatment for Kubernetes minor upgrades and node image updates
  • clear rollback/recovery runbook and communications process

Resource placement

Use ClusterResourcePlacement for cluster-scoped resources or entire namespaces, and use ResourcePlacement only after confirming namespace-scoped API support and preview/GA state. Prefer placement for centrally owned platform baselines, not for every app release.

Resource placement guardrails:

  • Avoid broad PickAll for production unless capacity and blast radius are understood.
  • Use PickFixed, PickN, or scoped label selection for controlled rollout.
  • Configure rollout strategy with conservative maxUnavailable, maxSurge, and unavailablePeriodSeconds.
  • Remember that placement status means resources were applied; it does not prove pods are ready or traffic is healthy.
  • Treat removing a selected cluster from placement as a potentially destructive change because placed resources can be removed.

Managed Fleet Namespaces

Managed Fleet Namespaces are useful for consistent namespace boundaries across clusters: quotas, labels, annotations, RBAC, and NetworkPolicy. Verify current GA/preview state before production. If still preview, use non-production by default unless there is explicit risk acceptance, support validation, rollback, and migration planning.

Cross-cluster networking and Cilium

Fleet Manager multi-cluster networking and upstream Cilium Cluster Mesh are advanced options but should not become defaults. Prefer ordinary regional ingress, API-level integration, externalized data stores, and clear failure domains unless cross-cluster service discovery is a real requirement.

Use this rule:

Prefer Microsoft-supported Fleet Manager and AKS networking capabilities first. Consider upstream Cilium Cluster Mesh only as an exception after supportability, routing, identity, policy, observability, and operational ownership are validated.

Cilium watchpoints:

  • Azure CNI Powered by Cilium on AKS is Microsoft-managed and does not imply full control over all upstream Cilium Cluster Mesh settings.
  • Cluster Mesh extends the datapath across clusters, supports policy enforcement, and can load-balance via Kubernetes annotations, but it adds operational and network failure modes.
  • KVStoreMesh improves scalability and isolation, but still requires explicit monitoring, upgrade, and incident ownership.
  • Hubble or equivalent network visibility is important for cross-cluster traffic validation.

Fleet RBAC and identity

Use Microsoft Entra ID and least-privilege RBAC across Fleet Manager ARM resources, the hub cluster, and member clusters. Keep platform-level fleet administration separate from application-team namespace administration.

Recommended guardrails:

  • Use Fleet Manager ARM contributor roles only for operators responsible for fleet resources, members, update strategies, and update runs.
  • Use hub read-only roles for inspection; reserve hub write/admin roles for platform operators responsible for placement and fleet CRDs.
  • Be careful with writer roles that can read Secrets because they can assume service-account credentials within the namespace.
  • Do not give application teams broad access to the fleet hub unless they own fleet-wide placement.
  • Audit RBAC role assignments at fleet, managed namespace, hub, and member-cluster scopes.

Fleet validation commands

# Fleet extension and resource inventory
az extension add --name fleet
az extension update --name fleet
az fleet show --resource-group <rg> --name <fleet> --output yaml
az fleet member list --resource-group <rg> --fleet-name <fleet> --output table

# Hub access for hub fleets
az fleet get-credentials --resource-group <rg> --name <fleet>
kubectl get memberclusters -o wide
kubectl get memberclusters --show-labels

# Update and placement evidence
az fleet updaterun list --resource-group <rg> --fleet-name <fleet> --output table
az fleet updatestrategy list --resource-group <rg> --fleet-name <fleet> --output table
kubectl get clusterresourceplacement
kubectl describe clusterresourceplacement <placement-name>
kubectl get resourceplacement -A

Fleet stop conditions

Stop before provisioning or production rollout if:

  • Fleet Manager is being introduced without a real multi-cluster requirement.
  • Hubless vs hub mode is not justified.
  • Public vs private hub access has not been approved.
  • Member cluster labels/taints are not authoritative.
  • Preview features are being proposed for production without explicit support and risk acceptance.
  • Resource placement could remove production resources and rollback is not tested.
  • Cross-cluster networking has not been validated for DNS, routing, NetworkPolicy, local/remote failover, observability, and incident response.

Multi-Region and Resilience

Choose data architecture first. AKS topology follows the data plane, not the other way around.

Pattern Use when Key risk
Single region, multi-zone Most production workloads Regional outage acceptance.
Warm standby Higher resilience with cost control Promotion and data lag.
Active-active Global/mission-critical workloads Data consistency and operational complexity.
Fleet governance Many clusters across regions/environments Feature maturity and management overhead.

Validation questions:

  • What is the RTO/RPO?
  • Where is the database primary?
  • How is DNS/traffic failover controlled?
  • Are images and secrets available in failover region?
  • Is the cluster rebuild automated?
  • Are firewall and private DNS dependencies duplicated?

Cost Optimisation

Main cost levers:

  • right-size requests and limits
  • use autoscaling/NAP for variable demand
  • use spot only for interruptible workloads
  • buy reserved capacity/savings plans only for proven baseline usage
  • control Log Analytics and Prometheus ingestion
  • monitor NAT Gateway, public IP, gateway, firewall, and egress costs
  • alert on idle GPU pools

Review monthly:

  • requested CPU/memory vs actual usage
  • top namespaces by cost
  • idle nodes and underutilised pools
  • autoscaler min/max settings
  • log ingestion spikes
  • GPU utilisation

Known Pitfalls and Mitigations

Pitfall Mitigation
SNAT exhaustion from load balancer egress NAT Gateway or central firewall with adequate capacity.
HPA with no resource requests Enforce requests in templates and policy.
CoreDNS customisation outage Test in non-prod and keep rollback ready.
CNI/CIDR conflict discovered late Do network ADR before provisioning.
Policy enforcement blocks release Roll out warning/audit first, then enforcement.
GitOps reverts emergency fix Document break-glass and reconcile back to Git.
Untested backup Schedule restore tests.
GPU quota not available Confirm quota before committing architecture.
Gateway API preview limitation Verify current feature status and limitations before production.
AKS Automatic unsupported feature Confirm support matrix before selecting Automatic.
Node OS lifecycle drift Track OS deadlines and use node pool replacement plans.

Drasi on AKS Notes

When Drasi sources, continuous queries, or reactions run on AKS:

  • Enable OIDC issuer and Workload Identity for Azure resource access.
  • Confirm Dapr sidecar requirements and health checks.
  • Kubernetes source integrations may require static service-account-token kubeconfig instead of az aks get-credentials exec auth.
  • Store kubeconfig material in Kubernetes Secrets and scope RBAC narrowly.
  • Monitor source, query, reaction, and Dapr component health.
  • Document recovery when source crashes deactivate Dapr actors.

ADR Template for Permanent Decisions

# ADR: <AKS decision>

## Status

Proposed | Accepted | Superseded

## Context

- Workload:
- Environment:
- Region:
- Compliance/network constraints:
- Availability target:

## Decision

<What was chosen.>

## Change difficulty

Permanent | Difficult | Reversible

## Options considered

1. <Option>
2. <Option>
3. <Option>

## Consequences

- Positive:
- Negative:
- Operational impact:
- Cost impact:

## Validation

- Commands run (cite the kubectl/az command lines used):
- Evidence captured (paste output snippet or link to artifact):
- Date verified:
- Verified by:
- Next re-verification date:

## Rollback or migration path

<What happens if the decision must change later.>

Production Readiness Checklist

Foundation

  • Region and data residency confirmed.
  • AKS Automatic vs Standard selected with rationale.
  • Kubernetes version verified in target region.
  • CIDRs checked for overlap.
  • CNI/IPAM/data plane selected and documented.
  • Availability zones confirmed.
  • API server exposure and DNS path confirmed.
  • Egress path and SNAT capacity validated.

Workload Platform

  • System/user workloads isolated.
  • Workload Identity enabled and used.
  • Azure RBAC/Kubernetes RBAC model documented.
  • Node pools sized and justified.
  • Namespace ResourceQuota, LimitRange, Pod Security labels, and RBAC boundaries defined.
  • Default-deny NetworkPolicy and explicit DNS/ingress/east-west/egress allow rules defined.
  • HPA/KEDA/VPA/NAP interactions understood.
  • Replica count, PDBs, readiness probes, rollout strategy, and topology spread defined for production apps.
  • GPU/Windows/spot pools have explicit rules if used.

Operations

  • Ingress/Gateway choice has support status and limitations documented.
  • TLS/DNS ownership documented.
  • ContainerLogV2 and Managed Prometheus configured.
  • Alerts and dashboards defined.
  • Maintenance windows configured.
  • Backup and restore tested.
  • GitOps or release process defined.
  • Fleet Manager decision documented when more than one cluster is in scope.
  • Hubless vs hub, hub access mode, member labels/taints, update stages, placement ownership, and preview gates documented when Fleet Manager is used.
  • Cost controls and budgets configured.
  • ADRs written for permanent/difficult decisions.

Related Skill Hooks

  • Production workload controls: use bundles/production-workload-controls/guide.md for NetworkPolicy, namespace, quota, replica, PDB, topology spread, HPA/KEDA, and VPA examples.
  • Fleet management: use bundles/fleet-management/guide.md for Azure Kubernetes Fleet Manager, update orchestration, placement, managed namespaces, cross-cluster networking, and Cilium Cluster Mesh watchpoints.
  • Kubernetes manifest authoring: use the repository Kubernetes instruction file when editing YAML.
  • Azure landing zone/networking: use the repository Azure/defaults/private-networking skill where available.
  • Observability: use the repository observability skill where available.
  • Drasi: use the Drasi skill for Sources, ContinuousQueries, and Reactions.
  • Azure Container Apps: consider ACA instead of AKS for simpler container platforms that do not need Kubernetes-level control.

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The skill is a comprehensive architecture and configuration guide for Azure Kubernetes Service (AKS). it emphasizes security best practices, including RBAC, NetworkPolicy, workload identity, and kernel-level isolation for AI agents. No malicious patterns or security risks were detected.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-26"
}
Other metadata
argument-hint
workload=<type>; region=<azure-region>; availability=<SLO>; network=<hub-spoke|standalone>; scope=<new-cluster|production-review|fleet>

README badge

README badge for lukemurraynz/hve-agent-skills/aks-cluster-architecture