Fleet Management Bundle
⚠️ Cross-cluster networking (Preview, March 2026): Azure Kubernetes Fleet Manager now supports cross-cluster service discovery and observability between member clusters. See Fleet cross-cluster networking for configuration and regional availability.
Other recent Fleet Manager additions: Multi-cluster auto-upgrade (GA, April 2025), automated GitHub-attached deployments (Preview, May 2025), placement drift detection with
applyStrategy(Preview, May 2025), resource placement (ClusterResourcePlacement) GA (July 2026); namespace-scoped ResourcePlacement remains preview (January 2026); Fleet Manager for Arc-enabled clusters (GA, Build 2026; a single fleet now supports up to 1,000 member clusters, up from 200), update run approval gates (Preview, September 2025), maximum allowed failures for update runs (Preview, July 2026), DNS-based public load balancing via Azure Traffic Manager (Preview, May 2025).
Multi-cluster governance: Azure Kubernetes Fleet Manager (hubless vs hub), update orchestration, resource placement, managed fleet namespaces, cross-cluster networking, Cilium Cluster Mesh watchpoints. Skip this bundle for single-cluster designs.
Load this bundle when an AKS architecture involves more than one cluster, multiple regions, multiple subscriptions, platform-team cluster governance, cross-cluster update orchestration, workload placement across clusters, managed namespace governance across clusters, or cross-cluster networking.
<!-- toc -->- When This Bundle Should Trigger
- Fleet Decision Rule
- Permanent / Difficult / Reversible Decisions
- Hubless vs Hub Fleet
- Member Cluster Model
- Update Orchestration
- Auto-Upgrade Profiles and Emergency CVE Patching
- Resource Placement
- Managed Fleet Namespaces
- Cross-Cluster Networking
- Cilium Cluster Mesh Watchpoints
- Security and Identity
- Observability and Operational Readiness
- Cost and Complexity Controls
- Stop Conditions
- ADRs Required
- Use With
Assume new AKS projects. Do not retrofit fleet-scale complexity into a simple single-cluster workload unless the customer has a credible near-term roadmap for multiple clusters.
When This Bundle Should Trigger
Use this bundle when the user mentions or implies:
- fleet, many clusters, multi-cluster, hub/spoke Kubernetes, central AKS governance, platform engineering at scale
- multiple AKS clusters across regions, subscriptions, environments, business units, or customer tenants
- safe staged upgrades across clusters, Kubernetes version alignment, node image upgrade rings, or maintenance-window coordination
- workload/resource placement across clusters, KubeFleet,
ClusterResourcePlacement, orResourcePlacement - managed fleet namespaces, namespace governance across clusters, shared quotas, RBAC, or NetworkPolicy baselines
- north-south or east-west cross-cluster traffic, multi-cluster service discovery, global services, DNS load balancing, or Cilium Cluster Mesh
- Arc-enabled Kubernetes clusters joining an Azure fleet
Do not load this bundle for ordinary single-cluster design unless multi-cluster governance is part of the stated target architecture.
Fleet Decision Rule
| Scenario | Preferred direction | Why |
|---|---|---|
| One AKS cluster only | Do not introduce Fleet Manager by default | Avoid governance overhead before it creates value. |
| Multiple AKS clusters needing coordinated Kubernetes or node image upgrades | Fleet Manager without hub cluster | Hubless fleets support update orchestration without hub-cluster cost or placement complexity. |
| Multiple clusters needing resource placement, namespace propagation, managed fleet namespaces, or DNS load balancing | Fleet Manager with hub cluster | Hub mode provides the Kubernetes API/control plane needed for placement and fleet namespace capabilities. |
| Production fleet with hub cluster | Prefer private hub access where operator network paths can support it | Public hub access is simpler but increases exposure. |
| Hybrid/multicloud member clusters | Arc-enabled member support is GA (Build 2026) | Capability parity with AKS members still differs — validate the capability matrix, support impact, regions, and Arc Gateway requirements before committing. |
| Cross-cluster networking or service discovery | Preview-gated pattern | Use only after confirming status, limits, network model, failure behaviour, and support boundary. |
| Upstream Cilium Cluster Mesh on AKS | Research/watchlist or explicitly validated exception | Prefer supported Azure/Fleet capabilities first; AKS-managed Cilium is not the same as unrestricted upstream Cilium. |
Permanent / Difficult / Reversible Decisions
| Decision | Difficulty | Review requirement |
|---|---|---|
| Whether a fleet is needed | Reversible to difficult | Easy before production, harder once clusters and governance workflows depend on it. |
| Fleet without hub vs with hub | Difficult | Hubless can be upgraded to hub mode; hub fleets cannot be downgraded to hubless. |
| Hub cluster public vs private access | Permanent | The hub network access type cannot be changed after creation. |
| Fleet resource region/subscription/tenant | Difficult | Member clusters must be in the same Microsoft Entra tenant; subscription and region choices affect ownership and operations. |
| Member cluster grouping strategy | Difficult | Drives upgrade rings, blast radius, placement, governance, ownership, and reporting. |
| Cluster label and taint taxonomy | Difficult | Placement, update stages, and governance rules depend on consistent labels. |
| Fleet update stages/groups | Reversible to difficult | Safe to adjust, but poor grouping can create rollout risk. |
| ClusterResourcePlacement for entire namespaces | Difficult | Removing a selected cluster can remove placed resources and affect traffic. |
| Namespace-scoped ResourcePlacement | Preview-gated/difficult | API/version support and production readiness must be confirmed. |
| Managed Fleet Namespaces | Preview-gated/difficult | Useful governance pattern, but production use requires explicit support validation and risk acceptance. |
| DNS load balancing / L4 multi-cluster services / cross-cluster networking | Preview-gated/difficult | Requires failure-mode testing, network validation, and clear support ownership. |
| Self-managed Cilium Cluster Mesh | Difficult | Requires explicit support, operations, identity, IP routing, policy, and troubleshooting model. |
Hubless vs Hub Fleet
Choose the smallest fleet shape that supports the requirement.
Fleet Manager without hub cluster
Use when the primary need is:
- grouping supported member clusters
- orchestrating Kubernetes version and node image updates across clusters
- aligning update runs with maintenance windows
- staging updates through non-prod, canary, production, and critical groups
Benefits:
- no hub cluster cost
- lower operational complexity
- can be upgraded later to a hub fleet if placement or managed namespace features become necessary
Limitations:
- no workload/resource placement
- no Managed Fleet Namespaces
- no DNS load balancing
- no hub Kubernetes API for KubeFleet-style resource propagation
Fleet Manager with hub cluster
Use when the design needs:
ClusterResourcePlacementorResourcePlacement- central staging of Kubernetes manifests on the hub for propagation to member clusters
- Managed Fleet Namespaces
- DNS load balancing or other hub-dependent multi-cluster networking capabilities
- placement policies based on member labels, location, cost, node count, or resource availability
Production guidance:
- Prefer private hub access when feasible.
- Do not mutate the managed hub AKS cluster or its underlying resources directly.
- Use
az fleet get-credentialsfor hub access. - Treat hub access, RBAC, and GitOps permissions as platform-admin capabilities, not application-team defaults.
Member Cluster Model
Define the fleet membership model before building clusters.
| Dimension | Recommended labels/tags | Design intent |
|---|---|---|
| Environment | `environment=dev | test |
| Criticality | `criticality=low | medium |
| Region | region=<azure-region> |
Residency, latency, failover, and placement. |
| Zone posture | `zones=single | multi` |
| Workload tier | `tier=platform | shared |
| Data class | `data-class=public | internal |
| Owner | owner=<team> |
Incident routing, cost, and change ownership. |
| Update group | `update-group=canary | early |
| Customer/tenant | tenant=<tenant-or-customer-code> |
Isolation, reporting, and delegated ownership. |
Rules:
- Labels used by placement or updates must be owned by the platform team and changed only through reviewed IaC or an approved platform workflow.
- Avoid labels that expose sensitive customer names or regulated data in places where cluster metadata is broadly visible.
- Use taints on member clusters for exclusion/exception patterns, then require explicit tolerations in placement rules.
- Do not place production workloads by broad
PickAllpolicies unless the blast radius and resource capacity are understood.
Update Orchestration
Use Fleet update orchestration when multiple member clusters need predictable Kubernetes or node image updates.
Recommended stage model:
| Stage | Example member selection | Purpose |
|---|---|---|
| Lab / platform validation | environment=dev, owner=platform |
Prove add-ons, policy, ingress, and observability still work. |
| Canary production | environment=prod, update-group=canary |
Validate real production conditions with small blast radius. |
| Standard production | environment=prod, update-group=standard |
Main rollout after canary passes. |
| Late / critical | criticality=mission-critical or update-group=late |
Final stage with explicit approval and runbook readiness. |
Update run requirements:
- Define update runs, update strategies, stages, groups, wait times, and maximum concurrency before production.
- Align update runs to member-cluster maintenance windows.
- Add manual approvals for production and mission-critical stages when supported and appropriate.
- Set maximum allowed failures (Preview, July 2026) when creating an update run so a small number of member-cluster failures can be tolerated instead of halting the fleet-wide run; document the chosen threshold and what it does not cover (e.g., failed health checks still stop the run).
- Separate Kubernetes version upgrades from node image updates where the risk profile differs.
- Stop or pause on failed health checks, pending pods, PDB violations, ingress failures, policy regressions, or application SLO breach.
- Do not use Fleet updates to compensate for weak per-cluster readiness. Each member cluster still needs normal AKS upgrade validation.
Validation commands:
# Azure fleet extension and fleet inventory
az extension add --name fleet
az extension update --name fleet
az fleet show --resource-group <rg> --name <fleet> --output yaml
az fleet member list --resource-group <rg> --fleet-name <fleet> --output table
# Hub access, if using hub mode
az fleet get-credentials --resource-group <rg> --name <fleet>
kubectl get memberclusters -o wide
kubectl get memberclusters --show-labels
# Update orchestration inventory
az fleet updaterun list --resource-group <rg> --fleet-name <fleet> --output table
az fleet updatestrategy list --resource-group <rg> --fleet-name <fleet> --output tableSee also: operations-resilience - Upgrade and Patch Strategy for the per-cluster upgrade machinery that staged rollouts invoke.
Auto-Upgrade Profiles and Emergency CVE Patching
Auto-upgrade profiles are the recommended way to keep a fleet on current Kubernetes and node image versions without hand-running an update run every time AKS publishes a new release. They are also the supported emergency response path for kernel/runtime CVE advisories such as the AKS-2026-0003 Copy Fail / Dirty Frag node-image class (illustrative - verify against current AKS Security Bulletins).
Channels
| Channel | What it tracks | When to use |
|---|---|---|
Stable |
Latest stable AKS Kubernetes minor (N-1) | Default cluster-version channel for most production fleets. |
Rapid |
Latest AKS Kubernetes minor (N) | Use only when the platform team actively tests new minors before they reach production stages. |
NodeImage |
AKS-published weekly node-image VHDs | Mandatory companion to a Kubernetes channel - keeps node OS patched between minor upgrades and is the primary CVE-response surface. |
Security patch (Preview; docs added 2026-08-21) |
In-place OS security patches between weekly VHD builds - the fleet analogue of the cluster-level nodeOSUpgradeChannel: SecurityPatch. Linux member-cluster nodes only; applies live patching when possible, otherwise deploys a newly patched machine image. CLI channel value: --channel SecurityPatch. |
Faster security-fix cadence than waiting for the next weekly image; Preview - verify current status against Automate upgrades of Kubernetes and node images before production adoption. |
TargetKubernetesVersion (preview) |
Pinned minor version | Only for organisations that need a specific minor and accept the preview gate. |
Creating an auto-upgrade profile:
# Create a NodeImage auto-upgrade profile bound to a phased update strategy.
# --channel accepts: NodeImage | Rapid | SecurityPatch | Stable | TargetKubernetesVersion
az fleet autoupgradeprofile create \
--resource-group <rg> \
--fleet-name <fleet> \
--name node-image-weekly \
--channel NodeImage \
--node-image-selection Consistent \
--update-strategy-id <update-strategy-resource-id>
# Inspect generated update runs from the profile
az fleet updaterun list --resource-group <rg> --fleet-name <fleet> --output table--node-image-selection accepts Latest (each member picks its newest available image at run time) or Consistent (every member converges to the same image build). Use Consistent for cross-region fleets so all members land on the same patched VHD.
IaC surface anchors.
Microsoft.ContainerService/fleets/autoUpgradeProfilesis the Bicep/ARM resource type for these profiles. Terraform exposes equivalents under theazurerm_kubernetes_fleet_*resource family - verify current azurerm provider support before relying on Terraform for fleet auto-upgrade profile creation (this resource set has churned across provider versions). The CLI is currently the most complete fleet surface.
Node image selection
| Setting | Behaviour | Recommended for |
|---|---|---|
Latest image |
Each member upgrades to the latest node image available in its own region; cross-region drift is possible | Single-region fleets, or fleets where members are intentionally independent. |
Consistent image |
Fleet Manager picks the latest VHD common to all member regions, so every cluster lands on the same node image | Default for multi-region production fleets and any rollout where image-version drift causes support or compliance problems. |
Snapshot caveat: agent pools created from a node-pool snapshot lose their creationData reference once a NodeImage channel run, or any Stable/Rapid/TargetKubernetesVersion run with Consistent image, picks them up. If a workload depends on a frozen snapshot, exclude that pool from auto-upgrade or accept that the snapshot reference will be removed.
Maintenance windows still apply
Auto-upgrade triggered runs honour the AKS-level maintenance windows on each member cluster:
aksManagedAutoUpgradeSchedule- Kubernetes minor/patch upgrades.aksManagedNodeOSUpgradeSchedule- node OS / VHD upgrades driven by theNodeImagechannel (or the cluster-level node OS auto-upgrade channel for non-fleet clusters).
For CVE response, design windows so they do not block patching for more than a single cycle. A four-hour weekly node OS window is a reasonable default; do not leave production on weekly off or shift the window so far out that the 90-day node-image validity period is breached.
To keep these per-cluster windows identical across the fleet, the Shared Maintenance Windows preview defines the schedule once as a standalone Microsoft.ContainerService/maintenanceWindows resource and links it into each member cluster's configurations via --maintenance-window-id; changing one resource retimes every linked cluster. It is complementary to Fleet Manager (windows = when, Fleet update runs = in what order) but is CLI/ARM-only and preview-gated; see operations-resilience before adopting.
90-day validity rule
Node image versions are valid for 90 days from publish. An update run that targets an image older than 90 days at the moment the run reaches a given member cluster can fail on that member. Treat this as a hard operating constraint:
- Run the
NodeImagechannel at minimum every 4 weeks; weekly is safer. - For staged fleets, size stage waits so the slowest stage still finishes within the 90-day window.
- Track
nodeImageVersionper pool and alert when any pool drifts past 60 days.
Validation commands
# Auto-upgrade profile inventory
az fleet autoupgradeprofile list --resource-group <rg> --fleet-name <fleet> --output table
az fleet autoupgradeprofile show --resource-group <rg> --fleet-name <fleet> --name <profile> --output yaml
# Current Kubernetes and node image state per member
az aks show --resource-group <rg> --name <cluster> --query "currentKubernetesVersion" --output tsv
az aks show --resource-group <rg> --name <cluster> --query "agentPoolProfiles[].{name:name,k8s:orchestratorVersion,nodeImage:nodeImageVersion}" --output tableEmergency CVE response runbook
Illustrative example using AKS-2026-0003 'Copy Fail'/Dirty Frag (verify the actual bulletin and patched VHD build against the AKS Security Bulletins feed before acting). Worked example: AKS-2026-0003 Copy Fail (algif_aead kernel LPE) and the Dirty Frag (CVE-2026-43284 / CVE-2026-43500) successor class. The same flow applies to any future AKS Security Bulletin that names a patched VHD build.
Emergency CLI examples:
# Generate an unscheduled NodeImage update run from an existing auto-upgrade profile
# Note: 'create' uses --name; 'generate-update-run' references the profile via --auto-upgrade-profile-name.
az fleet autoupgradeprofile generate-update-run \
--resource-group <rg> \
--fleet-name <fleet> \
--auto-upgrade-profile-name node-image-weekly
# For single-cluster emergency response, bypass the fleet and run the AKS-level command:
az aks nodepool upgrade \
--resource-group <rg> \
--cluster-name <cluster> \
--name <pool> \
--node-image-only \
--no-wait- Read the AKS Security Bulletin. Confirm the affected OS SKUs, the patched VHD build identifiers, and any interim host mitigation (for Copy Fail - illustrative; verify against current AKS Security Bulletins - the documented interim is
modprobe install algif_aead /bin/falsedeployed via DaemonSet). - Inventory exposure. Run
az aks nodepool list ... --query "[].{name:name,nodeImage:nodeImageVersion}"across every member cluster and compare against the patched VHD build. A Fleet-level loop driven from the member list is the fastest way to do this at scale. - Choose the response shape.
- Normal case: the fleet already has a
NodeImageauto-upgrade profile withConsistent image. The next scheduled run will land the patched VHD. Confirm the run will fit inside the 90-day window for every member. - Accelerated case: generate a one-off update run from the existing profile with
az fleet autoupgradeprofile generate-update-run --resource-group <rg> --fleet-name <fleet> --auto-upgrade-profile-name <profile>. This produces aNodeImageOnlyrun targeting the AKS-published image current at the moment of generation, honouring the existing strategy/stages. - Out-of-band case: create a manual update run with
--upgrade-type NodeImageOnlyand a tightened strategy that bypasses the normal canary wait for mission-critical members. Pair this with a temporary override ofaksManagedNodeOSUpgradeScheduleif the standing window would delay patching.
- Apply interim mitigation where the patched VHD is not yet reachable. For Copy Fail (illustrative - verify against current AKS Security Bulletins) this is the
algif_aeadmodprobe block delivered by DaemonSet to every affected node pool. Treat the mitigation as a placement-managed resource so it lands on every member cluster, and remove it once the patched VHD is confirmed. - Verify. Re-run the inventory query and confirm
nodeImageVersionmatches the patched build on every pool. Spot-check a node withkubectl debug node/<node> -it --image=<approved>oruname -avia DaemonSet logs. - Reconcile back to GitOps. Any temporary maintenance-window or strategy override must be reverted in source control and the auto-upgrade profile re-enabled. Capture the incident, runbook execution, and verification evidence in the platform change record.
See also: operations-resilience - CVE Response Decision Matrix for the urgency tiers and single-cluster fallback.
Stop conditions specific to auto-upgrade profiles
Stop and confirm before production rollout if:
- The fleet does not have a
NodeImageauto-upgrade profile. - A multi-region fleet is using
Latest imagerather thanConsistent imagewithout a documented reason. aksManagedNodeOSUpgradeScheduleis undefined, monthly-or-longer, or routinely overridden.- No process exists for monitoring AKS Security Bulletins and acting on
nodeImageVersiondrift. generate-update-runhas not been approved as the emergency CVE path, leaving operators to invent one under incident pressure.
Resource Placement
Status: GA (July 2026). Resource placement (ClusterResourcePlacement) for AKS fleets reached GA on 2026-07-29; namespace-scoped ResourcePlacement remains preview. Verify per-feature status in the target region before production.
Use resource placement only when central propagation is materially better than per-cluster GitOps or app-team deployment pipelines.
Good candidates:
- common namespaces and baseline namespace resources
- shared non-secret ConfigMaps where central ownership is appropriate
- platform add-on resources that must be identical across selected clusters
- workload placement where the platform, not individual app pipelines, decides eligible clusters
Poor candidates:
- secrets that should be sourced from Key Vault or workload-specific secret flows
- resources requiring cluster-specific mutation that is not modelled declaratively
- workloads with different release cadence or rollback ownership per cluster
- stateful workloads where placement changes can orphan data or break identity assumptions
Placement rules:
- Use explicit placement policies (
PickFixed,PickN, or carefully scopedPickAll) instead of broad propagation by default.PickFixed: Target specific clusters by name. Use for workload migration, cluster-specific deployments, or when cluster selection is deterministic.PickN: Select N clusters based on labels. Use for high availability or when placement should adapt to cluster state.PickAll: Place on all matching clusters. Use only for baseline resources (namespaces, shared ConfigMaps) where broad coverage is intentional.
- Prefer label-based member selection only after the label taxonomy is stable.
- Use
applyStrategywithwhenToTakeOver: IfNoDiffwhen taking over existing workloads from member clusters. See Workload Migration Between Clusters for the full migration pattern. - Use rollout strategies with conservative
maxUnavailable,maxSurge, andunavailablePeriodSecondsfor production resources. - Remember placement success means resources were applied; it does not automatically prove child pods are healthy or the app is serving traffic.
- Treat removal from placement as a destructive change because resources can be removed from member clusters.
- Use drift/diff status as operational evidence, but do not rely on it as the only production health signal.
Minimal production placement pattern:
apiVersion: placement.kubernetes-fleet.io/v1
kind: ClusterResourcePlacement
metadata:
name: platform-baseline-prod
spec:
resourceSelectors:
- group: ""
kind: Namespace
name: platform-baseline
version: v1
policy:
placementType: PickN
numberOfClusters: 2
affinity:
clusterAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
clusterSelectorTerms:
- labelSelector:
matchLabels:
environment: prod
update-group: canary
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 1
maxSurge: 1
unavailablePeriodSeconds: 60Validation commands:
az fleet get-credentials --resource-group <rg> --name <fleet>
kubectl get memberclusters --show-labels
kubectl get clusterresourceplacement
kubectl describe clusterresourceplacement <placement-name>
kubectl get resourceplacement -A
kubectl describe resourceplacement <placement-name> -n <namespace>Workload Migration Between Clusters
Use Fleet Manager resource placement to safely migrate workloads from one cluster to another without disruption. This pattern takes over existing workloads already running on member clusters and enables distribution to additional clusters.
When to use this pattern
- Migrating a workload from one cluster to another (e.g., during cluster maintenance or decommissioning)
- Taking over existing workloads not yet managed by Fleet Manager
- Distributing a workload across multiple clusters for high availability
- Moving workloads between regions or environments
Migration workflow
Step 1: Stage the workload on the Fleet Manager hub cluster
Apply the workload manifest to the Fleet Manager hub cluster. The workload is not scheduled on the hub but is ready for Fleet Manager to distribute.
# Get hub cluster credentials
az fleet get-credentials --resource-group <rg> --name <fleet>
# Apply workload to hub (not scheduled, just staged)
kubectl apply -f workload.yamlStep 2: Take over the existing workload on the source cluster
Use a PickFixed placement policy with applyStrategy to take over the workload running on a member cluster. The whenToTakeOver: IfNoDiff setting ensures Fleet Manager only takes over when there are no differences between the hub and member cluster workloads.
apiVersion: placement.kubernetes-fleet.io/v1beta1
kind: ClusterResourcePlacement
metadata:
name: workload-migration
spec:
resourceSelectors:
- group: ""
kind: Namespace
version: v1
name: "my-workload"
policy:
placementType: PickFixed
clusterNames:
- source-cluster
strategy:
applyStrategy:
whenToTakeOver: IfNoDiff
comparisonOption: PartialComparisonApply the manifest to the hub cluster:
kubectl apply -f placement-takeover.yamlMonitor the placement:
kubectl get clusterresourceplacement workload-migrationWhen the placement completes, Fleet Manager is now responsible for the workload on the source cluster.
Step 3: Add the target cluster to move the workload
Modify the placement to include the target cluster. Fleet Manager will roll out the workload to the new cluster.
apiVersion: placement.kubernetes-fleet.io/v1beta1
kind: ClusterResourcePlacement
metadata:
name: workload-migration
spec:
resourceSelectors:
- group: ""
kind: Namespace
version: v1
name: "my-workload"
policy:
placementType: PickFixed
clusterNames:
- source-cluster
- target-cluster
strategy:
applyStrategy:
whenToTakeOver: IfNoDiff
comparisonOption: PartialComparisonReapply the manifest to the hub cluster and watch as the workload rolls out to the target cluster.
applyStrategy options
| Option | Values | Purpose |
|---|---|---|
whenToTakeOver |
IfNoDiff, Always, Never |
Controls when Fleet Manager takes over existing workloads. IfNoDiff is safest for migration. |
comparisonOption |
PartialComparison, FullComparison |
Controls how thoroughly Fleet Manager compares hub vs member workloads. PartialComparison is faster. |
whenToTakeOver values:
IfNoDiff(recommended): Take over only when hub and member workloads are identical. Safest for existing workloads.Always: Take over regardless of differences. Use when you want Fleet Manager to overwrite the member workload.Never: Do not take over existing workloads. Only place on clusters with no existing workload.
comparisonOption values:
PartialComparison(recommended): Compare key fields only. Faster, suitable for most migration scenarios.FullComparison: Compare all fields. More thorough but slower.
PickFixed vs PickN for migration
| Policy | Use when |
|---|---|
PickFixed |
Migrating to/from specific clusters by name. Most common for workload migration. |
PickN |
Distributing workload across N clusters based on label selectors. Use for high availability. |
PickAll |
Placing workload on all clusters matching a selector. Use with caution for baseline resources. |
Validation commands
# Monitor placement status
kubectl get clusterresourceplacement <name> -o yaml
# Check placement conditions
kubectl describe clusterresourceplacement <name>
# Verify workload on member clusters
kubectl get pods -n <namespace> --context <member-context>
# Compare hub vs member state
kubectl diff -f workload.yaml --context <member-context>Managed Fleet Namespaces
Use Managed Fleet Namespaces to consider central namespace governance across member clusters when the organisation needs consistent quotas, labels, annotations, RBAC, or NetworkPolicy baselines.
Production gate:
- Verify current GA/preview status before recommending production use.
- If preview, state that it is non-production by default unless risk acceptance, support scope, rollback, and migration path are documented.
- Do not use managed namespaces to bypass application-team ownership. They define the boundary; workload teams still need clear release, incident, and cost ownership.
Recommended namespace baseline:
- owner, cost-centre, environment, app, data-class, criticality, and support-contact labels
- ResourceQuota and LimitRange
- Pod Security Admission labels
- default-deny NetworkPolicy plus explicit DNS/ingress/east-west/egress allow rules
- RBAC groups aligned to Entra ID
- audit evidence that the namespace exists and is governed consistently across selected member clusters
Cross-Cluster Networking
Treat Fleet Manager multi-cluster networking as an advanced, verification-driven pattern.
| Pattern | Use when | Gate |
|---|---|---|
| DNS load balancing | Public north-south entry across services exported from multiple AKS clusters | Verify preview/GA, Traffic Manager model, DNS ownership, health/failover, and app behaviour. |
| L4 multi-cluster load balancing | North-south load balancing across services inside a virtual network | Verify preview/GA, Azure Load Balancer behaviour, endpoint routing, and network reachability. |
| East-west cross-cluster networking | Service-to-service communication across clusters with policy enforcement | Verify current support, ACNS/Cilium requirements, flat/routed networking assumptions, and failure isolation. |
| Upstream Cilium Cluster Mesh | non-default exception | Validate support boundary, Cilium version, AKS CNI mode, identity, pod/service CIDRs, VNet peering/routing, and operational ownership. |
Rules:
- Do not recommend cross-cluster networking merely because multiple clusters exist.
- Prefer clear API contracts, externalised data stores, and regional ingress where they meet the requirement.
- Use cross-cluster service discovery only when application topology requires it and the failure modes are tested.
- Confirm how NetworkPolicy applies across local and remote endpoints.
- Document what happens when one cluster, one region, the hub, DNS, Traffic Manager, or the fleet network component is unavailable.
- Test local-preference/failover behaviour rather than assuming global service balancing matches application intent.
Cilium Cluster Mesh Watchpoints
Microsoft-managed Cilium on AKS does NOT enable upstream Cluster Mesh by default; Cluster Mesh is a self-managed extension and falls outside AKS supportability.
The New Stack fleet-scale theme is useful because it highlights a real industry pattern: fleet management and Cilium-based multi-cluster networking can help at very large scale. In this skill, convert that into guardrails rather than a blanket default.
Use this wording in recommendations:
For AKS, prefer Microsoft-supported Fleet Manager and AKS networking capabilities first. Consider upstream Cilium Cluster Mesh only as an exception after supportability, routing, identity, policy, and operational ownership are validated.
Watchpoints:
- AKS Azure CNI Powered by Cilium is Microsoft-managed; do not assume all upstream Cilium flags, Cluster Mesh behaviours, or Helm settings are supported.
- Cilium Cluster Mesh extends the datapath across clusters and supports policy enforcement and annotation-based load balancing, but it adds operational and network failure modes.
- KVStoreMesh improves scalability/isolation compared with each agent directly reading remote cluster state, but it still requires explicit operations, monitoring, upgrade, and incident ownership.
- Hubble or equivalent network visibility becomes more important when traffic crosses clusters.
- Do not mix Fleet multi-cluster networking and self-managed Cilium Cluster Mesh patterns without a written support and troubleshooting model.
See also: cluster-foundations - Azure CNI Overlay Powered by Cilium for the data-plane prerequisite, and production-workload-controls - Network Policy Strategy for namespace-level policy that interacts with cross-cluster traffic.
Security and Identity
- Use Microsoft Entra ID and least-privilege RBAC for Fleet Manager, hub access, member cluster access, and placement operations.
- Separate ARM-plane Fleet Manager contributors from Kubernetes data-plane users on the hub and member clusters.
- Use the Fleet Manager Hub Cluster User role for read-only hub inspection where appropriate; reserve hub-cluster admin/write roles for platform operators who own placement and fleet resources.
- Be careful with Fleet Manager RBAC Writer-style roles: writer access that can read Secrets can also assume service-account credentials in that namespace. Prefer reader roles for audit/review users and tightly scope writer/admin roles.
- Separate fleet administrators from application namespace contributors.
- Do not give application teams broad hub-cluster permissions unless they are explicitly responsible for fleet-level placement.
- Treat resource placement as a privileged deployment path because it can affect many clusters.
- Avoid distributing secrets through generic placement. Prefer workload identity plus Key Vault/External Secrets patterns where supported by the organisation.
- Audit changes to member labels, update strategies, update runs, auto-upgrade profiles, placement resources, Managed Fleet Namespaces, RBAC assignments, and cross-cluster networking CRDs.
Observability and Operational Readiness
Fleet readiness requires per-cluster and fleet-level evidence:
- member cluster health, version, and node image state
- update run status, failed stages/groups, and approvals
- placement scheduled/applied/available/drifted/diffed status
- namespace governance compliance
- cross-cluster networking health, DNS/Traffic Manager status, and service export/import health where used
- app-level SLOs per cluster and across the global service
- cost and capacity by member cluster, region, namespace, and owner
Do not mark a fleet rollout complete only because the hub reports resources as applied. Confirm workload readiness on member clusters.
Cost and Complexity Controls
- Do not create a hub fleet only for update orchestration.
- Account for hub cluster cost when hub mode is required.
- Keep the number of fleets small and aligned to clear operational boundaries such as tenant, platform, environment, or sovereignty boundary.
- Avoid one fleet per app unless isolation or ownership requires it.
- Use labels and placement to reduce duplication, but do not turn the fleet hub into an unreviewed central dumping ground for every manifest.
- Track Traffic Manager, Load Balancer, log/metric ingestion, cross-region egress, hub operations, and duplicated platform add-on costs.
Stop Conditions
Stop the design and ask for explicit confirmation if any of these are unknown:
- Is this multi-cluster, or are we adding Fleet before it has value?
- Are we using hubless or hub mode, and why?
- If hub mode is required, is public vs private hub access approved before creation?
- Which clusters can join the fleet, and are they in the same Microsoft Entra tenant?
- Which labels and taints are authoritative for update and placement decisions?
- Which features are GA and which are preview in the target region and member cluster type?
- What is the rollback plan for placement removing or changing resources across clusters?
- Who owns fleet hub access, placement policies, update approvals, and incidents?
- How are app health, DNS, ingress, and cross-cluster networking validated after rollout?
ADRs Required
Create ADRs for:
- Fleet Manager adoption or non-adoption.
- Hubless vs hub fleet.
- Hub public vs private access.
- Member cluster taxonomy, labels, taints, and update groups.
- Fleet update strategy, approvals, and maintenance-window model.
- Resource placement scope and ownership.
- Managed Fleet Namespace use and preview risk acceptance, if used.
- Cross-cluster networking and service discovery, if used.
- Cilium Cluster Mesh exception, if used.
Use With
- Top-level router: ../../SKILL.md
- Cluster foundations: ../cluster-foundations/guide.md
- Workload platform: ../workload-platform/guide.md
- Production workload controls: ../production-workload-controls/guide.md
- Operations resilience: ../operations-resilience/guide.md
- Full reference: ../../references/full-reference.md