Quick Reference - Non-Negotiables and Stop Conditions
Single-page summary for production gate reviews. Each item links back to its source bundle. Use this in tandem with the relevant bundle for context - this card is a checklist, not a substitute for the reasoning in the bundles.
How to use this card. For each Non-Negotiable, ask: (1) Does my design satisfy it as-stated, or do I have a documented ADR exception? (2) If exception, what is the expiry and named owner? (3) Which Stop Condition would my answer trip? If any Stop Condition fires, halt and resolve before production sign-off.
Non-Negotiables
- [Identity]: Use managed identities and Microsoft Entra Workload ID; no service principal secrets for new clusters. (SKILL.md)
- [CNI]: Prefer Azure CNI Overlay Powered by Cilium for new general-purpose clusters when workload and region support it. (SKILL.md)
- [Egress]: Do not rely on default load balancer SNAT for production; use NAT Gateway or UDR-through-firewall. (SKILL.md)
- [OIDC / Workload Identity]: Enable OIDC issuer and Workload Identity at cluster creation. (SKILL.md)
- [RBAC]: Use Azure RBAC for Kubernetes authorization unless a documented exception exists. (SKILL.md)
- [Pricing tier]: Use an AKS pricing tier suitable for production; do not silently default production to a free/dev posture. (SKILL.md)
- [Policy and platform telemetry]: Use Azure Policy / Deployment Safeguards, Managed Prometheus, ContainerLogV2, and a defined node OS patch channel from day one. (SKILL.md)
- [Node OS channel and window]: Set an explicit
--node-os-upgrade-channel(defaultNodeImage, neverNonein production) and anaksManagedNodeOSUpgradeSchedulethat does not block patching for more than one cycle. (SKILL.md) - [Fleet auto-upgrade]: For multi-cluster fleets, define a Fleet Manager
NodeImageauto-upgrade profile, chooseConsistent imagefor cross-region fleets, and pre-approveaz fleet autoupgradeprofile generate-update-runas the emergency CVE path. (SKILL.md) - [Node-image validity]: Treat AKS node images as having a 90-day validity window; designs must include a recurring patch cadence and a CVE-response runbook tied to AKS Security Bulletins. (SKILL.md)
- [NetworkPolicy]: Use default-deny NetworkPolicy in production namespaces, then add explicit label-based allow rules for DNS, ingress, east-west, and approved egress. (SKILL.md)
- [Namespace contract]: Every production namespace has ownership labels, ResourceQuota, LimitRange, RBAC boundary, and Pod Security Admission labels, or a documented exception. (SKILL.md)
- [Replicas and disruption]: Run production stateless services with at least two replicas (usually three across zones), plus readiness probes, PDBs, topology spread, and autoscaling limits. (SKILL.md)
- [Node-pool separation]: Separate system, user, spot, Windows, GPU, and stateful workloads into appropriate node pools or scheduling domains. (SKILL.md)
- [Storage]: For persistent data, choose the StorageClass and zone/replication model explicitly; Azure managed disks are zonal and pin a pod to one zone. Use
WaitForFirstConsumer+ cross-zone replication or ZRS, and size quotas against node allocatable. (cluster-foundations) - [Ingress / Gateway]: Avoid upstream/self-managed ingress-nginx as a new long-term production baseline; name AGC as the strategic target and document any NGINX bridge step with its migration path. (SKILL.md)
- [Preview features]: Do not recommend previews for production without calling out preview status, limitations, rollback approach, and support impact. (SKILL.md)
- [Fleet adoption]: Do not introduce Fleet Manager for a single-cluster design unless a multi-cluster operating model is approved with funded delivery plans. (SKILL.md)
- [Fleet shape]: Decide explicitly whether the fleet is hubless or hub-based and whether hub access must be private, before provisioning. (SKILL.md)
- [Verification-driven features]: Treat Fleet Manager cross-cluster networking, Managed Fleet Namespaces, Arc-enabled members, namespace-scoped placement, and Cilium Cluster Mesh as verification-driven with explicit support and preview gates. (SKILL.md)
Stop Conditions by Bundle
cluster-foundations
- Target region and data residency requirements not confirmed. (See cluster-foundations)
- Hub-spoke / VNet / on-prem CIDR ranges not known. (See cluster-foundations)
- Private API server requirement not decided. (See cluster-foundations)
- Outbound inspection requirement not decided. (See cluster-foundations)
- AKS Automatic vs Standard decision not made. (See cluster-foundations)
- Production availability target and zone support not confirmed. (See cluster-foundations)
- Identity / RBAC ownership model not agreed. (See cluster-foundations)
- GPU / Windows / confidential compute requirement not known. (See cluster-foundations)
workload-platform
- Node pool VM family, SKU, OS, or disk type chosen without performance evidence. (See workload-platform)
- HPA / KEDA / VPA enabled on a workload without CPU/memory requests or validated metrics/scaler source. (See workload-platform)
- HPA and VPA
Auto/Recreateconfigured on the same CPU/memory target without a tested design. (See workload-platform) - Workload Identity not enabled at cluster creation, or workloads still using AAD Pod Identity / SP secrets / shared platform MI. (See workload-platform)
- NAP / Karpenter introduced without confirmed networking compatibility and resource-request accuracy. (See workload-platform)
- Production stateless services at one replica without an explicit singleton design and ADR. (See workload-platform)
- GPU pools provisioned without verified regional quota, VM family, driver model, and a workload that requires GPU. (See workload-platform)
- Spot / Windows / stateful pools introduced without explicit scheduling rules and an upgrade plan. (See workload-platform)
- Scaling maximums not reconciled with downstream dependency capacity and budget. (See workload-platform)
- Workload-to-Azure paths bypass Workload Identity without a documented rotation owner. (See workload-platform)
production-workload-controls
- Namespace has no
ResourceQuota/LimitRange, or placeholder values not aligned to SLO and budget. (See production-workload-controls) pod-security.kubernetes.io/enforceunset, orrestrictedrequested for a namespace with known violations and no documented exception. (See production-workload-controls)- Default-deny NetworkPolicy about to be enforced without runtime connectivity tests in non-prod. (See production-workload-controls)
- Production stateless workloads at one replica, or with no PDB and no topology spread, without a documented exception. (See production-workload-controls)
- HPA / KEDA enabled without requests, accurate metrics, Workload-Identity-based scaler auth, and reconciled
maxReplicas. (See production-workload-controls) - Images referenced by mutable tags rather than immutable digests, unscanned, or with no registry allow-list at admission. (See production-workload-controls)
- CI pipeline does not validate manifests against the same policy set the cluster enforces. (See production-workload-controls)
- VPA
Autoenabled on a workload that also has HPA on CPU/memory without a tested design. (See production-workload-controls) - Workload writes its own Kubernetes
Secretwith Azure connection strings instead of using Workload Identity, with no rotation owner. (See production-workload-controls)
operations-resilience
- Ingress / Gateway controller chosen without confirmed support status, TLS/DNS ownership, or a migration path off any bridge controller. (See operations-resilience)
- Managed Prometheus, ContainerLogV2, baseline alerts, or audit-log forwarding to Sentinel/SIEM with immutable retention not configured. (See operations-resilience)
- Auto-upgrade channel set but
aksManagedAutoUpgradeSchedulemissing, or windows routinely skipped. (See operations-resilience) --node-os-upgrade-channelset to anything other thanNodeImagein production without a documented reason. (See operations-resilience)- Kubernetes minor upgrade planned without API-deprecation inventory, admission/conversion webhook checklist, or CRD storage-version migration check. (See operations-resilience)
- Add-on / operator compatibility (Flux, Argo CD, KEDA, Cilium/ACNS, NGINX, AGC, cert-manager, Istio, Kyverno, Gatekeeper, Ratify) not verified against target minor. (See operations-resilience)
- No recovery design for a failed minor upgrade (parallel cluster, restore-from-backup, node-pool snapshot) documented and rehearsed. (See operations-resilience)
- Backups exist but restore has never been tested end-to-end. (See operations-resilience)
- GitOps reconciliation in place but break-glass not time-bound, audited, and reconciled back into Git. (See operations-resilience)
- IaC pipeline missing pre-merge gates (tfsec/checkov, Azure Policy what-if, drift detection) or destructive-change protection. (See operations-resilience)
- CVE response runbook absent, or
az fleet autoupgradeprofile generate-update-run/az aks nodepool upgrade --node-image-onlynot pre-approved as the emergency path. (See operations-resilience) - Multi-region or DR pattern claimed but data plane and failover not validated. (See operations-resilience)
fleet-management
- Genuine multi-cluster need not established, or Fleet Manager added before it has value. (See fleet-management)
- Hubless vs hub mode not chosen with a documented reason. (See fleet-management)
- Public vs private hub access not approved before hub creation. (See fleet-management)
- Member clusters not in the same Microsoft Entra tenant, or eligibility not documented. (See fleet-management)
- Authoritative labels / taints for update and placement decisions not defined. (See fleet-management)
- GA vs preview status of selected features not confirmed for the target region and member cluster type. (See fleet-management)
- Rollback plan for placement removing or changing resources across clusters not defined. (See fleet-management)
- Ownership of fleet hub access, placement policies, update approvals, and incidents not agreed. (See fleet-management)
- Post-rollout validation for app health, DNS, ingress, and cross-cluster networking not designed. (See fleet-management)
- Fleet lacks a
NodeImageauto-upgrade profile or usesLatest imageacross a multi-region fleet without documented reason. (See fleet-management) aksManagedNodeOSUpgradeScheduleundefined, too infrequent, or routinely overridden. (See fleet-management)- No process for monitoring AKS Security Bulletins and acting on
nodeImageVersiondrift;generate-update-runnot pre-approved as the emergency CVE path. (See fleet-management)
This file is generated from authoritative content in the bundles - keep it in sync via the validator and the skill's release process.