All skills
microsoft avatar

/azure-kubernetes

@a8f19b4
by microsoftmicrosoft/skills3.1k stars
351

Plan, create, and configure production-ready Azure Kubernetes Service (AKS) clusters. Covers Day-0 checklist, SKU selection (Automatic vs Standard), networking options (private API server, Azure CNI Overlay, egress configuration), security, and operations (autoscaling, upgrade strategy, cost analysis). WHEN: create AKS environment, provision AKS, enable AKS observability, design AKS networking, choose AKS SKU, secure AKS, optimize AKS, AKS spot nodes, AKS cluster-autoscaler, rightsize AKS pod, pod rightsizing, over-provisioned AKS pod, pod resource requests and limits, Vertical Pod Autoscaler, VPA recommendations.

Use this Skill: https://skilld.dev/gh/microsoft/skills/azure-kubernetes

This session only. Nothing lands on disk.

azure-kubernetes-automatic-readinessreferencescommon-fixes.md

≈1.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Common Fix Patterns for AKS Automatic Compatibility

Loaded on demand when generating YAML fixes during assessment. Maps to constraint IDs in constraint-spec-v1.yaml.


safeguard-container-resource-requests — Add resource requests/limits

Before:

containers:
  - name: web
    image: myapp:v1.0.0

After:

containers:
  - name: web
    image: myapp:v1.0.0
    resources:
      requests:
        cpu: "250m"
        memory: "256Mi"
      limits:
        cpu: "500m"
        memory: "512Mi"

💡 Tip: Use safe minimums as starting values. VPA (auto-enabled on AKS Automatic) will tune these after deployment based on actual usage.


safeguard-container-capabilities — Drop all capabilities

Before:

securityContext:
  capabilities:
    add: ["NET_ADMIN"]

After:

securityContext:
  capabilities:
    drop: ["ALL"]

⚠️ Warning: If the app genuinely requires NET_ADMIN or similar, it is incompatible with AKS Automatic. Do not silently drop — explain the incompatibility and suggest redesign.


safeguard-allowed-seccomp-profiles — Add seccomp profile

Before:

spec:
  containers:
    - name: web

After:

spec:
  securityContext:
    seccompProfile:
      type: RuntimeDefault
  containers:
    - name: web

safeguard-allowed-seccomp-profiles — Remove 'Unconfined' seccomp profile

Before:

spec:
  securityContext:
    seccompProfile:
      type: Unconfined
  containers:
    - name: web

After:

spec:
  containers:
    - name: web

safeguard-enforce-apparmor — Add AppArmor annotation

Before:

metadata:
  name: my-deployment

After:

metadata:
  name: my-deployment
  annotations:
    container.apparmor.security.beta.kubernetes.io/web: runtime/default

💡 Tip: Replace web with the actual container name. Add one annotation per container.


safeguard-images-no-latest — Pin image tag (LLM-reasoned — ask user)

Before:

image: myapp:latest

After:

image: myapp:v1.2.3   # ← version confirmed with user

⚠️ Warning: Do not guess the version. Ask the user: "What specific version tag or SHA digest should I pin this image to?" If from a public registry, suggest checking Docker Hub or the registry for the latest stable tag.


safeguard-probes-configured — Add probes (best-practice recommendation — warning-only, not blocked at admission)

HTTP app (most common):

readinessProbe:
  httpGet:
    path: /healthz        # ← ask user for their health endpoint
    port: 8080            # ← ask user for port
  initialDelaySeconds: 5
  periodSeconds: 10
  failureThreshold: 3
livenessProbe:
  httpGet:
    path: /healthz
    port: 8080
  initialDelaySeconds: 15
  periodSeconds: 20
  failureThreshold: 3

TCP-only app (databases, Redis, etc.):

readinessProbe:
  tcpSocket:
    port: 6379           # ← service port
  initialDelaySeconds: 5
  periodSeconds: 10
livenessProbe:
  tcpSocket:
    port: 6379
  initialDelaySeconds: 15
  periodSeconds: 20

gRPC app:

readinessProbe:
  grpc:
    port: 50051
  initialDelaySeconds: 5
  periodSeconds: 10

safeguard-host-probes — Remove host field in probes and lifecycle hooks

Before:

spec:
  containers:
  - name: my-container
    image: nginx:v1.2.3
    livenessProbe:
      httpGet:
        host: "my-host"
        path: /healthz
        port: 8080
      initialDelaySeconds: 15
      periodSeconds: 20
      failureThreshold: 3

After: Remove the host field Example:

spec:
  containers:
  - name: my-container
    image: nginx:v1.2.3
    livenessProbe:
      httpGet:
        path: /healthz
        port: 8080
      initialDelaySeconds: 15
      periodSeconds: 20
      failureThreshold: 3

safeguard-pod-enforce-antiaffinity — Add topology spread (LLM-reasoned — ask user for label)

Ask user: "What label key/value identifies your workload's pods?"

spec:
  template:
    spec:
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: kubernetes.io/hostname
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels:
              app: <app-label>     # ← from user
      containers:
        - name: web

safeguard-csi-driver-storage-class — Migrate in-tree to CSI

Before (Azure Disk in-tree):

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: fast-storage
provisioner: kubernetes.io/azure-disk
parameters:
  skuName: Premium_LRS
reclaimPolicy: Delete
volumeBindingMode: Immediate

After (Azure Disk CSI):

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: fast-storage
provisioner: disk.csi.azure.com
parameters:
  skuName: Premium_LRS
reclaimPolicy: Delete
volumeBindingMode: WaitForFirstConsumer   # ← preferred for zonal disks
In-tree provisioner CSI replacement
kubernetes.io/azure-disk disk.csi.azure.com
kubernetes.io/azure-file file.csi.azure.com

PodDisruptionBudget — Add missing PDB

apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: <app-name>-pdb
  namespace: <namespace>
spec:
  maxUnavailable: 1
  selector:
    matchLabels:
      app: <app-label>

PodDisruptionBudget — Fix blocking maxUnavailable: 0

Before:

spec:
  maxUnavailable: 0

After:

spec:
  maxUnavailable: 1

⚠️ Warning: maxUnavailable: 0 completely blocks node drain during AKS Automatic upgrades. At least 1 pod must be allowed unavailable for upgrades to proceed.


safeguard-no-host-path-volumes — Replace hostPath (incompatible — suggest alternatives)

hostPath use case Recommended replacement
Log collection (/var/log) Azure Monitor Container Insights (auto-enabled on AKS Automatic)
Container runtime socket (/var/run/docker.sock) Use the AKS Automatic node observability features — direct socket access not supported
Shared config files configMap volume
Secrets / credentials Kubernetes secret volume or Azure Key Vault CSI Driver
Ephemeral scratch space emptyDir volume
Persistent app data Azure Disk CSI via PVC (disk.csi.azure.com)
Shared file storage across pods Azure Files CSI via PVC (file.csi.azure.com)

emptyDir example:

volumes:
  - name: scratch
    emptyDir: {}

Azure Files CSI PVC example:

apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: logs-pvc
spec:
  accessModes:
    - ReadWriteMany
  storageClassName: azurefile-csi
  resources:
    requests:
      storage: 10Gi

Source: SKILL.md on GitHub

No alerts9d3 checks · Risk SAFE
  • Gen Agent Trust Hub9d

    This skill provides comprehensive, production-grade guidance for Azure Kubernetes Service (AKS) with a strong emphasis on security best practices, including Workload Identity, hardened container configurations, and automated readiness assessments.

  • Socket9d

    No alerts

  • Snyk9d

    Risk: LOW · No issues

Signed by skilld at a8f19b4. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
metadata
{
  "author": "Microsoft",
  "version": "1.2.2"
}

README badge

README badge for microsoft/skills/azure-kubernetes