All skills
lukemurraynz avatar

/aks-cluster-architecture

@2cc2455

AKS cluster architecture decisions for new Azure Kubernetes Service projects: AKS Automatic vs Standard, networking topology, dual-stack (IPv4/IPv6), Kubernetes version and OS currency, node pool strategy, identity, production NetworkPolicy, namespaces, autoscaling, ingress and Gateway API, observability, operations, resilience, GPU and AI workloads, GPU partitioning (MIG, time-slicing, MPS), batch scheduling (Kueue), AKS on bare metal, AI Runway and KAITO model serving, AKS MCP server access, kars (Agent Reference Stack for Kubernetes) for agent isolation, Kata MicroVM pod sandboxing, Azure Kubernetes Fleet Manager, multi-cluster governance, update orchestration, resource placement, cross-cluster networking, and cost. WHEN: designing new AKS clusters, reviewing production readiness, choosing CNI or outbound connectivity, planning node pools, defining namespace, network and security controls, evaluating Fleet Manager, deploying AI agent runtimes on AKS, or making hard-to-reverse infrastructure decisions.

Use this Skill: https://skilld.dev/gh/lukemurraynz/hve-agent-skills/aks-cluster-architecture

This session only. Nothing lands on disk.

bundlesagent-runtimeguide.md

≈3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Agent Runtime Bundle

Deploying AI agent runtimes on AKS: kars (Agent Reference Stack for Kubernetes, an open-source Microsoft reference stack, not an officially supported product), Kata Containers pod sandboxing, KAITO / AI Runway model serving, AKS MCP server patterns, and security isolation for untrusted agent code.

Load this bundle for: deploying agent workloads that execute untrusted code, choosing between isolation primitives, designing agent-specific node pools, and wiring MCP servers to AKS-hosted agents.

kars. Agent Reference Stack for Kubernetes

kars (Azure/kars) is an open-source, Kubernetes-native Agent Reference Stack from Microsoft's Azure Cloud Native team for running AI agents on AKS. It is a reference implementation, not an officially supported Microsoft product (no SLA or support contract). It treats every agent as untrusted code with defense-in-depth isolation:

Isolation layer Mechanism
No-network agent container The agent runs as an unprivileged container (UID 1000) with no network access; all outbound traffic is forced through a co-located inference router
Rust inference router A separate container (UID 1001) is the security enforcement core: it holds all credentials and brokers every outbound call, enforcing identity, content safety, egress policy, and audit
Zero credentials in agent process The router exchanges a per-sandbox Microsoft Entra Agent ID (or cluster Workload Identity) for backend tokens via federated OIDC / IMDS; the agent process never sees a credential
L7 egress governance Every outbound CONNECT is checked against a per-sandbox allowlist plus a blocklist (OISD + URLhaus); token budgets, rate limits, and EgressApproval CRDs apply
Encrypted inter-agent mesh AgentMesh uses the Signal Protocol with a KNOCK trust handshake; the A2A Gateway handles cross-org peer-to-peer traffic
Belt-and-braces containment Default-deny NetworkPolicy and an egress-guard init container (iptables) contain blast radius if the router is bypassed

Microsoft Agent Framework (MAF) is one of kars's supported agent runtimes, and the router includes an MCP gateway that brokers calls to external MCP servers with OAuth and per-tool allowlists. Running kars node pools with AKS Pod Sandboxing (Kata Containers) optionally adds a per-pod guest kernel on top of the container-and-router model.

When to Use kars

Scenario Use kars?
Running third-party or user-supplied agent code Yes — the no-network agent container plus the credential-holding router is the baseline
Hosting MCP server-backed agents on AKS Yes — kars provides the hardened runtime; wire MCP servers as sidecars or cluster services
Multi-tenant agent platform (many agents per cluster) Yes — namespace-per-agent with resource quotas
Simple single-agent deployment with trusted code Consider ACA Express or Container Apps instead
Model inference/serving only (no agent tool execution) Use AI Runway / KAITO directly

Architecture

┌──────────────────────────────────────────────────┐
│  AKS Cluster (Automatic or Standard)               │
│                                                    │
│  ┌────────────────────────────────────────────┐    │
│  │  Sandbox pod (one agent per pod)           │    │
│  │  ┌───────────────┐   ┌───────────────────┐ │    │
│  │  │ Agent         │   │ Inference router  │ │    │
│  │  │ container     │─► │ (Rust)            │─┼──► governed egress
│  │  │ UID 1000      │   │ UID 1001          │ │    │  (L7 allow/blocklist)
│  │  │ no network    │   │ holds credentials │ │    │
│  │  └───────────────┘   └───────────────────┘ │    │
│  │  NetworkPolicy: default-deny               │    │
│  │  egress-guard init container (iptables)    │    │
│  └────────────────────────────────────────────┘    │
│                                                    │
│  AgentMesh (Signal Protocol) · A2A Gateway         │
│  MCP gateway (in router) ─► external MCP servers   │
└──────────────────────────────────────────────────┘

Deployment Checklist

Before deploying kars agents to production:

  • Node pool: Dedicated agent node pool with runtimeClassName: kata-vm-isolation or AKS Pod Sandboxing enabled.
  • VM SKU: Confidential VM (CVM) SKUs or v5+ family with nested virtualization support for Kata.
  • Identity: Per-agent Workload Identity (azure.workload.identity/client-id annotation on ServiceAccount). Zero static credentials.
  • NetworkPolicy: Default-deny per namespace; explicit allow rules for DNS, MCP endpoints, and approved egress.
  • ResourceQuota: CPU, memory, and ephemeral storage limits per agent namespace.
  • Pod Security: restricted Pod Security Admission with explicit securityContext.
  • Observability: ContainerLogV2, Managed Prometheus, and agent-specific metrics.
  • MCP transport: Streamable HTTP (preferred) or stdio via sidecar pattern. Do not expose MCP servers on public IPs.
  • Lifecycle: Agent pods should be stateless; checkpoint state externally. Use terminationGracePeriodSeconds for clean shutdown.

Agent Isolation Primitives Comparison

Primitive Isolation level Performance overhead Use case
Kata MicroVM / Pod Sandboxing Per-pod guest kernel Moderate (VM boot per pod) Untrusted code, multi-tenant agents
Confidential VM node pool Hardware-level (SEV-SNP/TDX) Low-moderate Data-sensitive agents, regulated workloads
gVisor (runsc) User-space kernel Low Not officially supported on AKS (GKE only); manual install carries the operational burden
Standard pod (no isolation) None None Trusted first-party agents only

kars's baseline isolation is the no-network agent container plus the inference router. To add a per-pod guest kernel on AKS, enable Pod Sandboxing (Kata Containers) with runtimeClassName: kata-vm-isolation, which is the AKS-native, officially supported (GA) way to get kernel isolation per agent. gVisor is not an AKS-supported runtime class, so treat that row as a general comparison, not an AKS option. Confidential Containers on AKS (kata-cc-isolation) is a separate, retiring preview, so do not build new work on it.

apiVersion: v1
kind: Pod
metadata:
  name: agent-pod
  namespace: agent-<id>
spec:
  runtimeClassName: kata-vm-isolation
  serviceAccountName: agent-sa
  containers:
  - name: agent
    image: <agent-image>
    resources:
      requests:
        cpu: "500m"
        memory: "512Mi"
      limits:
        cpu: "1"
        memory: "1Gi"
    securityContext:
      runAsNonRoot: true
      seccompProfile:
        type: RuntimeDefault
      allowPrivilegeEscalation: false
      capabilities:
        drop: ["ALL"]

MCP Server Sidecar Pattern

When the agent requires MCP tool access, run the MCP server as a sidecar in the same pod (or use a shared cluster service):

containers:
- name: agent
  image: <agent-image>
  env:
  - name: MCP_ENDPOINT
    value: "http://localhost:8080/mcp"
- name: mcp-sidecar
  image: <mcp-server-image>
  ports:
  - containerPort: 8080

For shared MCP servers (tool registry, knowledge base access), deploy as cluster services in a shared namespace with NetworkPolicy allowing agent namespaces.

AKS MCP Server Deployment

For MCP servers that are NOT part of kars but serve AKS-hosted agents:

  • Transport: Use Streamable HTTP (MCP 2026-07-28). SSE is deprecated for new deployments.

  • Auth: Workload Identity + Entra ID. Do not use API keys in Kubernetes Secrets.

  • Placement: System node pool (cluster-wide tools) or dedicated MCP node pool.

  • Scaling: HPA based on HTTP request rate. Cold-start aware if using scale-to-zero.

  • Session state vs replicas: FastMCP-style streamable-HTTP servers default to stateful sessions, which break behind a multi-replica Service with no session affinity - requests land on replicas that do not own the session. Run stateless mode (FastMCP(..., stateless_http=True)) so any replica serves any request, or pin to one replica / configure session affinity.

  • Networking: Internal cluster IP only (ClusterIP service). Do not expose via LoadBalancer or Ingress unless external agents need access, and then only via authenticated gateway.

  • Probes (required for all MCP server containers). Every MCP server container must declare both readinessProbe and livenessProbe. MCP server pods that lack probes will receive traffic before they are ready (causing tool call failures at startup) and will not be restarted when they hang or deadlock (causing silent agent failures). Minimum probe configuration:

    readinessProbe:
      httpGet:
        path: /health
        port: 8080
      initialDelaySeconds: 5
      periodSeconds: 10
      failureThreshold: 3
    livenessProbe:
      httpGet:
        path: /health
        port: 8080
      initialDelaySeconds: 15
      periodSeconds: 20
      failureThreshold: 3

    If the MCP server does not expose an HTTP health endpoint, use a tcpSocket or exec probe - many streamable-HTTP MCP endpoints require protocol headers on every request, so a plain httpGet probe returns 4xx even when the server is healthy. Absence of probes is a blocking finding in the readiness review.

Security Boundaries

  • Agent → MCP server: Authenticate via Workload Identity token. MCP servers validate the azp or sub claim against known agent identities.
  • Agent → external APIs: Route through egress firewall. Default-deny, explicit FQDN allow rules.
  • Agent → storage: Use Workload Identity + RBAC on Storage/Key Vault. No access keys.
  • Agent → other agents: NetworkPolicy isolation by default. Explicit cross-namespace allow rules only when agent-to-agent communication is designed.
  • One ServiceAccount per agent component (non-negotiable). Every agent container (main agent, MCP server sidecar, tool executor) must have its own Kubernetes ServiceAccount with its own federated identity credential. A shared ServiceAccount across multiple agent components collapses blast radius: compromise of any component grants the combined Azure RBAC scope of all components. Scope each ServiceAccount to the narrowest Azure roles required by that specific component only.

Observability

  • Agent health: Liveness/readiness probes on the agent pod's health endpoint.
  • Tool call metrics: Track tool invocation count, latency, and error rate via OpenTelemetry from the agent SDK.
  • MCP metrics: Request rate, latency, error rate on the MCP endpoint.
  • Isolation health: Monitor Kata VM boot times; pods stuck in ContainerCreating may indicate VM provisioning failures.

Related Skills

Source: SKILL.md on GitHub

No alerts8d3 checks · Risk SAFE
  • Gen Agent Trust Hub8d

    The skill is a comprehensive architecture and configuration guide for Azure Kubernetes Service (AKS). it emphasizes security best practices, including RBAC, NetworkPolicy, workload identity, and kernel-level isolation for AI agents. No malicious patterns or security risks were detected.

  • Socket8d

    No alerts

  • Snyk8d

    Risk: LOW · No issues

Signed by skilld at 2cc2455. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last month.

Steadyupdated last month
metadata
{
  "last_verified": "2026-08-26"
}
Other metadata
argument-hint
workload=<type>; region=<azure-region>; availability=<SLO>; network=<hub-spoke|standalone>; scope=<new-cluster|production-review|fleet>

README badge

README badge for lukemurraynz/hve-agent-skills/aks-cluster-architecture