Agent Runtime Bundle
Deploying AI agent runtimes on AKS: kars (Agent Reference Stack for Kubernetes, an open-source Microsoft reference stack, not an officially supported product), Kata Containers pod sandboxing, KAITO / AI Runway model serving, AKS MCP server patterns, and security isolation for untrusted agent code.
Load this bundle for: deploying agent workloads that execute untrusted code, choosing between isolation primitives, designing agent-specific node pools, and wiring MCP servers to AKS-hosted agents.
kars. Agent Reference Stack for Kubernetes
kars (Azure/kars) is an open-source, Kubernetes-native Agent Reference Stack from Microsoft's Azure Cloud Native team for running AI agents on AKS. It is a reference implementation, not an officially supported Microsoft product (no SLA or support contract). It treats every agent as untrusted code with defense-in-depth isolation:
| Isolation layer | Mechanism |
|---|---|
| No-network agent container | The agent runs as an unprivileged container (UID 1000) with no network access; all outbound traffic is forced through a co-located inference router |
| Rust inference router | A separate container (UID 1001) is the security enforcement core: it holds all credentials and brokers every outbound call, enforcing identity, content safety, egress policy, and audit |
| Zero credentials in agent process | The router exchanges a per-sandbox Microsoft Entra Agent ID (or cluster Workload Identity) for backend tokens via federated OIDC / IMDS; the agent process never sees a credential |
| L7 egress governance | Every outbound CONNECT is checked against a per-sandbox allowlist plus a blocklist (OISD + URLhaus); token budgets, rate limits, and EgressApproval CRDs apply |
| Encrypted inter-agent mesh | AgentMesh uses the Signal Protocol with a KNOCK trust handshake; the A2A Gateway handles cross-org peer-to-peer traffic |
| Belt-and-braces containment | Default-deny NetworkPolicy and an egress-guard init container (iptables) contain blast radius if the router is bypassed |
Microsoft Agent Framework (MAF) is one of kars's supported agent runtimes, and the router includes an MCP gateway that brokers calls to external MCP servers with OAuth and per-tool allowlists. Running kars node pools with AKS Pod Sandboxing (Kata Containers) optionally adds a per-pod guest kernel on top of the container-and-router model.
When to Use kars
| Scenario | Use kars? |
|---|---|
| Running third-party or user-supplied agent code | Yes — the no-network agent container plus the credential-holding router is the baseline |
| Hosting MCP server-backed agents on AKS | Yes — kars provides the hardened runtime; wire MCP servers as sidecars or cluster services |
| Multi-tenant agent platform (many agents per cluster) | Yes — namespace-per-agent with resource quotas |
| Simple single-agent deployment with trusted code | Consider ACA Express or Container Apps instead |
| Model inference/serving only (no agent tool execution) | Use AI Runway / KAITO directly |
Architecture
┌──────────────────────────────────────────────────┐
│ AKS Cluster (Automatic or Standard) │
│ │
│ ┌────────────────────────────────────────────┐ │
│ │ Sandbox pod (one agent per pod) │ │
│ │ ┌───────────────┐ ┌───────────────────┐ │ │
│ │ │ Agent │ │ Inference router │ │ │
│ │ │ container │─► │ (Rust) │─┼──► governed egress
│ │ │ UID 1000 │ │ UID 1001 │ │ │ (L7 allow/blocklist)
│ │ │ no network │ │ holds credentials │ │ │
│ │ └───────────────┘ └───────────────────┘ │ │
│ │ NetworkPolicy: default-deny │ │
│ │ egress-guard init container (iptables) │ │
│ └────────────────────────────────────────────┘ │
│ │
│ AgentMesh (Signal Protocol) · A2A Gateway │
│ MCP gateway (in router) ─► external MCP servers │
└──────────────────────────────────────────────────┘Deployment Checklist
Before deploying kars agents to production:
- Node pool: Dedicated agent node pool with
runtimeClassName: kata-vm-isolationor AKS Pod Sandboxing enabled. - VM SKU: Confidential VM (CVM) SKUs or v5+ family with nested virtualization support for Kata.
- Identity: Per-agent Workload Identity (
azure.workload.identity/client-idannotation on ServiceAccount). Zero static credentials. - NetworkPolicy: Default-deny per namespace; explicit allow rules for DNS, MCP endpoints, and approved egress.
- ResourceQuota: CPU, memory, and ephemeral storage limits per agent namespace.
- Pod Security:
restrictedPod Security Admission with explicit securityContext. - Observability: ContainerLogV2, Managed Prometheus, and agent-specific metrics.
- MCP transport: Streamable HTTP (preferred) or stdio via sidecar pattern. Do not expose MCP servers on public IPs.
- Lifecycle: Agent pods should be stateless; checkpoint state externally. Use
terminationGracePeriodSecondsfor clean shutdown.
Agent Isolation Primitives Comparison
| Primitive | Isolation level | Performance overhead | Use case |
|---|---|---|---|
| Kata MicroVM / Pod Sandboxing | Per-pod guest kernel | Moderate (VM boot per pod) | Untrusted code, multi-tenant agents |
| Confidential VM node pool | Hardware-level (SEV-SNP/TDX) | Low-moderate | Data-sensitive agents, regulated workloads |
| gVisor (runsc) | User-space kernel | Low | Not officially supported on AKS (GKE only); manual install carries the operational burden |
| Standard pod (no isolation) | None | None | Trusted first-party agents only |
kars's baseline isolation is the no-network agent container plus the inference router. To add a per-pod guest kernel on AKS, enable Pod Sandboxing (Kata Containers) with runtimeClassName: kata-vm-isolation, which is the AKS-native, officially supported (GA) way to get kernel isolation per agent. gVisor is not an AKS-supported runtime class, so treat that row as a general comparison, not an AKS option. Confidential Containers on AKS (kata-cc-isolation) is a separate, retiring preview, so do not build new work on it.
apiVersion: v1
kind: Pod
metadata:
name: agent-pod
namespace: agent-<id>
spec:
runtimeClassName: kata-vm-isolation
serviceAccountName: agent-sa
containers:
- name: agent
image: <agent-image>
resources:
requests:
cpu: "500m"
memory: "512Mi"
limits:
cpu: "1"
memory: "1Gi"
securityContext:
runAsNonRoot: true
seccompProfile:
type: RuntimeDefault
allowPrivilegeEscalation: false
capabilities:
drop: ["ALL"]MCP Server Sidecar Pattern
When the agent requires MCP tool access, run the MCP server as a sidecar in the same pod (or use a shared cluster service):
containers:
- name: agent
image: <agent-image>
env:
- name: MCP_ENDPOINT
value: "http://localhost:8080/mcp"
- name: mcp-sidecar
image: <mcp-server-image>
ports:
- containerPort: 8080For shared MCP servers (tool registry, knowledge base access), deploy as cluster services in a shared namespace with NetworkPolicy allowing agent namespaces.
AKS MCP Server Deployment
For MCP servers that are NOT part of kars but serve AKS-hosted agents:
Transport: Use Streamable HTTP (MCP 2026-07-28). SSE is deprecated for new deployments.
Auth: Workload Identity + Entra ID. Do not use API keys in Kubernetes Secrets.
Placement: System node pool (cluster-wide tools) or dedicated MCP node pool.
Scaling: HPA based on HTTP request rate. Cold-start aware if using scale-to-zero.
Session state vs replicas: FastMCP-style streamable-HTTP servers default to stateful sessions, which break behind a multi-replica Service with no session affinity - requests land on replicas that do not own the session. Run stateless mode (
FastMCP(..., stateless_http=True)) so any replica serves any request, or pin to one replica / configure session affinity.Networking: Internal cluster IP only (ClusterIP service). Do not expose via LoadBalancer or Ingress unless external agents need access, and then only via authenticated gateway.
Probes (required for all MCP server containers). Every MCP server container must declare both
readinessProbeandlivenessProbe. MCP server pods that lack probes will receive traffic before they are ready (causing tool call failures at startup) and will not be restarted when they hang or deadlock (causing silent agent failures). Minimum probe configuration:readinessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 5 periodSeconds: 10 failureThreshold: 3 livenessProbe: httpGet: path: /health port: 8080 initialDelaySeconds: 15 periodSeconds: 20 failureThreshold: 3If the MCP server does not expose an HTTP health endpoint, use a
tcpSocketorexecprobe - many streamable-HTTP MCP endpoints require protocol headers on every request, so a plainhttpGetprobe returns 4xx even when the server is healthy. Absence of probes is a blocking finding in the readiness review.
Security Boundaries
- Agent → MCP server: Authenticate via Workload Identity token. MCP servers validate the
azporsubclaim against known agent identities. - Agent → external APIs: Route through egress firewall. Default-deny, explicit FQDN allow rules.
- Agent → storage: Use Workload Identity + RBAC on Storage/Key Vault. No access keys.
- Agent → other agents: NetworkPolicy isolation by default. Explicit cross-namespace allow rules only when agent-to-agent communication is designed.
- One ServiceAccount per agent component (non-negotiable). Every agent container (main agent, MCP server sidecar, tool executor) must have its own Kubernetes ServiceAccount with its own federated identity credential. A shared ServiceAccount across multiple agent components collapses blast radius: compromise of any component grants the combined Azure RBAC scope of all components. Scope each ServiceAccount to the narrowest Azure roles required by that specific component only.
Observability
- Agent health: Liveness/readiness probes on the agent pod's health endpoint.
- Tool call metrics: Track tool invocation count, latency, and error rate via OpenTelemetry from the agent SDK.
- MCP metrics: Request rate, latency, error rate on the MCP endpoint.
- Isolation health: Monitor Kata VM boot times; pods stuck in
ContainerCreatingmay indicate VM provisioning failures.
Related Skills
- microsoft-agent-framework — Building agents with MAF, hosted agents, and Foundry integration.
- mcp-server-design — MCP server architecture, auth, and production hardening.
- azure-container-apps — ACA Express for simpler agent hosting (trusted code, lower isolation requirements).