All skills
microsoft avatar

/airunway-aks-setup

@d0c21c0 official

Set up AI Runway on AKS — from bare cluster to running model. Covers cluster verification, controller install, GPU assessment, provider setup, and first deployment. WHEN: "setup AI Runway", "onboard AKS cluster", "install AI Runway", "airunway setup", "deploy model to AKS", "GPU inference on AKS", "KAITO setup on AKS", "run LLM on AKS", "vLLM on AKS", "set up model serving on AKS", "AI Runway controller".

Use this Skill: https://skilld.dev/gh/microsoft/github-copilot-for-azure/airunway-aks-setup

This session only. Nothing lands on disk.

referencesstepsstep-5-deploy.md

≈912 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Step 5 — First Model Deployment

Goal: Deploy a starter model that fits the cluster's GPU capacity and reach Ready status.

Model recommendations (based on cluster VRAM from Step 3 — see gpu-profiles.md for full sizing guide):

Cluster Capacity Recommended Model Provider Notes
CPU-only google/gemma-3-1b-it-qat-q8_0-gguf KAITO (llama.cpp) GGUF Q8 quantized; runs on CPU
1× T4 (16 GB) microsoft/Phi-3-mini-4k-instruct KAITO (vLLM) Small, fast; fits T4 with headroom
1× A10G or L4 (24 GB) meta-llama/Llama-3.1-8B-Instruct KAITO (vLLM) Good general-purpose starter (gated — needs HF token)
1× A100 40 GB microsoft/Phi-3-medium-128k-instruct KAITO (vLLM) ~28 GB float16; non-gated; MIT license
1× A100 80 GB / H100 meta-llama/Llama-3.1-8B-Instruct KAITO (vLLM) Oversized for 8B; suggest 70B with tensor parallelism if more GPUs available
4× A100 80 GB meta-llama/Llama-3.1-70B-Instruct KAITO (vLLM, tensor parallel) Large model; requires tensor parallelism; gated — needs HF token

Present recommendation. Ask user to confirm or choose their own model.

Naming Convention

  • model-name: A DNS-safe Kubernetes resource name derived from the model ID (e.g., meta-llama/Llama-3.1-8B-Instruct → llama-3-1-8b-instruct)
  • model-id: The full HuggingFace model identifier (e.g., meta-llama/Llama-3.1-8B-Instruct)

Gated Models (HuggingFace Token)

Llama-family models are gated on HuggingFace and require an access token. Read the token interactively to avoid exposing it in shell history:

# Read token without echo, store in a securely permissioned temp file
hf_token_file="$(mktemp)"
trap 'rm -f "$hf_token_file"' EXIT

read -r -s -p "Enter HuggingFace token: " hf_token
printf '\n'
printf '%s' "$hf_token" > "$hf_token_file"
chmod 600 "$hf_token_file"
unset hf_token

kubectl create secret generic hf-token \
  --from-file=token="$hf_token_file" \
  -n <namespace> \
  --dry-run=client -o yaml | kubectl apply -f -

See powershell-notes.md for the PowerShell equivalent.

Apply the ModelDeployment CR

Use the appropriate manifest based on whether the model is gated.

For gated models (Llama etc. — requires HuggingFace token):

kubectl apply -f - <<EOF
apiVersion: airunway.ai/v1alpha1
kind: ModelDeployment
metadata:
  name: <model-name>
  namespace: <namespace>
spec:
  model:
    id: <model-id>
    huggingFaceTokenSecretRef:
      name: hf-token
      key: token
  provider:
    name: <provider-name>
EOF

For non-gated models (Phi-3, Gemma etc. — no token required):

kubectl apply -f - <<EOF
apiVersion: airunway.ai/v1alpha1
kind: ModelDeployment
metadata:
  name: <model-name>
  namespace: <namespace>
spec:
  model:
    id: <model-id>
  provider:
    name: <provider-name>
EOF

See powershell-notes.md for the PowerShell equivalent.

Monitor Deployment

kubectl get modeldeployment <model-name> -n <namespace> -w

Wait until STATUS = Ready. If not ready within 10 minutes:

  • kubectl describe modeldeployment <model-name> -n <namespace> — check events
  • kubectl get pods -l modeldeployment=<model-name> -n <namespace> — check pod status

Note: Large models (13B+) may take significantly longer than 10 minutes because weights must be downloaded first. For 70B models, expect 20–40+ minutes depending on network speed. Check pod logs to confirm download progress rather than assuming failure.

Source: SKILL.md on GitHub

No alerts5mo3 checks · Risk SAFE
  • Gen Agent Trust Hub5mo

    This skill facilitates the setup of AI Runway on Azure Kubernetes Service (AKS) using standard administrative tools. It incorporates security best practices, particularly regarding the handling of sensitive HuggingFace tokens during deployment. No security considerations were identified.

  • Socket5mo

    No alerts

  • Snyk5mo

    Risk: LOW · No issues

Signed by skilld at d0c21c0. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 19 hours ago.

Activeupdated 5 months ago
metadata
{
  "author": "Microsoft",
  "version": "0.0.0-placeholder"
}
argument-hint
[skip-to-step N]
  • kubernetes
  • aks
  • azure
  • llm
  • model-serving
  • gpu
  • inference
  • kaito
  • vllm
  • mlops

README badge

README badge for microsoft/github-copilot-for-azure/airunway-aks-setup

Installs and configures AI Runway on an existing AKS cluster, walking through controller setup, GPU assessment, inference provider selection, and model deployment. Targets users deploying large language models or other GPU-accelerated workloads to Kubernetes via KAITO, Dynamo, or KubeRay.

Generated from the current SKILL.md.

Do I need an existing AKS cluster to use this skill?
Yes. This skill assumes you already have an AKS cluster. If you don't, the skill will hand off to the azure-kubernetes skill to provision one first, then return here.
What inference providers does this skill support?
The skill covers KAITO, Dynamo, and KubeRay as inference provider options. Step 4 recommends and installs the appropriate provider for your setup.
Does this skill work with CPU-only clusters?
Yes, CPU-only inference is acceptable, though the skill includes GPU assessment and provider setup primarily designed for GPU workloads. A bare cluster without GPU resources is still supported.
What happens if I already have part of AI Runway set up?
You can use the `skip-to-step N` argument to resume from a specific phase. The skill will report the current state and skip already-completed steps.
What are the cost implications of using this skill?
GPU node pools incur significant charges — A100-80GB can cost $3–5+/hr. The skill will confirm you understand these costs before provisioning GPU resources.

Generated from the current SKILL.md. These answers refresh after source changes.