All skills
nvidia avatar

/physical-ai-infrastructure-setup-and-resilient-scaling

@6c6fc09
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling

This session only. Nothing lands on disk.

componentscluster-microk8sreference.md

≈452 tokens on demand. Your agent reads this file only when SKILL.md points to it.

MicroK8s Cluster

Prerequisites

  • Running on machine with GPU available and NVIDIA driver >= 525 installed
  • snapd >= 2.45.0
  • Ports 16443, 10250, 10255 available
  • 20 GB disk space
  • git >= 2.25.0 (for shallow clone of https://github.com/nvidia/osmo)

Deployment

Run as root (sudo) from repo root.

  1. Run preflight
REPO=$(git rev-parse --show-toplevel)
"$REPO/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/cluster-microk8s/scripts/preflight.sh"
  1. Clone https://github.com/nvidia/osmo - use main unless otherwise specified
OSMO_REF="${OSMO_REF:-main}"
OSMO_DIR="$HOME/.cache/physical-ai/osmo"
if [ -d "$OSMO_DIR/.git" ]; then
  git -C "$OSMO_DIR" fetch --depth 1 origin "$OSMO_REF"
  git -C "$OSMO_DIR" reset --hard FETCH_HEAD
else
  mkdir -p "$(dirname "$OSMO_DIR")"
  git clone --depth 1 --branch "$OSMO_REF" \
    https://github.com/NVIDIA/OSMO.git "$OSMO_DIR"
fi
  1. Run Microk8s bootstrap
sudo "$OSMO_DIR/deployments/scripts/microk8s/install.sh" --gpu

Verify

Check general Kubernetes state. Pods should be healthy and running.

kubectl get pods -A

Check GPUs are available and allocatable under nvidia.com/gpu.

kubectl describe node <node-name>

Ensure runtime class is marked as nvidia.

kubectl get runtimeclass nvidia -o jsonpath='{.handler}'

Troubleshooting

Symptom Fix
snap: command not found sudo apt-get install snapd
Node NotReady after install sudo microk8s status --wait-ready
GPU not visible nvidia-smi; verify driver ≥ 525
kubeconfig permission denied sudo chown $USER:$USER ~/.kube/config
Existing microk8s install in degraded state sudo snap remove microk8s --purge then re-run from step 2

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill from NVIDIA provides a comprehensive set of tools and instructions for setting up Physical AI infrastructure on Azure AKS or local MicroK8s clusters. The analysis found no malicious patterns; all external downloads are from trusted or well-known services (NVIDIA, Microsoft, HashiCorp, Hugging Face), and the orchestration design prioritizes user visibility through a partitioned multi-agent architecture.

  • Socket3mo

    1 alert: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 6c6fc09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
version
1.0.0
tools
[
  "Read",
  "Shell"
]
Other metadata
compatibility
Requires the selected component prerequisites, usually kubectl plus either MicroK8s or Azure CLI/Terraform, and OSMO or inference credentials for the chosen target.
metadata
{
  "author": "NVIDIA Physical AI",
  "tags": [
    "physical-ai",
    "infrastructure",
    "kubernetes",
    "azure",
    "microk8s",
    "osmo",
    "nim-operator",
    "scaling"
  ],
  "domain": "ai-ml",
  "languages": [
    "bash",
    "hcl",
    "yaml"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling