All skills
nvidia avatar

/physical-ai-infrastructure-setup-and-resilient-scaling

@6c6fc09
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling

This session only. Nothing lands on disk.

componentsazure-accessreference.md

≈775 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Azure Access

Use this before any selected Azure component preflight. It handles identity, PIM/RBAC, subscription, and region selection; Terraform outputs and cluster state are not expected yet.

Inputs

Get these from the user or org context before running Azure preflights:

Input Why
Tenant ID/domain Needed when az login lands in the wrong tenant.
Subscription ID/name All Azure components must target the same subscription.
Region (location) Drives quota, SKUs, and deploy.tfvars.
Caller CIDR (allowed_cidr) Required for Azure cluster deploy.tfvars; derive public IP/32 when possible.
PIM role Azure cluster deploy usually needs subscription Owner, or Contributor plus User Access Administrator.

If the user is unsure about PIM, tell them to open Azure Portal → PIM and activate the eligible role for the target subscription before continuing.

Login

Use browser login when available:

az login

For OpenClaw or any shell that cannot open a browser, use components/openclaw-azure-login/reference.md and keep the device-code command running until the user completes auth.

Subscription

List visible subscriptions and make the target explicit:

az account list --refresh \
  --query "[].{name:name,id:id,state:state,tenantId:tenantId,isDefault:isDefault}" \
  -o table
az account set --subscription <subscription-id-or-name>
az account show --query "{name:name,id:id,state:state,tenantId:tenantId,user:user.name}" -o table

Stop if the target subscription is missing or not Enabled. Ask the user to switch account/tenant or activate/access the subscription; do not infer another subscription.

Role And Provider Access

Run the selected Azure component preflight after login, subscription selection, and PIM activation. It checks subscription read access and required provider reads using az provider show.

If provider read fails, tell the user:

Activate the Azure PIM role for the target subscription, wait a minute for RBAC
propagation, then rerun the same preflight.

If a provider is readable but not Registered, ask for permission to register it before Terraform:

az provider register --namespace <provider> --subscription <subscription-id> --wait

Provider registration is a subscription mutation; do not run it as part of preflight.

Region

Choose the region before quota checks or Terraform. Confirm Azure exposes it:

az account list-locations \
  --query "[].{name:name,displayName:displayName}" \
  -o table

For Azure cluster deploys, write the selected subscription_id, location, and allowed_cidr to components/cluster-azure/scripts/deploy.tfvars, then rerun components/cluster-azure/scripts/preflight.sh. The preflight validates local tfvars when present.

Quota

After region selection, check quota for the planned SKUs before Terraform:

az vm list-usage -l <location> -o table

If CPU/GPU quota is insufficient, stop and ask for a quota increase or a different region/SKU before provisioning.

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill from NVIDIA provides a comprehensive set of tools and instructions for setting up Physical AI infrastructure on Azure AKS or local MicroK8s clusters. The analysis found no malicious patterns; all external downloads are from trusted or well-known services (NVIDIA, Microsoft, HashiCorp, Hugging Face), and the orchestration design prioritizes user visibility through a partitioned multi-agent architecture.

  • Socket3mo

    1 alert: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 6c6fc09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
version
1.0.0
tools
[
  "Read",
  "Shell"
]
Other metadata
compatibility
Requires the selected component prerequisites, usually kubectl plus either MicroK8s or Azure CLI/Terraform, and OSMO or inference credentials for the chosen target.
metadata
{
  "author": "NVIDIA Physical AI",
  "tags": [
    "physical-ai",
    "infrastructure",
    "kubernetes",
    "azure",
    "microk8s",
    "osmo",
    "nim-operator",
    "scaling"
  ],
  "domain": "ai-ml",
  "languages": [
    "bash",
    "hcl",
    "yaml"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling