All skills
nvidia avatar

/physical-ai-infrastructure-setup-and-resilient-scaling

@6c6fc09
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling

This session only. Nothing lands on disk.

componentsinference-azurereference.md

≈657 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Azure AI Foundry Inference

Docs: https://learn.microsoft.com/en-us/azure/ai-foundry/how-to/deploy-models-serverless

Prerequisites

Requirement Details
Azure CLI + ML extension Complete components/azure-access/reference.md first, then az extension add -n ml.
Foundry resource + project Provisioned by the Azure cluster component; consumed during install, not preflight.

Supporting files

Path Use When
scripts/preflight.sh Run first Checks Azure subscription/provider read access, CLI, jq, ML extension, and local Terraform root; Foundry outputs are install-time inputs.
scripts/install.sh Run Deploys or lists Azure AI Foundry serverless endpoints.

Capability catalog

Agent picks the endpoint name + Model ID when invoking install.sh:

Endpoint name Model ID Capabilities
llama-3-1-8b azureml://registries/azureml-meta/models/Meta-Llama-3.1-8B-Instruct text-llm, chat
llama-3-1-70b azureml://registries/azureml-meta/models/Meta-Llama-3.1-70B-Instruct text-llm, chat
phi-3-5-vision azureml://registries/azureml/models/Phi-3.5-vision-instruct vlm, image-qa, chat
deepseek-r1 azureml://registries/azureml-deepseek/models/DeepSeek-R1 text-llm, reasoning

Pattern for any Foundry-supported model: azureml://registries/<registry>/models/<model>.

Foundry has no video-generation / video-style-transfer — combine with NVCF or NIM Operator for video; root SKILL must reject unsatisfiable combos before submitting.

Install

skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/inference-azure/scripts/install.sh                         # deploy one endpoint (default llama-3-1-8b)
skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/inference-azure/scripts/install.sh -n <name> -m <model-id> # deploy a specific model from the catalog above
skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/inference-azure/scripts/install.sh --list                  # list deployed endpoints

Pipeline needs multiple endpoints → invoke install.sh per name. Reads RG + project from skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/cluster-azure/scripts TF outputs. --help for full flags.

Operations

az ml serverless-endpoint get-credentials -n <name>  # fetch URL + key for pipeline's *_URL env
az ml serverless-endpoint delete -n <name> --yes     # tear down one endpoint

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill from NVIDIA provides a comprehensive set of tools and instructions for setting up Physical AI infrastructure on Azure AKS or local MicroK8s clusters. The analysis found no malicious patterns; all external downloads are from trusted or well-known services (NVIDIA, Microsoft, HashiCorp, Hugging Face), and the orchestration design prioritizes user visibility through a partitioned multi-agent architecture.

  • Socket3mo

    1 alert: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 6c6fc09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
version
1.0.0
tools
[
  "Read",
  "Shell"
]
Other metadata
compatibility
Requires the selected component prerequisites, usually kubectl plus either MicroK8s or Azure CLI/Terraform, and OSMO or inference credentials for the chosen target.
metadata
{
  "author": "NVIDIA Physical AI",
  "tags": [
    "physical-ai",
    "infrastructure",
    "kubernetes",
    "azure",
    "microk8s",
    "osmo",
    "nim-operator",
    "scaling"
  ],
  "domain": "ai-ml",
  "languages": [
    "bash",
    "hcl",
    "yaml"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling