All skills
nvidia avatar

/physical-ai-infrastructure-setup-and-resilient-scaling

@6c6fc09
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling

This session only. Nothing lands on disk.

componentsinference-nvcfreference.md

≈816 tokens on demand. Your agent reads this file only when SKILL.md points to it.

NVCF Inference

Source docs: https://docs.nvidia.com/cloud-functions/

Capability catalog

Match pipeline needs against this table and export the *_URL env vars in Step 4.

Function Env var Capabilities Notes
Nurec (Llama 3.1 8B Instruct) LLAMA31_8B_URL / function ID text-llm, chat OpenAI-compat /v1/chat/completions
Cosmo (Cosmos Predict 1) COSMOS_PREDICT1_URL / function ID video-world-model, video-generation Used by Augmentation
Cosmos Transfer 2.5 COSMOS_TRANSFER25_URL video-style-transfer Optional
Qwen2.5 14B QWEN25_14B_URL text-llm, chat Alternative text LLM
Qwen3 VL 30B QWEN3_VL_30B_URL vlm, video-qa, chat Alternative VLM

Prerequisites

Requirement Details
Brev org Access to a Brev org whose linked NGC Org has Nurec/Cosmo provisioned
NGC Org API key One key per Brev org — obtain from NGC portal
curl For endpoint validation

Supporting files

Path Use When
scripts/preflight.sh Run first Checks local tools, loads repo .env, validates NGC_API_KEY, and probes NVCF unless network checks are skipped.

Setup

Brev org → NGC org is 1:1. One Org API key authenticates NVCF, nvcr.io pulls, and the NGC CLI.

  1. Get the Org-level NGC API key. If NGC_API_KEY is not in .env, prompt the user with these links (do NOT hardcode org IDs):

  2. Persist in repo-root .env:

    NGC_API_KEY=<your-ngc-org-api-key>
  3. Verify:

    curl -s -o /dev/null -w "%{http_code}" \
      -H "Authorization: Bearer ${NGC_API_KEY}" \
      https://api.nvcf.nvidia.com/v2/nvcf/functions

    Expect 200.

  4. List deployed functions and pick the ones the pipeline needs:

    curl -s -H "Authorization: Bearer ${NGC_API_KEY}" \
      https://api.nvcf.nvidia.com/v2/nvcf/functions \
      | jq '.functions[] | {name, id, status}'
  5. Export the matching *_URL env vars (from the catalog above) to repo-root .env. The URL format is https://api.nvcf.nvidia.com/v2/nvc/functions/<function-id>/versions/<version-id> — copy from the NVCF portal or derive from the list output.

Rotate the key

Replace NGC_API_KEY in .env and re-run every consuming stage's install script so pull secrets get recreated in each namespace.

Troubleshooting

Symptom Likely cause Fix
401 Unauthorized Wrong key type (personal vs org) Use the Org-level API key, not a personal token
403 Forbidden Key not associated with correct NGC org Verify org in NGC portal matches Brev org
Functions list empty No functions deployed to org Ask the NVCF function owner for your NGC org to provision Nurec/Cosmo
Image pull fails from nvcr.io nvcr-pull-secret missing/stale Re-run Step 2 to recreate the secret

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill from NVIDIA provides a comprehensive set of tools and instructions for setting up Physical AI infrastructure on Azure AKS or local MicroK8s clusters. The analysis found no malicious patterns; all external downloads are from trusted or well-known services (NVIDIA, Microsoft, HashiCorp, Hugging Face), and the orchestration design prioritizes user visibility through a partitioned multi-agent architecture.

  • Socket3mo

    1 alert: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 6c6fc09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
version
1.0.0
tools
[
  "Read",
  "Shell"
]
Other metadata
compatibility
Requires the selected component prerequisites, usually kubectl plus either MicroK8s or Azure CLI/Terraform, and OSMO or inference credentials for the chosen target.
metadata
{
  "author": "NVIDIA Physical AI",
  "tags": [
    "physical-ai",
    "infrastructure",
    "kubernetes",
    "azure",
    "microk8s",
    "osmo",
    "nim-operator",
    "scaling"
  ],
  "domain": "ai-ml",
  "languages": [
    "bash",
    "hcl",
    "yaml"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling