All skills
nvidia avatar

/physical-ai-infrastructure-setup-and-resilient-scaling

@6c6fc09
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling

This session only. Nothing lands on disk.

componentsosmo-clireferencesadvanced-patterns.md

≈622 tokens on demand. Your agent reads this file only when SKILL.md points to it.

OSMO Advanced Patterns Reference

Read this file only when the user's request clearly requires one of these specific capabilities: checkpointing, exit/retry behavior, or node exclusion. These are niche patterns not needed for most workflow generation tasks.


Checkpointing

Automatically upload a task's working directory to S3 at a fixed interval while the task runs. Useful for long-running training jobs where you want to preserve intermediate results if the job is interrupted.

tasks:
- name: train-with-checkpointing
  image: ubuntu:24.04
  command: [/bin/bash]
  args: [/tmp/run.sh]
  files:
  - path: /tmp/run.sh
    contents: |-
      #!/bin/bash
      set -ex
      mkdir -p /tmp/checkpoints
      python train.py --output /tmp/checkpoints
  checkpoint:
  - path: /tmp/checkpoints           # local directory to upload
    url: s3://my-bucket/checkpoints  # destination
    frequency: 60s                   # how often to sync

A final checkpoint is always uploaded when the task completes, regardless of interval.

Checkpoint only specific files

Use a regex to filter which files get uploaded:

checkpoint:
- path: /tmp/checkpoints
  url: s3://my-bucket/checkpoints
  frequency: 60s
  regex: .*\.(bin|pt)$   # only upload .bin and .pt files

Error Handling with Exit Actions

Control what happens when a task exits with a specific exit code. Useful for automatic retry logic.

tasks:
- name: resilient-task
  image: ubuntu:24.04
  command: ["bash", "-c", "python fetch_and_process.py"]
  exitActions:
    COMPLETE: 0       # exit code 0 → task completes normally
    RESCHEDULE: 1-255 # any non-zero exit → task is rescheduled (retried)

Available actions: COMPLETE, RESCHEDULE, FAIL. Ranges and comma-separated lists of exit codes are supported (e.g. 1,2,5 or 1-10).


Excluding Specific Nodes

Prevent a workflow from scheduling on known-problematic nodes using nodesExcluded in the resource spec:

workflow:
  name: exclude-nodes-demo
  resources:
    default:
      cpu: 4
      memory: 16Gi
      storage: 50Gi
      nodesExcluded:
      - worker-node-01
      - worker-node-02
  tasks:
  - name: my-task
    image: ubuntu:24.04
    command: ["bash", "-c", "echo Running on a healthy node"]

Warning: Excluding too many nodes can cause tasks to remain PENDING indefinitely. Only use this when specific nodes are confirmed to have hardware or network issues.

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill from NVIDIA provides a comprehensive set of tools and instructions for setting up Physical AI infrastructure on Azure AKS or local MicroK8s clusters. The analysis found no malicious patterns; all external downloads are from trusted or well-known services (NVIDIA, Microsoft, HashiCorp, Hugging Face), and the orchestration design prioritizes user visibility through a partitioned multi-agent architecture.

  • Socket3mo

    1 alert: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 6c6fc09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
version
1.0.0
tools
[
  "Read",
  "Shell"
]
Other metadata
compatibility
Requires the selected component prerequisites, usually kubectl plus either MicroK8s or Azure CLI/Terraform, and OSMO or inference credentials for the chosen target.
metadata
{
  "author": "NVIDIA Physical AI",
  "tags": [
    "physical-ai",
    "infrastructure",
    "kubernetes",
    "azure",
    "microk8s",
    "osmo",
    "nim-operator",
    "scaling"
  ],
  "domain": "ai-ml",
  "languages": [
    "bash",
    "hcl",
    "yaml"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling