All skills
nvidia avatar

/physical-ai-infrastructure-setup-and-resilient-scaling

@6c6fc09
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when the user wants to set up, scale, validate, or harden NVIDIA physical AI infrastructure for synthetic data generation workflows across local MicroK8s or Azure AKS, including Kubernetes clusters, inference endpoint deployment, OSMO deployment, workload submission readiness, and infrastructure failure recovery. Trigger keywords: physical ai infrastructure, resilient scaling, SDG infrastructure, microk8s, azure aks, NVCF deployment, NIM Operator, OSMO deploy, workflow scaling. Don't trigger for: OSMO log summarization or workload-only operations unless infrastructure setup, scaling, validation, or recovery is requested.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling

This session only. Nothing lands on disk.

componentsosmo-clireferencesworkflow-spec.md

≈2.2k tokens on demand. Your agent reads this file only when SKILL.md points to it.

OSMO Workflow Spec Reference

Complete schema for OSMO workflow YAML files.

Table of Contents


Top-Level Structure

version: 2              # optional; must be 2 if present (default)

workflow:
  name: <string>        # workflow name
  pool: <string>        # target pool (usually set via --pool flag instead)
  resources: ...        # named resource profiles
  tasks: [...]          # flat task list (mutually exclusive with groups)
  groups: [...]         # grouped tasks (mutually exclusive with tasks)
  timeout:
    exec_timeout: <duration>   # max execution time
    queue_timeout: <duration>  # max queue wait time

default-values:         # Jinja template defaults (top-level, outside workflow:)
  var_name: value

Rule: Exactly one of tasks: or groups: must be present — never both.


Resources

Named resource profiles referenced by tasks via the resource: field.

resources:
  default:              # every workflow has an implicit "default" profile
    cpu: 8
    gpu: 2
    memory: 32Gi        # must use binary units: Gi, Mi
    storage: 100Gi      # must use binary units: Gi, Mi
    platform: dgx-h100  # target hardware platform
    nodesExcluded:       # exclude specific nodes
    - bad-node-01
    topology:            # advanced placement constraints
    - key: <string>
      group: <string>
      requirementType: <string>

  gpu_heavy:            # custom named profile
    cpu: 16
    gpu: 8
    memory: 128Gi
    storage: 200Gi

Tasks use resource: gpu_heavy to select a profile. Default is "default".


Task Spec (TaskSpec)

Each task defines a container to run.

tasks:
- name: <string>                   # unique task name
  image: <string>                  # container image
  command: [<string>, ...]         # entrypoint (required, non-empty)
  args: [<string>, ...]            # additional arguments
  resource: <string>               # name of resource profile (default: "default")
  lead: <bool>                     # required in multi-task groups (one per group)

  # Data I/O
  inputs: [...]                    # data inputs (see below)
  outputs: [...]                   # data outputs (see below)

  # Configuration
  environment:                     # environment variables
    KEY: "value"
  files:                           # inline files created in the container
  - path: /tmp/script.sh
    contents: |
      #!/bin/bash
      echo "Hello"
  - path: /tmp/data.bin
    contents: <base64_string>
    base64: true

  # Credentials
  credentials:
    my_credential: /mnt/creds      # mount credential at path
    my_secret:                      # or map env vars to secret keys
      ENV_VAR: secret_key

  # Advanced
  privileged: <bool>
  hostNetwork: <bool>
  volumeMounts: [...]              # host volume mounts
  downloadType: <string>           # download behavior
  cacheSize: <string>              # cache size
  backend: <string>                # per-task backend override

  # Checkpointing
  checkpoint:
  - path: /tmp/checkpoints
    url: s3://bucket/checkpoints
    frequency: 60s
    regex: '.*\.(pt|bin)$'         # optional filter

  # Error handling
  exitActions:
    COMPLETE: 0                    # exit 0 = success
    RESCHEDULE: 1-255              # non-zero = retry
    FAIL: 137                      # specific code = fail

  # Monitoring
  kpis:
    index: <int>
    path: <string>

Inputs

Tasks can receive data from three sources:

Task-to-task (data dependency)

inputs:
- task: upstream_task_name
  regex: '.*\.csv$'        # optional: filter files

Creates a DAG dependency. Upstream task must complete before this task starts. Access via {{input:N}} (0-indexed by position in the inputs list) or {{input:upstream_task_name}}.

URL (cloud storage)

inputs:
- url: s3://bucket/data/
  regex: '.*\.png$'        # optional filter

Downloads from S3, GCS, Azure, Swift, or TOS at task start.

Dataset (OSMO-managed)

inputs:
- dataset:
    name: my_dataset
    path: /custom/mount    # optional
    regex: '.*'            # optional

Outputs

Dataset output

outputs:
- dataset:
    name: my_output_dataset
    # optional: metadata, labels

The task writes to {{output}}, which OSMO uploads as a dataset on completion. Fetch it with osmo dataset download <name> <local-path>.

URL output

outputs:
- url: s3://bucket/output/

Fetch URL outputs with osmo data list --no-pager <url> and osmo data download <url> <local-path>.

Update existing dataset

outputs:
- update_dataset:
    name: existing_dataset

Groups

Use groups when tasks need co-scheduling or network communication.

groups:
- name: training_group
  barrier: true              # default true; wait for all tasks in group
  ignoreNonleadStatus: true  # default true; group status follows lead only
  tasks:
  - name: coordinator
    lead: true               # exactly one lead per multi-task group
    image: my-image
    command: ["python", "coord.py"]
  - name: worker
    image: my-image
    command: ["python", "worker.py", "--coord={{host:coordinator}}"]

Group rules

  • Exactly one task must have lead: true in multi-task groups.
  • The group terminates when the lead task exits.
  • {{host:taskname}} resolves to the DNS name of a task in the same group.
  • {{host:taskname}} does NOT work across groups.

Cross-group dependencies

Groups depend on each other through task-level inputs::

groups:
- name: stage1
  tasks:
  - name: produce
    lead: true
    command: ["bash", "-c", "echo data > {{output}}/out.txt"]

- name: stage2
  tasks:
  - name: consume
    lead: true
    command: ["bash", "-c", "cat {{input:0}}/out.txt"]
    inputs:
    - task: produce     # stage2 waits for ALL of stage1 to complete

Special Tokens

Automatically set by OSMO — cannot be overridden with --set:

Token Resolves To
{{output}} Output directory path for this task
{{input:N}} Nth input path (0-indexed by inputs: list order)
{{input:taskname}} Input path by upstream task name
{{host:taskname}} DNS name of a task in the same group
{{workflow_id}} Unique workflow run ID

These tokens work in command, args, environment values, and files contents.


Jinja Templates

Make workflows configurable at submit time:

workflow:
  name: "{{workflow_name}}"
  resources:
    default:
      gpu: {{gpu_count}}
  tasks:
  - name: train
    image: "{{image}}"
    command: ["python", "train.py"]
    args: ["--epochs={{epochs}}"]

default-values:
  workflow_name: my-training
  gpu_count: 1
  image: nvcr.io/nvidia/pytorch:24.01-py3
  epochs: 10
# Submit with defaults
osmo workflow submit template.yaml --pool my-pool

# Override values
osmo workflow submit template.yaml --pool my-pool --set gpu_count=4 epochs=50

Jinja supports loops and conditionals:

tasks:
{% for i in range(num_workers) %}
- name: worker-{{i}}
  image: my-image
  command: ["python", "worker.py", "--id={{i}}"]
{% endfor %}

Cookbook Examples

Real-world examples in the OSMO repo under cookbook/:

Example Pattern
tutorials/hello_world.yaml Minimal single task
tutorials/template_hello_world.yaml Jinja template with defaults
tutorials/serial_workflow.yaml Serial task chain
tutorials/parallel_tasks.yaml Independent parallel tasks
tutorials/group_tasks.yaml Synchronized group
tutorials/group_tasks_communication.yaml {{host:...}} inter-task networking
tutorials/combination_workflow_complex.yaml Multi-group pipeline with data flow
tutorials/data_download.yaml S3 URL input
tutorials/dataset_upload.yaml Dataset output with {{output}}
tutorials/resources_platforms.yaml Multiple resource profiles + platforms
tutorials/resources_basic.yaml Basic resource configuration
dnn_training/torchrun_multinode/train.yaml Multi-node distributed training
reinforcement_learning/single_gpu/train_policy.yaml RL training with GPU
integration_and_tools/jupyterlab/jupyter.yaml Interactive Jupyter session
integration_and_tools/vscode/vscode.yaml VS Code remote session

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    This skill from NVIDIA provides a comprehensive set of tools and instructions for setting up Physical AI infrastructure on Azure AKS or local MicroK8s clusters. The analysis found no malicious patterns; all external downloads are from trusted or well-known services (NVIDIA, Microsoft, HashiCorp, Hugging Face), and the orchestration design prioritizes user visibility through a partitioned multi-agent architecture.

  • Socket3mo

    1 alert: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 6c6fc09. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
version
1.0.0
tools
[
  "Read",
  "Shell"
]
Other metadata
compatibility
Requires the selected component prerequisites, usually kubectl plus either MicroK8s or Azure CLI/Terraform, and OSMO or inference credentials for the chosen target.
metadata
{
  "author": "NVIDIA Physical AI",
  "tags": [
    "physical-ai",
    "infrastructure",
    "kubernetes",
    "azure",
    "microk8s",
    "osmo",
    "nim-operator",
    "scaling"
  ],
  "domain": "ai-ml",
  "languages": [
    "bash",
    "hcl",
    "yaml"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling