All skills
nvidia avatar

/physical-ai-video-data-augmentation

@0482ebc
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when running video data augmentation and auto-labeling workflows on OSMO: flow selection, preflight, submit-time interpolation, monitoring, and output retrieval. Trigger keywords: video data augmentation, data enrichment, auto labeling, VDA demo, OSMO workflow, pseudo labeling.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-video-data-augmentation

This session only. Nothing lands on disk.

referencesflowse2e_super_resolution.md

≈631 tokens on demand. Your agent reads this file only when SKILL.md points to it.

E2E (Super-Resolution Gated)

Table of Contents

Runs full VDA graph in sequential order where original auto-labeling (with SR enabled in setup .env) gates augmentation, then labels augmented outputs.

When to use

  • User requests SR-gated end-to-end execution.
  • User prefers deterministic sequencing over parallel throughput.
  • Hardware/resources favor sequential stage progression.

Graph

setup_group
  setup (SUPER_RESOLUTION_ENABLED=true)
      ▼
auto_labeling_original_group
  pl_original_worker_0
      ▼
augmentation_group
  cosmos_worker_0
      ▼
auto_labeling_augmented_group
  pl_augmented_worker_0

Inputs

Input Source Required by
Source video <storage_url>/datasets/<dataset>/<video>.mp4 setup + original AL + augmentation
Cosmos cache <storage_url>/data/models/cosmos_transfer cosmos_worker_0
Auto-labeling cache <storage_url>/data/models/auto_labeling pl_original_worker_0, pl_augmented_worker_0
VLM endpoint default in-cluster NIM or explicit override all workers
LLM endpoint default in-cluster NIM or explicit override all workers

Submit

Use the shared submit command and the common optional-overrides block in the SKILL.md "Submit (all flows)" section, with workflow YAML assets/configs/osmo/e2e_super_resolution.yaml. This flow runs the SR-gated sequential pipeline, so the full canonical single-flag submit shape applies (including both cache URL values).

Output layout

<storage_url>/datasets/<dataset>-outputs/<run_id>/
├─ setup_b0/
└─ outputs/
   ├─ pseudo_labeled/<video>/
   ├─ augmented/<video>_aug0/
   └─ pseudo_labeled_augmented/<video>_aug0/

Troubleshooting

  • Completion evidence requirement: include a side-by-side input vs augmented video artifact, augmentation summary (setup_b0/configs/manifest.yaml sampled_vars for <video>_aug0), and auto-labeling artifact summaries for both outputs/pseudo_labeled/<video> and outputs/pseudo_labeled_augmented/<video>_aug0 using the workspace-local run copy under media/vda/runs/<run_id>/.
  • If this mode appears parallel, confirm you submitted e2e_super_resolution.yaml (not e2e.yaml).
  • If augmentation starts before original AL completion, inspect group/task dependencies in rendered workflow.

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    The skill is a workflow orchestrator for video data augmentation and auto-labeling on the NVIDIA OSMO platform. It manages the end-to-end pipeline, including credential verification, configuration generation, worker execution, and result retrieval. No security issues or malicious patterns were detected; all external dependencies and network operations are associated with trusted vendors and the skill's primary functionality.

  • Socket3mo

    3 alerts: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 0482ebc. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
Other metadata
metadata
{
  "owner": "NVIDIA",
  "service": "data",
  "version": "1.0.0",
  "reviewed": "2026-05-26",
  "author": "NVIDIA",
  "tags": [
    "physical-ai",
    "video-data-augmentation",
    "auto-labeling",
    "cosmos"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-video-data-augmentation