All skills
nvidia avatar

/physical-ai-video-data-augmentation

@0482ebc
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when running video data augmentation and auto-labeling workflows on OSMO: flow selection, preflight, submit-time interpolation, monitoring, and output retrieval. Trigger keywords: video data augmentation, data enrichment, auto labeling, VDA demo, OSMO workflow, pseudo labeling.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-video-data-augmentation

This session only. Nothing lands on disk.

referencesflowsaugmentation_and_al.md

≈638 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Augmentation + Auto-Labeling

Table of Contents

Runs augmentation on source video(s) and then auto-labels augmented outputs. Original-video auto-labeling is not part of this flow.

When to use

  • User wants synthetic variants plus labels for those variants.
  • User does not need labels on source/original video path in the same run.
  • User wants a smaller graph than e2e while keeping augmentation + AL.

Graph

setup_group
  setup
    -> stages scripts + cookbook configs + generated per-video configs + .env
      ▼
augmentation_group
  cosmos_worker_0
    -> augmented outputs
      ▼
auto_labeling_augmented_group
  pl_augmented_worker_0
    -> pseudo-labeled augmented outputs

Inputs

Input Source Required by
Source video <storage_url>/datasets/<dataset>/<video>.mp4 setup, cosmos_worker_0
Cosmos cache <storage_url>/data/models/cosmos_transfer cosmos_worker_0
Auto-labeling cache <storage_url>/data/models/auto_labeling pl_augmented_worker_0
VLM endpoint default in-cluster NIM or explicit override augmentation + AL workers
LLM endpoint default in-cluster NIM or explicit override augmentation + AL workers

Submit

Use the shared submit command and the common optional-overrides block in the SKILL.md "Submit (all flows)" section, with workflow YAML assets/configs/osmo/augmentation_and_al.yaml. This flow runs augmentation first, then auto-labels the augmented outputs, so the full canonical single-flag submit shape applies (including both cache URL values).

Output layout

<storage_url>/datasets/<dataset>-outputs/<run_id>/
├─ setup_b0/
└─ outputs/
   ├─ augmented/<video>_aug0/
   └─ pseudo_labeled_augmented/<video>_aug0/

Troubleshooting

  • Completion evidence requirement: include a side-by-side input vs augmented video artifact, augmentation summary (setup_b0/configs/manifest.yaml sampled_vars for <video>_aug0), and augmented auto-labeling artifact summary from outputs/pseudo_labeled_augmented/<video>_aug0 using the workspace-local run copy under media/vda/runs/<run_id>/.
  • If submit fails with missing cache wiring, run setup_model_cache.yaml and rerun pre_submit_guard.py.
  • If workers stall on endpoints, verify vlm_url/llm_url health and /v1 availability before resubmitting.

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    The skill is a workflow orchestrator for video data augmentation and auto-labeling on the NVIDIA OSMO platform. It manages the end-to-end pipeline, including credential verification, configuration generation, worker execution, and result retrieval. No security issues or malicious patterns were detected; all external dependencies and network operations are associated with trusted vendors and the skill's primary functionality.

  • Socket3mo

    3 alerts: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 0482ebc. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
Other metadata
metadata
{
  "owner": "NVIDIA",
  "service": "data",
  "version": "1.0.0",
  "reviewed": "2026-05-26",
  "author": "NVIDIA",
  "tags": [
    "physical-ai",
    "video-data-augmentation",
    "auto-labeling",
    "cosmos"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-video-data-augmentation