All skills
nvidia avatar

/physical-ai-video-data-augmentation

@0482ebc
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use when running video data augmentation and auto-labeling workflows on OSMO: flow selection, preflight, submit-time interpolation, monitoring, and output retrieval. Trigger keywords: video data augmentation, data enrichment, auto labeling, VDA demo, OSMO workflow, pseudo labeling.

Use this Skill: https://skilld.dev/gh/nvidia/skills/physical-ai-video-data-augmentation

This session only. Nothing lands on disk.

assetscookbookswarehouseauto_labelingpromptsevent_analysis.md

≈717 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Warehouse-construction VLM event-analysis prompt.

Runtime file path: /workspace/configs/video_event_analysis_prompt_redid.md.

Formatting contract:

  • Emit two JSON objects only.
  • First object is metadata without an "events" key.
  • Second object contains the "events" array.
  • Keep output machine-readable; no commentary.

Video context:

  • Ground-level fixed camera inside an active warehouse buildout area.
  • Typical entities include workers, ladders, lifts, tools, materials, and floor cabling.

Assignment: Identify safety incidents, near misses, anomalies, and normal operations.

Allowed sub-categories:

  • worker_equipment_contact
  • worker_fall
  • near_miss_equipment
  • near_miss_falling_object
  • unsafe_ladder_use
  • cable_trip_hazard
  • normal_construction
  • equipment_operation
  • worker_transit

Metadata JSON requirements:

  • Mandatory keys: version, video_id, format, rectified, scenario_info, scene_description, event_summary, fps, duration, height, width, camera_id
  • scenario_info must equal "INDOOR_WAREHOUSE"
  • scene_description should summarize floor layout, active equipment, cable/material distribution, lighting mix, and hazard zones
  • event_summary should capture workforce activity level and timestamped outcomes

Event JSON requirements:

  • Root keys: version, events
  • Every event entry must include:
    • event_id
    • start_time
    • end_time
    • category (collision | near_miss | anomaly | normal_traffic)
    • sub_category (array)
    • instances (array)
    • event_caption

Category mapping:

  • collision -> worker_equipment_contact, worker_fall
  • near_miss -> near_miss_equipment, near_miss_falling_object
  • anomaly -> unsafe_ladder_use, cable_trip_hazard
  • normal_traffic -> normal_construction, equipment_operation, worker_transit

Validation constraints:

  • sub_category must always be a JSON array
  • event_caption includes severity (low/medium/high), actor/equipment references, and supporting timing
  • Prefer tracking IDs when present; otherwise use labels like worker/operator/crew member
  • Times are numeric seconds

No-worker scenario:

  • If no people are visible, return metadata plus an empty events array.

Warehouse-scene caveats:

  • Columns and material stacks can hide workers; short occlusion is not itself anomalous.
  • Floor cables are common background; classify cable_trip_hazard only when stretched across active walk paths.
  • Ladder and scaffold usage both occur; capture unsafe posture/placement under unsafe_ladder_use.
  • Mixed lighting may obscure detail; when uncertain, describe observable evidence instead of inferring unseen actions.
  • Missing visible PPE can be noted in captions, but keep primary category tied to the observed event class.
  • Slow lift repositioning is normal; near_miss_equipment requires close worker proximity during active motion.

Required output order: metadata object first, events object second.

Source: SKILL.md on GitHub

2 warnings3mo3 checks · Risk SAFE
  • Gen Agent Trust Hub3mo

    The skill is a workflow orchestrator for video data augmentation and auto-labeling on the NVIDIA OSMO platform. It manages the end-to-end pipeline, including credential verification, configuration generation, worker execution, and result retrieval. No security issues or malicious patterns were detected; all external dependencies and network operations are associated with trusted vendors and the skill's primary functionality.

  • Socket3mo

    3 alerts: gptAnomaly

  • Snyk3mo

    Risk: MEDIUM · 1 issue

Signed by skilld at 0482ebc. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 4 months ago
Other metadata
metadata
{
  "owner": "NVIDIA",
  "service": "data",
  "version": "1.0.0",
  "reviewed": "2026-05-26",
  "author": "NVIDIA",
  "tags": [
    "physical-ai",
    "video-data-augmentation",
    "auto-labeling",
    "cosmos"
  ]
}

README badge

README badge for nvidia/skills/physical-ai-video-data-augmentation