All skills
nvidia avatar

/deepstream-sop

@4cb1092
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification. Trigger even if the user does not name it: verify operator step sequence, detect missing or out-of-order SOP steps, score factory/work-cell video for procedure compliance, run VLM-based SOP checking on industrial cameras, or call /v1/chat/completions with a file, RTSP, or Basler camera. Also trigger for its internals: SOPVideoProcessor, DeepStream GEBD model (e.g. DDM) via Triton CAPI, nvds_custom_postprocess, Cosmos Reason 1/2 vLLM, SSE streaming, Kafka NvProto/JSON output, Basler/Pylon camera + emulation, Docker compose, chunk-level latency. Do NOT trigger for generic DeepStream pipelines, object detection/tracking, NIM imports, or video summarization.

Use this Skill: https://skilld.dev/gh/nvidia/skills/deepstream-sop

This session only. Nothing lands on disk.

referencesskill_13_verification_curl.md

≈1k tokens on demand. Your agent reads this file only when SKILL.md points to it.

§ 13 — Verification & curl Examples

Use these commands to verify the service after build and deploy (§ 9).


Build & Test

# 1. Build Docker image
docker compose -f deploy/compose.yaml build

# 2. Start service (requires GPU + DDM model + embedded vLLM)
docker compose -f deploy/compose.yaml up -d

# 3. Run full API test suite (24 tests)
TEST_VIDEO_PATH=test_video_whole_sop_h264.mp4 pytest tests/test_api_endpoints.py -v

curl Examples

BASE_URL="http://localhost:8300"

# --- Health checks ---
# NOTE: model loading takes several minutes on first run.
# Poll every 10s until HTTP 200 — see § 9 (skill_09) for the full readiness wait script.
curl $BASE_URL/v1/live
curl $BASE_URL/v1/ready
curl $BASE_URL/v1/startup

# --- List models ---
curl $BASE_URL/v1/models

# --- Upload a video file ---
curl -X POST $BASE_URL/v1/files \
  -F "file=@test_video.mp4" \
  -F "purpose=inference"

# --- List files ---
curl $BASE_URL/v1/files

# --- Chat completion (non-streaming, base64 video) ---
# Encodes video as base64 data URI → sent to Stage 1 pipeline.
# NOTE: Do NOT include a {"type":"text"} item in the message content.
#   The VLM prompt is controlled by VLM_PROMPT_PATH (e.g. configs/assy17_vlm_prompts.txt).
#   If a text item is present, it overrides the config-file prompt and breaks SOP classification.
VIDEO_B64=$(base64 -w0 test_video.mp4)
curl -X POST $BASE_URL/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ds_sop_model",
    "messages": [{"role": "user", "content": [
      {"type": "video_url", "video_url": {"url": "data:video/mp4;base64,'$VIDEO_B64'"}}
    ]}],
    "stream": false,
    "chunking_options": {"algorithm": "ddm-net", "threshold": 0.8}
  }'

# --- Streaming response ---
# Returns SSE chat.completion.chunk events as each chunk is processed.
# NOTE: Omit {"type":"text"} — VLM prompt comes from VLM_PROMPT_PATH config file.
curl -N -X POST $BASE_URL/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"ds_sop_model","messages":[{"role":"user","content":[{"type":"video_url","video_url":{"url":"data:video/mp4;base64,'$VIDEO_B64'"}}]}],"stream":true}'

# --- Uniform (fixed-length) chunking ---
# algorithm:"uniform" cuts fixed chunk_length_sec chunks and BYPASSES the DDM model (§ 3, § 6).
# extra="forbid": do NOT mix DDM fields (threshold / min_length_sec / max_length_sec) here → 422.
curl -N -X POST $BASE_URL/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"ds_sop_model","messages":[{"role":"user","content":[{"type":"video_url","video_url":{"url":"data:video/mp4;base64,'$VIDEO_B64'"}}]}],"stream":true,"chunking_options":{"algorithm":"uniform","chunk_length_sec":2.0}}'

# --- Physical Basler camera (serial 40748152, PFS optional) ---
# NOTE: Use max_length_sec <= 5.0 for responsive live streaming (default 60.0 is too slow)
curl -N -X POST $BASE_URL/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"ds_sop_model","messages":[{"role":"user","content":[
    {"type":"input_camera","input_camera":{
      "camera_id":"40748152","camera_vendor":"Basler",
      "camera_format":"RGB","camera_width":1280,"camera_height":720,
      "camera_fps_num":30,"camera_fps_den":1
    }}
  ]}],"stream":true,"chunking_options":{"algorithm":"ddm-net","threshold":0.8,"max_length_sec":2.0}}'

# --- Camera emulation (serial 0815-0000, PYLON_CAMEMU=1 in service env required) ---
curl -N -X POST $BASE_URL/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"ds_sop_model","messages":[{"role":"user","content":[
    {"type":"input_camera","input_camera":{
      "camera_id":"0815-0000","camera_vendor":"Basler",
      "config":"/opt/nvidia/nvds_sop/configs/Emulation_0815-0000.pfs",
      "camera_format":"RGB","camera_width":1280,"camera_height":720,
      "camera_fps_num":30,"camera_fps_den":1
    }}
  ]}],"stream":true,"chunking_options":{"algorithm":"ddm-net","threshold":0.8,"max_length_sec":2.0}}'

# --- Prometheus metrics ---
curl $BASE_URL/v1/metrics

Source: SKILL.md on GitHub

2 warnings27d3 checks · Risk MEDIUM
  • Gen Agent Trust Hub27d

    This skill provides a framework for building an industrial SOP monitoring microservice. It uses multi-stage processing, including Vision-Language Models and custom logic. Security observations include dynamic code generation based on configuration files, execution of build scripts via subprocess, and reliance on external research repositories for model weights and code.

  • Socket27d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk27d

    Risk: LOW · No issues

Signed by skilld at 4cb1092. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 months ago
owner
windy@nvidia.com
service
deepstream-sop
version
1.0.0
Other metadata
reviewed
2026-04-08
metadata
{
  "author": "Wind Yuan <windy@nvidia.com>",
  "tags": [
    "deepstream",
    "sop",
    "vlm",
    "triton",
    "gpu"
  ],
  "languages": [
    "python"
  ],
  "frameworks": [
    "deepstream",
    "triton",
    "fastapi"
  ],
  "domain": "video-analytics"
}

README badge

README badge for nvidia/skills/deepstream-sop