All skills
nvidia avatar

/deepstream-sop

@4cb1092
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification. Trigger even if the user does not name it: verify operator step sequence, detect missing or out-of-order SOP steps, score factory/work-cell video for procedure compliance, run VLM-based SOP checking on industrial cameras, or call /v1/chat/completions with a file, RTSP, or Basler camera. Also trigger for its internals: SOPVideoProcessor, DeepStream GEBD model (e.g. DDM) via Triton CAPI, nvds_custom_postprocess, Cosmos Reason 1/2 vLLM, SSE streaming, Kafka NvProto/JSON output, Basler/Pylon camera + emulation, Docker compose, chunk-level latency. Do NOT trigger for generic DeepStream pipelines, object detection/tracking, NIM imports, or video summarization.

Use this Skill: https://skilld.dev/gh/nvidia/skills/deepstream-sop

This session only. Nothing lands on disk.

referencesskill_02_pydantic_schemas.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

§ 2 — Pydantic API Schemas (api_types.py)

Generates: nvds_action_detector/api_types.py Critical Rules: UNIFORM_CHUNKING_BYPASSES_DDM Pipeline stage: Data contracts — defines request/response shapes for all endpoints (§ 1)


Request Models

from pydantic import BaseModel, Field
from typing import Literal, Optional, Union
from typing_extensions import Annotated

# --- Input content types ---
# The chat completion endpoint (§ 1) accepts multimodal content:
#   TextContent     — user's text prompt (e.g., "Analyze SOP compliance")
#   VideoURLContent — video via HTTP URL, RTSP stream, or base64 data URI
#   VideoFileContent — previously uploaded file (via POST /v1/files)
#   CameraInputContent — live Basler industrial camera feed (§ 8)

class TextContent(BaseModel):
    type: Literal["text"] = "text"
    text: str

class VideoURL(BaseModel):
    url: str  # HTTP(S), data:video/mp4;base64,... or rtsp://

class VideoURLContent(BaseModel):
    type: Literal["video_url"] = "video_url"
    video_url: VideoURL

class VideoFileContent(BaseModel):
    type: Literal["input_video"] = "input_video"
    file_id: str

class InputCamera(BaseModel):
    """Basler industrial camera configuration.
    See § 8 for physical camera and emulation usage patterns."""
    camera_id: str
    camera_vendor: Literal["Basler"] = "Basler"
    config: Optional[str] = None             # path to .pfs file
    camera_format: Optional[Literal["RGB", "UYVY", "YUY2"]] = None
    camera_width: Optional[int] = None
    camera_height: Optional[int] = None
    camera_fps_num: Optional[int] = Field(None, ge=1, le=1e8)
    camera_fps_den: Optional[int] = Field(None, ge=1, le=1e8)

class CameraInputContent(BaseModel):
    type: Literal["input_camera"] = "input_camera"
    input_camera: InputCamera

# Discriminated union: Pydantic uses the "type" field to pick the right model
ChatMessageContent = Annotated[
    Union[TextContent, VideoURLContent, VideoFileContent, CameraInputContent],
    Field(discriminator="type")
]

class ChatCompletionMessage(BaseModel):
    role: str
    content: Union[list[ChatMessageContent], str]

# --- Chunking options ---
# Two selectable algorithms, chosen via a discriminated union on `algorithm`:
#   "ddm-net" (default) — DDM event-boundary detection (Stage 1 CV + Stage 2)
#   "uniform"           — fixed-length chunks; bypasses the DDM model entirely (see § 3, § 6)
# Both set model_config extra="forbid", so a payload that mixes the two algorithms'
# fields (e.g. `threshold` under `algorithm:"uniform"`) is rejected with HTTP 422.
# Note: internal ChunkParams.max_length_sec defaults to 10.0s,
#        but API DdmNetChunkingOptions defaults to 60.0s

class DdmNetChunkingOptions(BaseModel):
    model_config = {"extra": "forbid"}
    algorithm: Literal["ddm-net"] = "ddm-net"
    threshold: float = Field(0.8, gt=0, lt=1.0)
    min_length_sec: float = Field(1.0, gt=0)
    max_length_sec: float = Field(60.0, gt=0)

class UniformChunkingOptions(BaseModel):
    model_config = {"extra": "forbid"}
    algorithm: Literal["uniform"] = "uniform"
    chunk_length_sec: float = Field(5.0, gt=0)    # fixed chunk duration in seconds

# Discriminated union — Pydantic picks the model from the `algorithm` literal.
ChunkingOptions = Annotated[
    Union[DdmNetChunkingOptions, UniformChunkingOptions],
    Field(discriminator="algorithm"),
]

class ChatCompletionRequest(BaseModel):
    messages: list[ChatCompletionMessage]
    model: Optional[str] = None
    stream: Optional[bool] = False
    temperature: Optional[float] = None
    max_completion_tokens: Optional[int] = None
    seed: Optional[int] = None
    top_p: Optional[float] = None
    # Defaults to DDM-net when omitted; pass {"algorithm":"uniform","chunk_length_sec":N} for fixed chunks.
    chunking_options: Optional[ChunkingOptions] = Field(default_factory=DdmNetChunkingOptions)

Response Models

# --- Health check response ---
class HealthSuccessResponse(BaseModel):
    object: Literal["health.response"] = "health.response"
    message: str

# --- File management responses ---
class FileObject(BaseModel):
    id: str                          # "file-<uuid>"
    object: Literal["file"] = "file"
    bytes: int
    created_at: int                  # Unix timestamp
    filename: str
    purpose: str

class FileList(BaseModel):
    object: Literal["list"] = "list"
    data: list[FileObject]

class DeletionStatus(BaseModel):
    id: str
    object: Literal["file.deleted"] = "file.deleted"
    deleted: bool = True

# --- Streaming (SSE) response models ---
# Used by the SSE generator (§ 7) to emit chat.completion.chunk events.
# Each chunk carries one action classification result from the 4-stage pipeline.

class DeltaMessage(BaseModel):
    role: Literal["user", "assistant", "system"] = "assistant"
    content: Optional[str] = None

class ChatCompletionResponseStreamChoice(BaseModel):
    index: int = 0
    delta: DeltaMessage
    finish_reason: Optional[str] = None
    chunk_metadata: Optional[dict] = None

class ChatCompletionStreamResponse(BaseModel):
    id: str                                          # "chatcmpl-<uuid>"
    object: Literal["chat.completion.chunk"] = "chat.completion.chunk"
    created: int
    model: str
    choices: list[ChatCompletionResponseStreamChoice]

# --- Non-streaming response models ---
# Used when stream=false; collects all chunk results before responding.

class ChatCompletionResponseChoice(BaseModel):
    index: int
    message: ChatCompletionMessage
    finish_reason: Optional[str] = None
    chunk_metadata_list: Optional[list[dict]] = None

class ChatCompletionResponse(BaseModel):
    id: str                                          # "chatcmpl-<uuid>"
    object: Literal["chat.completion"] = "chat.completion"
    created: int
    model: str
    choices: list[ChatCompletionResponseChoice]
    usage: Optional[dict] = None

Source: SKILL.md on GitHub

2 warnings27d3 checks · Risk MEDIUM
  • Gen Agent Trust Hub27d

    This skill provides a framework for building an industrial SOP monitoring microservice. It uses multi-stage processing, including Vision-Language Models and custom logic. Security observations include dynamic code generation based on configuration files, execution of build scripts via subprocess, and reliance on external research repositories for model weights and code.

  • Socket27d

    2 alerts: gptSecurity, gptAnomaly

  • Snyk27d

    Risk: LOW · No issues

Signed by skilld at 4cb1092. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 months ago
owner
windy@nvidia.com
service
deepstream-sop
version
1.0.0
Other metadata
reviewed
2026-04-08
metadata
{
  "author": "Wind Yuan <windy@nvidia.com>",
  "tags": [
    "deepstream",
    "sop",
    "vlm",
    "triton",
    "gpu"
  ],
  "languages": [
    "python"
  ],
  "frameworks": [
    "deepstream",
    "triton",
    "fastapi"
  ],
  "domain": "video-analytics"
}

README badge

README badge for nvidia/skills/deepstream-sop