Changelog
All notable changes to deepstream-import-vision-model are documented here. Format follows
Keep a Changelog and the skill uses
Semantic Versioning.
[1.5.2] — 2026-08-05
Changed
- Added
scripts/model/resolve-engine.shso Steps 6–7 and Step 8 share one engine resolver instead of each inlining the same glob, empty check, andMAX_BSparse.
Fixed
- Six defects found running the skill end-to-end on an H100 rather than by review.
[1.5.1] — 2026-08-04
Changed
- De-duplicated ONNX label extraction into a single helper shared by the HuggingFace and NGC routes.
[1.5.0] — 2026-08-04
Fixed
- Cleared the HIGH defects reported by SkillCritic across both DeepStream skills.
[1.4.3] — 2026-08-04
Added
- The missing
## Examplessection, with three concrete invocations: the default end-to-end run, a SafeTensors export showing the real dynamo → TorchScript fallback, and a revision-pinned build.
[1.4.2] — 2026-08-04
Fixed
- Cleared the remaining Tier-1 Code Risk Analysis findings, including B615
huggingface_unsafe_download— everyfrom_pretrained/snapshot_downloadcall now pins a revision.
[1.4.1] — 2026-08-04
Fixed
- Neutralised the NGC images' global
PIP_CONSTRAINT, whose tested pins conflicted with this skill's dependency set and made pip fail withResolutionImpossible. - Repaired transformers 5.x breakage in the export path.
[1.4.0] — 2026-08-04
Changed
- Replaced
optimum-cliwithtorch.onnx.export.optimum[exporters]pinned transformers below 4.54.0, but the two HIGH RCE advisories (GHSA-29pf-2h5f-8g72, GHSA-fgcw-684q-jj6r) are only fixed in 5.3.0 and 5.5.0, so the pin was unfixable while optimum stayed. optimum 2.1.0 had also dropped theonnxsubcommand, making that path a dead end regardless. The exporter now uses the dynamo backend with a TorchScript fallback and verifies the batch dimension stayed dynamic.
[1.3.6] — 2026-08-04
Fixed
- Cleared the NVSkills-Eval Tier-1 high-risk and Tier-2 findings that blocked the content
gate, including the Agent Snooping findings in
install.sh.
[1.3.5] — 2026-07-29
Added
- Explicit model-choice intake: the skill now always offers the validated default model and a custom object-detection model, and never silently substitutes one.
[1.3.4] — 2026-07-28
Fixed
- Updated the full workflow and container references to DeepStream 9.1 and CUDA 13.2.
- Made Bash and PowerShell reinstallation safe when source and destination are identical.
- Reconciled the README, phase references, tests, and eval expectations with the consolidated single-skill architecture.
- Pinned the Python dependency stack used by the in-container setup.
- Added standard Codex UI metadata and repository-required SPDX identifiers.
[1.3.3] — 2026-07-20
Fixed
bcdependency removed — timing/throughput math silently returned empty. The DeepStream container has nobc, so every$(echo "$A - $B" | bc)produced an empty string with no error, leaving all pipeline-timing (and some throughput) values blank. Installingbcviasetup.shwould not help — apt installs do not persist across the ephemeral--rmphase containers. All reference-doc timing now usespython3 -c "print(round(...))"; the helper scripts (benchmark-ds.sh,ds-sweep.sh,benchmark-trtexec.sh) useawk "BEGIN{printf ...}". Bothpython3andawkare always present in the image.- Step 8 PDF generation failed on a fresh run.
setup.shinstallswkhtmltopdfvia apt inside an ephemeral--rmcontainer, so the binary is gone by the time Step 8 runs in a later container (only the/work-mounted venv persists).scripts/report/md-to-html-pdf.pynow self-heals via anensure_wkhtmltopdf()helper that installs it if missing before rendering the PDF. The HTML (charts base64-inlined) was already unaffected.
Changed
- Real-time stream selection now converges instead of halving. When DS Run 2 came in marginally
under 30 fps/stream, the old fallback halved RT_STREAMS (e.g. 38 streams @ 29.6 fps → 19),
discarding ~half the GPU's real capacity and reporting a misleadingly low real-time count. Step 7
now recomputes the target from the measured throughput (
floor(TOTAL_FPS_RUN2 / 30)) and steps down one stream at a time, landing on the true ceiling (e.g. 37) in 1–2 short retries.
[1.3.2] — 2026-07-17
Removed
.claude-plugin/plugin.json— the skill now ships as a plain skill (like the siblingdeepstream-eval-and-finetune), consistent with how it is installed byinstall.sh(whole-directory copy into.claude/skills/and.cursor/skills/) and used both standalone and bundled with other skills. The manifest was not referenced byinstall.shand the skill was not registered in the repo marketplace, so removal has no effect on standalone or bundled use. This also lets NVCARPS nv-base classify the directory asType: skilland run its Tier-3 live agent-eval (producingBENCHMARK.md+ anAGENT_EVALresult), which the content gate requires and which theType: pluginpath skipped. To publish it later as a standalone marketplace plugin, re-add the manifest and run the signed marketplace flow.
[1.3.1] — 2026-07-16
Added
install.ps1— native-Windows (PowerShell) installer twin ofinstall.sh, with the identical sequence and flags (-Target=--target,-NoCursor=--no-cursor,-DryRun=--dry-run). Copies the skill into<project>\.claude\skills\(and.cursor\skills\). PowerShell 5.1+ compatible.
Changed
.gitattributesforces LF on*.ps1;references/windows.mdinstall note now points toinstall.ps1.
[1.3.0] — 2026-07-16
Changed
- Runs entirely through Docker — no host packages. Every step (venv/ONNX export, TensorRT engine build, nvinfer parser compile, DeepStream run, PDF report) now executes INSIDE the DeepStream container. The host needs only Docker + the NVIDIA driver, so the skill runs identically on Linux, Windows (Docker Desktop + WSL2 backend), and macOS.
- Removed the host-native toolchain assumptions (host
trtexec/nvidia-smi/dpkg/make/host venv/apt-get);wkhtmltopdf+ the export venv (build/.venv_optimum) are provisioned in-container by the newsetup.sh. - Reversed the "always build engines on the host" guidance — build and run now share one image, so there is no TensorRT build-vs-runtime version skew (the exact failure the old rule tried to avoid).
install.shnow installs the whole self-contained skill dir (SKILL.md + references + scripts + setup.sh) into.claude/skills/…, dropping the separatescripts/tree and theln -sfsymlink path.
Added
setup.sh(in-container bootstrap: venv + deps + wkhtmltopdf),scripts/preflight.sh(GPU + venv + trtexec, with container-mode),scripts/requirements.txt,scripts/dsrun.sh(docker wrapper),.gitattributes(LF), andreferences/windows.md(cross-platform runbook).
Fixed
scripts/model/safetensors-to-onnx.shno longer runspython3 -m venv(fails on the container python, which lacks ensurepip) — it reuses the virtualenv built bysetup.sh.
[1.2.2] — 2026-05-19
Changed
- Skill renamed from
deepstream-byovmtodeepstream-import-vision-modelacross all files:name:inSKILL.mdand.claude-plugin/plugin.json, package directory (team-skills/deepstream-sdk/deepstream-import-vision-model/), installed skill directories (.claude/skills/deepstream-import-vision-model,.cursor/skills/deepstream-import-vision-model), runtime scripts path (scripts/deepstream-import-vision-model/), invocation hints, eval prompts, tag (byovm→import-vision-model), README/title (DS BYOVM→DeepStream Import Vision Model), and cross-skill references inteam-skills/deepstream-sdk/README.md,team-skills/deepstream-sdk/deepstream-profile-pipeline/SKILL.md, andteam-skills/deepstream-sdk/deepstream-profile-pipeline/README.md. Body content (SKILL.md sections,references/*.md, and 5 differing scripts) was also resynced with the upstreamds-copilot/skills/deepstream-import-vision-modelsource - Encoder fallback: replaced
x264encfallback withtheoraenc + oggmux(LGPL, outputs.ogv).x264encandopenh264encare now prohibited (Rule 10). When neither NVENC northeoraenc/oggmuxis available, single-stream capture is skipped gracefully (DS_SINGLE_STREAM_MODE=skipped) - Video source: enforced
sample_720p.mp4(1280×720) as the mandatory default; custom paths only via explicitDS_VIDEO(Rule 11) - Performance measurement: switched DS multi-stream benchmark from
gst-launch-1.0 ! fpsdisplaysink(parsingCurrent FPS:) todeepstream-app -c … enable-perf-measurement=1(parsing**PERF:log lines) via the newscripts/deepstream/ds-perf-run.shwrapper. Removes runtime dependency ongstreamer1.0-plugins-bad - Media probing: replaced
ffprobe/gst-discoverercalls inbenchmark-ds.shandds-sweep.shwithmediainfo(with safe fallbacks) - Pipeline NVENC primary: switched
nvvideoconvertoutput format fromI420toNV12ahead ofnvv4l2h264enc - Puppeteer sandbox: split into two vetted configs —
mermaid-puppeteer.json(sandboxed; non-root) andmermaid-puppeteer-root.json(sandbox disabled; only selected whenuid == 0).render-mermaid-for-pdf.pyauto-selects the right one and refuses any user-supplied config that does not resolve to one of these two shipped files (blocks--remote-debugging-port,--load-extension, etc.) benchmark-trtexec.shinterface: replaced fixedb1 b16 b32 b64positional args with variadic<bs:engine> [<bs:engine> …] [duration]ds-sweep.sh: input shape and tensor name are now derived dynamically viainspect-onnx.pyinstead of hardcodedinputs+640×640(fixes YOLOv8 / RT-DETR / DETR / non-YOLOX models). Power-law batch-size prediction is guarded against α≈0 (flat curves)- Frame extraction:
extract-frame.shnow auto-detects.mp4vs.ogvand routes through the matching demux+decoder chain - Custom parser filenames: introduced
MODEL_NAME_SAFE = tr -c 'A-Za-z0-9' '_'for.cpp/.sofilenames so models likertdetr-lproduce a consistentlibnvdsinfer_rtdetr_l_parser.so - nvinfer config: moved
cluster-modeinline#comments to their own lines in both heredocs (GKeyFile rejects inline#) make-static-batch-onnx.py: useonnx.numpy_helper.to_array/from_arrayfor Reshape initializer patching (the old raw-bytes path silently skippedint64_datainitializers, leavingbatch=1baked in)- Report verification: replaced
>500 KBheuristic with deterministicgrep -o 'data:image/png' benchmark_report.html | wc -l == 5 - System tools: pre-flight now installs
mediainfoand checks fordeepstream-app(instead ofgstreamer1.0-plugins-bad)
Added
scripts/deepstream/ds-perf-run.sh— wrapsdeepstream-appwithenable-perf-measurement=1, emits**PERF:log lines for the report parserscripts/report/mermaid-puppeteer-root.json— vetted root-only Puppeteer config- New SKILL.md table rows:
ds-perf-run.sh,md-to-pdf.sh,mermaid-puppeteer-root.json
Fixed
ds-kitti-dump.sh: addedset -euo pipefail, replaced manualrm -fwith trap-based cleanup, guardedtimeoutpipeline withset +o pipefailto preservePIPESTATUSsafetensors-to-onnx.sh: added missingset -euo pipefailgenerate-benchmark-charts.py: removed unusedimport math
[1.2.1] — 2026-04-24
Changed
- Skill renamed from
ds-byovmtodeepstream-byovmacross all files:name:inSKILL.mdandplugin.json, installed skill directory (.claude/skills/deepstream-byovm), runtime scripts path (scripts/deepstream-byovm/), invocation hints, eval prompts, and all cross-references inreferences/*.md
[1.2.0] — 2026-04-24
Changed
- References pattern: removed 4 standalone sub-skills (
nv-model-acquire,nv-engine-build,ds-run-pipeline,nv-byovm-report); their content is now inskills/ds-byovm/references/(4 .md files) matching the ds-copilotdeepstream-devconvention of single skill + reference documents - Single skill dir:
SKILL.mdmoved from package root intoskills/ds-byovm/SKILL.md(lean ~170 lines); rootSKILL.mdremoved - plugin.json:
"skills": "./"→"skills": "skills/ds-byovm/"to point at the skill directory instead of the package root - install.sh: creates one symlink (
skills/ds-byovm/→.claude/skills/ds-byovmand.cursor/skills/ds-byovm) instead of 5; no sub-skill symlinks - Installed structure is now:
.claude/skills/ds-byovm/ SKILL.md references/ model-acquire.md engine-build.md pipeline-run.md report-generation.md scripts/ds-byovm/ (19 scripts, unchanged) - Tests: updated install dry-run assertions for single-skill structure;
sub-skill name assertions removed;
assertNotInfor sub-skill names added
[1.1.0] — 2026-04-24
Changed
- Skill-only architecture: removed
agents/deepstream-sdk/ds-byovm.md; top-levelSKILL.mdnow serves both Claude Code and Cursor (Cursor does not support agents — skill is the correct primitive for cross-tool compatibility) - Sub-skills renamed with
nv-/ds-prefix for namespace clarity:hf-model-acquire→nv-model-acquiretrt-engine-build→nv-engine-buildds-integration→ds-run-pipelinebenchmark-report→nv-byovm-report
- SKILL.md enhanced (version 1.0.1 → 1.1.0): merged pre-flight checks, mandatory model folder structure, engine naming convention, run budget table, pipeline timing pattern, and report output convention from the removed agent doc
- install.sh updated: removed agent installation block, updated symlink targets to new sub-skill names; invocation hints say "skill" not "agent"
- README.md updated: skill-only usage section for Claude Code + Cursor; sub-skill standalone invocation documented; all agent references removed
- evals.json: all prompts and assertion text updated from "agent" to "skill"
- Tests expanded: dry-run test now asserts all 5 skill names, Cursor skills
presence, and absence of
.claude/agents/directory - Shebang consistency: all shell scripts use
#!/usr/bin/env bashfor portability across container images and macOS environments - Local
.gitignore: added skill-level.gitignorefor portability when placed in repos that do not inherit team-mind-hub's root.gitignore
[1.0.1] — 2026-04-23
Fixed
ds-integrationStep 6g KITTI dump produced zero detection files.gie-kitti-output-diris adeepstream-app[application]key — it is not read bynvinfer, so appending it to the nvinfer config and running agst-launch-1.0 ... nvinfer ...pipeline silently wrote no files. Step 6g now invokesscripts/ds-byovm/deepstream/ds-kitti-dump.sh, which wrapsdeepstream-appwith the correct[application]section.
Changed
hf-model-acquireStep 2b now uses a single sharedbuild/.venv_optimumfor SafeTensors → ONNX export across all models, matching whatscripts/ds-byovm/model/safetensors-to-onnx.shalready does. The previous prose created a freshbuild/.venv_$MODEL_NAMEper model, which re-installedoptimum/transformers/torchevery run (~minutes + GBs wasted). New models that need extra packages (timmfor DETR,onnxsim, etc.) shouldpip installinto the shared venv.cleanup.shstill removes any legacy per-model venvs for backward compatibility, and explicitly preserves the sharedbuild/.venv_optimum.
[1.0.0] — 2026-04-22
Added
- Initial release of the DeepStream Bring Your Own Vision Model (BYOVM) skill
- End-to-end pipeline: HuggingFace / NGC model → ONNX → TensorRT engine → DeepStream → benchmark report
- Four orchestrated sub-skills under
skills/:hf-model-acquire— model download and format routing (ONNX vs SafeTensors)trt-engine-build— dynamic TRT engine build +trtexecbenchmarksds-integration— customnvinferparser, single-stream + multi-stream DS runsbenchmark-report— 5-chart Markdown → HTML → PDF report
- Runtime scripts under
scripts/:model/— HF/NGC list + download helpers, ONNX inspection, SafeTensors → ONNX export, scoped cleanupengine/—trtexecbenchmark helperdeepstream/— single-stream, sweep, KITTI dump, frame extraction helpersreport/— chart generation, Mermaid → PNG, Markdown → HTML → PDF
- Installer (
install.sh) with validated--targetand dry-run mode - Declared
permissions:block inSKILL.mdfrontmatter: tool allowlist, MCP scope, network egress allowlist (huggingface.co,api.ngc.nvidia.com,api-inference.huggingface.co), and filesystem read/write scoping - Test suite (
tests/test_hardened_scripts.py) covering input validation for all hardened shell scripts and the PDF image-embedding helper (24 tests)
Security
- All shell scripts validate inputs against
^[A-Za-z0-9._/-]+$or tighter before touching filesystem or network curlinvocations pinned to HTTPS + TLSv1.2 with bounded timeouts; optional$HF_TOKENhonored for gated HuggingFace reposinstall.shrejects--targetvalues that are empty,/, contain.., or don't exist; destructiverm -rfis scoped to paths under$TARGETscripts/report/md-to-html-pdf.pybase64-inlines images before rendering;wkhtmltopdfruns without--enable-local-file-accessscripts/report/render-mermaid-for-pdf.pyrefuses user-supplied Puppeteer configs; always uses the vettedmermaid-puppeteer.jsonshipped with the skill
Known limitations
- Nested sub-skills under
skills/surface a low-severity schema warning from some scanners; kept in place because the agent file preloads them by name and the installer symlinks them into.claude/skills/at the target - Engine build time depends on GPU, ONNX complexity, and requested batch size; the skill retries on OOM by halving batch size but gives up after reaching 1