HoloHub flow benchmarking
Use this reference only after the application passes its normal finite and visual acceptance checks.
Preconditions
- Read the checked-in flow-benchmark README, tutorial, runner source, and analyzer help from the same HoloHub revision.
- Reconcile examples with local
./holohubhelp. The tested revision uses--extra-scripts(plural). - Run the normal finite application once so compilation or engine generation is outside measured trials.
Instrumented build
Preview and perform the exact selected mode/language build:
./holohub build <app> <mode> --language <cpp-or-python> --benchmark \
--extra-scripts benchmarking --dryrun --verbose
./holohub build <app> <mode> --language <cpp-or-python> --benchmark \
--extra-scripts benchmarking --verboseRun and analyze
Run helpers inside the wrapper-managed container, not directly on the host.
Inspect each helper's -h, require the runner's scheduler, and preserve the
compound command as one quoted shell argument after --:
Choose a fresh, empty, writable <host-output-dir> whose basename is the
shell-safe <output-dir-name>. Keep restricted or canonical evidence outside
the checkout; if repository policy permits an in-checkout result, first prove
the exact directory is ignored. Mount the same host directory into both
container invocations and use only its translated container path:
./holohub run-container <app> <mode> --language <cpp-or-python> \
--extra-scripts benchmarking --add-volume <host-output-dir> \
--dryrun --verbose -- \
'set -e
./holohub run <app> <mode> --local --no-local-build \
--language <cpp-or-python> --dryrun --verbose
python benchmarks/holoscan_flow_benchmarking/benchmark.py \
-a <app> --language <cpp-or-python> \
--run-command "./holohub run <app> <mode> --local --no-local-build --language <cpp-or-python> --verbose" \
--sched greedy -r 3 -i 1 -m <messages> \
-d /workspace/volumes/<output-dir-name>
test -s /workspace/volumes/<output-dir-name>/benchmark.log
if grep -Eq "exited with code [1-9][0-9]*" \
/workspace/volumes/<output-dir-name>/benchmark.log; then
echo "A benchmark application trial failed" >&2
exit 1
fi
test "$(find /workspace/volumes/<output-dir-name> -maxdepth 1 \
-type f -name "logger_greedy_*.log" -size +0c | wc -l)" -eq 3'
./holohub run-container <app> <mode> --language <cpp-or-python> \
--extra-scripts benchmarking --add-volume <host-output-dir> --verbose -- \
'set -e
./holohub run <app> <mode> --local --no-local-build \
--language <cpp-or-python> --dryrun --verbose
python benchmarks/holoscan_flow_benchmarking/benchmark.py \
-a <app> --language <cpp-or-python> \
--run-command "./holohub run <app> <mode> --local --no-local-build --language <cpp-or-python> --verbose" \
--sched greedy -r 3 -i 1 -m <messages> \
-d /workspace/volumes/<output-dir-name>
test -s /workspace/volumes/<output-dir-name>/benchmark.log
if grep -Eq "exited with code [1-9][0-9]*" \
/workspace/volumes/<output-dir-name>/benchmark.log; then
echo "A benchmark application trial failed" >&2
exit 1
fi
test "$(find /workspace/volumes/<output-dir-name> -maxdepth 1 \
-type f -name "logger_greedy_*.log" -size +0c | wc -l)" -eq 3'The first command previews only the outer container launch, so inspect its
quoted child sequence before continuing. The second command keeps the nested
run --dryrun --verbose as a gate: only after that exact selected-mode preview
succeeds does the runner execute the matching custom --run-command. Always
pass the selected mode through --run-command; the runner's autogenerated
command omits mode and can silently select default_mode. The current runner
logs a nonzero application return code from a worker thread but can still
return zero, so require a nonempty benchmark.log, reject every recorded
exited with code, and require the expected runs * instances nonempty raw
logs before analysis. The example expects three logs. Then run analyze.py in
the same environment against the mounted directory. Check each parser
independently; runner and analyzer short options do not necessarily mean the
same thing.
Preserve raw logs and analyzer output. Report workload, scheduler, instances, runs, input, retained samples, warm-up/trim rules, operator path, aggregation, variability, hardware, image, SDK, driver, and exclusions.
Restore and prove normal state
Before instrumentation, record concise status, affected source hashes, and the
normal build shape. After the run, search the application tree for *.bak,
compare status and hashes with that baseline, and restore only
benchmark-attributable changes without overwriting pre-existing work. Automatic
restoration is not guaranteed after failure or interruption.
Preview and run the normal build without --benchmark, confirm from its
configuration/build evidence that benchmark flags are absent, and rerun the
finite smoke case. If normal source, build state, or smoke behavior is not
restored after this single pass, stop and hand off the exact state to
holohub-debug-build-run; do not clear caches or loop speculatively. A clean
source diff alone does not remove instrumented binaries or cached flags.
Do not turn a benchmark into an accuracy, safety, clinical, regulatory, or product-wide performance claim.