DOCA GPUNetIO ib_write_bw — Tasks
Where to start: The verbs that carry real workflow
content are ## configure, ## build, ## run, ## test,
and ## debug. The other verbs (install, modify, use)
carry preconditions, refuses-to-patch routing, and the
use-side decision shape respectively.
This file is loaded by SKILL.md after
CAPABILITIES.md. It walks the agent
through the task verbs every artifact in this bundle exposes
(install / configure / build / modify / run / test / debug / use), then defers task verbs that do not belong here.
install
Goal: confirm the user's hosts (client AND server) have every precondition the build + run sequence needs before any GPUNetIO-specific work begins.
This skill does not own DOCA installation; that path
lives in doca-setup. The
gpunetio_ib_write_bw-specific preconditions:
doca-gpunetio.pcis present. Runpkg-config --modversion doca-gpunetioon each build host. If the.pcdoes not resolve, the installed DOCA does not include GPUNetIO; route to../../libs/doca-gpunetio/TASKS.mdfor the package-name lookup and re-install.doca-common.pcanddoca-rdma.pcagree on the same DOCA semver. Per the four-way match indoca-version CAPABILITIES.md ## Version compatibility.- CUDA Toolkit +
nvccare on the build host. Theclient/kernel.cuis compiled bynvcc; confirmnvcc --versionresolves and the toolkit version is paired with the installed DOCA per the DOCA release notes (looked up viadoca-public-knowledge-map). nvidia_peermemkernel module is loaded. GPUDirect- style memory registration is the precondition for binding GPU memory to DOCA. Verify per../../libs/doca-gpunetio/CAPABILITIES.md#safety-policyanddoca-setup TASKS.md ## debug.- GPU visible to the host.
nvidia-smilists the GPU and reports its PCIe bus ID. - IB device visible to DOCA.
doca_caps --list-devs(per../doca-caps/TASKS.md#run) reports the IB device for-d <ibdev>. - OOB connectivity exists between client and server
for the TCP exchange the source's
oob_connection_*helpers drive. - The operator explicitly confirms the IB fabric is trusted and non-shared for the benchmark window. If the fabric is shared or its trust / tenancy is unknown, stop; do not run sustained WRITE traffic.
If any precondition fails, stop and route; tool-level diagnosis against a half-installed environment wastes time.
configure
Goal: pick the right GPU + IB device pair on each host, commit to the runtime-surface choice, and prepare the build tree.
Steps the agent walks the user through, in order:
- Commit to the runtime surface. Walk
CAPABILITIES.md ## Capabilities and modessurface-selection table. If the application has not committed to GPUNetIO, surface the GPI alternative and the CPU-initiated alternative; do not silently default to GPUNetIO. - Identify the GPU + IB device pair and check the
pairing. On each host:
nvidia-smi --query-gpu=pci.bus_idfor the GPU,ibdev2netdev -v(ordoca_caps --list-devs) for the IB device. Confirm the pairing perCAPABILITIES.md ## Capabilities and modesGPU-NIC pairing rule. The pairing is a precondition, not a tunable parameter. - Confirm
nvidia_peermem. Without it, the doca-gpu context cannot bind GPU memory; the symptom shows up as a layer-4 lifecycle failure perCAPABILITIES.md ## Error taxonomy. - Pick the GID index.
--gid-indexis optional but the right value is GID-routing-rule-specific per../../libs/doca-rdma/CAPABILITIES.md. The same index must be used on both client and server. - Pick the client / server roles and the OOB IP. The
server is started first (binds and waits on the TCP
socket); the client is started with
-c <server-ip>. - Confirm the build inputs.
PKG_CONFIG_PATHincludes the install'spkgconfigdirectory; the agent does not invent the literal path. - Re-confirm the fabric gate. Record the operator's explicit confirmation that the IB fabric is trusted and non-shared for this run. No confirmation means stop.
For the canonical DOCA universal lifecycle that underlies
program-side configuration (which the binary runs internally
per the GPUNetIO + RDMA libraries), see
doca-programming-guide TASKS.md ## configure.
build
The tool is not pre-built under /opt/mellanox/doca/;
the user builds it from source under
doca/tools/gpunetio_ib_write_bw/ against the installed
DOCA. The build pattern is the canonical meson flow:
- Set the right
PKG_CONFIG_PATHsopkg-configcan finddoca-gpunetio.pc,doca-rdma.pc, anddoca-common.pc. On a stock install the path lives under the DOCApkgconfigdirectory documented indoca-public-knowledge-map. - Set up the build directory.
meson setup <build-dir> doca/tools/gpunetio_ib_write_bw/from the workspace root. The top-levelmeson.buildwires together theclient/andserver/subtrees; each carries its ownmeson.buildand source files. - Compile.
meson compile -C <build-dir>produces the client and server binaries (the target names are declared in the per-subtreemeson.buildfiles; the agent re-reads them rather than quoting from memory). - Smoke the build artifacts. Run
<client-binary> --helpand<server-binary> --helpon each host before deploying. If--helpdoes not resolve, the build did not produce the expected artifacts — re-route to## debuglayer 2.
Routing for nearby "build" questions:
- "Can I build the tool against a different DOCA than the one I have installed?" → no, not through this skill; install a different DOCA first.
- "I want to build my own GPUNetIO-based BW benchmark
from scratch." → route to
../../libs/doca-gpunetio/TASKS.mdanddoca-programming-guide TASKS.md ## build.
This documented Meson flow is the supported build path for the shipped source tree. The skill forbids inventing wrappers or replacement source, not following the tool's documented Meson build.
modify
Do not patch the shipped tool source tree. The shipped
client/{main.c,common.c/h,kernel.cu,perftest.c} and
server/{main.c,common.c/h,perftest.c} files are the
verified worked example for this benchmark class; modifying
them puts the user in contributor-to-DOCA territory.
What the agent does modify is the build + invocation
environment — the PKG_CONFIG_PATH, the chosen GPU + IB
device pair, the GID index, the client / server roles, the
OOB IP, the run-time environment variables
(DOCA_LOG_LEVEL, CUDA_VISIBLE_DEVICES, NUMA pinning
via numactl). Treat "modify the environment, not the
source" as the operating mode.
Routing for nearby "modify" questions:
- "The reported columns are inconvenient — can I change them?" → source-level change; out of scope for this skill.
- "I want to add a new message-size sweep." → out of scope; this would be a contribution to the shipped tool.
- "I need a different metric than this tool reports." →
re-examine the runtime-surface choice in
## configurestep 1; if the user genuinely needs a bespoke benchmark, route todoca-programming-guideand the matching library skill.
run
The smoke-before-bulk flow — every session goes through it.
The detailed flag surface lives in the binary's --help on
the installed build.
Do-not-invent guard (paths + binary layout). Real downstream agents have hallucinated a
/opt/mellanox/doca/samples/gpunetio/subtree for this tool — it does not exist. The bundle's verbatim source path is/opt/mellanox/doca/tools/gpunetio_ib_write_bw/{client/,server/}(perSKILL.mdcompatibility block); discover withls /opt/mellanox/doca/tools/ | grep gpunetio_ib_write_bw, NOT under/opt/mellanox/doca/samples/. CUDA pairing rules live in../../libs/doca-gpunetio/CAPABILITIES.md ## Version compatibility; there is nocuda-toolkitskill in this bundle.
- Confirm the build artifacts and the environment. Per
## installand## configure. - Bring up the server first. On the host that hosts
the remote buffer, run the
serverbinary without-c. The server listens on the OOB TCP socket. It has no--gpuargument. - Confirm the server's pre-run echo. The tool logs the IB device name, the GID index, and the server role. Do not expect or synthesize a server GPU field. If the echo does not match intent, stop now; the client's connect will pin the wrong pairing.
- Bring up the client. On the second host, run the
clientbinary with-c <server-ip>, the chosen-dand--gpuon the client side, and the same--gid-indexif a non-default was used on the server. - Read the single smoke output. The result line
reports sustained throughput at the configured message
size and iteration count per
CAPABILITIES.md ## Observability. Verify the number is in a defensible order of magnitude for the GPU-NIC pair and the link's documented capacity. If anything looks off, loop back to## debug. - Decompose the throughput per
CAPABILITIES.md ## Capabilities and modesthroughput-decomposition table. Name the binding constraint (GPU compute occupancy / NIC issue rate / link saturation / wrong pairing) BEFORE quoting the number. - Plan the bulk / swept run only after the smoke is
green. The tool's message-size and iteration surface
lives in the binary's
--help; the agent does not invent sweep flags.
When recording the run for downstream consumers, write
down: pkg-config --modversion doca-gpunetio,
nvcc --version, the host OS / kernel / NUMA topology /
firmware on each side, the GPU model and PCIe address on
each side, the IB device model and PCIe address on each
side, the GID index, the exact client and server command
lines, the binding constraint identified in the
decomposition, and the full stdout for both halves. Before
persisting or sharing it, redact GPU-side handles,
memory-mmap exports, and OOB connection descriptors; retain
all other lines and ordering. The downstream
## test and ## debug workflows
depend on those fields.
test
gpunetio_ib_write_bw is a measurement tool, so its
## test verb is about testing the measurement —
confirming the numbers are sound and reproducible — not
unit-testing the tool.
## test is an iterative loop. A run that completes is
not the same as a run that produced a defensible number;
each iteration tightens one axis of measurement soundness.
The eval-loop overlay:
| Iteration trigger | What it looks like | What changes next iteration |
|---|---|---|
| Smoke completes; BW is well below link capacity | One of three candidates (or pairing) — name which before iterating | Re-walk the throughput decomposition in CAPABILITIES.md ## Capabilities and modes; confirm nvidia-smi dmon shows the SM at occupancy; confirm ibstat shows the link at expected rate. |
| BW variation fails the acceptance criterion supplied by the benchmark owner or operator | Steady-state not reached; system not at idle | Lengthen the run using the documented iteration control; confirm no background traffic on the link; confirm no concurrent CUDA workloads. Do not invent a percentage threshold. |
| BW grows with message size but does not move with kernel-side concurrency | NIC issue rate is the binding constraint | Quote the NIC-issue-rate hypothesis; confirm against the device's documented per-transport submission rate. |
| BW grows with kernel-side concurrency but does not move with message size | GPU compute occupancy is the binding constraint | Re-walk the persistent-kernel pattern in ../../libs/doca-gpunetio/CAPABILITIES.md#capabilities-and-modes. |
| BW does not move with either knob and sits well below link capacity | Likely PCIe crossover (wrong pairing) | Re-walk the GPU-NIC pairing precondition; the answer is platform-side, not benchmark-side. |
| Same invocation produces different numbers across hosts at the same DOCA version | NUMA / firmware / driver delta below DOCA | Capture the tuple on both hosts; route through doca-version TASKS.md ## test and doca-setup TASKS.md ## debug. |
| Same invocation produces different numbers on the same host across DOCA versions | This is a regression signal — provided both tuples are captured | Cross-link both baselines, name the changed fields, route to doca-version TASKS.md ## debug. |
The agent's rule: every change to the environment re-opens the loop. Re-running with a different GPU, NUMA pinning, or firmware level without re-walking the decomposition is exactly the failure mode this loop replaces.
Baseline-capture rule. The captured artifact includes
the multi-axis tuple per
CAPABILITIES.md ## Safety policy
alongside the full stdout from both halves, redacted only for
GPU-side handles, memory-mmap exports, and OOB connection
descriptors, AND the named binding constraint. Without all
of them, the baseline cannot be
regression-tested later.
debug
When gpunetio_ib_write_bw fails to build, bring up, or
produce defensible numbers, walk the
CAPABILITIES.md ## Error taxonomy
layers in order:
- Config-syntax. Confirm flags exist in
--helpon the installed binaries; do not infer. - Build-time. Re-route through
## buildand the GPUNetIO library skill's install verification. Common cause: missingdoca-gpunetio.pc, missingnvcc, GCC / GLIBC mismatch. - GPU-NIC pairing. Re-walk
CAPABILITIES.md ## Capabilities and modesGPU-NIC pairing rule; if the pairing is wrong, the fix is platform-side. - GPUNetIO-lifecycle. Most common:
nvidia_peermemnot loaded; CUDA buffer registration with DOCA happened afterdoca_ctx_start()instead of before. Route to../../libs/doca-gpunetio/CAPABILITIES.md#error-taxonomy. - RDMA-connection. Confirm GID index matches on
client and server; confirm RDMA permissions include
WRITE; confirm the mmap export was accepted. Route to
../../libs/doca-rdma/CAPABILITIES.md. - Measurement-soundness. Walk the
## testeval loop; confirm warm-up applied; name the binding constraint. - Version. Cross-cutting partial-install / mixed-
version. Walk
doca-version TASKS.md ## debug; confirm same DOCA + CUDA pair on both build hosts. - Cross-cutting. Cause is below DOCA. Hand off to
doca-debug TASKS.md ## debuganddoca-setup TASKS.md ## debug.
In every case: quote what the binaries reported. Keep full stdout and its ordering; redact only GPU-side handles, memory-mmap exports, and OOB connection descriptors before persisting or sharing it. Do not summarize a sweep into a single number.
use
Goal: turn a captured throughput from this benchmark into a class-of-workload decision — "is the GPUNetIO path the right runtime surface for my sustained-throughput workload?".
The decision shape this skill teaches:
- Quote the BW alongside the binding constraint. "X Gbit/s, NIC-issue-rate-bound" is a defensible quote; "X Gbit/s" alone is not.
- Compare against the right alternative. If the
workload class is "GPU-initiated WRITE for sustained
throughput", the alternative this benchmark answers is
"the same pattern on the CPU-initiated
perftestpath" (data from the upstream tool, captured separately); the agent does not synthesize the CPU-side number. - Apply the GPU-NIC pairing precondition to the downstream design. A throughput that wins on this benchmark only carries over to production if the production GPU-NIC pair sits on the same PCIe / NVLink topology as the test bed.
- Per-release re-verification. Every DOCA upgrade — and every CUDA Toolkit upgrade — requires re-running this benchmark before re-quoting the number. The agent does not assume a known-good number survives a DOCA-version bump or a CUDA-Toolkit bump.
- Hand off to the application's own benchmark for the final answer. This tool measures the throughput of a single RDMA WRITE stream through GPUNetIO; the user's application sits above one WRITE stream and has its own bottlenecks. The right final answer for "will my application sustain X Gbit/s in production" is an application-level benchmark.
Deferred task verbs
The verbs below are not gpunetio_ib_write_bw work and
should be routed out:
- install DOCA ⇒
doca-setup TASKS.md(and## no-installfor the NGC container path). - author a bespoke GPUNetIO-based BW benchmark ⇒
../../libs/doca-gpunetio/TASKS.mdanddoca-programming-guide TASKS.md ## build. - CPU-initiated WRITE BW ⇒ upstream
perftestib_write_bw(not in this bundle); route viadoca-public-knowledge-map. - GPU-initiated WRITE latency ⇒
../doca-gpunetio-ib-write-lat/SKILL.mdfor the GPUNetIO latency analog; for the GPI programming surface (no shipped GPI benchmark binary) ⇒doca-gpi. - hardware-touching changes the benchmark surfaced a
need for (NIC firmware burn, BFB reflash, kernel
command-line changes for IOMMU mode, hugepage
reservation changes) ⇒
doca-hardware-safety.