All skills
nvidia avatar

/doca-collectx-deployment

@f64fa0e
by NVIDIA Corporationnvidia/skills3.5k stars
424

Use this skill to deploy and operate a CollectX (clx) based DOCA telemetry collector on a host or BlueField — wiring providers / counters into the collector, running the collection daemon, and shaping its exporters (Prometheus pull, Fluent Bit push, NetFlow, file / IPC) so the metrics actually leave the box. Trigger even when the user never says CollectX or clx — implicit phrasings: {collector emits nothing downstream}, {add a provider to the clx collector}, {turn on the Prometheus endpoint}, {ship counters to Fluent Bit from the DPU}, {daemon starts but no schema rows appear}. This skill owns the CollectX collection mechanism plus the operator's own doca-telemetry / doca-telemetry-exporter usage; it ROUTES the productized DOCA Telemetry Service (DTS) to public docs (AGENTS.md Non-goal #7), the reader API to doca-telemetry, and the publisher API to doca-telemetry-exporter. Refuse to invent clx symbols, provider names, schema fields, flags, or config paths — describe the class and route to the live source.

Use this Skill: https://skilld.dev/gh/nvidia/skills/doca-collectx-deployment

This session only. Nothing lands on disk.

referencesdetails.md

≈2.3k tokens on demand. Your agent reads this file only when SKILL.md points to it.

doca-collectx-deployment — reference detail

Moved out of SKILL.md to keep the loader under the per-file size budget. This is supporting detail, not routing logic.

Example questions this skill answers well

The CLASSES of collector-deployment questions this skill is built to answer, each with one worked example. The class is the load-bearing piece; the worked example is one instance.

  • "People keep saying 'DOCA telemetry' as if it's one thing — which surface am I actually deploying?" — worked example: "I see clx, doca-telemetry, doca-telemetry-exporter, and 'DOCA Telemetry Service' all referenced; which one is the collector I'm standing up?". Answered by the four-surface decomposition in CAPABILITIES.md ## Capabilities and modes
    • the scope-boundary section in SKILL.md — the agent draws the clx-collector / reader-library / publisher-library / productized-DTS split BEFORE any config.
  • "My collector daemon starts but no schema rows appear for a provider." — worked example: "the collector is running but the provider I enabled produces nothing downstream". Answered by the no-provider-rows layer in CAPABILITIES.md ## Error taxonomy
    • the gate-before-commit rule in CAPABILITIES.md ## Safety policy: the device does not expose the counter under the chosen dimensions — the canonical silently-dropped metric.
  • "How do I turn on the Prometheus endpoint / ship to Fluent Bit / export NetFlow from my collector?" — worked example: "I want a Prometheus scrape target for my BlueField telemetry collector". Answered by the export-backend table in CAPABILITIES.md ## Capabilities and modes (class level: pull vs push vs local) + the TASKS.md ## configure step 5 / run-side verification — with the exact flag and config-path spelling routed to the live config and the public DTS / Telemetry guide.
  • "The collector is running but my downstream Prometheus / Grafana shows nothing — where do I start?" — worked example: "exporter looks up, consumer is empty". Answered by the layered ladder in CAPABILITIES.md ## Error taxonomy (exporter-silent → downstream-skew → transport) + the end-to-end smoke in TASKS.md ## test: a running daemon is not proof of deployment.
  • "Is the counter family I want even exposed on this device before I commit it to the collector config?" — worked example: "will this PCIe-diagnostic counter family work on my BlueField-3 under these property dimensions?". Answered by the gate-before-commit rule in CAPABILITIES.md ## Safety policy
  • "I think I actually need the DOCA Telemetry Service container, not a hand-rolled collector — where do I go?" — worked example: "I want the turnkey DTS service auto-started on my BlueField". Answered by the scope boundary in SKILL.md: DTS-as-deployed is externally productized (Non-goal #7); the agent routes to the public DTS guide rather than synthesizing DTS config.

What this skill deliberately does not ship

This skill is agent guidance, not a templates / sample-config bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:

  • The doca-telemetry library API. The hardware-counter reader API (per-domain doca_telemetry_<domain> contexts on a doca_dev) is owned by doca-telemetry. This skill routes there for any reader-API question; it does not re-document the per-domain lifecycle, the cap-query rule, or the reader error taxonomy.
  • The doca-telemetry-exporter library API. The application-side publisher API (schema / source / type, register-before-emit) is owned by doca-telemetry-exporter. This skill routes there for any "how do I emit a counter from my program" question.
  • The productized DTS container. The DOCA Telemetry Service as-deployed — its packaged config schema, its built-in provider set, its kubelet manifest, its NGC image tag — is externally productized (per AGENTS.md Non-goal #7). The agent refuses to synthesize DTS config file names, provider knob names, or paths and routes to the public DTS guide.
  • Invented clx symbols, provider names, schema field names, exporter flag names, or config paths. All of these are install-bound and docs-bound; the authoritative sources are the live collector config on the target and the public DOCA Telemetry / DTS guides reached through doca-public-knowledge-map. Quoting a clx provider name or a config path from memory as authoritative is the load-bearing hallucination failure for this skill.
  • Pre-baked collector configs, sample exporter configs, sample systemd units, or sample pod-specs. Collector deployment is site-specific (which providers, which export sink, which launch posture); the safe answer is to derive the config against the live install and the public docs, with the launch posture routed to doca-bare-metal-deployment or doca-container-deployment.
  • A samples/, templates/, config/, or reference/ subtree of any kind. A mock or incomplete artifact in this skill's tree, even one labeled "reference", is misleading: operators will read it as production-ready.

Related skills

  • doca-telemetry — the hardware-counter reader library. When a source the collector samples is the operator's own program reading device counters, the reader API lives there. This skill owns the collector runtime; that skill owns the reader API.
  • doca-telemetry-exporter — the application-side publisher library. A DOCA program that feeds the collector as an external source publishes through this API. This skill owns the collector / sink side; that skill owns the publisher / source side.
  • doca-telemetry-utils — the operator-side support CLI that enumerates the diagnostic-counter schema, translates name ↔ Data ID, and probes per-device counter support before a config commits. This skill reuses its gate-before-commit discipline for the collector config.
  • doca-setup — env preparation and install verification. This skill assumes its preconditions are satisfied (DOCA installed and healthy on the collection target).
  • doca-version — canonical DOCA version-handling rules. This skill's ## Version compatibility cross-links the four-way match and adds the collector-vs-source-vs-consumer schema-version alignment overlay.
  • doca-hardware-safety — the cross-cutting meta-policy for any change touching DPU / NIC hardware state. The collector is read-only against the device; any mutating step (a privileged-data daemon needing a hardware-state change, an mlxconfig set, a firmware burn) leaves this skill for that meta-policy.
  • doca-bare-metal-deployment and doca-container-deployment — the two launch-posture skills. The collector's launch shape (foreground / service-supervised on bare metal, or container / kubelet) follows those skills; this skill owns the collection pipeline that runs inside whichever posture the operator picks.
  • doca-public-knowledge-map — the routing table to the public DOCA Telemetry guide, the public DTS guide (externally-productized row), and the installed DOCA layout. This skill does not duplicate URLs; it points at the map and adds the collector-deployment overlay. The productized DTS service is reached through that map (Non-goal #7).
  • doca-debug — the cross-cutting layered debug ladder. Collector-deployment-specific debug (no rows, exporter silent, downstream skew) layers on top of the cross-cutting host / network ladder there.

Source: SKILL.md on GitHub

No alerts2mo3 checks · Risk SAFE
  • Gen Agent Trust Hub2mo

    The skill provides safe, documentation-driven instructions for deploying NVIDIA DOCA CollectX telemetry collectors. It correctly defines operational boundaries and routes out-of-scope requests to official NVIDIA resources without any malicious patterns.

  • Socket2mo

    No alerts

  • Snyk2mo

    Risk: LOW · No issues

Signed by skilld at f64fa0e. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 months ago
metadata
{
  "kind": "library"
}
Other metadata
compatibility
No DOCA install required to read this skill (it is a deployment / operation overlay over the DOCA telemetry libraries and the CollectX collection mechanism). The hands-on steps DO require a live DOCA install at /opt/mellanox/doca on a host or BlueField, an operator account that can run the collector and reach its exporter sinks, and the public DOCA Telemetry / DTS guides on docs.nvidia.com for any concrete provider name, schema field, flag, or config path (this skill never invents those).

README badge

README badge for nvidia/skills/doca-collectx-deployment