Official agent skill

Doca Gpunetio Ib Write Bw

by NVIDIA in NVIDIA/skills

A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests…

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Doca Gpunetio Ib Write Bw

skills CLI
$ npx skills add NVIDIA/skills --skill doca-gpunetio-ib-write-bw -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills doca-gpunetio-ib-write-bw --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/doca-gpunetio-ib-write-bw .claude/skills/doca-gpunetio-ib-write-bw && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
doca-gpunetio-ib-write-bw
GitHub stars
3.5k
Token cost
~4.2k tokens
SKILL.md length
1,864 words
Files
8
Skills in repo
386
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests…

  • Works in 3 steps: Read this SKILL.md first to confirm the… → **For what the tool measures, the… → **For step-by-step workflows — install,…
  • The user is building
  • SKILL.md covers Example questions this skill…, Audience, Language scope and When to load this skill, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Doca Gpunetio Ib Write Bw is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use this skill when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests through the doca-gpunetio device-side surface to measure sustained GPU-driven WRITE bandwidth on a GPU+IB-device pair. Trigger even when the user does not explicitly mention "doca-gpunetio-ib-write-bw" or "GPUNetIO" — typical implicit phrasings include "measure WRITE BW when the GPU posts the WRs", "BW swings between runs on the…

Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `BENCHMARK.md`, `CAPABILITIES.md` and `SKILLCARD.yaml`). Compatibility notes: Requires DOCA SDK on Linux with a BlueField DPU or ConnectX NIC, NVIDIA GPU, CUDA toolkit and nvcc, loaded nvidiapeermem, and an InfiniBand RNIC paired with…

It sits in AI & LLM Engineering, covering GPU and accelerator computing. It works with CUDA and NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user is building
  • Even when the user does not explicitly mention doca-gpunetio-ib-write-bw
  • GPUNetIO — typical implicit phrasings include measure WRITE BW when the GPU posts the WRs
  • BW swings between runs on the same flags

Example prompts

  • “doca-gpunetio-ib-write-bw”
  • “GPUNetIO”
  • “measure WRITE BW when the GPU posts the WRs”
  • “/doca-gpunetio-ib-write-bw”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires DOCA SDK on Linux with a BlueField DPU or ConnectX NIC, NVIDIA GPU, CUDA toolkit and nvcc, loaded `nvidia_peermem`, and an InfiniBand RNIC paired with the GPU. Uses `pkg-config` for doca-gpunetio, doca-rdma, and doca-common, plus the installed gpunetio_ib_write_bw sources. Run only on a trusted, non-shared IB fabric during the benchmark window.

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Read this SKILL.md first to confirm the user's
  2. **For what the tool measures, the surface-selection
  3. **For step-by-step workflows — install, configure,

What it can do on your machine

Read from SKILL.md and the folder at commit dfdd080. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires DOCA SDK on Linux with a BlueField DPU or ConnectX NIC, NVIDIA GPU, CUDA toolkit and nvcc, loaded `nvidia_peermem`, and an InfiniBand RNIC paired with the GPU. Uses `pkg-config` for doca-gpunetio, doca-rdma, and doca-common, plus the installed gpunetio_ib_write_bw sources. Run only on a trusted, non-shared IB fabric during the benchmark window.

    From compatibility in the SKILL.md frontmatter.

Context cost

Doca Gpunetio Ib Write Bw loads about 4.2k tokens when it runs. Until then it costs about 253 tokens; SKILL.md has 1,864 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~253
When it runs · the whole SKILL.md, loaded when a task matches
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit dfdd080, republished under its Apache-2.0 licence (© NVIDIA). 1,864 words, ~4,217 tokens.

Download SKILL.mdSave it as .claude/skills/doca-gpunetio-ib-write-bw/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
doca-gpunetio-ib-write-bw
description
Use this skill when the user is building, running, or interpreting the doca/tools/gpunetio_ib_write_bw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests through the doca-gpunetio device-side surface to measure sustained GPU-driven WRITE bandwidth on a GPU+IB-device pair. Trigger even when the user does not explicitly mention "doca-gpunetio-ib-write-bw" or "GPUNetIO" — typical implicit phrasings include "measure WRITE BW when the GPU posts the WRs", "BW swings between runs on the same flags", "is the NIC saturated or am I CPU-bound on the CUDA kernel", "meson compile fails for the GPUNetIO bw tool", "nvidia_peermem isn't picking up my GPU buffer", or "GPU-initiated WRITE throughput vs CPU-initiated perftest". Refuse and route elsewhere for general doca-gpunetio library work, DOCA install, the GPU-initiated WRITE latency analog, the CPU-initiated upstream perftest, or application-level end-to-end throughput — those belong to other skills.
compatibility
Requires DOCA SDK on Linux with a BlueField DPU or ConnectX NIC, NVIDIA GPU, CUDA toolkit and nvcc, loaded `nvidia_peermem`, and an InfiniBand RNIC paired with the GPU. Uses `pkg-config` for doca-gpunetio, doca-rdma, and doca-common, plus the installed gpunetio_ib_write_bw sources. Run only on a trusted, non-shared IB fabric during the benchmark window.
license
Apache-2.0
metadata.kind
tool

DOCA GPUNetIO ib_write_bw

Where to start: This is a tool skill for the GPUNetIO- flavored ib_write_bw benchmark shipped under doca/tools/gpunetio_ib_write_bw/ (a client + server pair, built from source against the installed DOCA via meson). It measures sustained RDMA WRITE bandwidth when the WRs are posted from a CUDA kernel through the doca-gpunetio device-side surface, with the GPU on the data path. Open TASKS.md and start at ## configure for the GPU-NIC pairing precondition and the build pattern; jump to ## run for the smoke-before-bulk flow. Open CAPABILITIES.md when the question is what this tool actually measures, how the result decomposes (GPU occupancy vs NIC issue rate vs link saturation), or how the result reads against the GPI sister tool and the upstream CPU-initiated perftest ib_write_bw. If DOCA is not installed yet, route to doca-setup first; if the user is still deciding between the GPI and GPUNetIO programming surfaces, the picture in ../../libs/doca-gpunetio/CAPABILITIES.md#capabilities-and-modes and ../../libs/doca-gpi/CAPABILITIES.md#capabilities-and-modes is the first stop.

Example questions this skill answers well

The CLASSES of doca-gpunetio-ib-write-bw questions this skill is built to answer, each with one worked example. The class is the load-bearing piece; the worked example is one instance.

  • "What sustained RDMA-WRITE bandwidth can the GPUNetIO path deliver on this GPU-NIC pair?" — worked example: "measure sustained WRITE BW between two hosts with an H100 + ConnectX-7 on each side". Answered by the GPU-NIC pairing precondition in CAPABILITIES.md ## Capabilities and modes
  • "Where is the bottleneck — GPU compute occupancy, NIC issue rate, or link saturation?" — worked example: "I see 120 Gbit/s on a 200 Gbit/s link; is the NIC saturated, am I CPU-bound on the client, or is the CUDA kernel not driving enough WRs in flight?". Answered by the throughput-decomposition rules in CAPABILITIES.md ## Observability
  • "How does the result differ from the classic CPU- initiated perftest ib_write_bw?" — worked example: "my team has a CPU-initiated WRITE BW number on this same NIC; should I expect the GPUNetIO number to match or be different?". Answered by the "GPU-initiated path adds (or removes) overhead vs the CPU-initiated path" rule in CAPABILITIES.md ## Capabilities and modes.
  • "Is the doca-gpunetio path the right surface for my sustained-throughput workload class?" — worked example: "my application streams sensor data from GPU memory at line rate to a remote consumer". Answered by the "when GPUNetIO is the right surface vs GPI vs CPU- initiated" rule in CAPABILITIES.md ## Capabilities and modes
  • "My BW number swings between runs. What do I check before quoting it?" — worked example: "three runs at the same flags gave 145, 187, and 160 Gbit/s; is the benchmark noisy or is my platform inconsistent?". Answered by the measurement-soundness rules in CAPABILITIES.md ## Error taxonomy layer 5 + the steady-state guidance in TASKS.md ## test.
  • "What version of DOCA + CUDA Toolkit do I need for this binary to build and run?" — worked example: "my install has DOCA at one semver and CUDA at another; will the ToT-shipped gpunetio_ib_write_bw even link?". Answered by the version overlay in CAPABILITIES.md ## Version compatibility which cross-links the canonical detection chain in doca-version.

Audience

This skill serves external developers and performance engineers who need a reproducible measurement of sustained RDMA WRITE bandwidth when the WRs are posted from a CUDA kernel through doca-gpunetio, on the user's actual install and GPU-NIC pair. Concretely:

  • A developer comparing the GPUNetIO path against the GPI path or the host-initiated perftest-style path before committing an application design to one of them.
  • A platform operator validating a tuning change (NUMA pinning, GPU PCIe placement, IB device choice, GID index, NIC firmware burn) by re-running this benchmark against the new state.
  • An SRE / performance engineer producing a "this is the GPUNetIO-driven WRITE BW on this GPU-NIC pair today" artifact downstream consumers can cite.
  • An AI agent answering "is the doca-gpunetio path a win for my sustained-throughput workload class" honestly — with a measured number, the build + invocation that produced it, and the GPU + NIC + DOCA version that scopes it — rather than guessing from datasheet headlines.

It is not for users debugging the doca-gpunetio library itself (route to ../../libs/doca-gpunetio/SKILL.md), and not a substitute for the perftest upstream ib_write_bw (which measures CPU-initiated WRITE BW).

Language scope

The doca-gpunetio-ib-write-bw tool is shipped as C plus a CUDA .cu translation unit under doca/tools/gpunetio_ib_write_bw/, split into a client/ subtree and a server/ subtree. The verified surface (per client/{main.c,common.h,common.c,kernel.cu,perftest.c} and server/{main.c,common.h,common.c,perftest.c}): host-side build via meson against the installed DOCA pkg-config modules (doca-gpunetio, doca-rdma, doca-common); the device-side build via nvcc against the DOCA GPU NetIO device-side header set; the OOB descriptor exchange via a TCP socket between client and server. There is no Python / Rust / Go binding — the tool is a pair of CLI binaries. The skill's job is to keep the operator-side workflow language-neutral; the device-side CUDA surface is not wrappable in another language.

When to load this skill

Load this skill when the user is — or the agent needs to — build and run the gpunetio_ib_write_bw client + server on real hosts with DOCA installed plus a CUDA Toolkit matched to the DOCA install, and a GPU + IB device pair on the host's PCIe topology. Concretely:

  • Measuring sustained kernel-initiated RDMA WRITE bandwidth between two hosts (or a host and a BlueField DPU) with the GPUNetIO surface.
  • Deciding whether the GPUNetIO path is the right runtime surface for a class of workload vs the GPI programming surface (the doca-gpi library — doca/tools/ ships no GPI benchmark binary) or the classic CPU-initiated perftest path.
  • Capturing a documented baseline (build + invocation + DOCA version + GPU + NIC + as-deployed environment + numbers) for later regression hunts.
  • Diagnosing a build / link / run failure that surfaces the GPUNetIO + RDMA bring-up sequence under this tool's shipped scaffolding.

Do not load this skill for general DOCA orientation, library API work, or installation. For those, use doca-public-knowledge-map, ../../libs/doca-gpunetio/SKILL.md, or doca-setup. Do not load it for application-level end-to-end throughput either — this benchmark measures the WR-submission path through GPUNetIO, not the user's full pipeline.

Show full SKILL.md (865 more words)Show less

What this skill provides

This is a thin loader. Substantive material lives in two companion files:

  • CAPABILITIES.md — what the tool measures (the sustained-WRITE-BW primitive driven by a client-side CUDA kernel through doca-gpunetio), the runtime-surface selection rule (GPUNetIO vs GPI vs CPU-initiated), the GPU-NIC pairing precondition, the throughput-decomposition guide (GPU compute occupancy vs NIC issue rate vs link saturation), the version overlay (DOCA .pc PLUS CUDA Toolkit), the layered error taxonomy (config-syntax / build-time / GPU-NIC- pairing / GPUNetIO-lifecycle / RDMA-connection / measurement-soundness / version / cross-cutting), the observability surface (stdout report, DOCA log levels, OOB-socket exchange), and the safety overlay (the "GPU-side handle is a credential" rule from doca-gpunetio; the cross-cutting hardware-safety meta-policy).
  • TASKS.md — step-by-step workflows for the in-scope task verbs: install (preconditions — DOCA install, CUDA Toolkit, GPU + NIC pair, OOB connectivity), configure (build-tree under doca/tools/gpunetio_ib_write_bw/ and the meson build wrapping the shipped DOCA), build (the meson setup + meson compile pattern from the public DOCA build documentation), modify (do not patch the shipped tool source; modify the invocation and the surrounding environment instead), run (smoke- before-bulk; client + server bring-up order; reading the per-iteration report), test (the eval loop — steady-state, NUMA placement, NIC saturation cross- check), debug (walk the error taxonomy layer by layer), use (how a BW result feeds a class-of- workload decision), plus a Deferred task verbs block routing out-of-scope questions.

The skill assumes a host where DOCA is already installed, a CUDA Toolkit matched to the install is present, and the operator has whatever privileges the public install profile expects for binding a doca_dev, a doca_gpu, and an OOB TCP socket.

What this skill deliberately does not ship

This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:

  • Specific flag strings or expected throughput numbers beyond what the tool's shipped --help and main.c ARGP registration establish. The flag surface is small (device name, GPU PCIe address, GID index, server IP on the client side); the agent re-reads the binary's --help on the installed version before quoting flag strings. Throughput numbers are device-, firmware-, version-, and topology-specific.
  • Pre-written DOCA GPUNetIO or CUDA kernel source code that would compete with the shipped tool tree. The shipped client/{main.c,kernel.cu,perftest.c,common.{c,h}} and server/{main.c,perftest.c,common.{c,h}} files are the verified worked example; the agent's job is to route the user there and prescribe minimum-diff modification per the universal modify-a-sample workflow in doca-programming-guide.
  • Wrappers, parsers, or scripts in any language that consume the tool's stdout. The output format is small and documented in CAPABILITIES.md ## Observability; if the user wants to script against it, the right answer is "read the live source, write the parser against your installed binary".
  • A samples/, bindings/, or reference/ subtree. This is a thin loader for a shipped tool tree; substantive material lives in the source tree and in the GPUNetIO library docs.

Loading order

  1. Read this SKILL.md first to confirm the user's question is in scope (the user actually wants to measure sustained kernel-initiated WRITE BW through GPUNetIO, not learn GPUNetIO as a library or do a CPU-initiated measurement).
  2. For what the tool measures, the surface-selection rule against the GPI sister tool and the CPU-initiated perftest, the throughput-decomposition guide, the version overlay, the error taxonomy, the observability surface, and the safety overlay, see CAPABILITIES.md.
  3. For step-by-step workflows — install, configure, build, modify, run, test, debug, use — see TASKS.md.
  • ../../libs/doca-gpunetio/SKILL.md — the library this tool wraps. The per-GPU doca_gpu context, the GPU-visible doca_gpu_eth_* and RDMA-side handles, the CUDA-side persistent-kernel pattern, the dual capability-discovery rule (DOCA cap-query AND cudaGetDeviceProperties), and the env preconditions (nvidia_peermem loaded, CUDA buffers registered with DOCA) live there.
  • ../../libs/doca-rdma/SKILL.md — the underlying RDMA library. The RDMA queue this tool binds is created and connected via doca-rdma; the queue lifecycle, transport type (RC vs UC vs UD), permission matrix, and connection method are owned there.
  • ../../libs/doca-verbs/SKILL.md — the raw-verbs escape hatch beneath doca-rdma / doca-gpunetio. This tool stays on the higher-level surfaces; doca-verbs is the right place only if the user needs a specific WR flag / QP attribute the GPUNetIO + RDMA surfaces do not expose.
  • ../doca-gpunetio-ib-write-lat/SKILL.md — the latency analog of this tool. Same physical operation; same runtime framework; different metric class (BW vs latency). The two together carry the full GPUNetIO-side throughput / latency picture.
  • doca-gpi — the GPI programming surface (CUDA-kernel-initiated RDMA), the alternative runtime framework for the same physical operation. doca/tools/ ships no GPI ib_write_lat / ib_write_bw benchmark binary, so the GPI comparison is against the library surface, not a sibling tool. The selection rule in CAPABILITIES.md ## Capabilities and modes is the decision aid.
  • doca-version — the canonical version-detection chain, four-way match rule, NGC container semantics, and headers-win-over-docs rule. The ## Version compatibility section in this skill is a thin overlay; the body lives there.
  • doca-setup — env preparation, install verification, GPU + CUDA Toolkit pairing, nvidia_peermem load, hugepages, NUMA, and the I have no install yet path with the public NGC DOCA container.
  • doca-public-knowledge-map — routing to the public DOCA documentation set (DOCA GPU NetIO, DOCA RDMA pages on docs.nvidia.com) and the docs.nvidia.com/cuda/ pointer for the CUDA Toolkit.
  • doca-debug — the cross-cutting debug ladder. The tool surfaces its own error taxonomy; when the cause is below DOCA, the taxonomy hands off here.
  • doca-hardware-safety — the bundle-wide hardware-safety meta-policy. The ## Safety policy overlay cross-links it.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files in skills/doca-gpunetio-ib-write-bw of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • CAPABILITIES.md
  • SKILLCARD.yaml
  • TASKS.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit dfdd080

Compare with similar skills

Doca Gpunetio Ib Write Bw next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Doca Gpunetio Ib Write Bw compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Doca Gpunetio Ib Write Bw this skillNVIDIA/skills3.5k—~4.2kAutomated safety check: PassApache-2.0
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone
Cv DeployLMIXR/CV_Deployment_skill170—~547Automated safety check: PassNone
Triton SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Hyperpod Version Checkerawslabs/agent-plugins915—~910Automated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    170 GitHub stars~547 tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Triton Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    915 GitHub stars~910 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    328 GitHub stars~3.5k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check: notes

More from NVIDIA/skills

All 386 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about Doca Gpunetio Ib Write Bw

What does Doca Gpunetio Ib Write Bw do?

A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests…. Doca Gpunetio Ib Write Bw is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use this skill when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests through the doca-gpunetio device-side surface to measure sustained GPU-driven WRITE bandwidth on a GPU+IB-device pair.

When should I use Doca Gpunetio Ib Write Bw?

Doca Gpunetio Ib Write Bw fits situations like: the user is building; even when the user does not explicitly mention doca-gpunetio-ib-write-bw; GPUNetIO — typical implicit phrasings include measure WRITE BW when the GPU posts the WRs; BW swings between runs on the same flags.

How do I install Doca Gpunetio Ib Write Bw in Claude Code?

Run `npx skills add NVIDIA/skills --skill doca-gpunetio-ib-write-bw -a claude-code`. Or copy the skill folder (skills/doca-gpunetio-ib-write-bw in NVIDIA/skills) into .claude/skills/doca-gpunetio-ib-write-bw in your project. Claude Code loads it when a task matches its description.

How do I install Doca Gpunetio Ib Write Bw in Codex?

Run `npx skills add NVIDIA/skills --skill doca-gpunetio-ib-write-bw -a codex`. Or copy the skill folder (skills/doca-gpunetio-ib-write-bw in NVIDIA/skills) into .agents/skills/doca-gpunetio-ib-write-bw in your project. Codex loads it when a task matches its description.

Can I use Doca Gpunetio Ib Write Bw in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill doca-gpunetio-ib-write-bw -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doca-gpunetio-ib-write-bw, .gemini/skills/doca-gpunetio-ib-write-bw, .github/skills/doca-gpunetio-ib-write-bw and .opencode/skills/doca-gpunetio-ib-write-bw in your project.

What does Doca Gpunetio Ib Write Bw need to run?

SKILL.md names no scripts, command-line tools or credentials: Doca Gpunetio Ib Write Bw is instructions for the agent only. Our summary lists: Python 3. Compatibility (from SKILL.md): Requires DOCA SDK on Linux with a BlueField DPU or ConnectX NIC, NVIDIA GPU, CUDA toolkit and nvcc, loaded `nvidia_peermem`, and an InfiniBand RNIC paired with the GPU. Uses `pkg-config` for doca-gpunetio, doca-rdma, and doca-common, plus the installed gpunetio_ib_write_bw sources. Run only on a trusted, non-shared IB fabric during the benchmark window. .

Does Doca Gpunetio Ib Write Bw access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Doca Gpunetio Ib Write Bw safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Doca Gpunetio Ib Write Bw use?

Doca Gpunetio Ib Write Bw is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Doca Gpunetio Ib Write Bw use?

About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Doca Gpunetio Ib Write Bw?

Skills that share tags, products or a category with Doca Gpunetio Ib Write Bw: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 925 stars), Cv Deploy (LMIXR/CV_Deployment_skill, 170 stars) and Triton Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Doca Gpunetio Ib Write Bw?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,546 GitHub stars. The repository holds 386 skills in this directory. The repository was last updated on October 9, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.