Aoti Debug
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
A skill your agent uses when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that…
$ npx skills add NVIDIA/skills --skill doca-bench-extension -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills doca-bench-extension --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/doca-bench-extension .claude/skills/doca-bench-extension && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "doca-bench-extension" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench-extension into .claude/skills/doca-bench-extension/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench-extension", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/doca-bench-extensionType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill doca-bench-extension -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills doca-bench-extension --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/doca-bench-extension .agents/skills/doca-bench-extension && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "doca-bench-extension" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench-extension into .agents/skills/doca-bench-extension/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench-extension", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill doca-bench-extension -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills doca-bench-extension --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/doca-bench-extension .cursor/skills/doca-bench-extension && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "doca-bench-extension" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench-extension into .cursor/skills/doca-bench-extension/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench-extension", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/doca-bench-extension--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill doca-bench-extension -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills doca-bench-extension --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/doca-bench-extension .gemini/skills/doca-bench-extension && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "doca-bench-extension" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench-extension into .gemini/skills/doca-bench-extension/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench-extension", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills doca-bench-extensionInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill doca-bench-extension -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/doca-bench-extension .github/skills/doca-bench-extension && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "doca-bench-extension" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench-extension into .github/skills/doca-bench-extension/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench-extension", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill doca-bench-extension -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills doca-bench-extension --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/doca-bench-extension .opencode/skills/doca-bench-extension && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "doca-bench-extension" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench-extension into .opencode/skills/doca-bench-extension/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench-extension", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
doca-bench-extensionA skill your agent uses when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that…
Doca Bench Extension is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use this skill when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that doca-bench loads to measure a workload class its built-in modes do not cover, with docabenchcuda as the shipped reference exemplar. Trigger even when the user does not say "doca-bench-extension" or "docabenchcuda" — typical implicit phrasings include "no built-in doca-bench mode fits my workload", "how do I benchmark a CUDA…
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `BENCHMARK.md`, `CAPABILITIES.md` and `SKILLCARD.yaml`). Compatibility notes: Requires DOCA SDK installed at /opt/mellanox/doca on Linux (Ubuntu 22.04/24.04 or RHEL/SLES) with a BlueField DPU or ConnectX NIC. Source tree…
It sits in Development. It works with CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires DOCA SDK installed at /opt/mellanox/doca on Linux (Ubuntu 22.04/24.04 or RHEL/SLES) with a BlueField DPU or ConnectX NIC. Source tree: `/opt/mellanox/doca/tools/bench_extension/` (underscored, NOT kebab-case); the built shared library `libdoca_bench_cuda_impl.so` lands in the platform libdir on a binary install. Also needs `pkg-config doca-common` and, for the GPU-side reference exemplar (DOCA GPUNetIO RX/TX kernels), an NVIDIA GPU + matching CUDA toolkit.
From compatibility in the SKILL.md frontmatter.
Doca Bench Extension loads about 4k tokens when it runs. Until then it costs about 250 tokens; SKILL.md has 1,743 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,743 words, ~3,978 tokens.
.claude/skills/doca-bench-extension/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Where to start: This is a tool skill for the extension /
plug-in framework that augments
doca-bench — NOT a workload-shape
skill on its own. Open TASKS.md and start at
## configure to commit to the three-axis
decision (workload class is genuinely outside doca-bench's
built-in modes × extension API surface fits × parent-tool
co-load is acceptable), then ## build for
how a custom extension is compiled and laid out, then
## run for how doca-bench discovers and
invokes the extension, then ## test for the
smoke-before-bulk loop the agent applies to every new
extension. Open CAPABILITIES.md when the
question is what an extension can do that built-in
doca-bench modes cannot, what the extension API surface
looks like in broad strokes (the DOCA_EXPERIMENTAL C entry
points the shipped reference exposes), how the
build / registration / discovery flow works, or how the
extension's lifetime is bounded by the parent doca-bench
invocation. If doca-bench itself is the question, route to
doca-bench. If the question is
"which built-in doca-bench mode do I pick?", that is also
doca-bench — extensions are the
exit ramp for workloads built-in modes do not cover.
<X> — does doca-bench measure it
natively, or do I need an extension?" — the
extension-vs-built-in decision question. The agent walks
the user back to doca-bench's
built-in mode inventory FIRST and only routes to the
extension framework when no built-in mode applies.doca_bench_cuda extension under
/opt/mellanox/doca/tools/bench_extension/doca_bench_cuda/ as the
reference exemplar and walks the operator through its
API surface and build shape.doca-bench actually discover and load my
custom extension at runtime? Is it a versioned shared
library? What does my entry-point need to look like?" —
the build / registration / discovery flow question. The
agent walks the Meson-built shared library shape, the
versioning, and the parent-tool's runtime discovery path
(which the agent does NOT invent from memory — the
shipped extension's meson.build and the public DOCA Bench
documentation on docs.nvidia.com are the source of
truth).DOCA_EXPERIMENTAL.
What does that mean for my extension's stability across
DOCA releases? Am I going to have to rebuild it every
release?" — the experimental-surface and version
compatibility question.doca-bench actually loaded it, called
into it, and that the call returned the data the parent
tool expected?" — the smoke-before-bulk question.doca-bench says it
cannot find / load / call it. Where do I look first?" —
the layered-debug question that distinguishes
build-failures, load-failures, registration-mismatches,
and runtime-call-failures.Experienced AI agents and platform / performance engineers
who already use doca-bench for
the built-in workload modes and now have a workload class
that the built-in modes do not cover. Readers are expected
to be comfortable with native build systems (Meson, in this
codebase), shared-library packaging on Linux, and the
DOCA_EXPERIMENTAL API stability contract. If the user
asks about GPU-side benchmarking via the shipped
doca_bench_cuda reference extension, the reader is also
expected to be familiar with DOCA GPUNetIO and CUDA toolchain
basics — those domains live in their own skills, not here.
This skill is NOT for:
doca-bench's built-in modes — that is
doca-bench;A doca-bench extension surfaces as:
.so with
soversion matching the DOCA release), built via the
doca-bench-extension Meson rules in the shipped
/opt/mellanox/doca/tools/bench_extension/meson.build and the
per-extension subdirectory (the reference exemplar is
doca_bench_cuda/).DOCA_EXPERIMENTAL-marked C entry
points that the parent doca-bench invokes — i.e. the
API surface declared in the extension's header file.
The shipped doca_bench_cuda/doca_bench_cuda.h is the
reference for what that surface shape looks like in
practice (*_init, *_device_query,
*_device_synchronize, and per-workload kernel-start
entry points such as *_start_nop_kernel,
*_start_eth_recv_kernel, *_start_eth_send_kernel,
*_start_eth_bidir_kernel).doca_bench_cuda_kernel_settings,
doca_bench_cuda_eth_rx_kernel_settings,
doca_bench_cuda_eth_tx_kernel_settings,
doca_bench_cuda_eth_bidir_kernel_settings carry block
counts, threads-per-block, RX / TX queues, buffer
address / mkey / size, a stop flag, and a stats
pointer).The skill itself is Markdown. The user's extension source is whatever language the workload requires (C / C++ / CUDA in the reference case). The agent does NOT prescribe a language beyond what the shipped reference demonstrates.
Load doca-bench-extension when ANY of the following is
true:
doca-bench-extension, the
doca_bench_cuda reference extension, the
doca_bench_cuda_impl shared library, or any of the
DOCA_EXPERIMENTAL extension entry points;doca-bench TASKS.md ## configure)
that none of doca-bench's built-in workload modes
measures the class they want, and an extension is the
exit ramp;doca_bench_cuda reference into a custom GPU-side
workload extension;doca-bench cannot find / load
/ call a custom extension they built.Co-load this skill with:
doca-bench (the parent tool —
ALWAYS co-loaded; extensions only have value as
plug-ins into doca-bench);doca-version (the
DOCA_EXPERIMENTAL surface is versioned with DOCA; the
extension's soversion is the DOCA soversion; the
four-way version match applies);doca-gpunetio when
the extension is GPU-side and uses GPUNetIO RX / TX
queues like the reference exemplar (route the GPUNetIO
semantics there, not here);doca-debug and
doca-setup for the
env-side debug ladder (driver, firmware, CUDA toolkit,
dynamic linker).Do NOT load this skill when the user's workload fits a
doca-bench built-in mode — extensions add cost (build
toolchain, version churn, the experimental-surface
contract); the built-in modes are always the first answer to
try.
Three companion files in this directory, each owning a different question shape:
SKILL.md — this file. Audience, scope,
loading order, related skills. Routes everything else.CAPABILITIES.md — what an extension
can do that the built-in modes cannot, what the API
surface looks like in broad strokes, how the
build / registration / discovery flow works, what
versions it ships in (including the
DOCA_EXPERIMENTAL-stability overlay on top of
doca-version), the layered error taxonomy,
observability, and the safety policy overlay.TASKS.md — the procedural verbs
(configure, build, run, test, debug, etc.) plus
a doca-bench-extension-specific command appendix and
the agent-side use workflow that consumes the captured
extension run.The combined skill teaches an AI agent to drive the
extension-author-and-wire-in class of doca-bench
questions: confirm an extension is needed at all; locate
the shipped reference exemplar
(/opt/mellanox/doca/tools/bench_extension/doca_bench_cuda/); copy
its build + API surface shape; build a versioned shared
library that matches the DOCA release; smoke that the
parent doca-bench actually loads it; diagnose layered
failures when it does not.
doca-bench's built-in workload modes.
That belongs to doca-bench.
This skill is the exit ramp for what the built-in modes
do not cover; it does not duplicate the parent's mode
inventory.DOCA_EXPERIMENTAL entry-point names beyond
what the shipped reference declares. The shipped
doca_bench_cuda/doca_bench_cuda.h on the user's
install is the reference for what the surface shape
looks like; the agent does not assert other extensions
exist with specific signatures.doca_bench_cuda reference IS the canonical layout;
rewriting it here would drift from the source of truth.
The agent points the operator at the shipped tree and
walks the operator through adapting it.doca-bench uses to
locate and load extensions (search path, naming
convention, registration call) lives in the public DOCA
Bench documentation on docs.nvidia.com and the
installed doca-bench binary. The agent points the
operator there rather than asserting a mechanism from
memory.doca-gpunetio;
this skill cross-links rather than duplicates.docs.nvidia.com; this skill does not duplicate it.doca-bench invocation details
unrelated to extensions. The parent's CLI flags,
pipeline shapes, and built-in workload classes belong
to doca-bench.When a doca-bench-extension question arrives:
doca-bench is reachable
on the user's install — if not, route to
doca-setup;doca-bench's built-in modes covers
the workload class — if any of them does, route back
to doca-bench TASKS.md ## configure
and stop. Extensions are the exit ramp, not the first
answer;CAPABILITIES.md to commit to
the three-axis decision and walk the reference
exemplar's API surface shape;TASKS.md and walk
## configure → ## build → ## run → ## test → ## debug
in that order; do NOT start with ## run without the
build precondition step.Cross-link conventions follow the bundle's relative path
contract from tools/<X>/:
doca-bench — the parent tool.
ALWAYS co-loaded. Extensions are plug-ins into
doca-bench; they do not replace it, they do not have a
standalone CLI, they do not measure anything without the
parent invoking them. Every question on this skill
presupposes the parent.doca-version — the
DOCA_EXPERIMENTAL surface is versioned with DOCA; the
extension's soversion matches the DOCA release per
the shipped meson.build. The four-way version match
applies; rebuilding the extension across DOCA upgrades
is the rule, not the exception.doca-gpunetio —
when the extension is GPU-side and uses GPUNetIO RX /
TX queues like the reference doca_bench_cuda. Route
the GPUNetIO semantics there.doca-setup — DOCA install
posture (does doca-bench exist? does the
doca_bench_cuda_impl reference library exist? is the
CUDA toolchain installed when needed?).doca-debug — the
cross-cutting debug ladder for env-side issues (dynamic
linker, library search path, CUDA driver / toolkit,
firmware).doca-public-knowledge-map
— routing to the public DOCA Bench / DOCA GPUNetIO pages
on docs.nvidia.com and the release notes for the
documented extension lifecycle / discovery mechanism.doca-structured-tools-contract
— the agent's detect → prefer → fall back → report
contract for the structured helpers
(doca-env --json, doca-capability-snapshot,
version-matrix.json) the build / load preconditions
rely on.doca-hardware-safety
— the canonical hardware-safety meta-policy that
CAPABILITIES.md ## Safety policy
overlays. Extensions are external code loaded into
doca-bench; the safety implications of loading
experimental code into a benchmark that touches the
dataplane / device are real.This skill assumes the user has built shared libraries on Linux before and knows what a Meson build is. Background material on those topics belongs in the toolchain docs, not in this skill.
© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files in skills/doca-bench-extension of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Doca Bench Extension next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Doca Bench Extension this skillNVIDIA/skills | 3.5k | — | ~4k | Automated safety check: Pass | Apache-2.0 | |
| Aoti Debugpytorch/pytorch | 104k | 1 repos | ~1.7k | Automated safety check: Pass | Custom licence | |
| Debug Distributed Hangsgl-project/sglang | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | |
| CUTLASS FMHA Incremental Rebuildmicrosoft/onnxruntime | 22k | — | ~1.3k | Automated safety check: Pass | MIT | |
| ONNX Runtime Source Buildmicrosoft/onnxruntime | 22k | — | ~1.4k | Automated safety check: Pass | MIT | |
| The Art of Debuggingstas00/the-art-of-debugging | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 |
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
sgl-project/sglang
Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).
microsoft/onnxruntime
Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.
microsoft/onnxruntime
Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
pytorch/executorch
Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
A skill your agent uses when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that…. Doca Bench Extension is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Use this skill when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that doca-bench loads to measure a workload class its built-in modes do not cover, with docabenchcuda as the shipped reference exemplar.
Doca Bench Extension fits situations like: the operator is authoring; with docabenchcuda as the shipped reference exemplar; even when the user does not say doca-bench-extension; docabenchcuda — typical implicit phrasings include no built-in doca-bench mode fits my workload.
Run `npx skills add NVIDIA/skills --skill doca-bench-extension -a claude-code`. Or copy the skill folder (skills/doca-bench-extension in NVIDIA/skills) into .claude/skills/doca-bench-extension in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill doca-bench-extension -a codex`. Or copy the skill folder (skills/doca-bench-extension in NVIDIA/skills) into .agents/skills/doca-bench-extension in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill doca-bench-extension -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doca-bench-extension, .gemini/skills/doca-bench-extension, .github/skills/doca-bench-extension and .opencode/skills/doca-bench-extension in your project.
SKILL.md names no scripts, command-line tools or credentials: Doca Bench Extension is instructions for the agent only. Compatibility (from SKILL.md): Requires DOCA SDK installed at /opt/mellanox/doca on Linux (Ubuntu 22.04/24.04 or RHEL/SLES) with a BlueField DPU or ConnectX NIC. Source tree: `/opt/mellanox/doca/tools/bench_extension/` (underscored, NOT kebab-case); the built shared library `libdoca_bench_cuda_impl.so` lands in the platform libdir on a binary install. Also needs `pkg-config doca-common` and, for the GPU-side reference exemplar (DOCA GPUNetIO RX/TX kernels), an NVIDIA GPU + matching CUDA toolkit. .
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Doca Bench Extension is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Doca Bench Extension: Aoti Debug (pytorch/pytorch, 104k stars), Debug Distributed Hang (sgl-project/sglang, 37k stars), CUTLASS FMHA Incremental Rebuild (microsoft/onnxruntime, 22k stars) and ONNX Runtime Source Build (microsoft/onnxruntime, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.