Harness Bench
ruvnet/ruflo
Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness.
Run docabench (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or…
$ npx skills add NVIDIA/skills --skill doca-bench -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills doca-bench --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/doca-bench .claude/skills/doca-bench && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "doca-bench" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench into .claude/skills/doca-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/doca-benchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill doca-bench -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills doca-bench --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/doca-bench .agents/skills/doca-bench && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "doca-bench" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench into .agents/skills/doca-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill doca-bench -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills doca-bench --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/doca-bench .cursor/skills/doca-bench && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "doca-bench" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench into .cursor/skills/doca-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/doca-bench--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill doca-bench -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills doca-bench --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/doca-bench .gemini/skills/doca-bench && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "doca-bench" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench into .gemini/skills/doca-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills doca-benchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill doca-bench -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/doca-bench .github/skills/doca-bench && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "doca-bench" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench into .github/skills/doca-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill doca-bench -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills doca-bench --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/doca-bench .opencode/skills/doca-bench && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "doca-bench" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/doca-bench into .opencode/skills/doca-bench/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "doca-bench", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
doca-benchRun docabench (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or…
Doca Bench is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Run docabench (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or BlueField Arm. Use it to discover enabled benchmark libraries, capture a reproducible command/version/device/environment baseline, compare stable runs against a declared tolerance, or diagnose configuration, device-binding, workload-precondition, and measurement failures. Trigger for requests such as measuring…
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files (for example `BENCHMARK.md`, `CAPABILITIES.md` and `SKILLCARD.yaml`). Compatibility notes: Requires DOCA SDK ≥ 2.7.0 installed at /opt/mellanox/doca on Linux (Ubuntu 22.04/24.04 or RHEL/SLES) with a BlueField DPU or ConnectX NIC attached and the…
The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Requires DOCA SDK ≥ 2.7.0 installed at /opt/mellanox/doca on Linux (Ubuntu 22.04/24.04 or RHEL/SLES) with a BlueField DPU or ConnectX NIC attached and the `doca_bench` binary present at /opt/mellanox/doca/tools/doca_bench. Companion app must run on the far side for remote-memory / RDMA / Eth scenarios; host and BlueField-Arm execution both supported.
From compatibility in the SKILL.md frontmatter.
Doca Bench loads about 3.5k tokens when it runs. Until then it costs about 181 tokens; SKILL.md has 1,553 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,553 words, ~3,461 tokens.
.claude/skills/doca-bench/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.doca_bench)Where to start: This is a tool skill for invoking doca_bench,
the cross-library micro-benchmark harness. Open
TASKS.md and start at
## configure for the three-axis decision
(target library × workload shape × measurement axis), then
## run for the smoke-before-bulk flow. Open
CAPABILITIES.md when the question is what
doca_bench can measure, which DOCA libraries it can drive, or
how to interpret throughput / latency / op-rate output without
fooling yourself on warm-up or steady-state. If DOCA is not
installed yet, route to
doca-setup first; if the install
version is < 2.7.0, doca_bench is not shipped on this host.
The CLASSES of doca_bench questions this skill is built to answer,
each with one worked example. The class is the load-bearing piece;
the worked example is one instance.
CAPABILITIES.md ## Capabilities and modesTASKS.md ## run. The same shape answers
"send-side throughput of DOCA RDMA" — doca_bench is
cross-library, not single-library.doca_bench actually drive on this
install?" — worked example: "is doca_sha enumerable on a
granular-build install". Answered by the built-in query system
surfaced in
CAPABILITIES.md ## Capabilities and modesTASKS.md ## configure step 2
(probe-before-bench). Empty enumeration = library not installed,
not bench failure.CAPABILITIES.md ## Error taxonomy
layer 5 + TASKS.md ## test (the eval-loop
overlay treats warm-up / steady-state / outliers as
re-iteration triggers, not one-shot facts).doca_bench shows
zero ops for AES-GCM but doca_caps says the device supports
it". Answered by the layered error taxonomy in
CAPABILITIES.md ## Error taxonomy
(config-syntax → device-binding → library-precondition →
workload-precondition → measurement-soundness → version →
cross-cutting) + TASKS.md ## debug.TASKS.md ## test (capture command line +
version + device + as-deployed environment alongside the
numbers; quoting numbers without the four-tuple is the
cross-version regression-hunt failure mode).doca_bench returns nothing for library X — what does that
mean?" — worked example: "empty output for DOCA SHA".
Answered by the empty-output interpretation rules in
TASKS.md ## debug +
CAPABILITIES.md ## Error taxonomy.
Re-route through
doca-caps for the coarse
per-device per-library capability ground truth, then back
into bench once the capability is confirmed present.This skill serves external operators, developers, and AI agents who need a reproducible, vendor-supported way to measure DOCA library performance on the user's actual install and device. Concretely:
doca_bench baseline against the new state.It is not for users debugging the doca_bench source code,
and not a substitute for the live public DOCA Bench guide on
docs.nvidia.com.
doca_bench is shipped as a tool (a single CLI binary plus a
companion app for the remote half of remote-memory / RDMA / Eth
scenarios), not a library you link against. The skill uses the
same kind: tool three-file shape as the rest of the bundle so
the agent's task-verb contract
(configure / build / modify / run / test / debug) is uniform
across libraries, services, and tools — even when individual
verbs collapse to a routing stub for a shipped binary.
Load this skill when the user is — or the agent needs to — invoke
doca_bench on a real host with DOCA ≥ 2.7.0 installed (or
inside the public NGC DOCA container with the equivalent version)
to measure performance of a DOCA library. Concretely:
tools/bench/doca_bench/configuration.hpp are not
interchangeable.TASKS.md ## debug).Do not load this skill for general DOCA orientation, library
API work, or installation. For those, use
doca-public-knowledge-map,
the matching libs/<library> skill, or
doca-setup. Do not load it for
application-level end-to-end benchmarking either — doca_bench
measures the DOCA library surface, not the user's application
above it.
This is a thin loader. Substantive material lives in two companion files:
CAPABILITIES.md — what doca_bench can measure (the
cross-library scope, the three-axis configuration model, the
documented operating modes, the warm-up / pipeline / multi-core
concepts that constrain measurement soundness), the version
overlay (doca-bench-specific facts on top of the canonical
doca-version rules), the layered error taxonomy
(config-syntax / device-binding / library-precondition /
workload-precondition / measurement-soundness / version /
cross-cutting), the observability surface (screen + CSV
output, real-time stats, query system), and the safety
posture (the public guide's "not for production" warning,
the host vs BlueField execution rule, the companion-app
attack surface).TASKS.md — step-by-step workflows for the in-scope task
verbs: configure (the three-axis decision + the
probe-before-bench step), build (route to install — the
binary is shipped, the companion app is shipped), modify
(refuse — do not patch the bench binary; modify the bench
invocation instead), run (the smoke-before-bulk flow),
test (the eval loop — warm-up, steady-state, outliers,
cross-version), debug (walk the error taxonomy layer by
layer), plus a Deferred task verbs block routing
out-of-scope questions and a Command appendix of
doca_bench-specific invocation classes.The skill assumes a host where DOCA ≥ 2.7.0 is already installed
(or the public NGC DOCA container is running at an equivalent
version) and the operator has whatever permissions the public
guide requires for doca_bench to bind devices and allocate
resources on their platform.
This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:
--help on the installed version are the
authoritative answer. Inventing a flag is the most common
hallucination failure for this skill.doca_bench CSV or stdout. The output formats are documented;
if a user wants to script against them, the right answer is
"read the live guide, write the parser against your installed
version".samples/ or reference/ subtree. This is a thin
loader for a documented CLI; substantive material lives on
the public page and in --help.SKILL.md first to confirm the user's question is
in scope (the user actually wants to invoke doca_bench for
measurement, not learn about a DOCA library in general).doca_bench measures, the three-axis model, the
version overlay, the error taxonomy, observability surface,
and safety posture, see CAPABILITIES.md.configure, build, modify, run, test,
debug — see TASKS.md.doca-public-knowledge-map
— routing to the public DOCA Bench page on docs.nvidia.com
and the rest of the public DOCA documentation set.doca-version — the canonical
version-detection chain, four-way match rule, NGC container
semantics, and headers-win-over-docs rule. The
## Version compatibility section in this skill is a thin
overlay on top of doca-version; the body lives there.doca-structured-tools-contract
— the bundle-wide contract for structured-output helper tools.
Bench-runner / bench-snapshot executables that satisfy the
detect-prefer-fallback-report loop are deferred to PR2; the
contract is consumed here in advance so the
## Command appendix in TASKS.md is infra-aware
from PR1.doca-setup — env preparation,
install verification, hugepages, NUMA awareness, and the I
have no install yet path with the public NGC DOCA container.doca-debug — the cross-cutting
debug ladder. Bench surfaces its own error taxonomy in
CAPABILITIES.md ## Error taxonomy;
when the cause turns out to be below DOCA (driver, firmware,
NUMA), the bench taxonomy hands off to doca-debug.doca-caps — the sibling DOCA tool
for the coarse per-device per-library capability snapshot.
Bench probes capability at finer grain via its own query
system; doca_caps is the cheaper first step to confirm the
device is even visible to DOCA.libs/<library> skill — e.g.
doca-comch,
doca-compress — for
the workload-side preconditions, capability-query rules, and
error-taxonomy overlays of the library under test. Bench
drives the library; the library skill explains what
"healthy" means for it.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files in skills/doca-bench of NVIDIA/skills.
Open the folder on GitHubat commit 0e0d506
Doca Bench next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Doca Bench this skillNVIDIA/skills | 3.5k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | |
| Harness Benchruvnet/ruflo | 74k | — | ~586 | Automated safety check: Notes | MIT | |
| Bench Readgithub/awesome-copilot | 40k | — | ~747 | Automated safety check: Pass | MIT | |
| Benchddalcu/mlx-serve | 1.8k | 1 repos | ~1.1k | Automated safety check: Pass | Custom licence | |
| Terminal Bench Looppaperclipai/paperclip | 98k | — | ~6.3k | Automated safety check: Pass | MIT | |
| Harness Security Benchruvnet/ruflo | 74k | — | ~1.1k | Automated safety check: Notes | MIT |
ruvnet/ruflo
Manage @metaharness/darwin bench suites — bench create <repo scaffolds a JSON suite from a repo's test corpus; bench verify <suite.json checks suite well-formedness.
github/awesome-copilot
Read artifacts from the shared bench — the workspace where desks leave findings, verdicts, and work products for each other and the operator.
ddalcu/mlx-serve
mlx-serve benchmarking methodology — bench.sh/llmprobe usage, comparison-trap rules (same-methodology cells only, spec-decode variance, thermal lies, engine naming), perf-claim etiquette.
paperclipai/paperclip
Run one Terminal-Bench task through a bounded Paperclip smoke/diagnosis/fix loop.
ruvnet/ruflo
Run @metaharness/darwin security bench (upstream "Darwin Shield" / ADR-155) — evolves a champion security-detection harness against a 10-vuln / 9-decoy corpus and grades it on…
autonomous-ai/openharness
Plan real experiments and analyze supplied continuous measurements in Lab Bench.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Run docabench (DOCA 2.7.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or…. Doca Bench is an agent skill from NVIDIA/skills, published by the product's own GitHub organization.0 or newer) to measure throughput, bulk latency, precision latency, or maximum bandwidth for RDMA, Compress, AES-GCM, SHA, DMA, EC, Ethernet, Comch, or GPUNetIO on a host or BlueField Arm.
Doca Bench fits situations like: discover enabled benchmark libraries; capture a reproducible command/version/device/environment baseline; compare stable runs against a declared tolerance; diagnose configuration.
Run `npx skills add NVIDIA/skills --skill doca-bench -a claude-code`. Or copy the skill folder (skills/doca-bench in NVIDIA/skills) into .claude/skills/doca-bench in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill doca-bench -a codex`. Or copy the skill folder (skills/doca-bench in NVIDIA/skills) into .agents/skills/doca-bench in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill doca-bench -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/doca-bench, .gemini/skills/doca-bench, .github/skills/doca-bench and .opencode/skills/doca-bench in your project.
SKILL.md names no scripts, command-line tools or credentials: Doca Bench is instructions for the agent only. Compatibility (from SKILL.md): Requires DOCA SDK ≥ 2.7.0 installed at /opt/mellanox/doca on Linux (Ubuntu 22.04/24.04 or RHEL/SLES) with a BlueField DPU or ConnectX NIC attached and the `doca_bench` binary present at /opt/mellanox/doca/tools/doca_bench. Companion app must run on the far side for remote-memory / RDMA / Eth scenarios; host and BlueField-Arm execution both supported. .
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Doca Bench is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Doca Bench: Harness Bench (ruvnet/ruflo, 74k stars), Bench Read (github/awesome-copilot, 40k stars), Bench (ddalcu/mlx-serve, 1.8k stars) and Terminal Bench Loop (paperclipai/paperclip, 98k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.