Graphsignal
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
Evaluate whether an existing hot path is a credible NVIDIA Warp candidate.
$ npx skills add NVIDIA/skills --skill warp-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills warp-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/warp-eval .claude/skills/warp-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "warp-eval" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/warp-eval into .claude/skills/warp-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "warp-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/warp-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill warp-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills warp-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/warp-eval .agents/skills/warp-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "warp-eval" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/warp-eval into .agents/skills/warp-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "warp-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill warp-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills warp-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/warp-eval .cursor/skills/warp-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "warp-eval" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/warp-eval into .cursor/skills/warp-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "warp-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/warp-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill warp-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills warp-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/warp-eval .gemini/skills/warp-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "warp-eval" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/warp-eval into .gemini/skills/warp-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "warp-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills warp-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill warp-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/warp-eval .github/skills/warp-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "warp-eval" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/warp-eval into .github/skills/warp-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "warp-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill warp-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills warp-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/warp-eval .opencode/skills/warp-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "warp-eval" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/warp-eval into .opencode/skills/warp-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "warp-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
warp-evalEvaluate whether an existing hot path is a credible NVIDIA Warp candidate.
Warp Eval is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Evaluate whether an existing hot path is a credible NVIDIA Warp candidate. Use for irregular or spatial queries, particle or geometry simulation, branch-heavy loops, many small launches, host fallbacks, or large intermediates. CPU-only code and absent GPU dependencies are normal unless NVIDIA is prohibited. Exclude required cross-vendor or CPU-only deployment, vendor-lowered dense or NN layers, general Warp API questions, and already-selected Warp kernels. Contribution policy alone is not exclusion.
Its SKILL.md is about 4.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 51 other files, including scripts, reference files and assets (for example `BENCHMARK.md`, `assets/warp-evaluation-report-template.md` and `config/skillspector-baseline.yaml`). Compatibility notes: Screening, static evaluation and reporting need no GPU. Measuring Warp requires an NVIDIA CUDA GPU, the target project's dependencies and a representative…
It sits in AI & LLM Engineering. It works with NVIDIA AI Platform and CUDA. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
9 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Screening, static evaluation and reporting need no GPU. Measuring Warp requires an NVIDIA CUDA GPU, the target project's dependencies and a representative workload; without them, abort before profiling.
From compatibility in the SKILL.md frontmatter.
Warp Eval loads about 4.9k tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 129 tokens; SKILL.md has 2,435 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 2,435 words, ~4,918 tokens.
.claude/skills/warp-eval/SKILL.md (or your agent's skills folder). This skill also uses 43 other files; get the full folder from GitHub.Collect reproducible evidence about how a narrow seam in an existing codebase would behave in NVIDIA Warp. Report facts; the user decides.
Name Warp as the option under evaluation in the first line, state that no adoption recommendation will follow, and do not treat the triggering performance request as authorization to experiment.
Measured evaluations produce warp-evaluation-report/: the report, one
independently applicable diff per solution, the drivers, and raw results. Never
modify production code. Exits before measured work create no directory.
These override any local reasoning.
AWAITING INTENT).Static screening requires the target repository and its stated product constraints. Measured work additionally requires explicit authorization, an NVIDIA CUDA GPU, target-project dependencies and a representative workload.
latest docs track development, not the shipped release. Pin to the
target project's own Warp version; if it has none, use the current stable and
say which.hasattr(wp, "mesh_query_point") is False on a release that has it. Probe
the stub file or that version's docs before concluding a builtin is absent.enable_backward=False (kernel, module or global) removes adjoint
codegen. If nothing differentiates through the seam, set it before measuring
compile cost.| State | Reached when | Report directory |
|---|---|---|
ABORT | Any gate establishes that the evaluated seam cannot satisfy the stated scope | No if nothing was measured; otherwise preserve the evidence already collected |
AWAITING INTENT | The environment gates turn on a fact only the user has | No — one question, both branches concrete |
AWAITING AUTHORIZATION | A candidate pattern survives stage 1 | No — early findings plus scope/resource preview |
INCOMPLETE | Authorized work cannot obtain representative evidence required for the scoped evaluation | Yes — preserve collected evidence and name the one missing artifact |
| Report delivered | Authorized work produced measurements, or stopped after the report directory existed | Yes — facts per seam and regime, with missing evidence explicit |
Use this exact shape for an early exit:
ABORT — Gate <letter>: <cited fact>; <why the scoped Warp evaluation cannot proceed>.
Do not name a preferred alternative.
Reporting rules:
references/evidence-and-reporting.md.
A delivered directory follows the template's fixed order: schema and provenance,
authorization and evaluation state, stage census, one B<n> evidence section
per seam/regime, caveats, then environment/reproduction. solutions/,
benchmarks/ and results/ contain every linked artifact.
Gate exit: ABORT — Gate A: deployment.md requires one implementation with AMD, Apple and NVIDIA parity; a Warp-specific path cannot satisfy this scope.
Surviving candidate: Name the seam and pattern, label inferred facts as assumptions, state that no gate has fired, preview the profile/baseline/Warp prototype/benchmark scope and its cost, then ask the separate intent and authorization questions from references/authorization-checkpoint.md.
Required: the target repository and a performance, memory or scale problem with a candidate seam. Optional: explicit deployment/packaging constraints, existing profiles or logs, representative datasets and acceptance criteria. Prompt constraints take precedence over repository policy/configuration, then existing logs. User corrections override inference. Never substitute an assumption for a stated fact or measurement.
| Script | Purpose | Arguments |
|---|---|---|
scripts/driver-template.py | Copy once per bottleneck; define workloads and variants | Edit placeholders, then run the copied driver |
scripts/measure.py | Import from drivers for synchronized timing, memory and isolated cases | Python API; do not execute directly |
scripts/validate_report_schema.py | Validate the delivered report and evidence links | <report-directory> |
Use run_script("scripts/validate_report_schema.py", args=["warp-evaluation-report"])
when supported; otherwise invoke the script with Python and the report directory.
INCOMPLETE and name the missing
artifact; do not invent data or fire Gate F.Every stage before the last can end the evaluation. Stop as soon as a gate fires; do not gather evidence that cannot change the scoped facts.
| Gate | Fires when |
|---|---|
| A | Production is stated CPU-only or to need non-NVIDIA portability, with no acceptable optional CUDA path |
| B | Data must cross the host/device boundary per small or infrequent call and the boundary cannot be widened |
| C | The region is dense tensor algebra already mapped to a tuned framework or vendor library |
| D | A mature CUDA implementation already meets the contract, and no non-performance objective was requested |
| E | A stated policy blocks Warp's dependency, compilation, cache or fallback obligations |
| F | Representative evidence proves the region too small a share of its requested metric for any backend to move it |
ABORT.Requires explicit authorization and a settled intent question.
ABORT the affected scope under Gate F, preserve the evidence already
collected, and stop. The existing authorization already covered this
materiality check; do not ask for authorization again.INCOMPLETE, and stop. This is missing
evidence, not Gate F and not ABORT. An invented workload cannot prove
materiality.Protocol: references/benchmark-protocol.md.
Record per candidate: source, bottleneck evidence, objective, narrow seam, mechanism Warp could change, strongest incumbent, risks, acceptance threshold, cheapest falsifying experiment. Screen against references/target-patterns.md; if none survives, write the report and stop.
Define values, dtypes, shapes, devices, errors, mutation, ordering, ties, capacity/overflow, topology/degeneracy, tolerances, required gradients, streams, ownership, aliasing, invalidation, concurrency, capture, teardown and fallback.
ABORT before prototyping if the proposed seam cannot satisfy a required
contract.Hazards and adversarial checks: references/semantic-contract.md.
Algorithm before backend: (1) a better or output-sensitive algorithm; (2) chunking, tiling, sparse output, layout, rematerialization; (3) the incumbent framework's compiler and native primitives; (4) what the project already depends on — its own accelerator backend, a parallel idiom it ships but leaves off, or a capability an existing dependency exposes and nobody wired up; (5) only then narrow Warp.
Close every in-project route before prototyping Warp:
| State | What it takes to claim it |
|---|---|
| measured | timed through the same boundary as the baseline |
| absent | a cited declaration, symbol table or missing flag proves it is unavailable |
| waived | you asked the user and they chose to skip it; record their words |
A capability present but unbound is reachable, not absent. If exposing it costs no more than the planned Warp seam, measure it first. Ladder details: references/baselines.md. Waived routes do not block stage 6; every route not explicitly waived must be measured or evidenced absent before the Warp prototype begins.
ABORT for the affected scope.not measured.null_test, and one-time costs both separately and amortized.
Serialize GPU measurements under an exclusive device lock. Below 1.5× is no
measured difference.ABORT for a seam and regime whose predeclared end-to-end performance or
memory requirement fails, after preserving the measurements.Full protocol: references/benchmark-protocol.md.
pass, fail, not measured, not available, no representative data,
unknown or n/a only where a stated criterion makes that status objective.
Missing evidence remains missing.complete only when every evaluated seam and regime includes
an end-to-end Warp measurement. Use aborted — <gate and scope> when a gate
fired, or incomplete — <missing evidence and scope> when representative
evidence was unavailable. Preserve everything collected. Never deliver an
incumbent-only report as a completed warp-eval.results/.uv run python scripts/validate_report_schema.py <report-directory>© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 43 other files (scripts, references, assets) in skills/warp-eval of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
Warp Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Warp Eval this skillNVIDIA/skills | 3.5k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 911 | — | ~2.8k | Automated safety check: Pass | None | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Setup Workshop Nemoclawbrevdev/workshop-build-an-agent | 144 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 |
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
brevdev/workshop-build-an-agent
Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.
LMIXR/CV_Deployment_skill
基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Evaluate whether an existing hot path is a credible NVIDIA Warp candidate. Warp Eval is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Evaluate whether an existing hot path is a credible NVIDIA Warp candidate.
Warp Eval fits situations like: spatial queries; geometry simulation; branch-heavy loops; many small launches.
Run `npx skills add NVIDIA/skills --skill warp-eval -a claude-code`. Or copy the skill folder (skills/warp-eval in NVIDIA/skills) into .claude/skills/warp-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill warp-eval -a codex`. Or copy the skill folder (skills/warp-eval in NVIDIA/skills) into .agents/skills/warp-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill warp-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/warp-eval, .gemini/skills/warp-eval, .github/skills/warp-eval and .opencode/skills/warp-eval in your project.
Going by SKILL.md and its folder, Warp Eval needs Python for the scripts in its folder and the command-line tools its instructions call (uv). Our summary lists: Python 3. Compatibility (from SKILL.md): Screening, static evaluation and reporting need no GPU. Measuring Warp requires an NVIDIA CUDA GPU, the target project's dependencies and a representative workload; without them, abort before profiling. .
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Warp Eval is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.9k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 20k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Warp Eval: Graphsignal (graphsignal/graphsignal, 257 stars), LLM Torch Profiler Trace Analysis (BBuf/AI-Infra-Auto-Driven-SKILLS, 911 stars), Optimize Op (CVCUDA/CV-CUDA, 2.7k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.