Embeddings via 9Router
decolua/9router
Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.
Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config.
$ npx skills add NVIDIA/skills --skill rag-perf -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills rag-perf --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-perf .claude/skills/rag-perf && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "rag-perf" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/rag-perf into .claude/skills/rag-perf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-perf", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/rag-perfType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill rag-perf -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills rag-perf --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/rag-perf .agents/skills/rag-perf && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "rag-perf" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/rag-perf into .agents/skills/rag-perf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-perf", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill rag-perf -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills rag-perf --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/rag-perf .cursor/skills/rag-perf && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "rag-perf" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/rag-perf into .cursor/skills/rag-perf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-perf", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/rag-perf--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill rag-perf -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills rag-perf --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/rag-perf .gemini/skills/rag-perf && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "rag-perf" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/rag-perf into .gemini/skills/rag-perf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-perf", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills rag-perfInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill rag-perf -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/rag-perf .github/skills/rag-perf && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "rag-perf" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/rag-perf into .github/skills/rag-perf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-perf", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill rag-perf -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills rag-perf --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/rag-perf .opencode/skills/rag-perf && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "rag-perf" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/rag-perf into .opencode/skills/rag-perf/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "rag-perf", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
rag-perfPerformance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config.
RAG Perf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. Not for accuracy / RAGAS scoring (use rag-eval) or for deploying / repairing services (use rag-blueprint).
Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `BENCHMARK.md`, `eval/h100.json` and `eval/nvidia_hosted.json`). Compatibility notes: Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml)…
It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Load testing. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
ReadGrepGlobBash(ls *)Bash(python3 *)Bash(uv *)Bash(cat *)Bash(curl *)WriteEditFrom allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvpipcurlFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, pip and curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
NVIDIA_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml); reachable RAG server (default http://localhost:8081); for synthetic queries an OpenAI-compatible chat-completions endpoint is required (default http://localhost:8999/v1/chat/completions); aiperf load-test phase uses the bundled nvidia_rag endpoint plugin, registered automatically when rag-perf is installed editable.
From compatibility in the SKILL.md frontmatter.
RAG Perf loads about 4.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 1,513 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,513 words, ~4,071 tokens.
.claude/skills/rag-perf/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Drive a deployed NVIDIA RAG Blueprint server with a YAML config, run a server-side profiling pass (per-stage timing, citation quality, bottleneck inference) and an optional aiperf load test (TTFT / E2E / token & request throughput / error rate), and write a unified report. The CLI is intentionally minimal: rag-perf -c <config> plus --help / --version. Behaviour is fully config-driven; field variations belong in YAML.
uv sync --project scripts/rag-perf.uv sync --project scripts/rag-perf --extra dev (otherwise pytest-asyncio is missing and async tests error out at collection time).http://localhost:8081). For the aiperf phase, the bundled nvidia_rag endpoint plugin must be installed — pip install -e ./scripts/rag-perf registers it via the aiperf.plugins entry point.synthetic.llm_url (default http://localhost:8999/v1/chat/completions).NVIDIA_API_KEY (unlike rag-eval). The synthetic LLM endpoint may require its own auth — that's the deployment's concern.Pick a preset. The three under scripts/rag-perf/configs/ are:
quick_profile.yaml — profile-only, ~30 s. Skips load test. For fast iteration on retrieval / reranker tuning.single_run.yaml — one concurrency level, profiling + aiperf, ~2 min. Regression checks.sweep.yaml — multi-axis sweep. load.concurrency, rag.vdb_top_k, rag.reranker_top_k are all int | list[int]; any of them as a list becomes a sweep axis (Cartesian product).Edit the preset. Required: replace rag.collection_names: ["<collection_name>"] with a real collection on the deployed ingestor server. Verify the collection exists via GET /v1/collections on the ingestor. The placeholder <collection_name> validates fine but every request will fail at retrieval. Use a copied YAML preset for variants; the CLI surface is intentionally config-only.
Run. From repo root:
uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/single_run.yamlSame form for the other presets. The CLI accepts only -c / --config (required), --help, --version.
Read stdout. Every invocation prints, in order: a startup banner, a one-line summary, the fully resolved config as YAML (so the run is reproducible from terminal output), per-grid-point progress with the shlex-joined aiperf command in copy-pastable form, a rich per-point summary table (stage breakdown with bars, citation quality, bottleneck, load-test block), and finally a side-by-side comparison table auto-labelled by whichever axis varied. See references/output-and-analysis.md.
Inspect artifacts. Layout depends on run shape — flat for single-point + iterations=1, nested under iter_<i>/<point>/... otherwise. See references/output-and-analysis.md for the full directory tree, file purposes, and how to parse results.json / results.csv / report.md.
Summarise for the user. When reporting back, follow the playbook in references/output-and-analysis.md#summarising-results-to-the-user: pick the canonical result file for the run shape, build a headline table (concurrency × top-k axes × TTFT × throughput × bottleneck × citation quality), compute scaling efficiency on sweeps, always flag zero citations / non-zero error rate / suspect llm_ttft_ms / small-sample p99, and propose a concrete next-experiment YAML.
Tune. Schema is fully documented in docs/performance-benchmarking.md and the deeper-dive references below. Common knobs: turn aiperf.enabled: false for profile-only mode, increase load.iterations for variance estimation, set load.sleep_between_points_s: 60 for overnight Cartesian sweeps.
Profile-only (quickest signal on retrieval / reranker tuning):
uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/quick_profile.yamlOutput: rag-perf-results/quick_profile/run_<ts>/{profile_report.md, profile_results.json, profiling/}. The aiperf_rag_on/ directory is omitted. Filenames are profile_* because aiperf.enabled: false.
Single benchmark point with full report:
uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/single_run.yamlOutput: flat run_<ts>/{report.md, results.json, results.csv, profiling/, aiperf_rag_on/}.
Concurrency sweep:
uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/sweep.yamlOutput: nested run_<ts>/iter_1/<CR:_VDB-K:_RERANKER-K:_…>/{profiling,aiperf_rag_on}/ per point, plus aggregate report.md / results.json / results.csv at the run root.
Run unit tests:
uv sync --project scripts/rag-perf --extra dev # one-time, installs pytest-asyncio
uv run --project scripts/rag-perf python -m pytest tests/unit/test_rag_perf/load.concurrency / rag.vdb_top_k / rag.reranker_top_k accept int | list[int]; the validator requires unique list values because each value names a unique point dir.input.file and input.synthetic follow an XOR rule — both set fails validation. When neither is set, synthetic auto-fills with defaults so a bare config still validates..jsonl or .csv); other extensions are rejected.synthetic.disable_thinking: true (the default). Without it the model exhausts the token budget on chain-of-thought and content returns empty — the generator now raises with a clear message instead of substituting reasoning_content for the answer.AiperfRunner._base_aiperf_cmd in scripts/rag-perf/rag_perf/runner.py.references/ to keep this file concise.| Error / signal | Likely cause | What to do |
|---|---|---|
Configuration errors in <yaml>: • input — ... XOR rule | Both input.file and input.synthetic set | Pick one. The XOR validator runs at YAML load time. |
input.file must end in .jsonl or .csv | Extension other than .jsonl / .csv | Rename or convert. |
load.concurrency has duplicate values | e.g. [2, 2, 4] | Each concurrency maps to a unique point dir; dedupe. |
warmup_requests must be >= 1 | YAML had warmup_requests: 0 | aiperf rejects warmup=0; minimum is 1. |
LLM returned empty content (reasoning_content was populated — model exhausted its budget on chain-of-thought; raise min_query_tokens or set synthetic.disable_thinking=true). | Reasoning model used CoT and ran out of tokens | Set synthetic.disable_thinking: true (the default) or raise min_query_tokens. |
✗ All N profiling requests failed across M point(s). + exit 1 | Bad URL, server down, wrong collection | Verify target.url, rag.collection_names (the <collection_name> placeholder will hit this). |
Per-iteration ⚠ N profiling requests failed warning, run continues | Some requests timed out / errored mid-run | Check rag-server logs, raise target.timeout_s, drop concurrency. |
RuntimeError: Random synthetic query generation failed at query N: ... | LLM endpoint rejected a request mid-generation | Partial JSONL is at synthetic.jsonl_output_path; fix endpoint and re-run with reduced num_queries, or point input.file at the partial file. |
Citation count (mean): 0 and Citation relevance score: N/A for a non-empty deployment | Collection mismatch between rag.collection_names and what's actually ingested | Run curl -s http://<ingestor>:8082/v1/collections to list real collections. |
Tests error with ModuleNotFoundError: No module named 'pytest_asyncio' | Dev extras missing | uv sync --project scripts/rag-perf --extra dev. |
CI: ModuleNotFoundError: No module named 'ruamel' from tests/unit/test_rag_perf/ | rag-perf package missing from CI venv | Add uv pip install -e ./scripts/rag-perf after the top-level install in the unit-tests job. |
scripts/rag-perf/examples/queries.jsonl and scripts/rag-perf/prompts/default_prompts.yaml with repo-root-relative paths. Running from inside scripts/rag-perf/ will fail those file lookups.rag.collection_names before the first run. The presets ship with ["<collection_name>"] as a deliberate placeholder. Validation passes, retrieval fails silently for every request — manifests as Citation count (mean): 0 everywhere.load.concurrency_list, rag.vdb_top_k_list, rag.reranker_top_k_list are read-only properties that normalise scalar-or-list to a list. Use them when reasoning about the grid; the underlying YAML field is whatever the user wrote.aiperf.enabled: false changes filenames. The top-level outputs become profile_report.md / profile_results.json / profile_results.csv. The aggregate sweep table also suppresses load-test rows and the "Optimal throughput" footer.\n $ python -m aiperf profile -m ... --endpoint-type nvidia_rag ... in stdout — copy-paste runnable for reproducing a single point outside rag-perf.--endpoint-type nvidia_rag comes from the bundled plugin at scripts/rag-perf/rag_perf/plugin/nvidia_rag.py. It teaches aiperf about the RAG /v1/generate request shape and parses citations + per-stage metrics out of the SSE stream. If aiperf can't resolve nvidia_rag, rag-perf needs editable installation in the venv — re-run uv sync --project scripts/rag-perf (or uv pip install -e ./scripts/rag-perf).[1, 4] × single vdb_top_k), the dir name encodes everything: CR:1_ISL:50_OSL:512_VDB-K:20_RERANKER-K:4_Model:.... Cluster / GPU / experiment_name (output.cluster, output.gpu, output.experiment_name) are appended too — useful for diff-friendly artifact paths across machines.load.iterations > 1 repeats the entire grid. Each repetition writes to its own iter_<i>/. Aggregate CSV row count = n_points × iterations.| Piece | Location |
|---|---|
| Driver | scripts/rag-perf/rag_perf/cli.py (main is the single Click command) |
| Schema | scripts/rag-perf/rag_perf/config.py (RunConfig and sub-models) |
| Orchestrator | scripts/rag-perf/rag_perf/runner.py (BenchmarkRunner.run, RagProfiler, AiperfRunner) |
| aiperf plugin | scripts/rag-perf/rag_perf/plugin/nvidia_rag.py |
| User-facing doc | docs/performance-benchmarking.md |
| Presets | scripts/rag-perf/configs/{quick_profile,single_run,sweep}.yaml |
| Sample queries | scripts/rag-perf/examples/queries.jsonl |
| Synthetic prompts | scripts/rag-perf/prompts/default_prompts.yaml |
| Config schema details | references/config-schema.md |
| Synthetic-query generation | references/synthetic-generation.md |
| Output layout & metric semantics | references/output-and-analysis.md |
uv sync --project scripts/rag-perf (one-time per checkout).scripts/rag-perf/configs/<preset>.yaml if you want a variant; always set rag.collection_names to a real collection.uv run --project scripts/rag-perf rag-perf -c <config> from repo root.output.dir/run_<ts>/ — see references/output-and-analysis.md. For multi-point runs, results.csv has one row per (point × iteration).references/output-and-analysis.md#summarising-results-to-the-user — headline table, scaling-efficiency math for sweeps, mandatory flags for zero citations / non-zero errors / suspect llm_ttft_ms / low sample size, and a concrete next-experiment YAML.quick_profile.yaml or aiperf.enabled: false for fast iteration, then return to single_run.yaml / sweep.yaml when characterising under load.references/output-and-analysis.md for empty-citation / bottleneck=N/A patterns.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (references) in skills/rag-perf of NVIDIA/skills.
Open the folder on GitHubat commit 67a13c0
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.
RAG Perf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| RAG Perf this skillNVIDIA/skills | 3.5k | 1 repos | ~4.1k | Automated safety check: Pass | Apache-2.0 | |
| Embeddings via 9Routerdecolua/9router | 30k | — | ~604 | Automated safety check: Pass | MIT | |
| Vss Deploy Detection Tracking 3DNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~5.1k | Automated safety check: Notes | Apache-2.0 | |
| Testing Prompt Injection In RAG Pipelinesmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Gke Inferencegoogle/skills | 21k | — | ~2k | Automated safety check: Pass | Apache-2.0 | |
| Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~2.3k | Automated safety check: Pass | MIT |
decolua/9router
Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC…
mukul975/Anthropic-Cybersecurity-Skills
Probes Retrieval-Augmented Generation pipelines for indirect prompt injection via poisoned retrieved documents and embedding-space manipulation, using NVIDIA garak, Promptfoo red-team plugins, and…
google/skills
Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
MoizIbnYousaf/ai-agent-skills
Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. RAG Perf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config.
RAG Perf fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Load testing.
Run `npx skills add NVIDIA/skills --skill rag-perf -a claude-code`. Or copy the skill folder (skills/rag-perf in NVIDIA/skills) into .claude/skills/rag-perf in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill rag-perf -a codex`. Or copy the skill folder (skills/rag-perf in NVIDIA/skills) into .agents/skills/rag-perf in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill rag-perf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-perf, .gemini/skills/rag-perf, .github/skills/rag-perf and .opencode/skills/rag-perf in your project.
Going by SKILL.md and its folder, RAG Perf needs the command-line tools its instructions call (uv, pip and curl) and credentials named NVIDIA_API_KEY. Our summary lists: Python 3; A credential in NVIDIA_API_KEY. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash(ls *), Bash(python3 *), Bash(uv *), Bash(cat *), Bash(curl *), Write, Edit. Compatibility (from SKILL.md): Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml); reachable RAG server (default http://localhost:8081); for synthetic queries an OpenAI-compatible chat-completions endpoint is required (default http://localhost:8999/v1/chat/completions); aiperf load-test phase uses the bundled nvidia_rag endpoint plugin, registered automatically when rag-perf is installed editable..
SKILL.md contains no URLs. Its commands use uv, pip and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
RAG Perf is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with RAG Perf: Embeddings via 9Router (decolua/9router, 30k stars), Vss Deploy Detection Tracking 3D (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars), Testing Prompt Injection In RAG Pipelines (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Gke Inference (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.