Official agent skill

RAG Perf

by NVIDIA in NVIDIA/skills

Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install RAG Perf

skills CLI
$ npx skills add NVIDIA/skills --skill rag-perf -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills rag-perf --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-perf .claude/skills/rag-perf && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-perf
GitHub stars
3.5k
Used in
1 other repo
Token cost
~4.1k tokens
SKILL.md length
1,513 words
Files
9 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config.

  • Works in 7 steps: Pick a preset. The three under… → Edit the preset. Required: replace… → Run. From repo root → …
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Purpose, Scope, Prerequisites and Instructions, plus 6 more sections
  • Calls uv, pip and curl; needs NVIDIA_API_KEY

What it does

RAG Perf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. Not for accuracy / RAGAS scoring (use rag-eval) or for deploying / repairing services (use rag-blueprint).

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including reference files (for example `BENCHMARK.md`, `eval/h100.json` and `eval/nvidia_hosted.json`). Compatibility notes: Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml)…

It sits in AI & LLM Engineering, covering Retrieval-augmented generation and Load testing. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Load testing

Example prompts

  • “/rag-perf”

Requirements

  • Python 3
  • A credential in NVIDIA_API_KEY
  • Compatibility (from SKILL.md): Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml); reachable RAG server (default http://localhost:8081); for synthetic queries an OpenAI-compatible chat-completions endpoint is required (default http://localhost:8999/v1/chat/completions); aiperf load-test phase uses the bundled nvidia_rag endpoint plugin, registered automatically when rag-perf is installed editable.
  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash(ls *), Bash(python3 *), Bash(uv *), Bash(cat *), Bash(curl *), Write, Edit

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Pick a preset. The three under scripts/rag-perf/configs/ are
  2. Edit the preset. Required: replace rag.collection_names: [""] with a real collection on the deployed ingestor server. Verify the…
  3. Run. From repo root
  4. Read stdout. Every invocation prints, in order: a startup banner, a one-line summary, the fully resolved config as YAML (so the run is…
  5. Inspect artifacts. Layout depends on run shape — flat for single-point + iterations=1, nested under iter_//... otherwise. See…
  6. Summarise for the user. When reporting back, follow the playbook in references/output-and-analysis.md#summarising-results-to-the-user…
  7. Tune. Schema is fully documented in docs/performance-benchmarking.md and the deeper-dive references below. Common knobs: turn…

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash(ls *)
    • Bash(python3 *)
    • Bash(uv *)
    • Bash(cat *)
    • Bash(curl *)
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, pip and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NVIDIA_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml); reachable RAG server (default http://localhost:8081); for synthetic queries an OpenAI-compatible chat-completions endpoint is required (default http://localhost:8999/v1/chat/completions); aiperf load-test phase uses the bundled nvidia_rag endpoint plugin, registered automatically when rag-perf is installed editable.

    From compatibility in the SKILL.md frontmatter.

Context cost

RAG Perf loads about 4.1k tokens when it runs, and up to ~11k if it reads all its reference files. Until then it costs about 63 tokens; SKILL.md has 1,513 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~11k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,513 words, ~4,071 tokens.

Download SKILL.mdSave it as .claude/skills/rag-perf/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
rag-perf
description
Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. Not for accuracy / RAGAS scoring (use rag-eval) or for deploying / repairing services (use rag-blueprint).
allowed-tools
Read, Grep, Glob, Bash(ls *), Bash(python3 *), Bash(uv *), Bash(cat *), Bash(curl *), Write, Edit
compatibility
Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml); reachable RAG server (default http://localhost:8081); for synthetic queries an OpenAI-compatible chat-completions endpoint is required (default http://localhost:8999/v1/chat/completions); aiperf load-test phase uses the bundled nvidia_rag endpoint plugin, registered automatically when rag-perf is installed editable.
version
2.6.0
license
Apache-2.0
metadata.tool-version
0.1.0
metadata.author
NVIDIA RAG <foundational-rag-dev@exchange.nvidia.com>
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/rag
metadata.endpoint-openapi-schemas
docs/api_reference/openapi_schema_rag_server.json
metadata.argument-hint
rag-perf | aiperf | TTFT | latency | throughput | concurrency sweep | bottleneck | retrieval / reranker tuning | profile-only | synthetic queries |…
metadata.tags
nvidia, blueprint, rag, performance, benchmarking, aiperf, nvidia-rag-blueprint
metadata.languages
python, shell
metadata.frameworks
aiperf, fastapi

RAG-Perf — config-driven perf benchmark CLI

Purpose

Drive a deployed NVIDIA RAG Blueprint server with a YAML config, run a server-side profiling pass (per-stage timing, citation quality, bottleneck inference) and an optional aiperf load test (TTFT / E2E / token & request throughput / error rate), and write a unified report. The CLI is intentionally minimal: rag-perf -c <config> plus --help / --version. Behaviour is fully config-driven; field variations belong in YAML.

Scope

  • Accuracy / RAGAS scoring of answer quality → use the rag-eval skill.
  • Deploying, repairing, or configuring services (compose, helm, NIM env vars) → use the rag-blueprint skill.
  • Production monitoring / alerting — rag-perf is a one-shot benchmark tool.
  • Runtime requirement: a deployed RAG server reachable on the network.

Prerequisites

  • Repo cloned; run commands from the repo root (config paths in the presets are repo-root-relative).
  • Python 3.11+ and uv on PATH.
  • Install rag-perf into its own uv-managed venv: uv sync --project scripts/rag-perf.
  • For unit tests: install dev extras as well — uv sync --project scripts/rag-perf --extra dev (otherwise pytest-asyncio is missing and async tests error out at collection time).
  • A reachable RAG server (default http://localhost:8081). For the aiperf phase, the bundled nvidia_rag endpoint plugin must be installed — pip install -e ./scripts/rag-perf registers it via the aiperf.plugins entry point.
  • For synthetic queries: an OpenAI-compatible chat-completions endpoint reachable at synthetic.llm_url (default http://localhost:8999/v1/chat/completions).
  • rag-perf itself runs without NVIDIA_API_KEY (unlike rag-eval). The synthetic LLM endpoint may require its own auth — that's the deployment's concern.

Instructions

  1. Pick a preset. The three under scripts/rag-perf/configs/ are:

    • quick_profile.yaml — profile-only, ~30 s. Skips load test. For fast iteration on retrieval / reranker tuning.
    • single_run.yaml — one concurrency level, profiling + aiperf, ~2 min. Regression checks.
    • sweep.yaml — multi-axis sweep. load.concurrency, rag.vdb_top_k, rag.reranker_top_k are all int | list[int]; any of them as a list becomes a sweep axis (Cartesian product).
  2. Edit the preset. Required: replace rag.collection_names: ["<collection_name>"] with a real collection on the deployed ingestor server. Verify the collection exists via GET /v1/collections on the ingestor. The placeholder <collection_name> validates fine but every request will fail at retrieval. Use a copied YAML preset for variants; the CLI surface is intentionally config-only.

  3. Run. From repo root:

    bash
    uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/single_run.yaml

    Same form for the other presets. The CLI accepts only -c / --config (required), --help, --version.

  4. Read stdout. Every invocation prints, in order: a startup banner, a one-line summary, the fully resolved config as YAML (so the run is reproducible from terminal output), per-grid-point progress with the shlex-joined aiperf command in copy-pastable form, a rich per-point summary table (stage breakdown with bars, citation quality, bottleneck, load-test block), and finally a side-by-side comparison table auto-labelled by whichever axis varied. See references/output-and-analysis.md.

  5. Inspect artifacts. Layout depends on run shape — flat for single-point + iterations=1, nested under iter_<i>/<point>/... otherwise. See references/output-and-analysis.md for the full directory tree, file purposes, and how to parse results.json / results.csv / report.md.

  6. Summarise for the user. When reporting back, follow the playbook in references/output-and-analysis.md#summarising-results-to-the-user: pick the canonical result file for the run shape, build a headline table (concurrency × top-k axes × TTFT × throughput × bottleneck × citation quality), compute scaling efficiency on sweeps, always flag zero citations / non-zero error rate / suspect llm_ttft_ms / small-sample p99, and propose a concrete next-experiment YAML.

  7. Tune. Schema is fully documented in docs/performance-benchmarking.md and the deeper-dive references below. Common knobs: turn aiperf.enabled: false for profile-only mode, increase load.iterations for variance estimation, set load.sleep_between_points_s: 60 for overnight Cartesian sweeps.

Examples

Profile-only (quickest signal on retrieval / reranker tuning):

bash
uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/quick_profile.yaml

Output: rag-perf-results/quick_profile/run_<ts>/{profile_report.md, profile_results.json, profiling/}. The aiperf_rag_on/ directory is omitted. Filenames are profile_* because aiperf.enabled: false.

Single benchmark point with full report:

bash
uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/single_run.yaml

Output: flat run_<ts>/{report.md, results.json, results.csv, profiling/, aiperf_rag_on/}.

Concurrency sweep:

bash
uv run --project scripts/rag-perf rag-perf -c scripts/rag-perf/configs/sweep.yaml

Output: nested run_<ts>/iter_1/<CR:_VDB-K:_RERANKER-K:_…>/{profiling,aiperf_rag_on}/ per point, plus aggregate report.md / results.json / results.csv at the run root.

Run unit tests:

bash
uv sync --project scripts/rag-perf --extra dev   # one-time, installs pytest-asyncio
uv run --project scripts/rag-perf python -m pytest tests/unit/test_rag_perf/

Limitations

  • The CLI is config-only: author or copy YAML to vary a parameter.
  • load.concurrency / rag.vdb_top_k / rag.reranker_top_k accept int | list[int]; the validator requires unique list values because each value names a unique point dir.
  • input.file and input.synthetic follow an XOR rule — both set fails validation. When neither is set, synthetic auto-fills with defaults so a bare config still validates.
  • File-based input format is inferred from extension only (.jsonl or .csv); other extensions are rejected.
  • Synthetic generation streams each query to disk as it completes (failure-resilient) but fails fast on the first LLM error — partial JSONL is preserved. Re-run after fixing the endpoint.
  • Reasoning models (Nemotron Omni, Qwen-Reasoning) require synthetic.disable_thinking: true (the default). Without it the model exhausts the token budget on chain-of-thought and content returns empty — the generator now raises with a clear message instead of substituting reasoning_content for the answer.
  • aiperf-specific knobs outside the YAML surface (request rate distribution, GPU telemetry config, etc.) require editing AiperfRunner._base_aiperf_cmd in scripts/rag-perf/rag_perf/runner.py.
  • Procedural detail lives under references/ to keep this file concise.

Troubleshooting

Error / signalLikely causeWhat to do
Configuration errors in <yaml>: • input — ... XOR ruleBoth input.file and input.synthetic setPick one. The XOR validator runs at YAML load time.
input.file must end in .jsonl or .csvExtension other than .jsonl / .csvRename or convert.
load.concurrency has duplicate valuese.g. [2, 2, 4]Each concurrency maps to a unique point dir; dedupe.
warmup_requests must be >= 1YAML had warmup_requests: 0aiperf rejects warmup=0; minimum is 1.
LLM returned empty content (reasoning_content was populated — model exhausted its budget on chain-of-thought; raise min_query_tokens or set synthetic.disable_thinking=true).Reasoning model used CoT and ran out of tokensSet synthetic.disable_thinking: true (the default) or raise min_query_tokens.
✗ All N profiling requests failed across M point(s). + exit 1Bad URL, server down, wrong collectionVerify target.url, rag.collection_names (the <collection_name> placeholder will hit this).
Per-iteration ⚠ N profiling requests failed warning, run continuesSome requests timed out / errored mid-runCheck rag-server logs, raise target.timeout_s, drop concurrency.
RuntimeError: Random synthetic query generation failed at query N: ...LLM endpoint rejected a request mid-generationPartial JSONL is at synthetic.jsonl_output_path; fix endpoint and re-run with reduced num_queries, or point input.file at the partial file.
Citation count (mean): 0 and Citation relevance score: N/A for a non-empty deploymentCollection mismatch between rag.collection_names and what's actually ingestedRun curl -s http://<ingestor>:8082/v1/collections to list real collections.
Tests error with ModuleNotFoundError: No module named 'pytest_asyncio'Dev extras missinguv sync --project scripts/rag-perf --extra dev.
CI: ModuleNotFoundError: No module named 'ruamel' from tests/unit/test_rag_perf/rag-perf package missing from CI venvAdd uv pip install -e ./scripts/rag-perf after the top-level install in the unit-tests job.
Show full SKILL.md (481 more words)Show less

Gotchas

  • Run from repo root. Preset configs reference scripts/rag-perf/examples/queries.jsonl and scripts/rag-perf/prompts/default_prompts.yaml with repo-root-relative paths. Running from inside scripts/rag-perf/ will fail those file lookups.
  • CLI is config-only. Edit the YAML or copy a preset for URL, concurrency, collection, and similar fields.
  • Always edit rag.collection_names before the first run. The presets ship with ["<collection_name>"] as a deliberate placeholder. Validation passes, retrieval fails silently for every request — manifests as Citation count (mean): 0 everywhere.
  • load.concurrency_list, rag.vdb_top_k_list, rag.reranker_top_k_list are read-only properties that normalise scalar-or-list to a list. Use them when reasoning about the grid; the underlying YAML field is whatever the user wrote.
  • aiperf.enabled: false changes filenames. The top-level outputs become profile_report.md / profile_results.json / profile_results.csv. The aggregate sweep table also suppresses load-test rows and the "Optimal throughput" footer.
  • Resolved-config dump is verbose (50+ lines) — expected. It's what makes terminal output a self-contained reproducer; don't filter it out in scripts.
  • The aiperf shell command is logged before each subprocess. Look for \n $ python -m aiperf profile -m ... --endpoint-type nvidia_rag ... in stdout — copy-paste runnable for reproducing a single point outside rag-perf.
  • --endpoint-type nvidia_rag comes from the bundled plugin at scripts/rag-perf/rag_perf/plugin/nvidia_rag.py. It teaches aiperf about the RAG /v1/generate request shape and parses citations + per-stage metrics out of the SSE stream. If aiperf can't resolve nvidia_rag, rag-perf needs editable installation in the venv — re-run uv sync --project scripts/rag-perf (or uv pip install -e ./scripts/rag-perf).
  • Sweep-mode point-name collision. When two points differ only in concurrency (e.g. [1, 4] × single vdb_top_k), the dir name encodes everything: CR:1_ISL:50_OSL:512_VDB-K:20_RERANKER-K:4_Model:.... Cluster / GPU / experiment_name (output.cluster, output.gpu, output.experiment_name) are appended too — useful for diff-friendly artifact paths across machines.
  • load.iterations > 1 repeats the entire grid. Each repetition writes to its own iter_<i>/. Aggregate CSV row count = n_points × iterations.

Source of truth

PieceLocation
Driverscripts/rag-perf/rag_perf/cli.py (main is the single Click command)
Schemascripts/rag-perf/rag_perf/config.py (RunConfig and sub-models)
Orchestratorscripts/rag-perf/rag_perf/runner.py (BenchmarkRunner.run, RagProfiler, AiperfRunner)
aiperf pluginscripts/rag-perf/rag_perf/plugin/nvidia_rag.py
User-facing docdocs/performance-benchmarking.md
Presetsscripts/rag-perf/configs/{quick_profile,single_run,sweep}.yaml
Sample queriesscripts/rag-perf/examples/queries.jsonl
Synthetic promptsscripts/rag-perf/prompts/default_prompts.yaml
Config schema detailsreferences/config-schema.md
Synthetic-query generationreferences/synthetic-generation.md
Output layout & metric semanticsreferences/output-and-analysis.md

Agent playbook

  1. Sync deps: uv sync --project scripts/rag-perf (one-time per checkout).
  2. Pick & customise a preset: copy scripts/rag-perf/configs/<preset>.yaml if you want a variant; always set rag.collection_names to a real collection.
  3. Run: uv run --project scripts/rag-perf rag-perf -c <config> from repo root.
  4. Read the per-point + aggregate tables on stdout. Bottleneck inference is in the per-point profiling section; comparison across points is the final aggregate table.
  5. Parse artifacts under output.dir/run_<ts>/ — see references/output-and-analysis.md. For multi-point runs, results.csv has one row per (point × iteration).
  6. Summarise for the user using the playbook in references/output-and-analysis.md#summarising-results-to-the-user — headline table, scaling-efficiency math for sweeps, mandatory flags for zero citations / non-zero errors / suspect llm_ttft_ms / low sample size, and a concrete next-experiment YAML.
  7. Tune retrieval / reranker: flip to quick_profile.yaml or aiperf.enabled: false for fast iteration, then return to single_run.yaml / sweep.yaml when characterising under load.
  8. Triage failures: see Troubleshooting above and references/output-and-analysis.md for empty-citation / bottleneck=N/A patterns.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/rag-perf of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • eval/h100.json
  • eval/nvidia_hosted.json
  • references/config-schema.md
  • references/output-and-analysis.md
  • references/synthetic-generation.md
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 67a13c0

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in NVIDIA/skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

RAG Perf next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Perf compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Perf this skillNVIDIA/skills3.5k1 repos~4.1kAutomated safety check: PassApache-2.0
Embeddings via 9Routerdecolua/9router30k—~604Automated safety check: PassMIT
Vss Deploy Detection Tracking 3DNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~5.1kAutomated safety check: NotesApache-2.0
Testing Prompt Injection In RAG Pipelinesmukul975/Anthropic-Cybersecurity-Skills34k—~3.3kAutomated safety check: PassApache-2.0
Gke Inferencegoogle/skills21k—~2kAutomated safety check: PassApache-2.0
Chroma Vector DatabaseOrchestra-Research/AI-Research-SKILLs13k8 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Embeddings via 9Router

    decolua/9router

    Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search.

    30k GitHub stars~604 tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Vss Deploy Detection Tracking 3D

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC…

    1.9k GitHub stars~5.1k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Testing Prompt Injection In RAG Pipelines

    mukul975/Anthropic-Cybersecurity-Skills

    Probes Retrieval-Augmented Generation pipelines for indirect prompt injection via poisoned retrieved documents and embedding-space manipulation, using NVIDIA garak, Promptfoo red-team plugins, and…

    34k GitHub stars~3.3k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Gke Inference

    google/skills

    Official

    Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers.

    21k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Application Dev

    MoizIbnYousaf/ai-agent-skills

    Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration.

    1.1k GitHub starsUsed in 2 repos~1.3k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about RAG Perf

What does RAG Perf do?

Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config. RAG Perf is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Performance benchmarking for a deployed NVIDIA RAG Blueprint server: profiling pass + aiperf load test driven by a single YAML config.

When should I use RAG Perf?

RAG Perf fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Load testing.

How do I install RAG Perf in Claude Code?

Run `npx skills add NVIDIA/skills --skill rag-perf -a claude-code`. Or copy the skill folder (skills/rag-perf in NVIDIA/skills) into .claude/skills/rag-perf in your project. Claude Code loads it when a task matches its description.

How do I install RAG Perf in Codex?

Run `npx skills add NVIDIA/skills --skill rag-perf -a codex`. Or copy the skill folder (skills/rag-perf in NVIDIA/skills) into .agents/skills/rag-perf in your project. Codex loads it when a task matches its description.

Can I use RAG Perf in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill rag-perf -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-perf, .gemini/skills/rag-perf, .github/skills/rag-perf and .opencode/skills/rag-perf in your project.

What does RAG Perf need to run?

Going by SKILL.md and its folder, RAG Perf needs the command-line tools its instructions call (uv, pip and curl) and credentials named NVIDIA_API_KEY. Our summary lists: Python 3; A credential in NVIDIA_API_KEY. Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash(ls *), Bash(python3 *), Bash(uv *), Bash(cat *), Bash(curl *), Write, Edit. Compatibility (from SKILL.md): Repository checkout with uv; Python 3.11+; run from repo root; uv sync --project scripts/rag-perf (perf deps live in scripts/rag-perf/pyproject.toml); reachable RAG server (default http://localhost:8081); for synthetic queries an OpenAI-compatible chat-completions endpoint is required (default http://localhost:8999/v1/chat/completions); aiperf load-test phase uses the bundled nvidia_rag endpoint plugin, registered automatically when rag-perf is installed editable..

Does RAG Perf access the network?

SKILL.md contains no URLs. Its commands use uv, pip and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is RAG Perf safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does RAG Perf use?

RAG Perf is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Perf use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.1k tokens, read only when the agent opens those files.

What are the alternatives to RAG Perf?

Skills that share tags, products or a category with RAG Perf: Embeddings via 9Router (decolua/9router, 30k stars), Vss Deploy Detection Tracking 3D (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars), Testing Prompt Injection In RAG Pipelines (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Gke Inference (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Perf?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.