Official agent skill

RAG Blueprint

by NVIDIA in NVIDIA/skills

NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage.

OfficialApache-2.0Auto-check: notesAI & LLM Engineering

Install RAG Blueprint

skills CLI
$ npx skills add NVIDIA/skills --skill rag-blueprint -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills rag-blueprint --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/rag-blueprint .claude/skills/rag-blueprint && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rag-blueprint
GitHub stars
3.5k
Token cost
~2.8k tokens
SKILL.md length
879 words
Files
39 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage.

  • Works in 4 steps: Match the user request to the intent… → Read the referenced playbook before… → Use repository docs and deployment… → …
  • Tasks that involve Retrieval-augmented generation
  • SKILL.md covers Purpose, Instructions, Prerequisites and Autonomy Principles, plus 6 more sections
  • Calls docker, kubectl and curl; needs NGC_API_KEY

What it does

RAG Blueprint is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. Handles any RAG action: deploy, install, start, enable, disable, toggle, change, configure, troubleshoot, debug, fix, shutdown, stop, or tear down any RAG feature or service (Agentic RAG, VLM, guardrails, query rewriting, models, search, ingestion, observability, summarization, reasoning, and more).

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 41 other files, including reference files (for example `BENCHMARK.md`, `eval/h100.json` and `eval/nvidia_hosted.json`). Compatibility notes: NVIDIA RAG Blueprint repository checkout; Docker/Compose or Kubernetes/Helm for deployments; Python 3.11+ for library workflows; NVIDIA GPU tooling for…

It sits in AI & LLM Engineering, covering Retrieval-augmented generation, Summarization and Observability. It works with NVIDIA AI Platform, CUDA and Docker. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Retrieval-augmented generation
  • Tasks that involve Summarization
  • Tasks that involve Observability

Example prompts

  • “/rag-blueprint”

Requirements

  • Python 3
  • Docker
  • A credential in NGC_API_KEY
  • Compatibility (from SKILL.md): NVIDIA RAG Blueprint repository checkout; Docker/Compose or Kubernetes/Helm for deployments; Python 3.11+ for library workflows; NVIDIA GPU tooling for self-hosted NIM services.
  • Pre-approved tools (allowed-tools): Bash(echo *), Bash(nvidia-smi *), Bash(curl --version *), Bash(docker ps *), Bash(docker info *), Bash(docker --version *), Bash(docker version *), Bash(docker logs *), Bash(docker inspect *), Bash(docker stats *), Bash(docker compose ps *), Bash(docker compose logs *), Bash(docker compose config *), Bash(docker compose version *), Bash(kubectl get *), Bash(kubectl describe *), Bash(kubectl version *), Bash(kubectl logs *), Bash(kubectl api-resources *), Bash(kubectl rollout status *), Bash(helm version *), Bash(helm list *), Bash(helm status *), Bash(oc get *), Bash(oc describe *), Bash(oc logs *), Bash(oc whoami *), Bash(oc version *), Bash(git rev-parse *), Bash(git describe *), Bash(git status *), Bash(python3 --version *), Bash(pip3 show *), Bash(df *), Bash(du *), Bash(cat /proc/*), Bash(cat /etc/os-release *), Bash(ss *), Bash(netstat *), Bash(ls *), Bash(grep *), Bash(lsof *), Bash(ps aux *), Read, Grep, Glob

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Match the user request to the intent routing table below.
  2. Read the referenced playbook before making changes.
  3. Use repository docs and deployment config files as the source of truth.
  4. Verify the affected service or workflow after changes.

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(echo *)
    • Bash(nvidia-smi *)
    • Bash(curl --version *)
    • Bash(docker ps *)
    • Bash(docker info *)
    • Bash(docker --version *)
    • Bash(docker version *)
    • Bash(docker logs *)
    • Bash(docker inspect *)
    • Bash(docker stats *)

    …and 36 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • docker
    • kubectl
    • curl
    • helm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use docker, kubectl, curl and helm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • NGC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    NVIDIA RAG Blueprint repository checkout; Docker/Compose or Kubernetes/Helm for deployments; Python 3.11+ for library workflows; NVIDIA GPU tooling for self-hosted NIM services.

    From compatibility in the SKILL.md frontmatter.

Context cost

RAG Blueprint loads about 2.8k tokens when it runs, and up to ~29k if it reads all its reference files. Until then it costs about 96 tokens; SKILL.md has 879 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~29k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:122
    r) | Any | Self-hosted | `deploy/compose/.env` |
  • NoteMentions a .env fileSKILL.md:129
    ployment. Config file is `deploy/compose/.env`. Correct?"

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 879 words, ~2,834 tokens.

Download SKILL.mdSave it as .claude/skills/rag-blueprint/SKILL.md (or your agent's skills folder). This skill also uses 38 other files; get the full folder from GitHub.
name
rag-blueprint
description
NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. Handles any RAG action: deploy, install, start, enable, disable, toggle, change, configure, troubleshoot, debug, fix, shutdown, stop, or tear down any RAG feature or service (Agentic RAG, VLM, guardrails, query rewriting, models, search, ingestion, observability, summarization, reasoning, and more).
allowed-tools
Bash(echo *), Bash(nvidia-smi *), Bash(curl --version *), Bash(docker ps *), Bash(docker info *), Bash(docker --version *), Bash(docker version *), Bash(docker logs *), Bash(docker inspect *), Bash(docker stats *), Bash(docker compose ps *), Bash(docker compose logs *), Bash(docker compose config *), Bash(docker compose version *), Bash(kubectl get *), Bash(kubectl describe *), Bash(kubectl version *), Bash(kubectl logs *), Bash(kubectl api-resources *), Bash(kubectl rollout status *), Bash(helm version *), Bash(helm list *), Bash(helm status *), Bash(oc get *), Bash(oc describe *), Bash(oc logs *), Bash(oc whoami *), Bash(oc version *), Bash(git rev-parse *), Bash(git describe *), Bash(git status *), Bash(python3 --version *), Bash(pip3 show *), Bash(df *), Bash(du *), Bash(cat /proc/*), Bash(cat /etc/os-release *), Bash(ss *), Bash(netstat *), Bash(ls *), Bash(grep *), Bash(lsof *), Bash(ps aux *), Read, Grep, Glob
compatibility
NVIDIA RAG Blueprint repository checkout; Docker/Compose or Kubernetes/Helm for deployments; Python 3.11+ for library workflows; NVIDIA GPU tooling for self-hosted NIM services.
version
2.6.0
license
Apache-2.0
metadata.author
NVIDIA RAG <foundational-rag-dev@exchange.nvidia.com>
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/rag
metadata.endpoint-openapi-schemas
docs/api_reference/openapi_schema_rag_server.json, docs/api_reference/openapi_schema_ingestor_server.json
metadata.argument-hint
deploy RAG | enable feature | disable feature | configure | troubleshoot | shutdown
metadata.tags
nvidia, blueprint, rag, deployment, configuration, troubleshooting
metadata.languages
python, typescript, shell
metadata.frameworks
fastapi, langchain, react, docker-compose, helm
metadata.domain
ai-ml

NVIDIA RAG Blueprint

Purpose

Use this skill for NVIDIA RAG Blueprint operations: deployment, configuration, troubleshooting, shutdown, and feature management across Docker, Helm, and library deployments.

Instructions

  1. Match the user request to the intent routing table below.
  2. Read the referenced playbook before making changes.
  3. Use repository docs and deployment config files as the source of truth.
  4. Verify the affected service or workflow after changes.

Prerequisites

  • NVIDIA RAG Blueprint repository checkout.
  • Docker/Compose or Kubernetes/Helm for deployments.
  • Python 3.11+ for library workflows.
  • NVIDIA GPU tooling for self-hosted NIM services.

Autonomy Principles

  • Auto-detect everything: GPU, VRAM, drivers, Docker, CUDA, disk, OS, ports, existing services, NGC key, repo state.
  • If it can be checked with a command, check it — don't ask the user.
  • Ask only when user action is required: providing an API key, confirming data deletion, or choosing between equally valid options.
  • Once analysis is done, route to the correct workflow and execute.

Intent Detection

Determine what the user wants and route immediately:

User IntentAction
Deploy, install, set up, start RAGRead and follow references/deploy.md
Configure, enable, change, toggle a featureUse the Configure section below
Troubleshoot, debug, fix, error, unhealthyRead and follow references/troubleshoot.md
Stop, shutdown, tear down, clean upRead and follow references/shutdown.md

If the intent is ambiguous, infer from context (e.g., "RAG isn't working" → troubleshoot; "get RAG running" → deploy). Only ask if genuinely unclear.


Configure

Requires a running RAG deployment. If services are not running, deploy first via references/deploy.md.

Match the user's request to a reference file, then read and follow it:

Feature KeywordsReference
VLM, VLM embeddings, image captioningreferences/configure/vlm.md
NeMo Guardrailsreferences/configure/guardrails.md
Agentic RAG, planning/execution agent, agentic streaming, stage eventsreferences/configure/agentic-rag.md
Query rewriting, decomposition, multi-turnreferences/configure/query-and-conversation.md
Ingestion (text-only, audio, Nemotron Parse, OCR, batch CLI, NV-Ingest, volume mount, performance)references/configure/ingestion.md
Search, retrieval, hybrid search, multi-collection, metadata, filters, Elasticsearch filters, reranker, topK, accuracy/performancereferences/configure/search-and-retrieval.md
LLM/embedding/ranking model changes, vector DB, Milvus/Elasticsearch auth, service keys, model profiles, ports/GPUreferences/configure/models-and-infrastructure.md
Reasoning, thinking mode, reasoning_content, self-reflection, prompts, generation params (tokens, temperature, citations), per-request LLM paramsreferences/configure/reasoning-and-generation.md
Summarizationreferences/configure/summarization.md
Observability (tracing, Zipkin, Grafana, Prometheus)references/configure/observability.md
Multimodal query (image + text)references/configure/multimodal-query.md
Data catalog (collection/document metadata)references/configure/data-catalog.md
User interface (UI settings, reasoning panel, metadata filters)references/configure/user-interface.md
API reference (endpoints, schemas)references/configure/api-reference.md
Evaluation (RAGAS metrics)references/configure/evaluation.md (and skill rag-eval)
MCP server & client, agent toolkitreferences/configure/mcp.md
Migration (version upgrades)references/configure/migration.md
Notebooks (setup and catalog)references/configure/notebooks.md
Configure Flow
  1. Match the user's request to a reference file from the table above.

  2. Detect what's running:

    bash
    echo "=== NIM ===" && docker ps --format '{{.Names}}' 2>/dev/null | grep -iE '(nim-llm|nemotron-(vlm-)?embedding|nemotron-ranking|nemotron-vlm|nemotron-3-nano-omni|page-elements|graphic-elements|table-structure|nemotron-ocr)' || echo "NO_LOCAL_NIMS"; echo "=== RAG ===" && docker ps --format '{{.Names}}' 2>/dev/null | grep -iE '(rag-server|ingestor-server|elasticsearch|milvus|seaweedfs|lancedb)' || echo "NO_DOCKER_RAG"; echo "=== K8S ===" && kubectl get pods -n rag 2>/dev/null | head -5 || echo "NO_K8S"; echo "=== LIBRARY ===" && ps aux 2>/dev/null | grep -E '(nvidia_rag|uvicorn.*rag)' | grep -v grep || echo "NO_LIBRARY"
  3. Use this table to determine platform, deployment type, and where config lives:

    Local NIMs running?RAG services running?Deployment TypeConfig Location
    Yes (Docker)AnySelf-hosteddeploy/compose/.env
    NoYes (Docker)NVIDIA-hosteddeploy/compose/nvdev.env
    Yes (K8s pods)AnySelf-hostedvalues.yaml (NIM sections)
    NoYes (K8s pods)NVIDIA-hostedvalues.yaml (envVars)
    —Library processesLibrary modenotebooks/config.yaml
    NoNoNot runningDeploy first via references/deploy.md

    Tell the user what you detected and ask to confirm. Example: "I see local NIM containers running (nim-llm-ms, nemotron-vlm-embedding-ms) — this is a self-hosted deployment. Config file is deploy/compose/.env. Correct?"

  4. Check current feature state before changing anything — read the config location from step 3, then cross-check the live service:

    • Docker: docker exec rag-server env 2>/dev/null | grep -E "<VAR_NAME>"
    • Helm: kubectl get pod -n rag -l app=rag-server -o jsonpath='{.items[0].spec.containers[0].env}' 2>/dev/null

    If the config file and live service disagree, tell the user the service has stale config and will need a restart.

  5. If the feature needs extra GPUs, check availability against hardware restrictions (see below):

    bash
    nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv,noheader 2>/dev/null || echo "NO_GPU"
  6. Read the reference file and apply changes:

    • Docker: edit the env file (uncomment to enable, re-comment to disable — the env file is the source of truth). Then restart the affected service:
      source <env-file> && docker compose -f deploy/compose/<compose-file> up -d
      ServiceCompose File
      rag-serverdocker-compose-rag-server.yaml
      ingestor-serverdocker-compose-ingestor-server.yaml
      Elasticsearch, Milvus, etcd, SeaweedFSvectordb.yaml
      NIM containers (LLM, embedding, ranking, VLM, OCR, parse, audio, extraction)nims.yaml
      guardrailsdocker-compose-nemo-guardrails.yaml
      observability (Grafana, Prometheus, Zipkin)observability.yaml
    • Helm: edit values.yaml, then upgrade: helm upgrade rag <chart> -n rag -f values.yaml
    • Library: edit notebooks/config.yaml, then restart the Python process
  7. Verify:

    • Docker: docker ps --format "table {{.Names}}\t{{.Status}}" | head -20; curl -s http://localhost:8081/v1/health?check_dependencies=true 2>/dev/null | head -1
    • Helm: kubectl get pods -n rag; kubectl rollout status deployment/rag-server -n rag --timeout=120s
    • Library: curl -s http://localhost:8081/v1/health 2>/dev/null | head -1
  8. If restart fails, read references/troubleshoot.md. If multiple features requested, repeat from step 1 for each.

Show full SKILL.md (162 more words)Show less

Examples

  • "Deploy RAG" -> route to references/deploy.md.
  • "Enable VLM" -> route to references/configure/vlm.md.
  • "RAG is unhealthy" -> route to references/troubleshoot.md.
  • "Stop RAG" -> route to references/shutdown.md.

Limitations

  • Operational guidance only applies to this RAG Blueprint repository.
  • Live deployment changes require a running Docker, Helm, or library target.
  • Secrets such as NGC_API_KEY must be supplied by the user environment.

Troubleshooting

Error / signalWhat to do
Services are not runningFollow references/deploy.md before configuring features.
Restart or health check failsFollow references/troubleshoot.md.
User requests teardownFollow references/shutdown.md and confirm destructive cleanup.
When User Says "Configure" Without Specifics

Run steps 2–3 above, then read the identified config file to list what's currently enabled:

bash
grep -E "^(export )?(ENABLE_|APP_)" <config-file> 2>/dev/null | sort

Summarize what's running and enabled, then ask which feature to change.


Hardware Restrictions

Read docs/support-matrix.md for current GPU requirements per deployment mode. Read docs/service-port-gpu-reference.md for port mappings and GPU assignments.

GPUFeature Restrictions
B200No VLM, No Guardrails, No Nemotron Parse. May need multi-GPU LLM (LLM_MS_GPU_ID).
RTX PRO 6000No Nemotron Parse. No Audio on Helm.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 38 other files (references) in skills/rag-blueprint of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • eval/h100.json
  • eval/nvidia_hosted.json
  • references/configure/agentic-rag.md
  • references/configure/api-reference.md
  • references/configure/data-catalog.md
  • references/configure/evaluation.md
  • references/configure/guardrails.md
  • references/configure/ingestion.md
  • references/configure/mcp.md
  • references/configure/migration.md
  • references/configure/models-and-infrastructure.md
  • references/configure/multimodal-query.md
  • references/configure/notebooks.md
  • references/configure/observability.md
  • references/configure/query-and-conversation.md
  • references/configure/reasoning-and-generation.md
  • … and 21 more

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

RAG Blueprint next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

RAG Blueprint compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
RAG Blueprint this skillNVIDIA/skills3.5k—~2.8kAutomated safety check: NotesApache-2.0
Vss Deploy Detection Tracking 3DNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~5.1kAutomated safety check: NotesApache-2.0
Nemotron Customizer AirgapNVIDIA-NeMo/Nemotron2.1k—~1.2kAutomated safety check: PassApache-2.0
Cv DeployLMIXR/CV_Deployment_skill126—~547Automated safety check: PassNone
Vllm Deploy Dockervllm-project/vllm-skills103—~2.5kAutomated safety check: NotesApache-2.0
Setup Workshopbrevdev/workshop-build-an-agent143—~2.3kAutomated safety check: NotesApache-2.0

Similar skills

  • Vss Deploy Detection Tracking 3D

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC…

    1.9k GitHub stars~5.1k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Nemotron Customizer Airgap

    NVIDIA-NeMo/Nemotron

    Prepare, validate, build, and use Nemotron Customizer airgap image bundles for offline clusters.

    2.1k GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Cv Deploy

    LMIXR/CV_Deployment_skill

    基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

    126 GitHub stars~547 tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Deploy Docker

    vllm-project/vllm-skills

    Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

    103 GitHub stars~2.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Setup Workshop

    brevdev/workshop-build-an-agent

    This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a.

    143 GitHub stars~2.3k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    143 GitHub stars~5.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated today
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated today
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated today
    Auto-check: notes

Questions about RAG Blueprint

What does RAG Blueprint do?

NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage. RAG Blueprint is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. NVIDIA RAG Blueprint — deploy, configure, troubleshoot, and manage.

When should I use RAG Blueprint?

RAG Blueprint fits situations like: tasks that involve Retrieval-augmented generation; tasks that involve Summarization; tasks that involve Observability.

How do I install RAG Blueprint in Claude Code?

Run `npx skills add NVIDIA/skills --skill rag-blueprint -a claude-code`. Or copy the skill folder (skills/rag-blueprint in NVIDIA/skills) into .claude/skills/rag-blueprint in your project. Claude Code loads it when a task matches its description.

How do I install RAG Blueprint in Codex?

Run `npx skills add NVIDIA/skills --skill rag-blueprint -a codex`. Or copy the skill folder (skills/rag-blueprint in NVIDIA/skills) into .agents/skills/rag-blueprint in your project. Codex loads it when a task matches its description.

Can I use RAG Blueprint in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill rag-blueprint -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rag-blueprint, .gemini/skills/rag-blueprint, .github/skills/rag-blueprint and .opencode/skills/rag-blueprint in your project.

What does RAG Blueprint need to run?

Going by SKILL.md and its folder, RAG Blueprint needs the command-line tools its instructions call (docker, kubectl, curl and helm) and credentials named NGC_API_KEY. Our summary lists: Python 3; Docker; A credential in NGC_API_KEY. Its frontmatter pre-approves these tools: Bash(echo *), Bash(nvidia-smi *), Bash(curl --version *), Bash(docker ps *), Bash(docker info *), Bash(docker --version *), Bash(docker version *), Bash(docker logs *), Bash(docker inspect *), Bash(docker stats *), Bash(docker compose ps *), Bash(docker compose logs *), Bash(docker compose config *), Bash(docker compose version *), Bash(kubectl get *), Bash(kubectl describe *), Bash(kubectl version *), Bash(kubectl logs *), Bash(kubectl api-resources *), Bash(kubectl rollout status *), Bash(helm version *), Bash(helm list *), Bash(helm status *), Bash(oc get *), Bash(oc describe *), Bash(oc logs *), Bash(oc whoami *), Bash(oc version *), Bash(git rev-parse *), Bash(git describe *), Bash(git status *), Bash(python3 --version *), Bash(pip3 show *), Bash(df *), Bash(du *), Bash(cat /proc/*), Bash(cat /etc/os-release *), Bash(ss *), Bash(netstat *), Bash(ls *), Bash(grep *), Bash(lsof *), Bash(ps aux *), Read, Grep, Glob. Compatibility (from SKILL.md): NVIDIA RAG Blueprint repository checkout; Docker/Compose or Kubernetes/Helm for deployments; Python 3.11+ for library workflows; NVIDIA GPU tooling for self-hosted NIM services..

Does RAG Blueprint access the network?

SKILL.md contains no URLs. Its commands use docker and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is RAG Blueprint safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does RAG Blueprint use?

RAG Blueprint is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does RAG Blueprint use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 26k tokens, read only when the agent opens those files.

What are the alternatives to RAG Blueprint?

Skills that share tags, products or a category with RAG Blueprint: Vss Deploy Detection Tracking 3D (NVIDIA-AI-Blueprints/video-search-and-summarization, 1.9k stars), Nemotron Customizer Airgap (NVIDIA-NeMo/Nemotron, 2.1k stars), Cv Deploy (LMIXR/CV_Deployment_skill, 126 stars) and Vllm Deploy Docker (vllm-project/vllm-skills, 103 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains RAG Blueprint?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.