Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Agent skill
by NVIDIA-AI-Blueprints in NVIDIA-AI-Blueprints/video-search-and-summarization
Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization benchmark-video-search --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/benchmark-video-search .claude/skills/benchmark-video-search && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-video-search" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/benchmark-video-search into .claude/skills/benchmark-video-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-video-search", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/benchmark-video-searchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization benchmark-video-search --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/benchmarking/benchmark-video-search .agents/skills/benchmark-video-search && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-video-search" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/benchmark-video-search into .agents/skills/benchmark-video-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-video-search", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization benchmark-video-search --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/benchmarking/benchmark-video-search .cursor/skills/benchmark-video-search && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-video-search" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/benchmark-video-search into .cursor/skills/benchmark-video-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-video-search", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git --path skills/benchmarking/benchmark-video-search--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization benchmark-video-search --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/benchmarking/benchmark-video-search .gemini/skills/benchmark-video-search && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-video-search" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/benchmark-video-search into .gemini/skills/benchmark-video-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-video-search", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization benchmark-video-searchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/benchmarking/benchmark-video-search .github/skills/benchmark-video-search && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-video-search" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/benchmark-video-search into .github/skills/benchmark-video-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-video-search", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization benchmark-video-search --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/benchmarking/benchmark-video-search .opencode/skills/benchmark-video-search && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-video-search" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/benchmark-video-search into .opencode/skills/benchmark-video-search/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-video-search", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-video-searchMeasure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.
Benchmark Video Search is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths. Use when the user asks to "benchmark video search", "compare search recall between builds", or "profile search latency" on a deployed VSS profile.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including scripts and reference files (for example `evals/evals.json`, `references/dataset-format.md` and `references/flags.md`).
It sits in AI & LLM Engineering. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 10 files in scripts/ (Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
uvcurlpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv and curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Video Search loads about 4.3k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 2,269 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 2,269 words, ~4,291 tokens.
.claude/skills/benchmark-video-search/SKILL.md (or your agent's skills folder). This skill also uses 20 other files; get the full folder from GitHub.Follow the routing table, then the numbered steps. Execute each step in order on
a first run; for a repeat run against an already-indexed deployment go straight
to Step 5. Reference material is in references/, the runner is in
scripts/.
Answers two questions about a search profile deployment: does retrieval find
the right footage, and where does a query's time go. It drives the vss CLI —
the same path the product uses, since the agent adapter shells out to
vss search run <mode> --raw for every search — so the numbers describe what
ships rather than a REST endpoint that is being retired.
vss-search-archive.run_eval.py noted in references/troubleshooting.md.| Situation | Action |
|---|---|
| User asks to benchmark the OpenClaw chat route | This runner does not measure it; explain the gap instead of presenting CLI timings as chat timings |
| User asks to benchmark the retired agent REST search endpoint | This runner does not measure it; for historical comparison, use the external legacy run_eval.py noted in references/troubleshooting.md |
| No search profile deployed in this session | Use vss-build-vision-ai to deploy the stock search profile with its in-stack agent REST API and in-stack LLM retained. Explicitly name both requirements in the build request so it skips the harness question (which would remove the agent and possibly the LLM); do not select a NemoClaw-only or CLI-only harness. Record the reachable agent endpoint and unified origin, then return here. Check agent /health and LLM /v1/models, run Step 3 to inspect indices, and complete Step 4 ingestion and its index probe before scoring |
| User did not give an endpoint | Ask for it. Do not guess, and do not default to localhost |
Endpoint uses localhost or 127.0.0.1 | This is valid when the runner is on the deployment host. Verify the actual VST clip URL is reachable from RT-VLM before trusting critic-filtered metrics |
| User did not give a dataset | Ask which dataset and where its --data-dir is. Do not invent one |
Elasticsearch has no mdx-* indices | Ingestion has not completed. Run Step 4; do not report zero scores as a quality result |
| User asks to reuse what is already ingested | Add --skip-ingest. It cannot be combined with --clear or --only-dataset |
| User asks to analyse an existing result file | Skip to Step 6; read the JSON, run nothing |
| Every query returns 0 hits | Stop. Diagnose with Step 3 before reporting anything — this is nearly always missing indices, not poor retrieval |
Critic-filtered metrics are all NA | The critic never ran. Check the clip-URL prerequisite below before concluding anything about verification quality |
Unless the user says otherwise, this is the run. Ask which dataset to use and, if missing, which deployment endpoint to target.
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
--data-dir /path/to/datasets --dataset DATASET \
--skip-download --skip-existing --name RUN_NAMEFive defaults, each load-bearing:
--skip-existing is non-destructive and avoids duplicate uploads on a
shared deployment. --only-dataset and --clear are explicit destructive
maintenance modes. Both require --confirm-delete in the same invocation;
there is no interactive prompt. Use either only when the user requested
deletion, and report the printed deletion inventory.
Live decomposition is on by default. The LLM origin is derived from
--endpoint (same host, port 30081), so there is no flag to forget. Every
query is decomposed the way the deployed agent decomposes it, and that choice
picks the retrieval path. Override with --llm-url / --llm-port; turn it
off with --no-decompose (replays stored routes) or --fixed-search-path
(ignores stored routes for a single-path baseline).
If the NIM is unreachable the run replays dataset routes when present;
otherwise it falls back to embed, which needs nothing but the query text.
The fallback records flow.query.live_decomposition.fell_back_to in the
result file. Check that field and Search paths: before quoting numbers.
--no-decompose also replays dataset routes when present; it does not force
one path. Use --fixed-search-path only for a deliberately single-path
baseline. Nothing else is inferred: a
devset_provenance.json sidecar is never loaded automatically. To use it as
an answer key, pass its full path with --decompositions; scoring against
perfect routing is a different experiment from scoring the live decomposer.
The pre-decomposition question is sent with every query as
--original-query. Retrieval ignores it; the critic verifies against it
instead of a question the library rebuilds out of --query and --attribute.
Without it the CLI's critic is asked a different question from the REST
flow's, so the two flows' rejection rates are not comparable — and on the
object path there is no question at all, so verification is skipped
outright. Turn it off only with --no-original-query, and only to reproduce
a result file captured before this existed.
A vss older than the flag is detected by one --help probe at startup and
the run continues without it, with a warning. Do not quote critic-filtered
metrics from a run that printed that warning.
Ingest is vst-direct: upload to VIOS, let its webhooks drive
perception. That is what the UI does — it stopped calling the agent's ingest
API — and it is the only flow that also triggers RT-VLM tagging. The
timestamp anchor rides in the upload metadata, not in anything /complete
does, so scores stay comparable with agent-3step baselines.
What agent-3step gave for free was proof: /complete returns
chunks_processed, and a zero failed the upload. vst-direct has no such
step, so a post-ingest index probe replaces it — one embed query, retried,
before the real run starts. Registered in VST is not the same as indexed in
Elasticsearch, and without this check an unindexed deployment scores 0.0
across the board and reads as a retrieval collapse.
If the probe aborts a run, check the deployed VIOS notification config and
that RTVI_EMBED_MODEL matches the webhook's model string — RT-Embed answers
a mismatch with HTTP 200 and inference: false. The Docker and Helm search
profiles enable webhooks; generic VIOS chart defaults may not. If the
deployed webhooks are disabled, use --ingest-flow agent-3step.
Never pass --skip-index-probe on a run whose numbers you intend to quote.
Concurrency stays at 1 (the script's default). Concurrent queries contend for the same VLM and embedding services, so per-stage latencies inflate and stop describing a single query.
Deviate only on request: --skip-ingest to reuse what is there, --clear to
wipe the whole deployment, --subset to narrow the slice, --concurrency N to
trade latency fidelity for wall-clock, --no-decompose to replay stored routes,
or --fixed-search-path for a single-path baseline.
Never add either decomposition override to a run the user called a live-routing eval. Both measure a different flow; the result records the routing mode.
| Requirement | How to check |
|---|---|
| Search profile agent reachable | curl -sf --connect-timeout 5 --max-time 10 ${ENDPOINT}/health returns 200 |
| Unified HAProxy origin routes the services | curl -sf --connect-timeout 5 --max-time 10 ${ORIGIN}/vst/api/v1/sensor/version returns 200. ORIGIN is usually the endpoint host on port 7777; use it for vss configure and the Step 3 Elasticsearch probe. The runner separately derives VST's direct ingress on port 30888 (--vst-port default); do not pass the HAProxy port as --vst-port |
vss CLI runs from this checkout | vss search run --help exits 0 |
| CLI points at the right origin | vss configure show reports the ORIGIN above |
| Dataset present locally | test -f ${DATA_DIR}/${DATASET}/dataset.json — layout in references/dataset-format.md |
Python 3.10+ for this runner; Python 3.13–3.14 for a separately installed vss CLI | python3 --version for the runner; use a Python 3.13 or 3.14 environment when installing libs/vss/core and libs/vss/cli (their requires-python is >=3.13,<3.15), then check vss search run --help there |
| LLM reachable (else the run falls back and says so) | curl -sf --connect-timeout 5 --max-time 10 http://HOST:30081/v1/models returns 200 |
| LLM model selected unambiguously | If /v1/models lists multiple IDs, pass --llm-model matching the deployed agent's model |
| Critic clip URL reachable from RT-VLM | Inspect a returned VST videoUrl and test that exact URL from the RT-VLM container. The CLI uses video_url_scope="internal"; CLI ORIGIN may legitimately be localhost on the host |
Agent /health confirms the process is reachable, not that VIOS and the search
indices are ready. Step 3 and the post-ingest index probe are the retrieval
readiness gates.
The CLI reads ~/.vss/config.json, which is separate state from --endpoint
and points at a different port: it discovers services by path prefix on the
unified origin, which the agent port does not route. Getting this wrong makes
every query exit 4.
vss configure --base-url http://HOST:7777
vss configure showlocalhost is valid here when the runner runs on the deployment host. This
URL controls where the CLI reaches the services; it does not determine the
clip URL sent to RT-VLM. If verification fails, inspect the VST-returned
videoUrl and test that exact address from the RT-VLM container.
Reports what would happen and contacts nothing.
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
--data-dir /path/to/datasets --dataset DATASET --dry-runRegistration with VST is not ingestion. A deployment can list every sensor and still have an empty Elasticsearch, in which case every query returns zero hits and every metric reads 0.0000 — which looks like catastrophic retrieval quality and is not.
curl -s --connect-timeout 5 --max-time 10 "http://HOST:7777/elasticsearch/_cat/indices?h=index,docs.count"Expect mdx-embed-filtered-*, mdx-behavior-* and mdx-raw-* with non-zero
counts. mdx-behavior-* is what the attribute and fusion paths read; if it
is missing, only embed can score. If the indices are absent, go to Step 4.
The default vst-direct flow uploads to VIOS and lets its webhooks trigger
perception. It does not call the agent's /complete endpoint or receive a
chunks_processed count. The runner waits for VST registration, then probes
the search index before scoring. Use --ingest-flow agent-3step only when the
deployment cannot run the webhook flow; see references/troubleshooting.md.
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
--data-dir /path/to/datasets --dataset DATASET --skip-download --skip-existingExpect this to be slow while the webhook pipeline indexes the videos:
--ingest-flow agent-3step explicitly. That
flow calls /complete, which can return a transient 502 and is retried.--skip-existing is the default: ingest what is missing and delete nothing.--only-dataset deletes foreign sources only when explicitly requested with
--confirm-delete; report its candidate inventory before it runs.--clear --confirm-delete deletes every source including other people's.
Use it only on explicit request.--clear and --only-dataset list sources via --vst-url but delete through
the agent endpoint. If those URLs have different hosts, verify they belong
to the same deployment before deletion; otherwise source IDs from one
stack could be sent to another.Then re-run Step 3. Indices are lazy; they appear after webhook processing,
not immediately after upload. See references/flags.md for ingest options.
uv run --with requests python3 scripts/run_eval_flows.py --endpoint http://HOST:8000 \
--data-dir /path/to/datasets --dataset DATASET --subset SUBSET \
--skip-download --skip-ingest --name RUN_NAME--skip-ingest reuses an already-indexed dataset; it cannot be combined with
--clear or --only-dataset. See references/flags.md before changing flags
that affect run meaning.
Decomposition needs no flag. An LLM turns the sentence into
{query, attributes, has_action, ...} and that decides which path runs — the
same call the deployed agent makes. The script derives the NIM origin from
--endpoint; if it cannot reach it the run falls back rather than stopping, and
says which fallback it took.
fusion cannot be a fallback. It needs an attribute per query and there is
nothing to supply one without decomposition — and feeding it the whole query as
its attribute is the failure this eval already measured: mAP −25%, HIT@1 halved,
latency +59%.
Leave --concurrency alone unless asked. It defaults to 1 because concurrent
queries contend for the same VLM and embedding services, which inflates every
per-stage latency in the report.
--name makes the result file findable; without it the name is
<dataset>_<subset>_<timestamp>.
Results land in scripts/cli_eval_result/<name>.json.
Report Recall and HIT@k as the quality signal. Precision and mAP are
dominated by --top-k — five results against roughly one relevant segment keeps
them low however good retrieval is — so they compare runs; they do not grade a
deployment.
State these alongside any number, because each one changes what it means:
Search paths: in the summary. A run that was all
embed says nothing about attribute.Routed to: — a routing shift changes what was measured.NA when the
response carried no verification block. NA is not 0.Read references/reading-results.md when interpreting stage latencies or
metrics. Read references/flags.md when a flag might change comparability.
Read references/troubleshooting.md when ingest fails or the CLI exits nonzero.
Read references/dataset-format.md when validating a dataset's layout.
A non-zero CLI exit aborts the run. A missing result envelope marks that
query unanswered and aborts by default; --tolerate-unanswered permits an
explicit partial run that excludes those queries from scoring. Never count
them as misses. Diagnose an all-zero run with Step 3 before reporting quality.
Critic-filtered NA means verification was not available, not zero quality. Use
references/troubleshooting.md for the specific failure and recovery steps.
Routed to: before comparing scores.© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 20 other files (scripts, references) in skills/benchmarking/benchmark-video-search of NVIDIA-AI-Blueprints/video-search-and-summarization.
Open the folder on GitHubat commit fdb6a7a
Benchmark Video Search next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Video Search this skillNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 4 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.9k | 14 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
jarrodwatts/claude-code-config
Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…
NVIDIA-AI-Blueprints/video-search-and-summarization
Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.
NVIDIA-AI-Blueprints/video-search-and-summarization
Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…
NVIDIA-AI-Blueprints/video-search-and-summarization
Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when operating VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, verification verdicts, on-demand verification, always-on operation, Slack…
Categories
Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths. Benchmark Video Search is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.
Benchmark Video Search fits situations like: the user asks to benchmark video search; compare search recall between builds; profile search latency on a deployed VSS profile.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a claude-code`. Or copy the skill folder (skills/benchmarking/benchmark-video-search in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/benchmark-video-search in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a codex`. Or copy the skill folder (skills/benchmarking/benchmark-video-search in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/benchmark-video-search in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill benchmark-video-search -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-video-search, .gemini/skills/benchmark-video-search, .github/skills/benchmark-video-search and .opencode/skills/benchmark-video-search in your project.
Going by SKILL.md and its folder, Benchmark Video Search needs Python for the scripts in its folder and the command-line tools its instructions call (uv, curl and python3). Our summary lists: Python 3; Docker.
SKILL.md contains no URLs. Its commands use uv and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Benchmark Video Search is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Benchmark Video Search: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.
Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.