LLM Benchmarking with lm-evaluation-harness
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
Agent skill
by NVIDIA-AI-Blueprints in NVIDIA-AI-Blueprints/video-search-and-summarization
Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-evaluate-caption-accuracy --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/vss-evaluate-caption-accuracy .claude/skills/vss-evaluate-caption-accuracy && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "vss-evaluate-caption-accuracy" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/vss-evaluate-caption-accuracy into .claude/skills/vss-evaluate-caption-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vss-evaluate-caption-accuracy", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/vss-evaluate-caption-accuracyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-evaluate-caption-accuracy --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/benchmarking/vss-evaluate-caption-accuracy .agents/skills/vss-evaluate-caption-accuracy && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "vss-evaluate-caption-accuracy" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/vss-evaluate-caption-accuracy into .agents/skills/vss-evaluate-caption-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vss-evaluate-caption-accuracy", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-evaluate-caption-accuracy --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/benchmarking/vss-evaluate-caption-accuracy .cursor/skills/vss-evaluate-caption-accuracy && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "vss-evaluate-caption-accuracy" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/vss-evaluate-caption-accuracy into .cursor/skills/vss-evaluate-caption-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vss-evaluate-caption-accuracy", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git --path skills/benchmarking/vss-evaluate-caption-accuracy--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-evaluate-caption-accuracy --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/benchmarking/vss-evaluate-caption-accuracy .gemini/skills/vss-evaluate-caption-accuracy && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "vss-evaluate-caption-accuracy" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/vss-evaluate-caption-accuracy into .gemini/skills/vss-evaluate-caption-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vss-evaluate-caption-accuracy", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-evaluate-caption-accuracyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/benchmarking/vss-evaluate-caption-accuracy .github/skills/vss-evaluate-caption-accuracy && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "vss-evaluate-caption-accuracy" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/vss-evaluate-caption-accuracy into .github/skills/vss-evaluate-caption-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vss-evaluate-caption-accuracy", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization vss-evaluate-caption-accuracy --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/benchmarking/vss-evaluate-caption-accuracy .opencode/skills/vss-evaluate-caption-accuracy && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "vss-evaluate-caption-accuracy" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/vss-evaluate-caption-accuracy into .opencode/skills/vss-evaluate-caption-accuracy/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "vss-evaluate-caption-accuracy", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
vss-evaluate-caption-accuracyMeasure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…
Vss Evaluate Caption Accuracy is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and emit an accuracy and processing-time table. Use when changing frame selection, decode, or model settings and you need evidence there is no accuracy regression.
Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `evals/evals.json`, `scripts/aggregate_table.py` and `scripts/multi_judge.py`).
It sits in AI & LLM Engineering, covering LLM evaluation. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3a75661. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 5 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
bashdockerpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use docker, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYANTHROPIC_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Vss Evaluate Caption Accuracy loads about 2.1k tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 865 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
ts | Set `MODEL_PATH` in the deployment `.env`. Runs here used `Qwen3-VL-32B-Instruct`; any vLLM-compatible VLM works |SE=vllm-compatible` | In the deployment `.env` || `OPENAI_API_KEY` | In the deployment `.env`. Only needed for the `gt` stage (ground truth is gpt-4.1) |rms. Unset inherits the container's own `.env`. When set, it is also used as `RTVI_MODEL_PATH_ALLOWLIST`, which `VLM_TRUAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit 3a75661, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 865 words, ~2,103 tokens.
.claude/skills/vss-evaluate-caption-accuracy/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.Answer one question with evidence: did this RT-VLM change make captions worse?
A configuration change that saves processing time is only useful if caption quality holds. This skill captures captions twice over the same videos — once with the change (HYP) and once without (REF) — scores both against a ground truth using an LLM judge, and reports accuracy delta alongside time saved.
Not for deploying RT-VLM, general LVS/RTVI benchmarking (see vss-benchmark-video-summarization), or one-off caption spot checks with no paired baseline.
Everything except the judge runs inside the RT-VLM container. The container
needs a GPU, the model on disk, and the nvdsframeselector DeepStream plugin
(shipped with DeepStream in the RT-VLM image).
| Requirement | Notes |
|---|---|
| Running RT-VLM container | Default name rtvi_vlm-$USER. Start it before running any stage |
| Model weights | Set MODEL_PATH in the deployment .env. Runs here used Qwen3-VL-32B-Instruct; any vLLM-compatible VLM works |
VLM_MODEL_TO_USE=vllm-compatible | In the deployment .env |
| Source videos | A directory of .mp4 files. Point DEDUP_DIR at it |
OPENAI_API_KEY | In the deployment .env. Only needed for the gt stage (ground truth is gpt-4.1) |
claude CLI on the host | The judge shells out to it. It is not installed in the container |
Deployment inside requires-vss | >=3.2.0,<4.0.0, from this skill's front matter. Not enforced automatically — this skill has no preflight.sh, and it drives RT-VLM's OpenAI-compatible HTTP surface directly rather than a VSS deployment origin, so there is no origin here to check. Verify by hand if a deployment's version is in doubt: python3 <repo>/services/agent/scripts/check_vss_version.py <deployment-origin> --skill <this SKILL.md> (0 = compatible, 3 = incompatible, 1 = indeterminate) |
scripts/run_captioning.py maps a filename to a short scene name in its SCENES
dict. Add your own videos there:
SCENES = {
"warehouse.mp4": "warehouse",
"GoPro5_10min_compressed.mp4": "new_warehouse",
# "<your-file>.mp4": "<scene-name>",
}Scene names are what you pass on the command line; files are resolved inside
DEDUP_DIR. A scene whose file is missing is skipped with a warning.
The skill does not choose a model — it inherits whatever the container is configured with, and sets only the knobs under test. Both arms run the same model so the comparison isolates the configuration change, not the checkpoint.
They run in different places, so invoke them separately:
| Stage | Where | What |
|---|---|---|
gt | container | Ground truth from gpt-4.1. Expensive — run once per video set and reuse via GT_SRC |
capture | container | Paired REF + HYP per scene, back to back |
judge | host | claude-opus-4-8 scores REF and HYP against GT, per chunk |
table | either | Accuracy + time-saved markdown, optionally against a baseline run |
# 0. ground truth (once per video set)
docker exec -e DESC=my-run -w /workspace rtvi_vlm-$USER \
bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh gt scene-a scene-b
# 1. capture — paired REF + HYP
docker exec -e DESC=my-run -w /workspace rtvi_vlm-$USER \
bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh capture scene-a scene-b
# 2. judge — on the host, not in the container
DESC=my-run bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh judge scene-a scene-b
# 3. table
DESC=my-run bash skills/benchmarking/vss-evaluate-caption-accuracy/scripts/run_eval.sh table scene-a scene-bReplace -w /workspace with the path the repository is mounted at in your
container.
| Var | Default | Meaning |
|---|---|---|
DESC | eval-<date> | Run name. Everything lands in results/<DESC>/ |
VSS_REPO_DIR | /workspace | Path the repository is mounted at inside the container |
RTVI_CONTAINER | rtvi_vlm-$USER | Container name to exec into |
MODEL_PATH | unset | Pin a checkpoint for both arms. Unset inherits the container's own .env. When set, it is also used as RTVI_MODEL_PATH_ALLOWLIST, which VLM_TRUST_REMOTE_CODE=true requires |
DEDUP_DIR | <repo>/videos | Directory holding the source videos |
SFC | unset | NVDS_FSELECT_STATIC_FRAME_COUNT — frames emitted for a chunk classified STATIC. Unset leaves the plugin default |
GT_SRC | $DESC | Run to copy gt.txt from, so one ground truth serves many runs |
BASELINE | none | Run to diff against in table |
RESULTS_ROOT | <skill>/results | Where run folders live |
HYP_VER | v1 | Which hyp_<desc>_vN.txt to judge |
MAX_WORKERS | 32 | Judge concurrency |
REF caption generation is nondeterministic run to run. Judging HYP against a REF
captured in a different session moved per-scene deltas by up to 0.05 — larger than
most effects being measured. capture therefore runs HYP and REF back to back per
scene in one session. Only GT is reused, because it is expensive and comes from a
different model.
table writes results/<DESC>/summary.md: accuracy and processing time per scene,
LLM-judge entity/event F1 detail, and — with BASELINE — the incremental effect
versus that run. Totals are chunk-weighted, so a 60-chunk scene does not carry the
same weight as a 14-chunk one.
Two cautions:
combined_score_macro_0_1
averages entity, event, critical-event and interaction F1. Interaction F1 is often
a handful of samples and carries little signal, so it can mask a real entity-F1
move. Read the F1 detail block, not just the combined column.Frame-level provenance is stronger evidence than matching totals. When frame selection is active the plugin logs one line per chunk:
grep -oE 'EOS OF-only -> [0-9]+' results/<DESC>/server_logs/hyp_<scene>.log \
| grep -oE '[0-9]+$' | tail -n <chunks> | sort -n | uniq -cIdentical per-chunk frame counts across two runs mean the selector chose the same
frames. Matching chunk counts alone do not. The first values belong to pipeline
warmup rather than the scene, so take the trailing <chunks> entries.
scripts/run_eval.sh — the four stagesscripts/run_captioning.py — capture engine (gt / ref / hyp)scripts/multi_judge.py — judge engine; uses the local claude CLI, so no
ANTHROPIC_API_KEY is neededscripts/score.py — per-chunk scoring into summary.csvscripts/aggregate_table.py — the report© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (scripts) in skills/benchmarking/vss-evaluate-caption-accuracy of NVIDIA-AI-Blueprints/video-search-and-summarization.
Open the folder on GitHubat commit 3a75661
Vss Evaluate Caption Accuracy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Vss Evaluate Caption Accuracy this skillNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~2.1k | Automated safety check: Notes | Apache-2.0 | |
| LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs | 13k | 8 repos | ~3k | Automated safety check: Pass | MIT | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| Looperksimback/looper | 710 | — | ~2.7k | Automated safety check: Notes | MIT | |
| Agent Eval Engineeringlangchain-ai/langchain-skills | 1.3k | — | ~4k | Automated safety check: Pass | MIT | |
| Quality FlywheelGoogleCloudPlatform/vertex-ai-samples | 792 | — | ~2k | Automated safety check: Pass | Apache-2.0 |
Orchestra-Research/AI-Research-SKILLs
Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
ksimback/looper
Scaffold a well-designed agent loop with best-practice coaching and a cross-model review council.
langchain-ai/langchain-skills
Builds agent evaluations in stages: inspect the repository and traces, agree a Task Spec with you, then build, audit and run a Harbor task with an independent verifier.
GoogleCloudPlatform/vertex-ai-samples
Evaluate and improve GenAI models and agents using the Google GenAI Evaluation SDK.
cloudnative-co/claude-code-starter-kit
Formal evaluation framework for Claude Code sessions implementing eval-driven development (EDD) principles.
NVIDIA-AI-Blueprints/video-search-and-summarization
Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when a user wants to search archived VSS video or ingest or delete a source for search.
NVIDIA-AI-Blueprints/video-search-and-summarization
Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.
NVIDIA-AI-Blueprints/video-search-and-summarization
Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when operating VSS alert workflows — real-time monitoring, Alert-Bridge subscriptions, verification verdicts, on-demand verification, always-on operation, Slack…
Categories
Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…. Vss Evaluate Caption Accuracy is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and emit an accuracy and processing-time table.
Vss Evaluate Caption Accuracy fits situations like: changing frame selection; model settings and you need evidence there is no accuracy regression.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a claude-code`. Or copy the skill folder (skills/benchmarking/vss-evaluate-caption-accuracy in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/vss-evaluate-caption-accuracy in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a codex`. Or copy the skill folder (skills/benchmarking/vss-evaluate-caption-accuracy in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/vss-evaluate-caption-accuracy in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill vss-evaluate-caption-accuracy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/vss-evaluate-caption-accuracy, .gemini/skills/vss-evaluate-caption-accuracy, .github/skills/vss-evaluate-caption-accuracy and .opencode/skills/vss-evaluate-caption-accuracy in your project.
Going by SKILL.md and its folder, Vss Evaluate Caption Accuracy needs Python and a shell for the scripts in its folder, the command-line tools its instructions call (bash, docker and python3) and credentials named OPENAI_API_KEY and ANTHROPIC_API_KEY. Our summary lists: Python 3; A Bash shell; Docker; A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY.
SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Vss Evaluate Caption Accuracy is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Vss Evaluate Caption Accuracy: LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), Looper (ksimback/looper, 710 stars) and Agent Eval Engineering (langchain-ai/langchain-skills, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,917 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 9, 2026.
Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.