SageMaker Serving Image Selection
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
Agent skill
by NVIDIA-AI-Blueprints in NVIDIA-AI-Blueprints/video-search-and-summarization
Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization profile-vllm-performance --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/profile-vllm-performance .claude/skills/profile-vllm-performance && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "profile-vllm-performance" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/profile-vllm-performance into .claude/skills/profile-vllm-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile-vllm-performance", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/profile-vllm-performanceType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization profile-vllm-performance --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/benchmarking/profile-vllm-performance .agents/skills/profile-vllm-performance && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "profile-vllm-performance" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/profile-vllm-performance into .agents/skills/profile-vllm-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile-vllm-performance", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization profile-vllm-performance --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/benchmarking/profile-vllm-performance .cursor/skills/profile-vllm-performance && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "profile-vllm-performance" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/profile-vllm-performance into .cursor/skills/profile-vllm-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile-vllm-performance", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git --path skills/benchmarking/profile-vllm-performance--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization profile-vllm-performance --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/benchmarking/profile-vllm-performance .gemini/skills/profile-vllm-performance && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "profile-vllm-performance" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/profile-vllm-performance into .gemini/skills/profile-vllm-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile-vllm-performance", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization profile-vllm-performanceInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/benchmarking/profile-vllm-performance .github/skills/profile-vllm-performance && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "profile-vllm-performance" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/profile-vllm-performance into .github/skills/profile-vllm-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile-vllm-performance", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization profile-vllm-performance --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/benchmarking/profile-vllm-performance .opencode/skills/profile-vllm-performance && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "profile-vllm-performance" agent skill from https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization/tree/develop/skills/benchmarking/profile-vllm-performance into .opencode/skills/profile-vllm-performance/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "profile-vllm-performance", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
profile-vllm-performanceProfile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.
Profile Vllm Performance is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure. Use this skill when GPU utilization is unexpectedly low or bursty, capacity trails another runtime, TTFT or ITL regresses, or multimodal prefill and decode appear serialized. Not for running a first benchmark without a reproducible fixed-load signal.
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/identify-serialized-vision.json`, `evals/reject-early-eos-osl-comparison.json` and `evals/reject-offered-load-gap.json`).
It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Profile Vllm Performance loads about 1.5k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 732 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 732 words, ~1,526 tokens.
.claude/skills/profile-vllm-performance/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.Find the dominant loss mechanism, prove it with a correlated timeline, make the smallest causal change, and validate the result with a matched A/B run. A capacity number or average GPU-utilization sample is not bubble attribution.
Record the exact code, container, model revision, precision, hardware, scheduler settings, workload shape, media, prompt, token budget, sampling, cache state, concurrency, and success criteria. Prove semantic correctness before optimizing. Do not compare runtimes when any of these dimensions differ unless that dimension is the declared variable.
nvidia-smi dmon for the resource envelope and a
short Nsight Systems or repository profiler capture for attribution. Profilers
perturb timing; never use a profiled run as the capacity result.Define a performance bubble as GPU-idle or low-occupancy time while eligible work is queued. Idle time caused by insufficient offered load is not a runtime bubble. If the scheduler queue is empty throughout every observed idle interval, first check whether eligible work is waiting upstream in media arrival, decode, preprocessing, or engine submission. Upstream waiting work is frontend or media starvation, not an offered-load gap. When neither the scheduler nor upstream stages have eligible work waiting, classify the result as insufficient offered load and make the first experiment a higher-fixed-concurrency replay with every other workload dimension unchanged. Do not recommend scheduler, kernel, batching, cache, prefill, or decode optimization until that replay shows GPU idle while eligible work remains queued.
Use stable request, sequence, stream, and chunk identifiers across:
Capture only enough NVTX or structured timing to align these phases with CUDA kernels, memory copies, CPU wakeups, and scheduler state. Avoid high-volume logging in the measured path.
Read references/bubble-playbook.md for the evidence matrix and the shortest discriminating experiment for each bubble class.
For each candidate, state the precise idle or under-occupancy interval, whether eligible work was queued, the event immediately preceding it, its frequency and share of measured time, and the observation that rules out the nearest competing explanation. Rank candidates by recoverable critical-path time, not visual prominence in one trace.
Do not infer a decode bottleneck from an OSL capacity gap alone. Compare matched fixed-concurrency prefill and decode timelines plus tokens per scheduler step.
For serialized vision, first compare an existing batched preprocessing or vision- encoding path using the same fixed visual shapes and request set. State its acceptance criteria explicitly: require output parity and multimodal visual-token count, order, and position parity before attributing any performance improvement to the change.
Stop at the first rung that removes the proven bubble:
Change one causal variable at a time. Repeat the short trace, then run an unprofiled matched A/B or A/B/B/A comparison at fixed concurrency. Rerun the capacity boundary only when capacity is the claim. Preserve output parity, ordering, cancellation, KV lifetime, and multimodal token and position invariants.
Report confirmed, improved, inconclusive, or invalid comparison, followed by
the dominant bubble, recoverable-time estimate, evidence, ruled-out alternatives,
fix, before/after distributions, artifact paths, remaining bottleneck, and smallest
next experiment.
Do not install profilers, launch costly GPU work, push changes, or mutate a shared machine without authorization.
© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/benchmarking/profile-vllm-performance of NVIDIA-AI-Blueprints/video-search-and-summarization.
Open the folder on GitHubat commit fdb6a7a
Profile Vllm Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Profile Vllm Performance this skillNVIDIA-AI-Blueprints/video-search-and-summarization | 1.9k | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | |
| SageMaker Serving Image Selectionhuggingface/skills | 11k | 1 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | |
| Hugging Face Local Model Evalshuggingface/skills | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | |
| CI Fails Buildkiteguqiong96/Lvllm | 465 | 2 repos | ~349 | Automated safety check: Pass | Apache-2.0 | |
| Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel | 1.3k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Add Diffusion Modelvllm-project/vllm-omni | 7.1k | — | ~7k | Automated safety check: Pass | Apache-2.0 |
huggingface/skills
Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.
huggingface/skills
Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.
guqiong96/Lvllm
Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.
ModelCloud/GPTQModel
Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.
vllm-project/vllm-omni
Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…
MetaX-MACA/vLLM-metax
Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.
NVIDIA-AI-Blueprints/video-search-and-summarization
Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…
NVIDIA-AI-Blueprints/video-search-and-summarization
Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.
NVIDIA-AI-Blueprints/video-search-and-summarization
Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…
NVIDIA-AI-Blueprints/video-search-and-summarization
Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…
NVIDIA-AI-Blueprints/video-search-and-summarization
A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…
Works with
Categories
Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure. Profile Vllm Performance is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.
Profile Vllm Performance fits situations like: GPU utilization is unexpectedly low; capacity trails another runtime; multimodal prefill and decode appear serialized.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a claude-code`. Or copy the skill folder (skills/benchmarking/profile-vllm-performance in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/profile-vllm-performance in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a codex`. Or copy the skill folder (skills/benchmarking/profile-vllm-performance in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/profile-vllm-performance in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/profile-vllm-performance, .gemini/skills/profile-vllm-performance, .github/skills/profile-vllm-performance and .opencode/skills/profile-vllm-performance in your project.
SKILL.md names no scripts, command-line tools or credentials: Profile Vllm Performance is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Profile Vllm Performance is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Profile Vllm Performance: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.
Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.