Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Profile Vllm Performance

skills CLI
$ npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA-AI-Blueprints/video-search-and-summarization profile-vllm-performance --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmarking/profile-vllm-performance .claude/skills/profile-vllm-performance && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
profile-vllm-performance
GitHub stars
1.9k
Token cost
~1.5k tokens
SKILL.md length
732 words
Files
5 (incl. references)
Skills in repo
22
Repo updated
First seen
Licence
Apache-2.0

At a glance

Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.

  • Works in 4 steps: Verify actual output-token counts,… → Reproduce at fixed load below the knee,… → Record TTFT, inter-token latency,… → …
  • GPU utilization is unexpectedly low
  • SKILL.md covers Freeze The Comparison, Establish The Signal, Correlate The Timeline and Prove The Dominant Bubble, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Profile Vllm Performance is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure. Use this skill when GPU utilization is unexpectedly low or bursty, capacity trails another runtime, TTFT or ITL regresses, or multimodal prefill and decode appear serialized. Not for running a first benchmark without a reproducible fixed-load signal.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `evals/identify-serialized-vision.json`, `evals/reject-early-eos-osl-comparison.json` and `evals/reject-offered-load-gap.json`).

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM. The repository describes itself as: NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts… The licence is Apache-2.0.

When your agent uses it

  • GPU utilization is unexpectedly low
  • Capacity trails another runtime
  • Multimodal prefill and decode appear serialized

Example prompts

  • “/profile-vllm-performance”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Verify actual output-token counts, successful work, fresh-stream coverage,
  2. Reproduce at fixed load below the knee, near the highest stable point, and just
  3. Record TTFT, inter-token latency, throughput, queue delay, batch size, scheduled
  4. Use scenario-scoped DCGM or nvidia-smi dmon for the resource envelope and a

What it can do on your machine

Read from SKILL.md and the folder at commit fdb6a7a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Profile Vllm Performance loads about 1.5k tokens when it runs, and up to ~3k if it reads all its reference files. Until then it costs about 111 tokens; SKILL.md has 732 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA-AI-Blueprints/video-search-and-summarization at commit fdb6a7a, republished under its Apache-2.0 licence (© NVIDIA-AI-Blueprints). 732 words, ~1,526 tokens.

Download SKILL.mdSave it as .claude/skills/profile-vllm-performance/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
profile-vllm-performance
description
Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure. Use this skill when GPU utilization is unexpectedly low or bursty, capacity trails another runtime, TTFT or ITL regresses, or multimodal prefill and decode appear serialized. Not for running a first benchmark without a reproducible fixed-load signal.
license
Apache-2.0
metadata.version
3.3.0-rc0
metadata.requires-vss
>=3.2.0,<4.0.0
metadata.author
NVIDIA Video Search and Summarization Team
metadata.github-url
https://github.com/NVIDIA-AI-Blueprints/video-search-and-summarization
metadata.tags
nvidia rt-vlm vllm performance profiling gpu

Profile vLLM Performance

Find the dominant loss mechanism, prove it with a correlated timeline, make the smallest causal change, and validate the result with a matched A/B run. A capacity number or average GPU-utilization sample is not bubble attribution.

Freeze The Comparison

Record the exact code, container, model revision, precision, hardware, scheduler settings, workload shape, media, prompt, token budget, sampling, cache state, concurrency, and success criteria. Prove semantic correctness before optimizing. Do not compare runtimes when any of these dimensions differ unless that dimension is the declared variable.

Establish The Signal

  1. Verify actual output-token counts, successful work, fresh-stream coverage, dropped chunks, cancellations, and cleanup. Early EOS or silent failure is not throughput.
  2. Reproduce at fixed load below the knee, near the highest stable point, and just beyond it when practical. Keep capacity search separate from profiling.
  3. Record TTFT, inter-token latency, throughput, queue delay, batch size, scheduled tokens, active sequences, preemptions, KV occupancy, GPU memory, and phase durations available in the checked-out version.
  4. Use scenario-scoped DCGM or nvidia-smi dmon for the resource envelope and a short Nsight Systems or repository profiler capture for attribution. Profilers perturb timing; never use a profiled run as the capacity result.

Define a performance bubble as GPU-idle or low-occupancy time while eligible work is queued. Idle time caused by insufficient offered load is not a runtime bubble. If the scheduler queue is empty throughout every observed idle interval, first check whether eligible work is waiting upstream in media arrival, decode, preprocessing, or engine submission. Upstream waiting work is frontend or media starvation, not an offered-load gap. When neither the scheduler nor upstream stages have eligible work waiting, classify the result as insufficient offered load and make the first experiment a higher-fixed-concurrency replay with every other workload dimension unchanged. Do not recommend scheduler, kernel, batching, cache, prefill, or decode optimization until that replay shows GPU idle while eligible work remains queued.

Correlate The Timeline

Use stable request, sequence, stream, and chunk identifiers across:

  • media arrival and decode;
  • CPU frame selection, resize, normalization, tokenization, and processor work;
  • host-to-device transfer and synchronization;
  • vision encoder and projector;
  • text and visual prefill;
  • scheduler admission, batch formation, KV allocation, preemption, and dispatch;
  • decode, sampling, detokenization, and output delivery.

Capture only enough NVTX or structured timing to align these phases with CUDA kernels, memory copies, CPU wakeups, and scheduler state. Avoid high-volume logging in the measured path.

Read references/bubble-playbook.md for the evidence matrix and the shortest discriminating experiment for each bubble class.

Show full SKILL.md (319 more words)Show less

Prove The Dominant Bubble

For each candidate, state the precise idle or under-occupancy interval, whether eligible work was queued, the event immediately preceding it, its frequency and share of measured time, and the observation that rules out the nearest competing explanation. Rank candidates by recoverable critical-path time, not visual prominence in one trace.

Do not infer a decode bottleneck from an OSL capacity gap alone. Compare matched fixed-concurrency prefill and decode timelines plus tokens per scheduler step.

For serialized vision, first compare an existing batched preprocessing or vision- encoding path using the same fixed visual shapes and request set. State its acceptance criteria explicitly: require output parity and multimodal visual-token count, order, and position parity before attributing any performance improvement to the change.

Fix And Validate

Stop at the first rung that removes the proven bubble:

  1. Correct benchmark or configuration defects such as early EOS, mismatched shapes, cache policy, scheduler limits, or insufficient offered load.
  2. Tune an existing runtime knob supported by the checked-out version.
  3. Reuse an existing batching, preprocessing, cache, CUDA-graph, tensor-IPC, or scheduler path that is configured incorrectly or bypassed.
  4. Fix an integration boundary that serializes or copies work unnecessarily.
  5. Change scheduling or model-runner code only when traces prove the existing path cannot express the required overlap.
  6. Change a kernel only after ruling out launch, scheduling, shape, and data movement.

Change one causal variable at a time. Repeat the short trace, then run an unprofiled matched A/B or A/B/B/A comparison at fixed concurrency. Rerun the capacity boundary only when capacity is the claim. Preserve output parity, ordering, cancellation, KV lifetime, and multimodal token and position invariants.

Report confirmed, improved, inconclusive, or invalid comparison, followed by the dominant bubble, recoverable-time estimate, evidence, ruled-out alternatives, fix, before/after distributions, artifact paths, remaining bottleneck, and smallest next experiment.

Do not install profilers, launch costly GPU work, push changes, or mutate a shared machine without authorization.

© NVIDIA-AI-Blueprints, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/benchmarking/profile-vllm-performance of NVIDIA-AI-Blueprints/video-search-and-summarization.

  • SKILL.md
  • evals/identify-serialized-vision.json
  • evals/reject-early-eos-osl-comparison.json
  • evals/reject-offered-load-gap.json
  • references/bubble-playbook.md

Open the folder on GitHubat commit fdb6a7a

Compare with similar skills

Profile Vllm Performance next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Profile Vllm Performance compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Profile Vllm Performance this skillNVIDIA-AI-Blueprints/video-search-and-summarization1.9k—~1.5kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
CI Fails Buildkiteguqiong96/Lvllm4652 repos~349Automated safety check: PassApache-2.0
Gptqmodel Tokenizer NormalizationModelCloud/GPTQModel1.3k—~1.1kAutomated safety check: PassCustom licence
Add Diffusion Modelvllm-project/vllm-omni7.1k—~7kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • CI Fails Buildkite

    guqiong96/Lvllm

    Fetch and diagnose vLLM Buildkite CI failure logs. An agent skill from guqiong96/Lvllm.

    465 GitHub starsUsed in 2 repos~349 tokens
    AI & LLM EngineeringAuto-check passed
  • Diagnose and correct GPT-QModel tokenizer initialization, tokenization normalization, special-token handling, prompt rendering, and chat-template problems.

    1.3k GitHub stars~1.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Add Diffusion Model

    vllm-project/vllm-omni

    Add a new diffusion model (text-to-image, text-to-video, image-to-video, text-to-audio, image editing) to vLLM-Omni, including native non-Diffusers ports, reference-parity validation, Cache-DiT…

    7.1k GitHub stars~7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Metax Model Upgrade

    MetaX-MACA/vLLM-metax

    Review and upgrade MetaX model support against a target vLLM revision and installed MACA components, recursively including model-dependent attention and kernels.

    180 GitHub stars~3.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from NVIDIA-AI-Blueprints/video-search-and-summarization

All 22 skills in this repo
  • Benchmark Video Search

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure retrieval quality and latency of a deployed VSS search profile by ingesting a labelled dataset and running the vss CLI across retrieval paths.

    1.9k GitHub stars~4.3k tokensUpdated yesterday
    Auto-check passed
  • Vss Search Archive

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when a user wants to search archived VSS video that is already registered in a configured deployment — by natural-language, similarity, attribute, object-ID, or lexical tag…

    1.9k GitHub stars~3.3k tokensUpdated yesterday
    Auto-check passed
  • Rtvi Vlm Perf Testing

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Plan, run, and diagnose reproducible RT-VLM GPU performance canaries and benchmarks.

    1.9k GitHub stars~8.6k tokensUpdated yesterday
    Auto-check: notes
  • Vss Build Vision AI

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Add agent-ready vision capabilities — dense captioning, detection, search, alerting, summarization — to an agent or application through a customizable, self-contained vision stack built on the…

    1.9k GitHub stars~15k tokensUpdated yesterday
    Auto-check: notes
  • Vss Evaluate Caption Accuracy

    NVIDIA-AI-Blueprints/video-search-and-summarization

    Measure whether an RT-VLM configuration change altered caption quality — capture paired baseline and candidate captions for a set of videos, score both against a ground truth with an LLM judge, and…

    1.9k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check: notes
  • Rtvi Byom Porting

    NVIDIA-AI-Blueprints/video-search-and-summarization

    A skill your agent uses when adding, debugging, or validating a bring-your-own VLM in VSS RT-VLM, including custom Hugging Face or NGC checkpoints, vLLM adapters or plugins, model shims, and…

    1.9k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Profile Vllm Performance

What does Profile Vllm Performance do?

Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure. Profile Vllm Performance is an agent skill from NVIDIA-AI-Blueprints/video-search-and-summarization. Profile vLLM and RT-VLM inference to identify and remove GPU-idle gaps, underfilled batches, transfer stalls, serialized multimodal work, scheduler gaps, or KV pressure.

When should I use Profile Vllm Performance?

Profile Vllm Performance fits situations like: GPU utilization is unexpectedly low; capacity trails another runtime; multimodal prefill and decode appear serialized.

How do I install Profile Vllm Performance in Claude Code?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a claude-code`. Or copy the skill folder (skills/benchmarking/profile-vllm-performance in NVIDIA-AI-Blueprints/video-search-and-summarization) into .claude/skills/profile-vllm-performance in your project. Claude Code loads it when a task matches its description.

How do I install Profile Vllm Performance in Codex?

Run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a codex`. Or copy the skill folder (skills/benchmarking/profile-vllm-performance in NVIDIA-AI-Blueprints/video-search-and-summarization) into .agents/skills/profile-vllm-performance in your project. Codex loads it when a task matches its description.

Can I use Profile Vllm Performance in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA-AI-Blueprints/video-search-and-summarization --skill profile-vllm-performance -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/profile-vllm-performance, .gemini/skills/profile-vllm-performance, .github/skills/profile-vllm-performance and .opencode/skills/profile-vllm-performance in your project.

What does Profile Vllm Performance need to run?

SKILL.md names no scripts, command-line tools or credentials: Profile Vllm Performance is instructions for the agent only.

Does Profile Vllm Performance access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Profile Vllm Performance safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Profile Vllm Performance use?

Profile Vllm Performance is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Profile Vllm Performance use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.4k tokens, read only when the agent opens those files.

What are the alternatives to Profile Vllm Performance?

Skills that share tags, products or a category with Profile Vllm Performance: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars), CI Fails Buildkite (guqiong96/Lvllm, 465 stars) and Gptqmodel Tokenizer Normalization (ModelCloud/GPTQModel, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Profile Vllm Performance?

NVIDIA-AI-Blueprints (a GitHub organization) maintains it in NVIDIA-AI-Blueprints/video-search-and-summarization, which has 1,919 GitHub stars. The repository holds 22 skills in this directory. The repository was last updated on October 10, 2026.

Source: NVIDIA-AI-Blueprints/video-search-and-summarization on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.