Official agent skill

Deepstream Profile Pipeline

by NVIDIA in NVIDIA/skills

Profile a DeepStream pipeline with Nsight Systems and derive its configs from the measurement.

OfficialApache-2.0Auto-check passedAI & LLM Engineering

Install Deepstream Profile Pipeline

skills CLI
$ npx skills add NVIDIA/skills --skill deepstream-profile-pipeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills deepstream-profile-pipeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/deepstream-profile-pipeline .claude/skills/deepstream-profile-pipeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
deepstream-profile-pipeline
GitHub stars
3.5k
Token cost
~4.1k tokens
SKILL.md length
1,467 words
Files
16 (incl. scripts, references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Profile a DeepStream pipeline with Nsight Systems and derive its configs from the measurement.

  • The user asks for an efficient
  • SKILL.md covers When to trigger, The 6-stage flow, Stage 0 — Preset-apply (at… and The verification flow (Stages…, plus 4 more sections
  • Runs Python scripts from its folder
  • Profiled pipeline —

What it does

Deepstream Profile Pipeline is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Profile a DeepStream pipeline with Nsight Systems and derive its configs from the measurement. Use when the user asks for an efficient, performant, or profiled pipeline — or to benchmark, tune, or measure FPS.

Its SKILL.md is about 4.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts and reference files (for example `BENCHMARK.md`, `README.md` and `evals/evals.json`). Compatibility notes: DeepStream SDK 9.0 on Ubuntu 22.04 or 24.04, run from the nvcr.io/nvidia/deepstream:9.0-triton-multiarch container (the dev image; the slimmer…

It sits in AI & LLM Engineering. It works with NVIDIA AI Platform. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • The user asks for an efficient
  • Profiled pipeline —

Example prompts

  • “/deepstream-profile-pipeline”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): DeepStream SDK 9.0 on Ubuntu 22.04 or 24.04, run from the `nvcr.io/nvidia/deepstream:9.0-triton-multiarch` container (the dev image; the slimmer `samples-multiarch` variant strips the nsys NVTX injector and produces empty per-plugin NVTX traces — do not use it for profiling). Requires `nsys` (Nsight Systems 2024+) and `nvidia-smi` on PATH. No GUI dependency — the skill runs fully headless and uses only `nsys profile` + `nsys stats`.

What it can do on your machine

Read from SKILL.md and the folder at commit 67a13c0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    DeepStream SDK 9.0 on Ubuntu 22.04 or 24.04, run from the `nvcr.io/nvidia/deepstream:9.0-triton-multiarch` container (the dev image; the slimmer `samples-multiarch` variant strips the nsys NVTX injector and produces empty per-plugin NVTX traces — do not use it for profiling). Requires `nsys` (Nsight Systems 2024+) and `nvidia-smi` on PATH. No GUI dependency — the skill runs fully headless and uses only `nsys profile` + `nsys stats`.

    From compatibility in the SKILL.md frontmatter.

Context cost

Deepstream Profile Pipeline loads about 4.1k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 59 tokens; SKILL.md has 1,467 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~4.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 67a13c0, republished under its Apache-2.0 licence (© NVIDIA). 1,467 words, ~4,066 tokens.

Download SKILL.mdSave it as .claude/skills/deepstream-profile-pipeline/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
deepstream-profile-pipeline
description
Profile a DeepStream pipeline with Nsight Systems and derive its configs from the measurement. Use when the user asks for an efficient, performant, or profiled pipeline — or to benchmark, tune, or measure FPS.
compatibility
DeepStream SDK 9.0 on Ubuntu 22.04 or 24.04, run from the `nvcr.io/nvidia/deepstream:9.0-triton-multiarch` container (the dev image; the slimmer `samples-multiarch` variant strips the nsys NVTX injector and produces empty per-plugin NVTX traces — do not use it for profiling). Requires `nsys` (Nsight Systems 2024+) and `nvidia-smi` on PATH. No GUI dependency — the skill runs fully headless and uses only `nsys profile` + `nsys stats`.
metadata.author
NVIDIA CORPORATION
metadata.tags
deepstream, profiling, nsight-systems, nvtx, nvidia-smi, benchmarking
metadata.languages
bash, python, yaml
metadata.domain
video-analytics
metadata.team
deepstream-sdk
owner
NVIDIA CORPORATION
service
deepstream
version
0.1.0
reviewed
2026-04-24
license
CC-BY-4.0 AND Apache-2.0

DeepStream Profiling Skill

Profile-driven pipeline creation. When the user indicates they want an efficient DeepStream pipeline, this skill replaces guesswork with two measured numbers — inference plateau batch and HW ceiling — and derives every other config from them. Then it profiles the E2E pipeline with Nsight Systems and reports per-plugin NVTX timings.

Model- and pipeline-agnostic. The skill assumes only that the inference element is nvinfer or nvinferserver (so model dims, precision, and batch knobs are settable through the standard config). It works for detection (with or without tracker), classification, segmentation, VLM, and embedding pipelines. Source can be file, RTSP, USB camera, or any mix. The skill reads the user's actual config to discover model dims / target FPS / source properties — it does NOT assume any particular model, codec, or resolution.

Constraint. Terminal only. Use nsys profile to capture and nsys stats to extract. Do not depend on Nsight Lens or any GUI.

When to trigger

Activate this skill at pipeline creation time when the user's ask carries efficiency intent. Concrete triggers:

  • "build an efficient / fast / performant / optimized pipeline"
  • "give me a pipeline that runs well on this GPU"
  • "benchmark / profile / measure / tune / optimize this pipeline"
  • "I want to run N streams at M FPS"
  • "how many streams can this GPU handle"
  • user explicitly asks for nsys or Nsight

For plain "build a pipeline" / "display this video" / "save this stream" with no perf intent, hand off to the deepstream-generate-pipeline skill instead.

The 6-stage flow

Run the stages in order. Stage 0 fires before the pipeline is generated, so the user starts from a perf-tuned skeleton. Stages 1–5 measure and verify.

Stage 0 — Preset-apply (at pipeline-creation time)

Trigger: any time the coding agent is about to generate a new DS pipeline AND the user's prompt carries efficiency intent (see "When to trigger" above).

Action: pre-apply these defaults without prompting. The user does not need to know any of them; they just get a pipeline that's already in the right shape.

KnobDefault valueSkip when
nvinfer.network-mode1 (INT8) if a calibration file is present at int8-calib-file=<path>, else 2 (FP16). Never FP32.Model has no INT8 calibration AND the user explicitly says "FP32".
nvinfer.model-engine-filePre-built .engine pathAlways set. Force a one-shot prebuild before measurement.
nvinfer.infer-dims3;<H>;<W> matching the model's native inputAlways set, even for static-shape ONNX (harmless).
nvstreammux.batch-sizemin(N_streams, 16) until microbench refines it—
nvstreammux.width / heightmodel's native input dims (read from the nvinfer config's infer-dims=3;H;W)User explicitly asks for native source resolution at the muxer.
nvstreammux.batched-push-timeout1e6 / source_fps µs (33333 for 30 fps)—
nvstreammux.nvbuf-memory-type0 (NVMM)—
Decoder num-extra-surfacesmin(batch_size, 5)—
Decoder cudadec-memtype0 (NVMM)—
Sinkfakesink sync=False for the benchmark variantUser asked for on-screen display or on-disk recording (then keep OSD/tiler/encoder/sink and produce TWO variants).
OSD + tileromitUser asked for visible output.
Tracker ll-config-fileconfig_tracker_NvDCF_max_perf.yml (perf-tuned NvDCF preset shipped with DS 9.0)Tracker not present.
Tracker tracker-width / height480 / 288—
Tracker enable-batch-process (in linked YAML)1—
Queue between source and pgiemax-size-buffers = batch_size × 4No queue requested (rare).
Kafka/message queuemax-size-buffers=2, leaky=2No Kafka.
Decode-side PerfMonitorattach (in addition to pgie-side)Pipeline is nvurisrcbin → pgie direct without intermediate queue.

Why Stage 0 exists: without it, every newly generated pipeline starts from display-first defaults and Stages 1–5 spend cycles fixing avoidable issues. Stage 0 is the "don't write a bad pipeline in the first place" gate.

The student / API user never sees these knobs. The skill's response back to the user is in plain English (FPS, stream count, observed bottleneck), not knob names.

The verification flow (Stages 1–5)

Run the stages in order. Do not skip a stage — later stages depend on earlier ones' outputs.

Stage 1 — NVTX coverage check

DeepStream plugins emit NVTX ranges natively; custom plugins and plain GStreamer-core elements (queue, tee, h264parse, etc.) do not. Before profiling, list the elements the pipeline uses and classify each.

  • Read the pipeline definition (gst-launch string or pipeline.py).
  • For each element, look it up in references/nvtx-coverage.md.
  • Classify COVERED (emits NVTX in this DS / image / nsys combo) or UNINSTRUMENTED.
  • MVP rule: the skill prefers per-plugin NVTX as confirmation but does not require it. Decode-bound diagnosis works from microbench shape + nvidia-smi dmon; compute-bound from CUDA kernel mix; memcpy from cuda_gpu_mem_time_sum. NVTX is a bonus.
  • For UNINSTRUMENTED elements, the skill reports "not directly measurable in this build" and still applies the closed-form R1–R6 knobs (which are derived from inputs, not from per-plugin profile data).
  • Auto-injecting NVTX for uninstrumented elements is out of scope for this version — flag it as follow-up in the final report.

Output of Stage 1: a short coverage table, e.g.

text
nvurisrcbin       COVERED
nvstreammux       COVERED
nvinfer           COVERED
nvtracker         COVERED
queue_src         UNINSTRUMENTED — not re-tuned
fakesink          UNINSTRUMENTED — not re-tuned
Stage 2 — HW discovery

Run nvidia-smi and derive theoretical ceilings for the host GPU. Minimum queries:

bash
# Identity + memory + compute
nvidia-smi --query-gpu=name,compute_cap,memory.total,memory.free,\
clocks.max.sm,clocks.max.memory,utilization.gpu \
--format=csv,noheader,nounits

# NVDEC / NVENC utilization (per-engine)
nvidia-smi --query-gpu=utilization.decoder,utilization.encoder \
--format=csv,noheader,nounits

# PCIe link width/gen (for H2D memcpy ceiling)
nvidia-smi --query-gpu=pcie.link.gen.current,pcie.link.width.current \
--format=csv,noheader,nounits

Derive from those numbers:

  • Decode ceiling (fps): NVDEC_count × per-unit H265/H264 fps for the source resolution (table in references/hw-ceiling-formulas.md).
  • Compute ceiling (TOPS): SM count × clock × ops-per-clock at the target precision. Gives an upper bound — real models hit 30–60% of this.
  • Memory-bandwidth ceiling (GB/s): memory clock × bus width. Model weight reads + activations should fit well under this.
  • Memcpy ceiling (GB/s): PCIe gen × width × 0.8 practical. Only relevant if NVMM is broken and H2D/D2H transfers appear in Stage 5.

Store the derived ceilings — they drive the Stage 5 "actual vs. theoretical" section.

Full formulas and the per-codec NVDEC throughput table: references/hw-ceiling-formulas.md.

Show full SKILL.md (621 more words)Show less
Stage 3 — Inference-only micro-benchmark

Run only the inference stage (source → streammux → nvinfer → fakesink), sweeping batch-size to find the plateau. This isolates the model's true peak FPS from everything else, and answers "how many streams fit into a single batch without FPS dropping?".

Sweep: batch-size ∈ {1, 2, 4, 8, 16, 32} (cap at N_streams and at GPU memory).

For each batch size:

  • Set nvstreammux.batch-size = nvinfer.batch-size = B.
  • Set nvstreammux.width/height = the model's native infer-dims (read from the nvinfer config).
  • fakesink sync=False as the only branch.
  • Run 30 s; measure FPS from measure_fps_probe (console) or DS PerfMonitor.
  • Record (B, fps).

Plateau batch = the smallest B where increasing to 2×B yields < 5% FPS gain. That is the target batch for the full pipeline.

If the user's N_streams ≤ plateau batch, set final batch = N_streams. Otherwise set final batch = plateau batch and note that the pipeline will process streams in multiple batches per tick.

Stage 4 — Derive configs

From (plateau_batch, HW_ceilings, N_streams, source_res, source_fps), set every tunable knob at once. Do not tune one knob at a time — the derivation rules are closed-form.

Knobs to set, in order:

  1. Streammux: batch-size = final_batch, width/height = min(source_res, infer_dims), batched-push-timeout = 1e6 / source_fps µs, nvbuf-memory-type = 0.
  2. Inference: batch-size = final_batch, network-mode = 1 (INT8) if calib file exists else 2 (FP16), interval = 0, infer-dims = model's native dims, model-engine-file = pre-built .engine path.
  3. Decoder (on nvurisrcbin / nvmultiurisrcbin / nvv4l2decoder): num-extra-surfaces = min(final_batch, 5), cudadec-memtype = 0, nvbuf-memory-type = 0.
  4. Tracker (if present): enable-batch-process = 1, tracker res 480×288, point ll-config-file at config_tracker_NvDCF_max_perf.yml.
  5. Queues (if present between decoder and streammux, or streammux and nvinfer): max-size-buffers = final_batch × 2. Kafka/message branches: leaky=2, max-size-buffers=2.

Full derivation table with each formula and a one-line "why": references/config-derivation-rules.md.

Write the derived values into the user's config files (pgie_config.yml, tracker_config.yml, pipeline.py source properties, any deepstream-app .txt). Always Read before Edit. Keep edits surgical — do not reformat unrelated lines.

Stage 5 — E2E profile + report

Run the E2E pipeline under nsys profile and extract per-plugin timings via nsys stats.

Capture:

bash
TS=$(date +%Y%m%d_%H%M%S)
OUT=/tmp/ds_profile_${TS}
nsys profile \
  --trace=cuda,nvtx,osrt \
  --gpu-metrics-devices=all \
  --cuda-memory-usage=true \
  --force-overwrite=true \
  --duration=30 \
  --output=${OUT} \
  <your-pipeline-launch-command>

Extract:

bash
# Per-kernel GPU time (top 10)
nsys stats --report cuda_gpu_kern_sum --format csv ${OUT}.nsys-rep | head -20

# Per-NVTX-range time (top 10) — this is the DS per-plugin breakdown
nsys stats --report nvtx_sum --format csv ${OUT}.nsys-rep | head -20

# Memcpy totals
nsys stats --report cuda_gpu_mem_time_sum --format csv ${OUT}.nsys-rep

# GPU metrics (SM activity, DRAM throughput) — requires --gpu-metrics-devices
nsys stats --report gpu_metric_gpu_util_sum --format csv ${OUT}.nsys-rep

Full command reference: references/nsys-cli-recipes.md.

Report (Markdown, to stdout — no external UI):

markdown
## Profile summary

**Hardware**: <name>, <mem_total> GB, SM x<sm>, NVDEC x<nvdec>, PCIe Gen<g> x<w>
**Ceilings**: decode <X> fps, compute ~<Y> TOPS @ INT8, memory <Z> GB/s

**Inference plateau**: batch=<B>, peak=<F> fps per batch → <F × B> fps aggregate

**E2E measured**: <actual> fps  (=<pct>% of inference plateau)

### Per-plugin time (from NVTX) — only for plugins emitting NVTX in this build

| Plugin          | Share of wall time | GPU / CPU | Notes |
|-----------------|--------------------|-----------|-------|
| nvinfer         | <pct>%             | GPU       | (always emitted; if absent, NVTX injection is broken) |
| nvdsosd         | <pct>%             | GPU       | (when in pipeline) |
| ...             | ...                | ...       | (other plugins as the verification probe shows) |

(Numbers above are illustrative — fill in from `nsys stats --report nvtx_sum`. Plugins
that don't emit NVTX in your DS / image combo simply don't appear; that's not a bug, it's
the limit of what NVTX captures here. See `references/nvtx-coverage.md`.)

### Applied configs (sample shape; values come from R1–R6 + Stage 3 measurements)

- `nvstreammux.batch-size = <plateau_batch>`
- `nvinfer.network-mode = 1 (INT8)` if calibration available, else `2 (FP16)`
- decoder `num-extra-surfaces = min(plateau_batch, 5)`
- queue between source and pgie, `max-size-buffers = plateau_batch × 4`
- ... (full list per the user's pipeline shape)

### Uninstrumented (skipped re-tune)

List the elements that didn't emit NVTX in this build (typically the closed-source binary
plugins — see `references/nvtx-coverage.md`) plus plain GStreamer-core helpers. Report
them so the user knows what wasn't directly measurable.

Keep the summary terse. Raw nsys stats CSV goes into the temp file, not the response.

Reference documents

DocumentUse when
references/nvtx-coverage.mdStage 1 — classifying each pipeline element as COVERED or UNINSTRUMENTED.
references/hw-ceiling-formulas.mdStage 2 — turning nvidia-smi output into decode / compute / memory ceilings.
references/config-derivation-rules.mdStage 4 — per-knob formula keyed to (plateau_batch, HW, N_streams, source_res, source_fps).
references/nsys-cli-recipes.mdStages 3 & 5 — exact nsys profile / nsys stats invocations.

Non-goals (this version)

  • No Nsight Lens / no GUI. Terminal only.
  • No NVTX auto-injection for uninstrumented plugins. MVP skips their knobs. Future work.
  • No iterative tune-measure-tune loop. Stage 4 derives configs once from closed-form rules; Stage 5 measures and reports. If the user wants to keep tuning, they can re-invoke the skill with updated inputs.
  • deepstream-generate-pipeline — upstream pipeline generation. This skill assumes a pipeline already exists or is about to be generated.
  • deepstream-byovm — HF → TensorRT engine building. Run first if the user brought a new model; come here after.

Notes

  • Lives in skills/deepstream-profile-pipeline/ alongside the other DS skills, per the repo convention in CLAUDE.md.
  • For ground-truth on any plugin's properties (types, defaults, ranges) and pad caps, query the loaded binary inside the DS container:
    bash
    gst-inspect-1.0 nvinfer
    gst-inspect-1.0 nvstreammux
    gst-inspect-1.0 nvurisrcbin   # works on closed-source binary plugins too
    gst-inspect-1.0 | grep ^nv    # list every NVIDIA-specific element this build ships
    Plugin naming convention: any element prefixed nv* is NVIDIA DeepStream-specific (NVMM-capable, may emit NVTX); everything else is upstream GStreamer-core (no NVMM, never emits DS NVTX). Use this prefix as the first-pass classifier when triaging an unfamiliar pipeline.
  • The open-source subset of plugin code lives under /opt/nvidia/deepstream/deepstream/sources/gst-plugins/ if you need to read the implementation (only some plugins are open — closed ones must be inspected via gst-inspect-1.0 and behaviour observed at runtime).
<!-- Signing refresh marker. -->

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (scripts, references) in skills/deepstream-profile-pipeline of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • README.md
  • evals/evals.json
  • references/boundedness-rules.md
  • references/config-derivation-rules.md
  • references/hw-ceiling-formulas.md
  • references/nsys-cli-recipes.md
  • references/nvtx-coverage.md
  • requirements.txt
  • scripts/capacity_report.py
  • skill-card.md
  • skill.oms.sig
  • tests/README.md
  • tests/__init__.py
  • tests/test_capacity_report.py

Open the folder on GitHubat commit 67a13c0

Compare with similar skills

Deepstream Profile Pipeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Deepstream Profile Pipeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Deepstream Profile Pipeline this skillNVIDIA/skills3.5k—~4.1kAutomated safety check: PassApache-2.0
Fla Triton To Gluonfla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMIT
DGX Spark Memory and Thermal Opswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
DGX Spark Training Gotchaswshobson/agents40k1 repos~2kAutomated safety check: PassMIT
Nemotron Add StepNVIDIA-NeMo/Nemotron2.1k—~1.7kAutomated safety check: PassApache-2.0
Yolo Detection 2026SharpAI/DeepCamera3.1k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Fla Triton To Gluon

    fla-org/flash-linear-attention

    Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

    5.8k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

    40k GitHub starsUsed in 1 repo~2k tokens
    AI & LLM EngineeringAuto-check passed
  • Nemotron Add Step

    NVIDIA-NeMo/Nemotron

    Add a new step under src/nemotron/steps/<category/<stepid/ — manifest (step.toml), runner glue, configs, and per-step README.md.

    2.1k GitHub stars~1.7k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Yolo Detection 2026

    SharpAI/DeepCamera

    YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera.

    3.1k GitHub stars~1.5k tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion.

    37k GitHub starsUsed in 2 repos~5k tokens
    AI & LLM EngineeringAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Questions about Deepstream Profile Pipeline

What does Deepstream Profile Pipeline do?

Profile a DeepStream pipeline with Nsight Systems and derive its configs from the measurement. Deepstream Profile Pipeline is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Profile a DeepStream pipeline with Nsight Systems and derive its configs from the measurement.

When should I use Deepstream Profile Pipeline?

Deepstream Profile Pipeline fits situations like: the user asks for an efficient; profiled pipeline —.

How do I install Deepstream Profile Pipeline in Claude Code?

Run `npx skills add NVIDIA/skills --skill deepstream-profile-pipeline -a claude-code`. Or copy the skill folder (skills/deepstream-profile-pipeline in NVIDIA/skills) into .claude/skills/deepstream-profile-pipeline in your project. Claude Code loads it when a task matches its description.

How do I install Deepstream Profile Pipeline in Codex?

Run `npx skills add NVIDIA/skills --skill deepstream-profile-pipeline -a codex`. Or copy the skill folder (skills/deepstream-profile-pipeline in NVIDIA/skills) into .agents/skills/deepstream-profile-pipeline in your project. Codex loads it when a task matches its description.

Can I use Deepstream Profile Pipeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill deepstream-profile-pipeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/deepstream-profile-pipeline, .gemini/skills/deepstream-profile-pipeline, .github/skills/deepstream-profile-pipeline and .opencode/skills/deepstream-profile-pipeline in your project.

What does Deepstream Profile Pipeline need to run?

Going by SKILL.md and its folder, Deepstream Profile Pipeline needs Python for the scripts in its folder. Our summary lists: Python 3. Compatibility (from SKILL.md): DeepStream SDK 9.0 on Ubuntu 22.04 or 24.04, run from the `nvcr.io/nvidia/deepstream:9.0-triton-multiarch` container (the dev image; the slimmer `samples-multiarch` variant strips the nsys NVTX injector and produces empty per-plugin NVTX traces — do not use it for profiling). Requires `nsys` (Nsight Systems 2024+) and `nvidia-smi` on PATH. No GUI dependency — the skill runs fully headless and uses only `nsys profile` + `nsys stats`. .

Does Deepstream Profile Pipeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Deepstream Profile Pipeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Deepstream Profile Pipeline use?

Deepstream Profile Pipeline is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Deepstream Profile Pipeline use?

About 4.1k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.4k tokens, read only when the agent opens those files.

What are the alternatives to Deepstream Profile Pipeline?

Skills that share tags, products or a category with Deepstream Profile Pipeline: Fla Triton To Gluon (fla-org/flash-linear-attention, 5.8k stars), DGX Spark Memory and Thermal Ops (wshobson/agents, 40k stars), DGX Spark Training Gotchas (wshobson/agents, 40k stars) and Nemotron Add Step (NVIDIA-NeMo/Nemotron, 2.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Deepstream Profile Pipeline?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,539 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.