Optimize Op
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
DALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks.
$ npx skills add NVIDIA/skills --skill dali-dynamic-mode -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVIDIA/skills dali-dynamic-mode --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/dali-dynamic-mode .claude/skills/dali-dynamic-mode && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "dali-dynamic-mode" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/dali-dynamic-mode into .claude/skills/dali-dynamic-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dali-dynamic-mode", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVIDIA/skills/tree/main/skills/dali-dynamic-modeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVIDIA/skills --skill dali-dynamic-mode -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVIDIA/skills dali-dynamic-mode --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/dali-dynamic-mode .agents/skills/dali-dynamic-mode && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "dali-dynamic-mode" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/dali-dynamic-mode into .agents/skills/dali-dynamic-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dali-dynamic-mode", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill dali-dynamic-mode -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVIDIA/skills dali-dynamic-mode --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/dali-dynamic-mode .cursor/skills/dali-dynamic-mode && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "dali-dynamic-mode" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/dali-dynamic-mode into .cursor/skills/dali-dynamic-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dali-dynamic-mode", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVIDIA/skills.git --path skills/dali-dynamic-mode--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVIDIA/skills --skill dali-dynamic-mode -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVIDIA/skills dali-dynamic-mode --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/dali-dynamic-mode .gemini/skills/dali-dynamic-mode && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "dali-dynamic-mode" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/dali-dynamic-mode into .gemini/skills/dali-dynamic-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dali-dynamic-mode", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVIDIA/skills dali-dynamic-modeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVIDIA/skills --skill dali-dynamic-mode -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/dali-dynamic-mode .github/skills/dali-dynamic-mode && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "dali-dynamic-mode" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/dali-dynamic-mode into .github/skills/dali-dynamic-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dali-dynamic-mode", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVIDIA/skills --skill dali-dynamic-mode -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVIDIA/skills dali-dynamic-mode --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/dali-dynamic-mode .opencode/skills/dali-dynamic-mode && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "dali-dynamic-mode" agent skill from https://github.com/NVIDIA/skills/tree/main/skills/dali-dynamic-mode into .opencode/skills/dali-dynamic-mode/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "dali-dynamic-mode", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
dali-dynamic-modeDALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks.
Dali Dynamic Mode is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. DALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks.
Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts (for example `BENCHMARK.md`, `agents/openai.yaml` and `evals/evals.json`).
It sits in AI & LLM Engineering. It works with NVIDIA AI Platform, CUDA and Python. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Dali Dynamic Mode loads about 3.7k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,306 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,306 words, ~3,731 tokens.
.claude/skills/dali-dynamic-mode/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Guide AI agents in writing, reviewing, and migrating code that uses DALI's imperative dynamic-mode API, nvidia.dali.experimental.dynamic (ndd).
nvidia.dali.experimental.dynamic as ndd and write code as direct ndd calls in ordinary Python; do not use pipeline-mode APIs such as Pipeline, @pipeline_def, pipe.build(), or pipe.run().batch_size to next_epoch(...).batch_size to random ops; there is no pipeline-level batch size to inherit.device="gpu" instead of pipeline-mode "mixed", Batch.tensors[...] for sample selection, and Batch.slice[...] for per-sample slicing..torch() to convert a tensor or batch to a PyTorch tensor. Use pad=True for batches with variable shapes.nvidia.dali.experimental.dynamic..torch().Dynamic mode is DALI's imperative Python API. It lets code call DALI operators directly from normal Python control flow instead of building and running a pipeline graph.
t = ndd.tensor(data) # copy
t = ndd.as_tensor(data) # wrap, no copy if possible
t.cpu() # move to CPU
t.gpu() # move to GPU
t.torch(copy=False) # conversion to PyTorch tensor with no copy (default)
t[1:3] # slicing supported
np.asarray(t) # NumPy via __array__ (CPU only)Supports __dlpack__, __cuda_array_interface__, __array__, arithmetic operators.
b = ndd.batch([arr1, arr2]) # copy
b = ndd.as_batch(data) # wrap, no copy if possibleBatch has no __getitem__ -- batch[i] raises TypeError because indexing is ambiguous (sample selection vs. per-sample slicing). Use the explicit APIs instead:
| Intent | Method | Returns |
|---|---|---|
| Get sample i | batch.tensors[i] | Tensor |
| Get subset of samples | batch.tensors[slice_or_list] | Batch |
| Slice within each sample | batch.slice[...] | Batch (same batch_size) |
| Sample-wise slicing | batch.slice[batch_of_indices] | Batch (same batch_size) |
.tensors[] picks which samples. .slice indexes inside each sample.
xy = ndd.random.uniform(batch_size=16, range=[0, 1], shape=2)
crop_x = xy.slice[0] # Batch of 16 scalars, first element from each sample
crop_y = xy.slice[1] # Batch of 16 scalars, second element from each sample
sample_0 = xy.tensors[0] # Tensor, the entire first sample [x, y]The .slice[] API accepts batches of indices, allowing the user to mix and match batches and
scalar values, e.g.:
imgs = ndd.imread(filenames) # a batch of images, if `filenames` is a list
sliced = imgs.slice[
42 : # the range start is broadcast to all samples
ndd.batch(imgs.shape).slice[0] // 2 # per-sample range stop (half of each image)
]PyTorch conversion:
batch.torch() -- works for uniform shapes; raises for ragged batchesbatch.torch(pad=True) -- zero-pads ragged batches to max shape (use for variable-length audio, detection boxes, etc.)batch.torch(copy=None) is the default (avoids copy if possible)__dlpack__ -- use ndd.as_tensor(batch) first for DLPack consumers. ndd.as_tensor supports pad as well.Tensor.torch(copy=False) is default (no copy)Iteration: for sample in batch: yields Tensors.
Readers are stateful objects -- create once, reuse across epochs. This matters because readers track internal state like shuffle order and shard position.
reader = ndd.readers.File(file_root=image_dir, random_shuffle=True)
for epoch in range(num_epochs):
for jpegs, labels in reader.next_epoch(batch_size=64):
# jpegs, labels are Batch objects
...Key points:
labels.torch().to(device)).ndd.readers.File(...), ndd.readers.COCO(...), ndd.readers.TFRecord(...)batch_size goes to next_epoch(), not to the reader constructornext_epoch(batch_size=N) yields tuples of Batch; next_epoch() without batch_size yields tuples of Tensornext_epoch() must be fully consumed before calling next_epoch() againSharded reading for distributed training:
reader = ndd.readers.File(
file_root=image_dir,
shard_id=rank, num_shards=world_size,
stick_to_shard=True,
pad_last_batch=True,
)device="gpu" (NOT "mixed"). The "mixed" keyword is a pipeline-mode concept for implicit CPU-to-GPU transfer; in dynamic mode, passing device="gpu" triggers the same hardware-accelerated decode path..cpu() before passing to a GPU model -- .torch() gives you a GPU tensor directly. .cpu() is only needed for consumers requiring host memory (numpy, __array__).Default mode is eager -- async execution in a background thread, returns immediately.
No .evaluate() needed in most cases. Any data consumption (.torch(), __dlpack__, __array__, .shape, property access, iteration) triggers evaluation automatically.
For debugging, switch to synchronous mode so errors surface at the exact call site rather than later in the async queue:
with ndd.EvalMode.sync_cpu:
images = ndd.decoders.image(jpegs, device="gpu")
images = ndd.resize(images, size=[224, 224])
# Any error surfaces here, at the exact op that failedModes (increasing synchronicity): deferred < eager < sync_cpu < sync_full
Use EvalMode.sync_full for debugging instead of scattering .evaluate() calls -- it's cleaner and catches all issues at once. sync_cpu is often sufficient and lighter than sync_full.
ndd.set_num_threads(4) # Call once at startup, only if necessary to override the defaultsControls DALI's internal worker threads for CPU operators. Defaults to CPU affinity count or DALI_NUM_THREADS env var. Unrelated to Python-level threading.
Two approaches (use one, not both):
# Approach 1: set the thread-local default seed (simple, good enough for most cases)
ndd.random.set_seed(42)
angles = ndd.random.uniform(batch_size=64, range=(-30, 30))
# Approach 2: explicit RNG object (finer control, pass rng= to each op)
rng = ndd.random.RNG(seed=42)
values = ndd.random.uniform(batch_size=64, range=[0, 1], shape=2, rng=rng)When rng= is passed to a random op, the explicit RNG overrides the default seed. Thread-local: each thread has independent random state.
Random ops need an explicit batch_size when working with batches -- there is no pipeline-level batch size to inherit.
Dynamic mode has no pipeline-level checkpoint. Checkpoints aggregate the state of individual stateful objects: readers and RNG instances. Stateless ops (decoders, resize, rotate, normalize, ...) are not part of a checkpoint.
ckpt = ndd.checkpoint.Checkpoint()
ckpt.register(reader, "my_reader")
ckpt.register(rng, "rng")
# ... iterate for a while ...
ckpt.collect() # snapshot the registered objects
ckpt.save("ckpt_{seq:04d}.json") # writes ckpt_0000.json, ckpt_0001.json, ...Restoring is the symmetric operation -- build a fresh reader and RNG, then load + register. The loaded state is applied to each object at register time:
reader = ndd.readers.File(file_root=..., enable_checkpointing=True, name="my_reader")
rng = ndd.random.RNG()
ckpt = ndd.checkpoint.Checkpoint()
ckpt.load("ckpt_{seq:04d}.json") # picks the highest sequence number
ckpt.register(reader, "my_reader") # state applied here
ckpt.register(rng, "rng") # ditto
for batch in reader.next_epoch(batch_size=N):
... # produces the next batch after the checkpointed iterationKey rules:
enable_checkpointing=True. Registering an already-iterated reader without it raises RuntimeError; if the reader has not been iterated yet, register enables it retroactively.next_epoch call. The prefetch thread starts on first iteration and the snapshot queue is locked after that. set_state (or a register from a loaded checkpoint) on an already-iterated reader raises RuntimeError.enable_checkpointing=True is incompatible with compile=True. Calling reader.next_epoch(..., compile=True) on a checkpointing-enabled reader raises NotImplementedError.register(op) uses sequential keys (__op_0, __op_1, ...) so the registration order must match between save and restore. Type tags catch cross-type swaps but not reorders of compatible types. Prefer register(op, name).ndd.checkpoint.current() returns the Checkpoint bound to the current thread-local EvalContext. It's shared across calls -- call ckpt.clear() if reusing the default context for unrelated runs.save/load take a Python format string with a single {seq} placeholder (e.g. "ckpt_{seq:04d}.json"). save picks the next free sequence; load picks the highest matching one on disk.deserialize rejects payloads from a different checkpoint format version -- no automatic upgrade.Checkpoint per thread.Manual get_state / set_state is also available directly on each Reader and RNG -- the Checkpoint aggregator is built on top of it. Use the manual API only when integrating with an external checkpoint system.
import nvidia.dali.experimental.dynamic as ndd
reader = ndd.readers.File(file_root="/data/imagenet/train", random_shuffle=True)
for epoch in range(num_epochs):
for jpegs, labels in reader.next_epoch(batch_size=64):
images = ndd.decoders.image(jpegs, device="gpu")
images = ndd.resize(images, size=[224, 224])
images = ndd.crop_mirror_normalize(
images,
mean=[0.485 * 255, 0.456 * 255, 0.406 * 255],
std=[0.229 * 255, 0.224 * 255, 0.225 * 255],
)
train_step(images.torch(), labels.torch())| Wrong | Right | Why |
|---|---|---|
device="mixed" | device="gpu" | "mixed" is pipeline mode only |
batch[i] | batch.tensors[i] | Batch has no __getitem__ |
batch.tensors[0] for per-sample slicing | batch.slice[0] | .tensors pick samples; .slice slices within each sample |
.evaluate() after every op | Let consumption trigger eval | .torch(), .shape, etc. trigger it automatically |
.cpu() before GPU model | .torch() directly | Avoids wasteful D2H + H2D round-trip |
| Recreate reader each epoch | reader.next_epoch() | Readers are stateful -- create once, reuse |
ndd.readers.file(...) | ndd.readers.File(...) | Reader classes are PascalCase |
break from next_epoch() loop | Exhaust iterator or create new reader | Iterator must be fully consumed before next next_epoch() |
No batch_size to random ops | ndd.random.uniform(batch_size=N, ...) | No pipeline-level batch size to inherit |
register(reader) after first next_epoch to restore | Register the freshly built reader before the first iteration | Reader state can only be applied before the prefetch thread starts |
Restoring into a reader built without enable_checkpointing=True after iteration | Pass enable_checkpointing=True at construction (or register before first iteration) | Backend doesn't keep snapshots otherwise |
| Spelling out default argument values | Skip default argument values | Very high Python-side overhead, especially when the argument accepts Tensors/Batches. Skipping arguments uses a fast path, actually passing a sentinel value. |
| Pipeline Mode | Dynamic Mode |
|---|---|
@pipeline_def / pipe.build() / pipe.run() | Direct function calls in a loop |
fn.readers.file(...) | ndd.readers.File(...) (PascalCase, stateful) |
fn.decoders.image(jpegs, device="mixed") | ndd.decoders.image(jpegs, device="gpu") |
fn.op_name(...) | ndd.op_name(...) |
Pipeline-level batch_size=64 | reader.next_epoch(batch_size=64) + random ops batch_size=64 |
Pipeline-level seed=42 | ndd.random.set_seed(42) or ndd.random.RNG(seed=42) |
Pipeline-level num_threads=4 | ndd.set_num_threads(4) at startup |
output.at(i) | batch.tensors[i] |
output.as_cpu() | batch.cpu() |
pipe.run() returns tuple of TensorList | reader.next_epoch(batch_size=N) yields tuples of Batch |
Pipeline(..., enable_checkpointing=True) + pipe.checkpoint() / pipeline(checkpoint=...) | ndd.checkpoint.Checkpoint + per-object register / collect / save / load; readers opt in with enable_checkpointing=True |
Dynamic mode is more flexible than pipeline mode, but can have slightly worse performance. For maximum throughput, prefer pipeline mode.
EvalMode.sync_cpu or EvalMode.sync_full.next_epoch() iterator is fully consumed.© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts) in skills/dali-dynamic-mode of NVIDIA/skills.
Open the folder on GitHubat commit 0e0d506
Dali Dynamic Mode next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Dali Dynamic Mode this skillNVIDIA/skills | 3.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Optimize OpCVCUDA/CV-CUDA | 2.7k | — | ~834 | Automated safety check: Pass | Custom licence | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Triton SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Make Op ScaffoldCVCUDA/CV-CUDA | 2.7k | — | ~306 | Automated safety check: Pass | Custom licence | |
| Vllm Deploy Simplevllm-project/vllm-skills | 103 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 |
CVCUDA/CV-CUDA
Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
slowlyC/agent-gpu-skills
Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.
CVCUDA/CV-CUDA
Scaffold a new CV-CUDA operator — a complete, wired, building skeleton — and delegate the implementation to a human or another AI.
vllm-project/vllm-skills
Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
NVIDIA/skills
A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.
NVIDIA/skills
Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.
NVIDIA/skills
Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.
NVIDIA/skills
Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.
NVIDIA/skills
Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.
NVIDIA/skills
Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.
Works with
Categories
DALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks. Dali Dynamic Mode is an agent skill from NVIDIA/skills, published by the product's own GitHub organization.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks.
Dali Dynamic Mode fits situations like: working on ndd code; migrating pipelines; skip pipeline-only tasks.
Run `npx skills add NVIDIA/skills --skill dali-dynamic-mode -a claude-code`. Or copy the skill folder (skills/dali-dynamic-mode in NVIDIA/skills) into .claude/skills/dali-dynamic-mode in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVIDIA/skills --skill dali-dynamic-mode -a codex`. Or copy the skill folder (skills/dali-dynamic-mode in NVIDIA/skills) into .agents/skills/dali-dynamic-mode in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill dali-dynamic-mode -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/dali-dynamic-mode, .gemini/skills/dali-dynamic-mode, .github/skills/dali-dynamic-mode and .opencode/skills/dali-dynamic-mode in your project.
Going by SKILL.md and its folder, Dali Dynamic Mode needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Dali Dynamic Mode is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Dali Dynamic Mode: Optimize Op (CVCUDA/CV-CUDA, 2.7k stars), Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars), Triton Skill (slowlyC/agent-gpu-skills, 169 stars) and Make Op Scaffold (CVCUDA/CV-CUDA, 2.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.