Agent skill

Cuda Cpp Kernel

by vipshop in vipshop/cache-dit

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Cuda Cpp Kernel

skills CLI
$ npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vipshop/cache-dit cuda-cpp-kernel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cuda-cpp-kernel .claude/skills/cuda-cpp-kernel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cuda-cpp-kernel
GitHub stars
1.3k
Token cost
~2.3k tokens
SKILL.md length
1,097 words
Files
897 (incl. references)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

  • Works in 4 steps: Identify the exact instruction, API,… → Search the narrowest relevant… → Read only the matching file or a short… → …
  • Optimizing CUDA C++
  • SKILL.md covers Goal, When to Use, Reference Style Rule and Bundled Reference Map, plus 7 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cuda Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems or Nsight Compute; or reasoning about Tensor Core instructions, shared memory, bank conflicts, occupancy, async copy, TMA, WGMMA, and architecture-specific behavior on Ampere, Hopper, or Blackwell.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 898 other files, including reference files (for example `kernel-templates.md`, `references/best-practices-guide/1-overview.md` and `references/best-practices-guide/10-memory-optimizations.md`).

It sits in AI & LLM Engineering, covering Deep learning, GPU and accelerator computing and Agent memory. It works with CUDA and C++. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.

When your agent uses it

  • Optimizing CUDA C++
  • Investigating CUDA Runtime
  • Driver API behavior
  • Profiling kernels with Nsight Systems

Example prompts

  • “/cuda-cpp-kernel”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Identify the exact instruction, API, metric, or concept.
  2. Search the narrowest relevant subdirectory first.
  3. Read only the matching file or a short relevant range.
  4. Translate the documentation into the specific kernel or operator constraint you are implementing.

What it can do on your machine

Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cuda Cpp Kernel loads about 2.3k tokens when it runs, and up to ~2.6M if it reads all its reference files. Until then it costs about 98 tokens; SKILL.md has 1,097 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.6M

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 1,097 words, ~2,324 tokens.

Download SKILL.mdSave it as .claude/skills/cuda-cpp-kernel/SKILL.md (or your agent's skills folder). This skill also uses 896 other files; get the full folder from GitHub.
name
cuda-cpp-kernel
description
Use when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems or Nsight Compute; or reasoning about Tensor Core instructions, shared memory, bank conflicts, occupancy, async copy, TMA, WGMMA, and architecture-specific behavior on Ampere, Hopper, or Blackwell.
argument-hint
Describe the kernel or CUDA problem, target GPU architecture, dtype, shapes, current bottleneck, and whether the task is implementation, debugging…
user-invocable
true

CUDA C++ and PTX Kernel Development

Goal

Use the bundled CUDA, PTX, and profiling references in this skill to implement, debug, and optimize CUDA kernels without relying on agent-specific install paths or ad hoc web searches.

When to Use

Use this skill when you need to:

  • write or review CUDA C++ kernels or supporting host code
  • reason about PTX instructions, inline PTX, Tensor Core instructions, or memory model details
  • debug CUDA Runtime API or Driver API failures
  • profile or optimize a kernel with Nsight Systems, Nsight Compute, compute-sanitizer, or cuda-gdb
  • investigate shared memory bank conflicts, memory coalescing, occupancy, register pressure, async copy, TMA, or cluster behavior
  • compare a custom operator or kernel against a PyTorch baseline for correctness or performance

Do not use this skill for:

  • CUTLASS or CuTe template design as the primary task; use cutlass-cpp-kernel
  • CuTe DSL Python kernel authoring as the primary task; use cute-dsl-kernel
  • cache-dit operator registration, packaging, or public API migration work by itself; pair with operator-migration

Reference Style Rule

Use skill-local relative paths for bundled references, for example:

  • references/ptx-docs/
  • references/cuda-runtime-docs/
  • references/ncu-docs/ProfilingGuide.md

Do not write agent-specific install paths into follow-up notes or generated docs.

Bundled Reference Map

The following directories are bundled under references/ inside this skill:

  • references/ptx-docs/ — full PTX ISA reference
  • references/ptx-simple/ — condensed PTX quick reference
  • references/cuda-runtime-docs/ — CUDA Runtime API reference
  • references/cuda-driver-docs/ — CUDA Driver API reference
  • references/cuda-guide/ — CUDA Programming Guide
  • references/best-practices-guide/ — CUDA C++ Best Practices Guide
  • references/ncu-docs/ — Nsight Compute docs
  • references/nsys-docs/ — Nsight Systems docs
  • references/debugging-tools.md — debugging workflow notes
  • references/performance-traps.md — common optimization traps

The following architecture and kernel reference files are also bundled at the top level of this skill:

  • sm89-optimization-guide.md — Ada-specific optimization and profiling guidance
  • sm90-optimization-guide.md — Hopper-specific optimization and profiling guidance
  • sm100-optimization-guide.md — Blackwell datacenter optimization and profiling guidance
  • sm103-optimization-guide.md — Blackwell Ultra optimization and profiling guidance
  • sm120-optimization-guide.md — Blackwell desktop or workstation optimization and profiling guidance
  • kernel-templates.md — low-level CUDA kernel templates and implementation patterns
  • troubleshooting.md — debugging, compute-sanitizer, and profiling troubleshooting notes

How to Search the Bundle

Prefer narrow text search over loading large reference files into context.

Suggested workflow:

  1. Identify the exact instruction, API, metric, or concept.
  2. Search the narrowest relevant subdirectory first.
  3. Read only the matching file or a short relevant range.
  4. Translate the documentation into the specific kernel or operator constraint you are implementing.

Typical search targets:

  • PTX instruction syntax: references/ptx-docs/9-instruction-set/
  • Quick PTX lookup: references/ptx-simple/
  • CUDA Runtime APIs: references/cuda-runtime-docs/modules/
  • CUDA Driver APIs: references/cuda-driver-docs/modules/
  • architecture and programming-model behavior: references/cuda-guide/
  • profiling metrics and sections: references/ncu-docs/ProfilingGuide.md
  • timeline and launch behavior: references/nsys-docs/UserGuide.md

For architecture-specific tuning and profiling, read the matching optimization guide first:

  • sm89-optimization-guide.md
  • sm90-optimization-guide.md
  • sm100-optimization-guide.md
  • sm103-optimization-guide.md
  • sm120-optimization-guide.md

Architecture-Specific Profiling Workflow

When using Nsight Systems or Nsight Compute, interpret the results in the context of the target architecture instead of treating all GPUs the same.

Use the bundled architecture guides as the first reference for cross-architecture analysis:

  1. For sm89, focus on memory throughput, L2 hit rate, kernel fusion opportunity, and the lack of TMA or cluster features.
  2. For sm90, focus on TMA overlap, warpgroup behavior, shared-memory staging, and whether the timeline shows good load or compute overlap.
  3. For sm100 and sm103, focus on tcgen05 usage, TMEM behavior, TMA v2 overlap, cluster behavior, and whether the kernel actually benefits from Blackwell datacenter features.
  4. For sm120, treat it closer to Ada than to datacenter Blackwell for profiling purposes: watch memory throughput, L2 hit rate, shared-memory limits, and the lack of TMEM or cluster features while deciding explicitly whether TMA or cp.async is the better staging path.

Recommended order:

  1. Read the matching smXX-optimization-guide.md file.
  2. Use nsys to identify launch gaps, overlap, copy or compute concurrency, and end-to-end bottlenecks.
  3. Use ncu to inspect architecture-specific limits such as occupancy, memory throughput, L2 hit rate, tensor core utilization, stall reasons, register pressure, or shared-memory pressure.
  4. Compare the observations against the architecture guide before changing tile shapes, pipelines, or memory movement.
Show full SKILL.md (459 more words)Show less

Implementation Checklist

Before changing code, answer these questions:

  1. What is the exact shape, dtype, and layout contract?
  2. What architectural assumptions exist, such as SM target, shared memory budget, alignment, or Tensor Core mode?
  3. Is the bottleneck compute, memory, launch overhead, or synchronization?
  4. Which CUDA Runtime, Driver, or PTX rules must be preserved?
  5. What verification will prove the kernel is correct and faster enough to justify the change?

If the task is a migration into cache-dit, keep the kernel work separate from registration and packaging decisions and use operator-migration for the repository-integration layer.

Debugging Workflow

  1. Reproduce the issue with the smallest input that still fails.
  2. Confirm the failure mode: wrong value, launch error, illegal memory access, race, hang, or performance regression.
  3. When the kernel uses shared memory, cp.async, or any other asynchronous staging path, treat data-synchronization bugs as a first-line hypothesis. If only some shapes or stage-count cases fail, suspect missing barriers, premature shared-memory slot reuse, or incomplete predicate protection before assuming the math is wrong.
  4. Use compute-sanitizer or cuda-gdb for correctness problems.
  5. Use Nsight Systems first for end-to-end bottlenecks, then Nsight Compute for per-kernel root cause.
  6. After each change, rerun the focused correctness test before doing broader benchmarks.

Performance Workflow

Never optimize by intuition alone.

  1. Establish a baseline wall-clock measurement.
  2. Use Nsight Systems to see where time is spent.
  3. Use Nsight Compute to explain why the kernel is slow.
  4. Change one dimension at a time: tile shape, memory movement, synchronization, vectorization, or epilogue structure.
  5. Re-measure and compare against the previous version, not just the current absolute time.

If the result differs across GPU generations, consult the matching smXX-optimization-guide.md file before generalizing the bottleneck diagnosis.

Validation Requirements

Every operator or kernel task completed under this skill must include validation.

Minimum requirements:

  1. Add or update unit tests for correctness.
  2. Compare numerical accuracy against a PyTorch baseline or another trusted high-level reference when applicable.
  3. Compare performance against that baseline when the task claims a performance benefit or replaces a baseline path.
  4. Record the exact benchmark setup: shapes, dtypes, device, warmup, iterations, and timing method.

Additional requirement for rewrites or migrations:

  1. If you rewrite an existing operator, such as moving from one CUDA implementation to another, or from a handwritten kernel to a new implementation style, compare the new implementation against the pre-rewrite operator on both accuracy and performance.
  2. Treat a PyTorch baseline and the pre-rewrite operator as separate comparisons when both exist.

Output Expectations

When you finish a task using this skill, report:

  • the implementation scope
  • the main architectural or API constraints
  • the tests added or run
  • the PyTorch-baseline accuracy result
  • the performance result
  • if applicable, the old-versus-new operator comparison

© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 896 other files (references) in .github/skills/cuda-cpp-kernel of vipshop/cache-dit.

  • SKILL.md
  • kernel-templates.md
  • references/best-practices-guide/1-overview.md
  • references/best-practices-guide/10-memory-optimizations.md
  • references/best-practices-guide/10.1-data-transfer-between-host-and-device.md
  • references/best-practices-guide/10.2-device-memory-spaces.md
  • references/best-practices-guide/10.3-allocation.md
  • references/best-practices-guide/10.4-numa-best-practices.md
  • references/best-practices-guide/11-execution-configuration-optimizations.md
  • references/best-practices-guide/11.1-occupancy.md
  • references/best-practices-guide/11.2-hiding-register-dependencies.md
  • references/best-practices-guide/11.3-thread-and-block-heuristics.md
  • references/best-practices-guide/11.4-effects-of-shared-memory.md
  • references/best-practices-guide/11.5-concurrent-kernel-execution.md
  • references/best-practices-guide/11.6-multiple-contexts.md
  • references/best-practices-guide/12-instruction-optimization.md
  • references/best-practices-guide/12.1-arithmetic-instructions.md
  • references/best-practices-guide/12.2-memory-instructions.md
  • references/best-practices-guide/13-control-flow.md
  • … and 878 more

Open the folder on GitHubat commit a7898aa

Compare with similar skills

Cuda Cpp Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cuda Cpp Kernel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cuda Cpp Kernel this skillvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0
At Dispatch V2intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0
Qualcomm QNN Backend Developmentpytorch/executorch5.1k—~1.8kAutomated safety check: PassCustom licence
Pt2 Bug Basherpytorch/pytorch104k—~3.5kAutomated safety check: PassCustom licence
Debug Cuda Crashsgl-project/sglang37k2 repos~4.9kAutomated safety check: PassApache-2.0
ONNX Runtime CUDA Attention Patternsmicrosoft/onnxruntime22k—~6.5kAutomated safety check: PassMIT

Similar skills

  • At Dispatch V2

    intel/torch-xpu-ops

    Official

    Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

    115 GitHub starsUsed in 3 repos~2.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging.

    5.1k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Pt2 Bug Basher

    pytorch/pytorch

    Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches.

    104k GitHub stars~3.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Debug Cuda Crash

    sgl-project/sglang

    Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

    37k GitHub starsUsed in 2 repos~4.9k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

    22k GitHub stars~6.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed

More from vipshop/cache-dit

  • High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…

    1.3k GitHub stars~11k tokensUpdated 8 days ago
    Auto-check passed
  • Cute Dsl Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

    1.3k GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check passed
  • Cutlass Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

    1.3k GitHub stars~2.2k tokensUpdated 8 days ago
    Auto-check passed
  • Operator Migration

    vipshop/cache-dit

    A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…

    1.3k GitHub stars~3.8k tokensUpdated 8 days ago
    Auto-check passed
  • Ptq Workflow Integration

    vipshop/cache-dit

    A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…

    1.3k GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check passed
  • Triton Kernel

    vipshop/cache-dit

    Write optimized Triton GPU kernels for deep learning operations.

    1.3k GitHub stars~1.1k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about Cuda Cpp Kernel

What does Cuda Cpp Kernel do?

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…. Cuda Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems or Nsight Compute; or reasoning about Tensor Core instructions, shared memory, bank conflicts, occupancy, async copy, TMA, WGMMA, and architecture-specific behavior on Ampere, Hopper, or Blackwell.

When should I use Cuda Cpp Kernel?

Cuda Cpp Kernel fits situations like: optimizing CUDA C++; investigating CUDA Runtime; driver API behavior; profiling kernels with Nsight Systems.

How do I install Cuda Cpp Kernel in Claude Code?

Run `npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a claude-code`. Or copy the skill folder (.github/skills/cuda-cpp-kernel in vipshop/cache-dit) into .claude/skills/cuda-cpp-kernel in your project. Claude Code loads it when a task matches its description.

How do I install Cuda Cpp Kernel in Codex?

Run `npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a codex`. Or copy the skill folder (.github/skills/cuda-cpp-kernel in vipshop/cache-dit) into .agents/skills/cuda-cpp-kernel in your project. Codex loads it when a task matches its description.

Can I use Cuda Cpp Kernel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-cpp-kernel, .gemini/skills/cuda-cpp-kernel, .github/skills/cuda-cpp-kernel and .opencode/skills/cuda-cpp-kernel in your project.

What does Cuda Cpp Kernel need to run?

SKILL.md names no scripts, command-line tools or credentials: Cuda Cpp Kernel is instructions for the agent only. Our summary lists: Python 3.

Does Cuda Cpp Kernel access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cuda Cpp Kernel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cuda Cpp Kernel use?

Cuda Cpp Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cuda Cpp Kernel use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6M tokens, read only when the agent opens those files.

What are the alternatives to Cuda Cpp Kernel?

Skills that share tags, products or a category with Cuda Cpp Kernel: At Dispatch V2 (intel/torch-xpu-ops, 115 stars), Qualcomm QNN Backend Development (pytorch/executorch, 5.1k stars), Pt2 Bug Basher (pytorch/pytorch, 104k stars) and Debug Cuda Crash (sgl-project/sglang, 37k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cuda Cpp Kernel?

vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.