At Dispatch V2
intel/torch-xpu-ops
Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
$ npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vipshop/cache-dit cuda-cpp-kernel --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cuda-cpp-kernel .claude/skills/cuda-cpp-kernel && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cuda-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cuda-cpp-kernel into .claude/skills/cuda-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cpp-kernel", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vipshop/cache-dit/tree/main/.github/skills/cuda-cpp-kernelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vipshop/cache-dit cuda-cpp-kernel --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/cuda-cpp-kernel .agents/skills/cuda-cpp-kernel && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cuda-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cuda-cpp-kernel into .agents/skills/cuda-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cpp-kernel", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vipshop/cache-dit cuda-cpp-kernel --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/cuda-cpp-kernel .cursor/skills/cuda-cpp-kernel && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cuda-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cuda-cpp-kernel into .cursor/skills/cuda-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cpp-kernel", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vipshop/cache-dit.git --path .github/skills/cuda-cpp-kernel--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vipshop/cache-dit cuda-cpp-kernel --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/cuda-cpp-kernel .gemini/skills/cuda-cpp-kernel && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cuda-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cuda-cpp-kernel into .gemini/skills/cuda-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cpp-kernel", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vipshop/cache-dit cuda-cpp-kernelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/cuda-cpp-kernel .github/skills/cuda-cpp-kernel && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cuda-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cuda-cpp-kernel into .github/skills/cuda-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cpp-kernel", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vipshop/cache-dit cuda-cpp-kernel --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/cuda-cpp-kernel .opencode/skills/cuda-cpp-kernel && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cuda-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cuda-cpp-kernel into .opencode/skills/cuda-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cuda-cpp-kernel", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cuda-cpp-kernelA skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
Cuda Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems or Nsight Compute; or reasoning about Tensor Core instructions, shared memory, bank conflicts, occupancy, async copy, TMA, WGMMA, and architecture-specific behavior on Ampere, Hopper, or Blackwell.
Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 898 other files, including reference files (for example `kernel-templates.md`, `references/best-practices-guide/1-overview.md` and `references/best-practices-guide/10-memory-optimizations.md`).
It sits in AI & LLM Engineering, covering Deep learning, GPU and accelerator computing and Agent memory. It works with CUDA and C++. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cuda Cpp Kernel loads about 2.3k tokens when it runs, and up to ~2.6M if it reads all its reference files. Until then it costs about 98 tokens; SKILL.md has 1,097 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 1,097 words, ~2,324 tokens.
.claude/skills/cuda-cpp-kernel/SKILL.md (or your agent's skills folder). This skill also uses 896 other files; get the full folder from GitHub.Use the bundled CUDA, PTX, and profiling references in this skill to implement, debug, and optimize CUDA kernels without relying on agent-specific install paths or ad hoc web searches.
Use this skill when you need to:
Do not use this skill for:
cutlass-cpp-kernelcute-dsl-kerneloperator-migrationUse skill-local relative paths for bundled references, for example:
references/ptx-docs/references/cuda-runtime-docs/references/ncu-docs/ProfilingGuide.mdDo not write agent-specific install paths into follow-up notes or generated docs.
The following directories are bundled under references/ inside this skill:
references/ptx-docs/ — full PTX ISA referencereferences/ptx-simple/ — condensed PTX quick referencereferences/cuda-runtime-docs/ — CUDA Runtime API referencereferences/cuda-driver-docs/ — CUDA Driver API referencereferences/cuda-guide/ — CUDA Programming Guidereferences/best-practices-guide/ — CUDA C++ Best Practices Guidereferences/ncu-docs/ — Nsight Compute docsreferences/nsys-docs/ — Nsight Systems docsreferences/debugging-tools.md — debugging workflow notesreferences/performance-traps.md — common optimization trapsThe following architecture and kernel reference files are also bundled at the top level of this skill:
sm89-optimization-guide.md — Ada-specific optimization and profiling guidancesm90-optimization-guide.md — Hopper-specific optimization and profiling guidancesm100-optimization-guide.md — Blackwell datacenter optimization and profiling guidancesm103-optimization-guide.md — Blackwell Ultra optimization and profiling guidancesm120-optimization-guide.md — Blackwell desktop or workstation optimization and profiling guidancekernel-templates.md — low-level CUDA kernel templates and implementation patternstroubleshooting.md — debugging, compute-sanitizer, and profiling troubleshooting notesPrefer narrow text search over loading large reference files into context.
Suggested workflow:
Typical search targets:
references/ptx-docs/9-instruction-set/references/ptx-simple/references/cuda-runtime-docs/modules/references/cuda-driver-docs/modules/references/cuda-guide/references/ncu-docs/ProfilingGuide.mdreferences/nsys-docs/UserGuide.mdFor architecture-specific tuning and profiling, read the matching optimization guide first:
sm89-optimization-guide.mdsm90-optimization-guide.mdsm100-optimization-guide.mdsm103-optimization-guide.mdsm120-optimization-guide.mdWhen using Nsight Systems or Nsight Compute, interpret the results in the context of the target architecture instead of treating all GPUs the same.
Use the bundled architecture guides as the first reference for cross-architecture analysis:
sm89, focus on memory throughput, L2 hit rate, kernel fusion opportunity, and the lack of TMA or cluster features.sm90, focus on TMA overlap, warpgroup behavior, shared-memory staging, and whether the timeline shows good load or compute overlap.sm100 and sm103, focus on tcgen05 usage, TMEM behavior, TMA v2 overlap, cluster behavior, and whether the kernel actually benefits from Blackwell datacenter features.sm120, treat it closer to Ada than to datacenter Blackwell for profiling purposes: watch memory throughput, L2 hit rate, shared-memory limits, and the lack of TMEM or cluster features while deciding explicitly whether TMA or cp.async is the better staging path.Recommended order:
smXX-optimization-guide.md file.nsys to identify launch gaps, overlap, copy or compute concurrency, and end-to-end bottlenecks.ncu to inspect architecture-specific limits such as occupancy, memory throughput, L2 hit rate, tensor core utilization, stall reasons, register pressure, or shared-memory pressure.Before changing code, answer these questions:
If the task is a migration into cache-dit, keep the kernel work separate from registration and packaging decisions and use operator-migration for the repository-integration layer.
cp.async, or any other asynchronous staging path, treat data-synchronization bugs as a first-line hypothesis. If only some shapes or stage-count cases fail, suspect missing barriers, premature shared-memory slot reuse, or incomplete predicate protection before assuming the math is wrong.Never optimize by intuition alone.
If the result differs across GPU generations, consult the matching smXX-optimization-guide.md file before generalizing the bottleneck diagnosis.
Every operator or kernel task completed under this skill must include validation.
Minimum requirements:
Additional requirement for rewrites or migrations:
When you finish a task using this skill, report:
© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 896 other files (references) in .github/skills/cuda-cpp-kernel of vipshop/cache-dit.
Open the folder on GitHubat commit a7898aa
Cuda Cpp Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cuda Cpp Kernel this skillvipshop/cache-dit | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| At Dispatch V2intel/torch-xpu-ops | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Qualcomm QNN Backend Developmentpytorch/executorch | 5.1k | — | ~1.8k | Automated safety check: Pass | Custom licence | |
| Pt2 Bug Basherpytorch/pytorch | 104k | — | ~3.5k | Automated safety check: Pass | Custom licence | |
| Debug Cuda Crashsgl-project/sglang | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | |
| ONNX Runtime CUDA Attention Patternsmicrosoft/onnxruntime | 22k | — | ~6.5k | Automated safety check: Pass | MIT |
intel/torch-xpu-ops
Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.
pytorch/executorch
Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging.
pytorch/pytorch
Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches.
sgl-project/sglang
Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging
microsoft/onnxruntime
Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.
pytorch/pytorch
Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.
vipshop/cache-dit
High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…
vipshop/cache-dit
A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…
vipshop/cache-dit
A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…
vipshop/cache-dit
A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…
vipshop/cache-dit
Write optimized Triton GPU kernels for deep learning operations.
Categories
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…. Cuda Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems or Nsight Compute; or reasoning about Tensor Core instructions, shared memory, bank conflicts, occupancy, async copy, TMA, WGMMA, and architecture-specific behavior on Ampere, Hopper, or Blackwell.
Cuda Cpp Kernel fits situations like: optimizing CUDA C++; investigating CUDA Runtime; driver API behavior; profiling kernels with Nsight Systems.
Run `npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a claude-code`. Or copy the skill folder (.github/skills/cuda-cpp-kernel in vipshop/cache-dit) into .claude/skills/cuda-cpp-kernel in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a codex`. Or copy the skill folder (.github/skills/cuda-cpp-kernel in vipshop/cache-dit) into .agents/skills/cuda-cpp-kernel in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill cuda-cpp-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cuda-cpp-kernel, .gemini/skills/cuda-cpp-kernel, .github/skills/cuda-cpp-kernel and .opencode/skills/cuda-cpp-kernel in your project.
SKILL.md names no scripts, command-line tools or credentials: Cuda Cpp Kernel is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cuda Cpp Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.6M tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Cuda Cpp Kernel: At Dispatch V2 (intel/torch-xpu-ops, 115 stars), Qualcomm QNN Backend Development (pytorch/executorch, 5.1k stars), Pt2 Bug Basher (pytorch/pytorch, 104k stars) and Debug Cuda Crash (sgl-project/sglang, 37k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.
Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.