ExecuTorch Build Guide
pytorch/executorch
Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…
$ npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vipshop/cache-dit cutlass-cpp-kernel --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cutlass-cpp-kernel .claude/skills/cutlass-cpp-kernel && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cutlass-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cutlass-cpp-kernel into .claude/skills/cutlass-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutlass-cpp-kernel", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vipshop/cache-dit/tree/main/.github/skills/cutlass-cpp-kernelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vipshop/cache-dit cutlass-cpp-kernel --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/cutlass-cpp-kernel .agents/skills/cutlass-cpp-kernel && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cutlass-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cutlass-cpp-kernel into .agents/skills/cutlass-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutlass-cpp-kernel", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vipshop/cache-dit cutlass-cpp-kernel --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/cutlass-cpp-kernel .cursor/skills/cutlass-cpp-kernel && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cutlass-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cutlass-cpp-kernel into .cursor/skills/cutlass-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutlass-cpp-kernel", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vipshop/cache-dit.git --path .github/skills/cutlass-cpp-kernel--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vipshop/cache-dit cutlass-cpp-kernel --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/cutlass-cpp-kernel .gemini/skills/cutlass-cpp-kernel && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cutlass-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cutlass-cpp-kernel into .gemini/skills/cutlass-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutlass-cpp-kernel", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vipshop/cache-dit cutlass-cpp-kernelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/cutlass-cpp-kernel .github/skills/cutlass-cpp-kernel && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cutlass-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cutlass-cpp-kernel into .github/skills/cutlass-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutlass-cpp-kernel", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vipshop/cache-dit cutlass-cpp-kernel --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/cutlass-cpp-kernel .opencode/skills/cutlass-cpp-kernel && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cutlass-cpp-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cutlass-cpp-kernel into .opencode/skills/cutlass-cpp-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cutlass-cpp-kernel", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cutlass-cpp-kernelA skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…
Cutlass Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM schedules, or CuTe headers; or analyzing template configuration, tiling, memory movement, and kernel structure for Hopper or Blackwell GPUs.
Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `kernel-templates.md`, `sm100-optimization-guide.md` and `sm103-optimization-guide.md`).
It sits in Development. It works with C++, CUDA and Python. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cutlass Cpp Kernel loads about 2.2k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,044 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 1,044 words, ~2,193 tokens.
.claude/skills/cutlass-cpp-kernel/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Use the workspace CUTLASS checkout to understand, implement, and optimize CUTLASS- or CuTe-based C++ kernels while keeping CuTe DSL Python authoring in a separate dedicated workflow.
Use this skill when you need to:
layout.hpp, tensor.hpp, swizzle.hpp, and atom/*Do not use this skill for:
cute-dsl-kernelcuda-cpp-kerneloperator-migrationThis skill is the main home for:
For CuTe DSL Python kernels, JIT flows, and generated CuTe DSL API reference files, switch to cute-dsl-kernel.
In this workspace, the CUTLASS checkout is:
vipshop/cutlass/workspace/dev/vipshop/cutlassDo not rely on skill-local repository mirrors, update scripts, or any agent-local install path.
When citing CUTLASS sources, prefer workspace-relative paths such as:
vipshop/cutlass/include/cutlass/gemm/collective/vipshop/cutlass/include/cute/layout.hppvipshop/cutlass/examples/49_hopper_gemm_with_collective_builder/Use absolute shell paths only inside literal command examples when required.
This skill also bundles low-level CUDA architecture profiling references so CUTLASS tuning can be interpreted in the right hardware context.
Use these files when reading Nsight Systems or Nsight Compute results for CUTLASS or CuTe C++ kernels:
sm89-optimization-guide.mdsm90-optimization-guide.mdsm100-optimization-guide.mdsm103-optimization-guide.mdsm120-optimization-guide.mdtroubleshooting.mdThese guides are especially useful when a CUTLASS kernel behaves differently across Ada, Hopper, Blackwell datacenter, and Blackwell desktop targets.
Primary areas to inspect:
vipshop/cutlass/include/cutlass/ — CUTLASS library headersvipshop/cutlass/include/cutlass/gemm/collective/ — collective mainloop and epilogue building blocksvipshop/cutlass/include/cutlass/pipeline/ — pipeline abstractionsvipshop/cutlass/include/cute/ — CuTe C++ core headersvipshop/cutlass/include/cute/arch/ — architecture-specific copies and MMA helpersvipshop/cutlass/include/cute/atom/ — MMA and copy atomsvipshop/cutlass/examples/ — executable reference implementationsvipshop/cutlass/examples/cute/tutorial/ — CuTe C++ tutorial kernelsvipshop/cutlass/media/docs/pythonDSL/ — CuTe DSL conceptual and workflow docs useful when reviewing C++ to CuTe DSL rewritesPrefer targeted search rather than reading large header trees end-to-end.
Start from the most likely source family:
examples/ for a runnable pattern close to your target.include/cutlass/gemm/collective/ for collective-builder and mainloop configuration.include/cutlass/epilogue/ for fusion and output transforms.include/cutlass/pipeline/ for stage and async-copy structure.include/cute/ for layout algebra, tensor partitioning, swizzle, and atom semantics.For performance diagnosis, pair source study with the bundled architecture guides:
sm89 and sm120, prioritize memory throughput, L2 hit rate, occupancy, and the cost of not having cluster or TMEM-backed datacenter features; on sm120, decide explicitly whether TMA or cp.async is the better staging path.sm90, check whether the design is actually exploiting Hopper-specific staging and overlap opportunities.sm100 and sm103, verify that the kernel structure aligns with tcgen05, TMEM, TMA v2, and cluster-capable execution rather than only recompiling an older design.When using this skill for optimization or rewrites, do not read nsys or ncu output in isolation.
Recommended workflow:
smXX-optimization-guide.md first.nsys to determine whether the CUTLASS kernel has launch gaps, poor overlap, or a fusion opportunity at the operator level.ncu to determine whether the issue is occupancy, memory throughput, L2 reuse, register pressure, shared-memory pressure, or architecture-specific execution features.This matters most when comparing the same CUTLASS-style operator across multiple architectures.
Before editing code, answer these questions:
If the kernel uses shared memory, async-copy pipelines, TMA-like staging, or multi-stage buffering, explicitly audit synchronization before blaming layout algebra or MMA semantics. When only some shapes, stage counts, or schedule variants fail, prioritize checking barrier placement, stage-slot reuse, and predicate guards for partial tiles.
If the task becomes repository integration work, move that part to operator-migration and keep this skill focused on kernel structure and source study.
For architecture-specific bottlenecks or Nsight interpretation questions, use the bundled optimization guides as supporting reference material instead of assuming the same diagnosis applies across Ada, Hopper, and Blackwell.
When the task is a rewrite, such as moving an operator from handwritten C++ to CUTLASS, or replacing one CUTLASS design with another:
If the target implementation is CuTe DSL Python rather than C++ templates, use this skill for source study and switch to cute-dsl-kernel for the authoring workflow.
Every operator or kernel task completed under this skill must include validation.
Minimum requirements:
Additional requirement for rewrites or ports:
When you finish a task using this skill, report:
© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files in .github/skills/cutlass-cpp-kernel of vipshop/cache-dit.
Open the folder on GitHubat commit a7898aa
Cutlass Cpp Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cutlass Cpp Kernel this skillvipshop/cache-dit | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| ExecuTorch Build Guidepytorch/executorch | 5.1k | — | ~2.3k | Automated safety check: Notes | Custom licence | |
| AI ReviewPaddlePaddle/Paddle | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | |
| Mpk Internalsmirage-project/mirage | 2.5k | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | |
| Project StructurespiriMirror/libuipc | 335 | — | ~823 | Automated safety check: Pass | Apache-2.0 | |
| Paddle BuildPaddlePaddle/Paddle | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 |
pytorch/executorch
Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.
PaddlePaddle/Paddle
使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。
mirage-project/mirage
Reference guide for the MPK compilation-to-runtime pipeline.
spiriMirror/libuipc
Overview of the main directories and important files in the repository.
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
vipshop/cache-dit
High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
vipshop/cache-dit
A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…
vipshop/cache-dit
A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…
vipshop/cache-dit
A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…
vipshop/cache-dit
Write optimized Triton GPU kernels for deep learning operations.
Categories
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…. Cutlass Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM schedules, or CuTe headers; or analyzing template configuration, tiling, memory movement, and kernel structure for Hopper or Blackwell GPUs.
Cutlass Cpp Kernel fits situations like: optimizing CUTLASS; cuTe C++ kernels and templates; navigating CUTLASS examples; analyzing template configuration.
Run `npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a claude-code`. Or copy the skill folder (.github/skills/cutlass-cpp-kernel in vipshop/cache-dit) into .claude/skills/cutlass-cpp-kernel in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a codex`. Or copy the skill folder (.github/skills/cutlass-cpp-kernel in vipshop/cache-dit) into .agents/skills/cutlass-cpp-kernel in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cutlass-cpp-kernel, .gemini/skills/cutlass-cpp-kernel, .github/skills/cutlass-cpp-kernel and .opencode/skills/cutlass-cpp-kernel in your project.
SKILL.md names no scripts, command-line tools or credentials: Cutlass Cpp Kernel is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cutlass Cpp Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cutlass Cpp Kernel: ExecuTorch Build Guide (pytorch/executorch, 5.1k stars), AI Review (PaddlePaddle/Paddle, 24k stars), Mpk Internals (mirage-project/mirage, 2.5k stars) and Project Structure (spiriMirror/libuipc, 335 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.
Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.