Paddle Build
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…
$ npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vipshop/cache-dit cute-dsl-kernel --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cute-dsl-kernel .claude/skills/cute-dsl-kernel && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "cute-dsl-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cute-dsl-kernel into .claude/skills/cute-dsl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cute-dsl-kernel", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vipshop/cache-dit/tree/main/.github/skills/cute-dsl-kernelType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vipshop/cache-dit cute-dsl-kernel --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.github/skills/cute-dsl-kernel .agents/skills/cute-dsl-kernel && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "cute-dsl-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cute-dsl-kernel into .agents/skills/cute-dsl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cute-dsl-kernel", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vipshop/cache-dit cute-dsl-kernel --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.github/skills/cute-dsl-kernel .cursor/skills/cute-dsl-kernel && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "cute-dsl-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cute-dsl-kernel into .cursor/skills/cute-dsl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cute-dsl-kernel", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vipshop/cache-dit.git --path .github/skills/cute-dsl-kernel--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vipshop/cache-dit cute-dsl-kernel --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.github/skills/cute-dsl-kernel .gemini/skills/cute-dsl-kernel && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "cute-dsl-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cute-dsl-kernel into .gemini/skills/cute-dsl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cute-dsl-kernel", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vipshop/cache-dit cute-dsl-kernelInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .github/skills && cp -r skills-src/.github/skills/cute-dsl-kernel .github/skills/cute-dsl-kernel && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "cute-dsl-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cute-dsl-kernel into .github/skills/cute-dsl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cute-dsl-kernel", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vipshop/cache-dit cute-dsl-kernel --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.github/skills/cute-dsl-kernel .opencode/skills/cute-dsl-kernel && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "cute-dsl-kernel" agent skill from https://github.com/vipshop/cache-dit/tree/main/.github/skills/cute-dsl-kernel into .opencode/skills/cute-dsl-kernel/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "cute-dsl-kernel", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
cute-dsl-kernelA skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…
Cute Dsl Kernel is an agent skill from vipshop/cache-dit. Use when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or rewriting an existing CUDA or C++ operator into CuTe DSL while preserving correctness and performance expectations.
Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files (for example `cute.md`, `cute_arch.md` and `cute_nvgpu.md`).
It sits in AI & LLM Engineering. It works with C++, Python and CUDA. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Cute Dsl Kernel loads about 2.8k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,269 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 1,269 words, ~2,824 tokens.
.claude/skills/cute-dsl-kernel/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.Use the bundled CuTe DSL API snapshots in this skill and the workspace CUTLASS checkout to design, implement, debug, and integrate CuTe DSL GPU kernels in a way that is reusable across projects, including cache-dit.
Use this skill when you need to:
Do not use this skill for:
cutlass-cpp-kernelcuda-cpp-kerneloperator-migrationRead the relevant API reference files before writing kernel code.
Do not guess CuTe DSL APIs or architecture helpers from memory when the bundled docs or workspace CUTLASS examples can answer the question precisely.
Use Copilot-friendly sibling-file references for bundled docs in this skill, for example:
cute.mdcute_runtime.mdutils.mdcute_nvgpu_tcgen05.mdpipeline.mdUse workspace-relative paths for CUTLASS sources, for example:
vipshop/cutlass/python/CuTeDSL/vipshop/cutlass/examples/python/CuTeDSL/vipshop/cutlass/python/pycute/vipshop/cutlass/include/cute/vipshop/cutlass/media/docs/pythonDSL/Do not use agent-specific skill paths or placeholder-driven argument text in the final skill content.
Core API references:
cute.md — core CuTe DSL types and tensor or layout operationscute_runtime.md — runtime helpers and data interoputils.md — helper utilities and hardware infoArchitecture-specific references:
cute_nvgpu.md — architecture API indexcute_nvgpu_warp.md — warp-level APIs for SM80 to SM89cute_nvgpu_warpgroup.md — warpgroup APIs for SM90cute_nvgpu_tcgen05.md — tcgen05 and SM100+ APIscute_nvgpu_cpasync.md — async-copy APIscute_arch.md — low-level architecture primitivesutils_sm90.md and utils_sm100.md — architecture helpersPipeline and overview:
pipeline.mdintro.mdAdditional workflow and concept references from the workspace CUTLASS docs:
vipshop/cutlass/media/docs/pythonDSL/overview.rst — high-level positioning of CUTLASS DSLs and how CuTe DSL relates to CUTLASS C++vipshop/cutlass/media/docs/pythonDSL/quick_start.rst — environment, install, and setup assumptionsvipshop/cutlass/media/docs/pythonDSL/functionality.rst — supported dtypes, architectures, and current feature scopevipshop/cutlass/media/docs/pythonDSL/limitations.rst — current CuTe DSL limitations and unsupported casesvipshop/cutlass/media/docs/pythonDSL/faqs.rst — common issues and expected behaviorvipshop/cutlass/media/docs/pythonDSL/cute_dsl.rst — CuTe DSL workflow overviewvipshop/cutlass/media/docs/pythonDSL/cute_dsl_api.rst — API documentation entrypointvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_introduction.rst — DSL programming model and mental modelvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_control_flow.rst — control-flow semantics and restrictionsvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_dynamic_layout.rst — static vs dynamic layout handlingvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_jit_arg_generation.rst — JIT argument typing and signature generationvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_jit_caching.rst — JIT cache behaviorvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_jit_compilation_options.rst — compilation flags and debugging optionsvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/framework_integration.rst — framework interop patternsvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_ahead_of_time_compilation.rst — AOT compilation and export flowvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/debugging.rst — debugging workflow and generated-artifact inspectionvipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/autotuning_gemm.rst — autotuning guidance for GEMM kernelsThese workspace docs are especially valuable when the bundled API snapshots are too terse for workflow, compilation, debugging, or integration questions.
CUDA architecture and profiling references bundled in this skill:
sm89-optimization-guide.mdsm90-optimization-guide.mdsm100-optimization-guide.mdsm103-optimization-guide.mdsm120-optimization-guide.mdtroubleshooting.mdUse these files when interpreting nsys and ncu results for generated CuTe DSL kernels on different GPU families.
Use the workspace CUTLASS checkout for source examples and implementation patterns.
Key locations:
vipshop/cutlass/python/CuTeDSL/ — CuTe DSL implementation sourcesvipshop/cutlass/examples/python/CuTeDSL/ — CuTe DSL examples by architecture and topicvipshop/cutlass/python/pycute/ — pycute helpers and layout utilitiesvipshop/cutlass/include/cute/ — CuTe C++ headers for semantic groundingUse the shell path /workspace/dev/vipshop/cutlass only when you need a literal command path.
CuTe DSL kernels often need architecture-aware profiling because the generated kernel structure can look similar while the best bottleneck diagnosis differs by GPU generation.
Use the bundled optimization guides as follows:
sm89 and sm120, prioritize memory throughput, L2 hit rate, occupancy, and fusion opportunity; these targets do not have TMA, TMEM, or cluster features.sm90, inspect whether TMA-style overlap, warpgroup execution, and shared-memory staging are actually visible in the timeline and counters.sm100 and sm103, inspect whether tcgen05 or WGMMA, TMEM, TMA v2, and cluster-capable execution are being used effectively.Recommended profiling order:
smXX-optimization-guide.md file.nsys to identify launch gaps, missing overlap, copy or compute imbalance, and end-to-end bottlenecks.ncu to inspect occupancy, memory throughput, L2 hit rate, register pressure, shared-memory pressure, tensor core utilization, and stall reasons.Before writing code, answer these questions:
Then work in this order:
vipshop/cutlass/media/docs/pythonDSL/ when the question is about control flow, JIT behavior, debugging, AOT, integration, or limitations.vipshop/cutlass/examples/python/CuTeDSL/.When tuning the generated kernel, treat the bundled smXX-optimization-guide.md files as first-line references for interpreting profiling output rather than relying only on generic CUDA advice.
Keep integration guidance generic unless the target repository requires a specific loader or manifest format.
For cache-dit or other repositories:
operator-migrationWhen rewriting an existing operator into CuTe DSL:
Use cutlass-cpp-kernel alongside this skill when you need C++ CUTLASS or CuTe source study to understand the original design.
cp.async, pipeline stages, or other asynchronous movement, treat synchronization as a primary suspect. When only specific shapes or pipeline configurations produce bad outputs, first inspect barrier placement, shared-stage reuse, and predicate coverage on partial-tile loads or stores.Every operator or kernel task completed under this skill must include validation.
Minimum requirements:
Additional requirement for rewrites or migrations:
When you finish a task using this skill, report:
© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 21 other files in .github/skills/cute-dsl-kernel of vipshop/cache-dit.
Open the folder on GitHubat commit a7898aa
Cute Dsl Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Cute Dsl Kernel this skillvipshop/cache-dit | 1.3k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Paddle BuildPaddlePaddle/Paddle | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Ako4allTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT | |
| Paddle Op DevPaddlePaddle/Paddle | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Cutlass SkillslowlyC/agent-gpu-skills | 169 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Add Jit Kernelguqiong96/Lsglang | 143 | 1 repos | ~10k | Automated safety check: Pass | Apache-2.0 |
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
TongmingLAIC/AKO4ALL
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
slowlyC/agent-gpu-skills
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.
guqiong96/Lsglang
Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module
sgl-project/sglang
Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)
vipshop/cache-dit
High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…
vipshop/cache-dit
A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…
vipshop/cache-dit
A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…
vipshop/cache-dit
A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…
vipshop/cache-dit
Write optimized Triton GPU kernels for deep learning operations.
Categories
A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…. Cute Dsl Kernel is an agent skill from vipshop/cache-dit. Use when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or rewriting an existing CUDA or C++ operator into CuTe DSL while preserving correctness and performance expectations.
Cute Dsl Kernel fits situations like: optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; rewriting an existing CUDA.
Run `npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a claude-code`. Or copy the skill folder (.github/skills/cute-dsl-kernel in vipshop/cache-dit) into .claude/skills/cute-dsl-kernel in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a codex`. Or copy the skill folder (.github/skills/cute-dsl-kernel in vipshop/cache-dit) into .agents/skills/cute-dsl-kernel in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cute-dsl-kernel, .gemini/skills/cute-dsl-kernel, .github/skills/cute-dsl-kernel and .opencode/skills/cute-dsl-kernel in your project.
SKILL.md names no scripts, command-line tools or credentials: Cute Dsl Kernel is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Cute Dsl Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Cute Dsl Kernel: Paddle Build (PaddlePaddle/Paddle, 24k stars), Ako4all (TongmingLAIC/AKO4ALL, 369 stars), Paddle Op Dev (PaddlePaddle/Paddle, 24k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.
Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.