Agent skill

Cutlass Cpp Kernel

by vipshop in vipshop/cache-dit

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

Apache-2.0Auto-check passedDevelopment

Install Cutlass Cpp Kernel

skills CLI
$ npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vipshop/cache-dit cutlass-cpp-kernel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cutlass-cpp-kernel .claude/skills/cutlass-cpp-kernel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cutlass-cpp-kernel
GitHub stars
1.3k
Token cost
~2.2k tokens
SKILL.md length
1,044 words
Files
8
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

  • Works in 5 steps: examples/ for a runnable pattern close… → include/cutlass/gemm/collective/ for… → include/cutlass/epilogue/ for fusion and… → …
  • Optimizing CUTLASS
  • SKILL.md covers Goal, When to Use, Scope Split and Workspace Source Location, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cutlass Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM schedules, or CuTe headers; or analyzing template configuration, tiling, memory movement, and kernel structure for Hopper or Blackwell GPUs.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `kernel-templates.md`, `sm100-optimization-guide.md` and `sm103-optimization-guide.md`).

It sits in Development. It works with C++, CUDA and Python. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.

When your agent uses it

  • Optimizing CUTLASS
  • CuTe C++ kernels and templates
  • Navigating CUTLASS examples
  • Analyzing template configuration

Example prompts

  • “/cutlass-cpp-kernel”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. examples/ for a runnable pattern close to your target.
  2. include/cutlass/gemm/collective/ for collective-builder and mainloop configuration.
  3. include/cutlass/epilogue/ for fusion and output transforms.
  4. include/cutlass/pipeline/ for stage and async-copy structure.
  5. include/cute/ for layout algebra, tensor partitioning, swizzle, and atom semantics.

What it can do on your machine

Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cutlass Cpp Kernel loads about 2.2k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 1,044 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 1,044 words, ~2,193 tokens.

Download SKILL.mdSave it as .claude/skills/cutlass-cpp-kernel/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
cutlass-cpp-kernel
description
Use when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM schedules, or CuTe headers; or analyzing template configuration, tiling, memory movement, and kernel structure for Hopper or Blackwell GPUs.
argument-hint
Describe the CUTLASS or CuTe C++ task, target architecture, data types, kernel family, and whether the work is source navigation, implementation…
user-invocable
true

CUTLASS and CuTe C++ Kernel Development

Goal

Use the workspace CUTLASS checkout to understand, implement, and optimize CUTLASS- or CuTe-based C++ kernels while keeping CuTe DSL Python authoring in a separate dedicated workflow.

When to Use

Use this skill when you need to:

  • study or modify CUTLASS GEMM, epilogue, pipeline, or collective-builder code
  • navigate CuTe C++ headers such as layout.hpp, tensor.hpp, swizzle.hpp, and atom/*
  • inspect CUTLASS example kernels for Hopper, Blackwell, FP8, FP4, grouped GEMM, sparse GEMM, or MoE patterns
  • reason about template choices such as tile shapes, kernel schedules, epilogue schedules, stage count, and copy or MMA atoms
  • debug CUTLASS template errors, layout mismatches, or schedule/configuration issues
  • review a rewrite plan from an existing C++ kernel to another CUTLASS- or CuTe-based implementation

Do not use this skill for:

  • CuTe DSL Python kernel authoring as the primary task; use cute-dsl-kernel
  • generic CUDA/PTX debugging with no CUTLASS or CuTe angle; use cuda-cpp-kernel
  • cache-dit integration plumbing by itself; pair with operator-migration

Scope Split

This skill is the main home for:

  • CUTLASS C++ source navigation
  • CuTe C++ header semantics
  • collective and pipeline configuration
  • example-based C++ implementation patterns

For CuTe DSL Python kernels, JIT flows, and generated CuTe DSL API reference files, switch to cute-dsl-kernel.

Workspace Source Location

In this workspace, the CUTLASS checkout is:

  • workspace-relative: vipshop/cutlass
  • shell path: /workspace/dev/vipshop/cutlass

Do not rely on skill-local repository mirrors, update scripts, or any agent-local install path.

Reference Style Rule

When citing CUTLASS sources, prefer workspace-relative paths such as:

  • vipshop/cutlass/include/cutlass/gemm/collective/
  • vipshop/cutlass/include/cute/layout.hpp
  • vipshop/cutlass/examples/49_hopper_gemm_with_collective_builder/

Use absolute shell paths only inside literal command examples when required.

Bundled CUDA Architecture References

This skill also bundles low-level CUDA architecture profiling references so CUTLASS tuning can be interpreted in the right hardware context.

Use these files when reading Nsight Systems or Nsight Compute results for CUTLASS or CuTe C++ kernels:

  • sm89-optimization-guide.md
  • sm90-optimization-guide.md
  • sm100-optimization-guide.md
  • sm103-optimization-guide.md
  • sm120-optimization-guide.md
  • troubleshooting.md

These guides are especially useful when a CUTLASS kernel behaves differently across Ada, Hopper, Blackwell datacenter, and Blackwell desktop targets.

Key Source Map

Primary areas to inspect:

  • vipshop/cutlass/include/cutlass/ — CUTLASS library headers
  • vipshop/cutlass/include/cutlass/gemm/collective/ — collective mainloop and epilogue building blocks
  • vipshop/cutlass/include/cutlass/pipeline/ — pipeline abstractions
  • vipshop/cutlass/include/cute/ — CuTe C++ core headers
  • vipshop/cutlass/include/cute/arch/ — architecture-specific copies and MMA helpers
  • vipshop/cutlass/include/cute/atom/ — MMA and copy atoms
  • vipshop/cutlass/examples/ — executable reference implementations
  • vipshop/cutlass/examples/cute/tutorial/ — CuTe C++ tutorial kernels
  • vipshop/cutlass/media/docs/pythonDSL/ — CuTe DSL conceptual and workflow docs useful when reviewing C++ to CuTe DSL rewrites

Search Workflow

Prefer targeted search rather than reading large header trees end-to-end.

Start from the most likely source family:

  1. examples/ for a runnable pattern close to your target.
  2. include/cutlass/gemm/collective/ for collective-builder and mainloop configuration.
  3. include/cutlass/epilogue/ for fusion and output transforms.
  4. include/cutlass/pipeline/ for stage and async-copy structure.
  5. include/cute/ for layout algebra, tensor partitioning, swizzle, and atom semantics.

For performance diagnosis, pair source study with the bundled architecture guides:

  1. On sm89 and sm120, prioritize memory throughput, L2 hit rate, occupancy, and the cost of not having cluster or TMEM-backed datacenter features; on sm120, decide explicitly whether TMA or cp.async is the better staging path.
  2. On sm90, check whether the design is actually exploiting Hopper-specific staging and overlap opportunities.
  3. On sm100 and sm103, verify that the kernel structure aligns with tcgen05, TMEM, TMA v2, and cluster-capable execution rather than only recompiling an older design.

Architecture-Specific Nsight Guidance

When using this skill for optimization or rewrites, do not read nsys or ncu output in isolation.

Recommended workflow:

  1. Read the relevant smXX-optimization-guide.md first.
  2. Use nsys to determine whether the CUTLASS kernel has launch gaps, poor overlap, or a fusion opportunity at the operator level.
  3. Use ncu to determine whether the issue is occupancy, memory throughput, L2 reuse, register pressure, shared-memory pressure, or architecture-specific execution features.
  4. Only then decide whether to change tile shapes, stage count, epilogue schedule, kernel schedule, or data-movement strategy.

This matters most when comparing the same CUTLASS-style operator across multiple architectures.

Show full SKILL.md (409 more words)Show less

Implementation Workflow

Before editing code, answer these questions:

  1. Which existing CUTLASS example is the closest semantic starting point?
  2. Which layout, copy, and MMA atoms define the kernel's data movement?
  3. Which collective, schedule, and stage decisions matter for the target architecture?
  4. What public operator contract or wrapper must remain stable?
  5. What tests and benchmarks will prove the rewrite is valid?

If the kernel uses shared memory, async-copy pipelines, TMA-like staging, or multi-stage buffering, explicitly audit synchronization before blaming layout algebra or MMA semantics. When only some shapes, stage counts, or schedule variants fail, prioritize checking barrier placement, stage-slot reuse, and predicate guards for partial tiles.

If the task becomes repository integration work, move that part to operator-migration and keep this skill focused on kernel structure and source study.

For architecture-specific bottlenecks or Nsight interpretation questions, use the bundled optimization guides as supporting reference material instead of assuming the same diagnosis applies across Ada, Hopper, and Blackwell.

Rewrite Guidance

When the task is a rewrite, such as moving an operator from handwritten C++ to CUTLASS, or replacing one CUTLASS design with another:

  1. Preserve the operator contract first.
  2. Keep shape, dtype, alignment, and epilogue semantics explicit.
  3. Verify the new implementation against the original operator before claiming success.
  4. Only optimize after correctness and parity are established.

If the target implementation is CuTe DSL Python rather than C++ templates, use this skill for source study and switch to cute-dsl-kernel for the authoring workflow.

Validation Requirements

Every operator or kernel task completed under this skill must include validation.

Minimum requirements:

  1. Add or update unit tests.
  2. Compare numerical accuracy against a PyTorch baseline or another trusted eager reference when applicable.
  3. Compare performance against that baseline when the work replaces or claims to improve a baseline path.
  4. Record benchmark setup details clearly.

Additional requirement for rewrites or ports:

  1. Compare the new implementation against the pre-rewrite operator on both accuracy and performance.
  2. Treat "PyTorch baseline" and "previous operator implementation" as separate comparisons when both are available.
  3. If a rewrite changes schedules, stages, or layouts, isolate whether the change improved throughput, latency, or both.

Output Expectations

When you finish a task using this skill, report:

  • which CUTLASS or CuTe source pattern was used as reference
  • what layout, stage, or epilogue choices mattered
  • what tests were added or run
  • the PyTorch-baseline accuracy and performance result
  • the old-versus-new operator comparison when a rewrite was involved

© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files in .github/skills/cutlass-cpp-kernel of vipshop/cache-dit.

  • SKILL.md
  • kernel-templates.md
  • sm100-optimization-guide.md
  • sm103-optimization-guide.md
  • sm120-optimization-guide.md
  • sm89-optimization-guide.md
  • sm90-optimization-guide.md
  • troubleshooting.md

Open the folder on GitHubat commit a7898aa

Compare with similar skills

Cutlass Cpp Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cutlass Cpp Kernel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cutlass Cpp Kernel this skillvipshop/cache-dit1.3k—~2.2kAutomated safety check: PassApache-2.0
ExecuTorch Build Guidepytorch/executorch5.1k—~2.3kAutomated safety check: NotesCustom licence
AI ReviewPaddlePaddle/Paddle24k—~303Automated safety check: PassApache-2.0
Mpk Internalsmirage-project/mirage2.5k—~5.4kAutomated safety check: PassApache-2.0
Project StructurespiriMirror/libuipc335—~823Automated safety check: PassApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0

Similar skills

  • ExecuTorch Build Guide

    pytorch/executorch

    Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.

    5.1k GitHub stars~2.3k tokensUpdated today
    DevelopmentAuto-check: notes
  • AI Review

    PaddlePaddle/Paddle

    使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。

    24k GitHub stars~303 tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Mpk Internals

    mirage-project/mirage

    Reference guide for the MPK compilation-to-runtime pipeline.

    2.5k GitHub stars~5.4k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Project Structure

    spiriMirror/libuipc

    Overview of the main directories and important files in the repository.

    335 GitHub stars~823 tokensUpdated 4 days ago
    DevelopmentAuto-check passed
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 7 days ago
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed

More from vipshop/cache-dit

  • High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…

    1.3k GitHub stars~11k tokensUpdated 8 days ago
    Auto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 8 days ago
    Auto-check passed
  • Cute Dsl Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

    1.3k GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check passed
  • Operator Migration

    vipshop/cache-dit

    A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…

    1.3k GitHub stars~3.8k tokensUpdated 8 days ago
    Auto-check passed
  • Ptq Workflow Integration

    vipshop/cache-dit

    A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…

    1.3k GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check passed
  • Triton Kernel

    vipshop/cache-dit

    Write optimized Triton GPU kernels for deep learning operations.

    1.3k GitHub stars~1.1k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about Cutlass Cpp Kernel

What does Cutlass Cpp Kernel do?

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…. Cutlass Cpp Kernel is an agent skill from vipshop/cache-dit. Use when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM schedules, or CuTe headers; or analyzing template configuration, tiling, memory movement, and kernel structure for Hopper or Blackwell GPUs.

When should I use Cutlass Cpp Kernel?

Cutlass Cpp Kernel fits situations like: optimizing CUTLASS; cuTe C++ kernels and templates; navigating CUTLASS examples; analyzing template configuration.

How do I install Cutlass Cpp Kernel in Claude Code?

Run `npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a claude-code`. Or copy the skill folder (.github/skills/cutlass-cpp-kernel in vipshop/cache-dit) into .claude/skills/cutlass-cpp-kernel in your project. Claude Code loads it when a task matches its description.

How do I install Cutlass Cpp Kernel in Codex?

Run `npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a codex`. Or copy the skill folder (.github/skills/cutlass-cpp-kernel in vipshop/cache-dit) into .agents/skills/cutlass-cpp-kernel in your project. Codex loads it when a task matches its description.

Can I use Cutlass Cpp Kernel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill cutlass-cpp-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cutlass-cpp-kernel, .gemini/skills/cutlass-cpp-kernel, .github/skills/cutlass-cpp-kernel and .opencode/skills/cutlass-cpp-kernel in your project.

What does Cutlass Cpp Kernel need to run?

SKILL.md names no scripts, command-line tools or credentials: Cutlass Cpp Kernel is instructions for the agent only. Our summary lists: Python 3.

Does Cutlass Cpp Kernel access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cutlass Cpp Kernel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cutlass Cpp Kernel use?

Cutlass Cpp Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cutlass Cpp Kernel use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cutlass Cpp Kernel?

Skills that share tags, products or a category with Cutlass Cpp Kernel: ExecuTorch Build Guide (pytorch/executorch, 5.1k stars), AI Review (PaddlePaddle/Paddle, 24k stars), Mpk Internals (mirage-project/mirage, 2.5k stars) and Project Structure (spiriMirror/libuipc, 335 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cutlass Cpp Kernel?

vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.