Agent skill

Cutlass Skill

by slowlyC in slowlyC/agent-gpu-skills

Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

MITAuto-check passedAI & LLM Engineering

Install Cutlass Skill

skills CLI
$ npx skills add slowlyC/agent-gpu-skills --skill cutlass-skill -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install slowlyC/agent-gpu-skills cutlass-skill --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/slowlyC/agent-gpu-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/cutlass-skill .claude/skills/cutlass-skill && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cutlass-skill
GitHub stars
169
Token cost
~1.3k tokens
SKILL.md length
436 words
Files
2
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

  • Works in 5 steps: Identify the programming surface:… → Find the closest example for the target… → Trace the example into the… → …
  • The task explicitly involves CUTLASS/CuTe/CuTeDSL
  • SKILL.md covers Locate the checkout, Choose the source surface, Query workflow and Implementation discipline, plus 1 more section
  • Calls rg, bash and cmake

What it does

Cutlass Skill is an agent skill from slowlyC/agent-gpu-skills. Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. Use when the task explicitly involves CUTLASS/CuTe/CuTeDSL, cute::Layout, cute::Tensor, TiledMMA, TiledCopy, CollectiveBuilder, CollectiveMainloop, CollectiveEpilogue, GemmUniversal, KernelSchedule, EpilogueSchedule, CUTLASS pipelines, EVT, pycute, or CUTLASS template errors. Use cuda-skill for raw CUDA/PTX or NVIDIA architecture facts, and triton-skill for Triton or Gluon implementation work.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `quick-reference.md`).

It sits in AI & LLM Engineering. It works with CUDA, NVIDIA AI Platform, Python and C++. The licence is MIT.

When your agent uses it

  • The task explicitly involves CUTLASS/CuTe/CuTeDSL
  • CollectiveBuilder
  • CollectiveMainloop
  • CollectiveEpilogue

Example prompts

  • “/cutlass-skill”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Identify the programming surface: CuTeDSL Python, CuTe C++, or CUTLASS C++.
  2. Find the closest example for the target architecture and operation.
  3. Trace the example into the implementation and headers it instantiates.
  4. Check dtype, layout, alignment, architecture target, schedule, and epilogue together.
  5. Build or run the smallest representative example before adapting it to the target repository.

What it can do on your machine

Read from SKILL.md and the folder at commit ae02d07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg
    • bash
    • cmake
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cutlass Skill loads about 1.3k tokens when it runs. Until then it costs about 132 tokens; SKILL.md has 436 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from slowlyC/agent-gpu-skills at commit ae02d07, republished under its MIT licence (© slowlyC). 436 words, ~1,347 tokens.

Download SKILL.mdSave it as .claude/skills/cutlass-skill/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
cutlass-skill
description
Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. Use when the task explicitly involves CUTLASS/CuTe/CuTeDSL, cute::Layout, cute::Tensor, TiledMMA, TiledCopy, CollectiveBuilder, CollectiveMainloop, CollectiveEpilogue, GemmUniversal, KernelSchedule, EpilogueSchedule, CUTLASS pipelines, EVT, pycute, or CUTLASS template errors. Use cuda-skill for raw CUDA/PTX or NVIDIA architecture facts, and triton-skill for Triton or Gluon implementation work.

CUTLASS and CuTe development

Use the local CUTLASS checkout as the primary implementation reference. Prefer current source and examples over remembered APIs because CuTeDSL paths and interfaces change frequently.

Locate the checkout

Resolve the directory containing this SKILL.md, then use its repos/cutlass/ child. The installer links that path to agent-gpu-skills/third_party/cutlass/ or to the checkout supplied through CUTLASS_REPO.

In commands below, replace CUTLASS_REPO with the resolved absolute path:

bash
CUTLASS_REPO=/absolute/path/to/cutlass-skill/repos/cutlass

If the checkout is missing, run this from the agent-gpu-skills repository and reinstall the Skill:

bash
bash update-repos.sh cutlass
bash install.sh --skill cutlass-skill

Choose the source surface

TaskStart here
CuTeDSL API and implementationpython/CuTeDSL/cutlass/
All CuTeDSL examples and tutorialsexamples/python/CuTeDSL/
CUTLASS C++ examplesexamples/
CUTLASS C++ kernels and buildersinclude/cutlass/
CuTe C++ layout, copy, and MMAinclude/cute/
Python layout and swizzle utilitiespython/pycute/
Python manifest and generator utilitiespython/cutlass_library/

Read quick-reference.md when mapping an operation or architecture to a concrete file in the validated checkout.

Query workflow

  1. Identify the programming surface: CuTeDSL Python, CuTe C++, or CUTLASS C++.
  2. Find the closest example for the target architecture and operation.
  3. Trace the example into the implementation and headers it instantiates.
  4. Check dtype, layout, alignment, architecture target, schedule, and epilogue together.
  5. Build or run the smallest representative example before adapting it to the target repository.

Discover files before loading a large source file:

bash
find "$CUTLASS_REPO/examples/python/CuTeDSL" -type f | sort
find "$CUTLASS_REPO/examples" -mindepth 1 -maxdepth 1 -type d | sort

Search CuTeDSL definitions and usage together:

bash
rg -n 'TiledMMA|tiled_mma' \
  "$CUTLASS_REPO/python/CuTeDSL/cutlass/cute" \
  "$CUTLASS_REPO/examples/python/CuTeDSL"

rg -n 'Pipeline|pipeline' \
  "$CUTLASS_REPO/python/CuTeDSL/cutlass/pipeline" \
  "$CUTLASS_REPO/examples/python/CuTeDSL/cute"

Trace CUTLASS C++ builders and kernels:

bash
rg -n 'CollectiveBuilder' \
  "$CUTLASS_REPO/examples/49_hopper_gemm_with_collective_builder"

rg -n 'CollectiveMainloop' \
  "$CUTLASS_REPO/include/cutlass/gemm/collective"

rg -n 'GemmUniversal' \
  "$CUTLASS_REPO/include/cutlass/gemm"

Trace CuTe layout and architecture primitives:

bash
rg -n 'make_layout|composition|complement' \
  "$CUTLASS_REPO/include/cute/layout.hpp"

rg -n 'TiledCopy|make_tiled_copy' "$CUTLASS_REPO/include/cute"
rg -n 'SM90_TMA|SM100_TMA' "$CUTLASS_REPO/include/cute/arch"

For a C++ example, inspect its CMakeLists.txt and build only the selected target. Replace the architecture and target below for the workload:

bash
cmake -S "$CUTLASS_REPO" -B /tmp/cutlass-build -DCUTLASS_NVCC_ARCHS=100a
cmake --build /tmp/cutlass-build \
  --target 71_blackwell_gemm_with_collective_builder

Before running a CuTeDSL example, inspect the current python/CuTeDSL/pyproject.toml and requirements files instead of assuming a compatible CUDA or Python version.

Show full SKILL.md (160 more words)Show less

Implementation discipline

Keep three layers separate when diagnosing a CUTLASS kernel:

text
problem shape and tensor layout
  → selected collective and schedule
  → generated CUDA/PTX behavior on the target architecture

Do not infer the selected schedule from an example directory name. Follow the instantiated types and builder parameters. When an architecture or instruction detail determines correctness, add cuda-skill and verify it against the CUDA or PTX reference.

For correctness work:

  • preserve the original problem shapes, strides, layouts, dtypes, scaling format, and epilogue;
  • compare against an independent reference implementation;
  • test boundary shapes and alignment cases before performance tuning;
  • treat template compilation success as necessary but not sufficient.

For performance work:

  • establish a stable baseline and measurement method;
  • change one schedule, tile, stage count, or epilogue choice at a time;
  • record compiler resource usage and the exact architecture target;
  • profile only after confirming that the measured dispatch path uses the intended kernel.

Updating the source

From the agent-gpu-skills repository:

bash
bash update-repos.sh cutlass
python3 scripts/validate_repo.py --require-sources

The checkout follows CUTLASS main, while third_party/UPSTREAMS.toml records the commit last accepted by this Skill. Review source-map drift before updating that record.

© slowlyC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/cutlass-skill of slowlyC/agent-gpu-skills.

  • SKILL.md
  • quick-reference.md

Open the folder on GitHubat commit ae02d07

Compare with similar skills

Cutlass Skill next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cutlass Skill compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cutlass Skill this skillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Make Op VerifyCVCUDA/CV-CUDA2.7k—~433Automated safety check: PassCustom licence
Review Op SupportCVCUDA/CV-CUDA2.7k—~248Automated safety check: PassCustom licence
Cudaq GuideNVIDIA/skills3.5k—~1.3kAutomated safety check: PassApache-2.0
Cuopt DeveloperNVIDIA/skills3.5k—~3.2kAutomated safety check: NotesApache-2.0
Holoscan Install CondaNVIDIA/skills3.5k—~2kAutomated safety check: PassApache-2.0

Similar skills

  • Make Op Verify

    CVCUDA/CV-CUDA

    Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).

    2.7k GitHub stars~433 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Review Op Support

    CVCUDA/CV-CUDA

    Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix.

    2.7k GitHub stars~248 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • Cudaq Guide

    NVIDIA/skills

    Official

    A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance.

    3.5k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuopt Developer

    NVIDIA/skills

    Official

    Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI).

    3.5k GitHub stars~3.2k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Official

    Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment.

    3.5k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Official

    Install Holoscan SDK natively on Ubuntu via apt. An agent skill from NVIDIA/skills.

    3.5k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes

More from slowlyC/agent-gpu-skills

  • Tilelang Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

    169 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Triton Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Cuda Skill

    slowlyC/agent-gpu-skills

    Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references.

    169 GitHub stars~1.8k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Cutlass Skill

What does Cutlass Skill do?

Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. Cutlass Skill is an agent skill from slowlyC/agent-gpu-skills. Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

When should I use Cutlass Skill?

Cutlass Skill fits situations like: the task explicitly involves CUTLASS/CuTe/CuTeDSL; collectiveBuilder; collectiveMainloop; collectiveEpilogue.

How do I install Cutlass Skill in Claude Code?

Run `npx skills add slowlyC/agent-gpu-skills --skill cutlass-skill -a claude-code`. Or copy the skill folder (skills/cutlass-skill in slowlyC/agent-gpu-skills) into .claude/skills/cutlass-skill in your project. Claude Code loads it when a task matches its description.

How do I install Cutlass Skill in Codex?

Run `npx skills add slowlyC/agent-gpu-skills --skill cutlass-skill -a codex`. Or copy the skill folder (skills/cutlass-skill in slowlyC/agent-gpu-skills) into .agents/skills/cutlass-skill in your project. Codex loads it when a task matches its description.

Can I use Cutlass Skill in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add slowlyC/agent-gpu-skills --skill cutlass-skill -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cutlass-skill, .gemini/skills/cutlass-skill, .github/skills/cutlass-skill and .opencode/skills/cutlass-skill in your project.

What does Cutlass Skill need to run?

Going by SKILL.md and its folder, Cutlass Skill needs the command-line tools its instructions call (rg, bash, cmake and python3). Our summary lists: Python 3.

Does Cutlass Skill access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cutlass Skill safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cutlass Skill use?

Cutlass Skill is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cutlass Skill use?

About 1.3k tokens (SKILL.md is roughly 5.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cutlass Skill?

Skills that share tags, products or a category with Cutlass Skill: Make Op Verify (CVCUDA/CV-CUDA, 2.7k stars), Review Op Support (CVCUDA/CV-CUDA, 2.7k stars), Cudaq Guide (NVIDIA/skills, 3.5k stars) and Cuopt Developer (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cutlass Skill?

slowlyC (a GitHub user) maintains it in slowlyC/agent-gpu-skills, which has 169 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on August 8, 2026.

Source: slowlyC/agent-gpu-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.