Agent skill

Cute Dsl Kernel

by vipshop in vipshop/cache-dit

A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Cute Dsl Kernel

skills CLI
$ npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vipshop/cache-dit cute-dsl-kernel --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vipshop/cache-dit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/cute-dsl-kernel .claude/skills/cute-dsl-kernel && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
cute-dsl-kernel
GitHub stars
1.3k
Token cost
~2.8k tokens
SKILL.md length
1,269 words
Files
22
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

  • Works in 3 steps: On sm89 and sm120, prioritize memory… → On sm90, inspect whether TMA-style… → On sm100 and sm103, inspect whether…
  • Optimizing CuTe DSL GPU kernels in Python
  • SKILL.md covers Goal, When to Use, Core Rule and Reference Style Rule, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Cute Dsl Kernel is an agent skill from vipshop/cache-dit. Use when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or rewriting an existing CUDA or C++ operator into CuTe DSL while preserving correctness and performance expectations.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files (for example `cute.md`, `cute_arch.md` and `cute_nvgpu.md`).

It sits in AI & LLM Engineering. It works with C++, Python and CUDA. The repository describes itself as: A PyTorch-native inference engine with cache, parallelism, quantization and cpu offload for DiTs. The licence is Apache-2.0.

When your agent uses it

  • Optimizing CuTe DSL GPU kernels in Python
  • Reading CuTe DSL API reference material
  • Integrating a CuTe DSL kernel into a project
  • Rewriting an existing CUDA

Example prompts

  • “/cute-dsl-kernel”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. On sm89 and sm120, prioritize memory throughput, L2 hit rate, occupancy, and fusion opportunity; these targets do not have TMA, TMEM, or…
  2. On sm90, inspect whether TMA-style overlap, warpgroup execution, and shared-memory staging are actually visible in the timeline and…
  3. On sm100 and sm103, inspect whether tcgen05 or WGMMA, TMEM, TMA v2, and cluster-capable execution are being used effectively.

What it can do on your machine

Read from SKILL.md and the folder at commit a7898aa. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Cute Dsl Kernel loads about 2.8k tokens when it runs. Until then it costs about 76 tokens; SKILL.md has 1,269 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vipshop/cache-dit at commit a7898aa, republished under its Apache-2.0 licence (© vipshop). 1,269 words, ~2,824 tokens.

Download SKILL.mdSave it as .claude/skills/cute-dsl-kernel/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.
name
cute-dsl-kernel
description
Use when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or rewriting an existing CUDA or C++ operator into CuTe DSL while preserving correctness and performance expectations.
argument-hint
Describe the target kernel, tensor shapes, dtypes, target GPU architecture, required fusion, integration target, and whether this is a new kernel or a rewrite…
user-invocable
true

Write a CuTe DSL GPU Kernel

Goal

Use the bundled CuTe DSL API snapshots in this skill and the workspace CUTLASS checkout to design, implement, debug, and integrate CuTe DSL GPU kernels in a way that is reusable across projects, including cache-dit.

When to Use

Use this skill when you need to:

  • write or modify a CuTe DSL GPU kernel in Python
  • study CuTe DSL types, runtime helpers, architecture APIs, or pipeline abstractions
  • port or rewrite an existing CUDA or C++ operator into CuTe DSL
  • use CuTe DSL examples from the workspace CUTLASS checkout as reference material
  • debug CuTe DSL compilation, runtime behavior, layout issues, or integration problems

Do not use this skill for:

  • CUTLASS or CuTe C++ template work as the primary task; use cutlass-cpp-kernel
  • generic CUDA/PTX documentation lookup with no CuTe DSL angle; use cuda-cpp-kernel
  • repository integration plumbing by itself; pair with operator-migration

Core Rule

Read the relevant API reference files before writing kernel code.

Do not guess CuTe DSL APIs or architecture helpers from memory when the bundled docs or workspace CUTLASS examples can answer the question precisely.

Reference Style Rule

Use Copilot-friendly sibling-file references for bundled docs in this skill, for example:

  • cute.md
  • cute_runtime.md
  • utils.md
  • cute_nvgpu_tcgen05.md
  • pipeline.md

Use workspace-relative paths for CUTLASS sources, for example:

  • vipshop/cutlass/python/CuTeDSL/
  • vipshop/cutlass/examples/python/CuTeDSL/
  • vipshop/cutlass/python/pycute/
  • vipshop/cutlass/include/cute/
  • vipshop/cutlass/media/docs/pythonDSL/

Do not use agent-specific skill paths or placeholder-driven argument text in the final skill content.

Read These Files First

Core API references:

  • cute.md — core CuTe DSL types and tensor or layout operations
  • cute_runtime.md — runtime helpers and data interop
  • utils.md — helper utilities and hardware info

Architecture-specific references:

  • cute_nvgpu.md — architecture API index
  • cute_nvgpu_warp.md — warp-level APIs for SM80 to SM89
  • cute_nvgpu_warpgroup.md — warpgroup APIs for SM90
  • cute_nvgpu_tcgen05.md — tcgen05 and SM100+ APIs
  • cute_nvgpu_cpasync.md — async-copy APIs
  • cute_arch.md — low-level architecture primitives
  • utils_sm90.md and utils_sm100.md — architecture helpers

Pipeline and overview:

  • pipeline.md
  • intro.md

Additional workflow and concept references from the workspace CUTLASS docs:

  • vipshop/cutlass/media/docs/pythonDSL/overview.rst — high-level positioning of CUTLASS DSLs and how CuTe DSL relates to CUTLASS C++
  • vipshop/cutlass/media/docs/pythonDSL/quick_start.rst — environment, install, and setup assumptions
  • vipshop/cutlass/media/docs/pythonDSL/functionality.rst — supported dtypes, architectures, and current feature scope
  • vipshop/cutlass/media/docs/pythonDSL/limitations.rst — current CuTe DSL limitations and unsupported cases
  • vipshop/cutlass/media/docs/pythonDSL/faqs.rst — common issues and expected behavior
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl.rst — CuTe DSL workflow overview
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_api.rst — API documentation entrypoint
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_introduction.rst — DSL programming model and mental model
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_control_flow.rst — control-flow semantics and restrictions
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_dynamic_layout.rst — static vs dynamic layout handling
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_jit_arg_generation.rst — JIT argument typing and signature generation
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_jit_caching.rst — JIT cache behavior
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_jit_compilation_options.rst — compilation flags and debugging options
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/framework_integration.rst — framework interop patterns
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/dsl_ahead_of_time_compilation.rst — AOT compilation and export flow
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/debugging.rst — debugging workflow and generated-artifact inspection
  • vipshop/cutlass/media/docs/pythonDSL/cute_dsl_general/autotuning_gemm.rst — autotuning guidance for GEMM kernels

These workspace docs are especially valuable when the bundled API snapshots are too terse for workflow, compilation, debugging, or integration questions.

CUDA architecture and profiling references bundled in this skill:

  • sm89-optimization-guide.md
  • sm90-optimization-guide.md
  • sm100-optimization-guide.md
  • sm103-optimization-guide.md
  • sm120-optimization-guide.md
  • troubleshooting.md

Use these files when interpreting nsys and ncu results for generated CuTe DSL kernels on different GPU families.

Workspace Source Map

Use the workspace CUTLASS checkout for source examples and implementation patterns.

Key locations:

  • vipshop/cutlass/python/CuTeDSL/ — CuTe DSL implementation sources
  • vipshop/cutlass/examples/python/CuTeDSL/ — CuTe DSL examples by architecture and topic
  • vipshop/cutlass/python/pycute/ — pycute helpers and layout utilities
  • vipshop/cutlass/include/cute/ — CuTe C++ headers for semantic grounding

Use the shell path /workspace/dev/vipshop/cutlass only when you need a literal command path.

Architecture-Specific Profiling Guidance

CuTe DSL kernels often need architecture-aware profiling because the generated kernel structure can look similar while the best bottleneck diagnosis differs by GPU generation.

Use the bundled optimization guides as follows:

  1. On sm89 and sm120, prioritize memory throughput, L2 hit rate, occupancy, and fusion opportunity; these targets do not have TMA, TMEM, or cluster features.
  2. On sm90, inspect whether TMA-style overlap, warpgroup execution, and shared-memory staging are actually visible in the timeline and counters.
  3. On sm100 and sm103, inspect whether tcgen05 or WGMMA, TMEM, TMA v2, and cluster-capable execution are being used effectively.

Recommended profiling order:

  1. Read the relevant smXX-optimization-guide.md file.
  2. Use nsys to identify launch gaps, missing overlap, copy or compute imbalance, and end-to-end bottlenecks.
  3. Use ncu to inspect occupancy, memory throughput, L2 hit rate, register pressure, shared-memory pressure, tensor core utilization, and stall reasons.
  4. Only then decide whether to change tiling, pipelining, copy strategy, or fusion structure.
Show full SKILL.md (585 more words)Show less

Implementation Workflow

Before writing code, answer these questions:

  1. What are the input and output shapes, dtypes, and memory-layout constraints?
  2. What is the target architecture: SM80, SM89, SM90, SM100, or newer?
  3. Is this a brand-new kernel or a rewrite of an existing operator?
  4. Which CuTe DSL APIs and source examples are the closest match?
  5. How will the compiled kernel integrate into the target project?

Then work in this order:

  1. Read the relevant bundled API docs.
  2. Read the relevant conceptual or workflow docs under vipshop/cutlass/media/docs/pythonDSL/ when the question is about control flow, JIT behavior, debugging, AOT, integration, or limitations.
  3. Read the closest source example in vipshop/cutlass/examples/python/CuTeDSL/.
  4. Decide the kernel structure: elementwise, reduction, tiled GEMM, fused kernel, or another pattern.
  5. Implement the kernel and launch path.
  6. Integrate the compiled artifact or launcher into the target repository.
  7. Run correctness tests before performance tuning.

When tuning the generated kernel, treat the bundled smXX-optimization-guide.md files as first-line references for interpreting profiling output rather than relying only on generic CUDA advice.

Integration Guidance

Keep integration guidance generic unless the target repository requires a specific loader or manifest format.

For cache-dit or other repositories:

  1. preserve the external operator contract first
  2. keep the integration layer explicit rather than burying it inside the kernel definition
  3. document any generated artifact layout, launcher assumptions, or runtime dependencies
  4. if repository-level registration or packaging changes are needed, pair this skill with operator-migration

Rewrite Guidance

When rewriting an existing operator into CuTe DSL:

  1. preserve behavior before optimizing
  2. preserve shape, dtype, and numerics explicitly
  3. keep the old implementation available long enough to benchmark and compare
  4. do not claim success based only on compilation or a smoke test

Use cutlass-cpp-kernel alongside this skill when you need C++ CUTLASS or CuTe source study to understand the original design.

Debugging Workflow

  1. Use compile-time inspection for layouts, tiling, and static shapes.
  2. Use runtime printing sparingly for GPU-side debugging.
  3. Save PTX or IR when you need to inspect code generation.
  4. Reduce the problem to the smallest shape that still reproduces the failure.
  5. If the kernel relies on shared memory, cp.async, pipeline stages, or other asynchronous movement, treat synchronization as a primary suspect. When only specific shapes or pipeline configurations produce bad outputs, first inspect barrier placement, shared-stage reuse, and predicate coverage on partial-tile loads or stores.
  6. Once correctness is stable, profile before tuning.

Validation Requirements

Every operator or kernel task completed under this skill must include validation.

Minimum requirements:

  1. Add or update unit tests.
  2. Compare numerical accuracy against a PyTorch baseline or another trusted eager reference when applicable.
  3. Compare performance against that baseline when the kernel is meant to replace or outperform it.
  4. Record benchmark setup details clearly.

Additional requirement for rewrites or migrations:

  1. If the task rewrites an existing operator, such as replacing a C++ implementation with CuTe DSL, compare the new kernel against the pre-rewrite implementation on both accuracy and performance.
  2. Treat the PyTorch baseline and the previous implementation as separate validation targets when both are available.
  3. Explain any gap that remains after the rewrite instead of masking it with only a single favorable benchmark.

Output Expectations

When you finish a task using this skill, report:

  • which bundled docs and source examples were used
  • what integration assumptions were introduced
  • what tests were added or run
  • the PyTorch-baseline accuracy and performance result
  • the old-versus-new operator comparison when the task was a rewrite

© vipshop, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 21 other files in .github/skills/cute-dsl-kernel of vipshop/cache-dit.

  • SKILL.md
  • cute.md
  • cute_arch.md
  • cute_nvgpu.md
  • cute_nvgpu_common.md
  • cute_nvgpu_cpasync.md
  • cute_nvgpu_tcgen05.md
  • cute_nvgpu_warp.md
  • cute_nvgpu_warpgroup.md
  • cute_runtime.md
  • intro.md
  • kernel-templates.md
  • pipeline.md
  • sm100-optimization-guide.md
  • sm103-optimization-guide.md
  • sm120-optimization-guide.md
  • sm89-optimization-guide.md
  • sm90-optimization-guide.md
  • troubleshooting.md
  • utils.md
  • utils_sm100.md
  • … and 1 more

Open the folder on GitHubat commit a7898aa

Compare with similar skills

Cute Dsl Kernel next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Cute Dsl Kernel compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Cute Dsl Kernel this skillvipshop/cache-dit1.3k—~2.8kAutomated safety check: PassApache-2.0
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Paddle Op DevPaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0
Cutlass SkillslowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT
Add Jit Kernelguqiong96/Lsglang1431 repos~10kAutomated safety check: PassApache-2.0

Similar skills

  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Cutlass Skill

    slowlyC/agent-gpu-skills

    Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

    169 GitHub stars~1.3k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Add Jit Kernel

    guqiong96/Lsglang

    Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

    143 GitHub starsUsed in 1 repo~10k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Sgl Kernel

    sgl-project/sglang

    Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

    37k GitHub starsUsed in 2 repos~3.4k tokens
    AI & LLM EngineeringAuto-check passed

More from vipshop/cache-dit

  • High-level guide for integrating a new DiT model into cache-dit: Cache (BlockAdapter/ForwardPattern), Context Parallelism, Tensor Parallelism, Text Encoder Parallelism (TE-P), VAE Parallelism…

    1.3k GitHub stars~11k tokensUpdated 9 days ago
    Auto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 9 days ago
    Auto-check passed
  • Cutlass Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

    1.3k GitHub stars~2.2k tokensUpdated 9 days ago
    Auto-check passed
  • Operator Migration

    vipshop/cache-dit

    A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…

    1.3k GitHub stars~3.8k tokensUpdated 9 days ago
    Auto-check passed
  • Ptq Workflow Integration

    vipshop/cache-dit

    A skill your agent uses when integrating a new PTQ workflow into cache-dit; designing quantize/load API shape, backend-specific config validation, save/load manifests, benchmark and regression…

    1.3k GitHub stars~2.8k tokensUpdated 9 days ago
    Auto-check passed
  • Triton Kernel

    vipshop/cache-dit

    Write optimized Triton GPU kernels for deep learning operations.

    1.3k GitHub stars~1.1k tokensUpdated 9 days ago
    Auto-check passed

Works with

Questions about Cute Dsl Kernel

What does Cute Dsl Kernel do?

A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…. Cute Dsl Kernel is an agent skill from vipshop/cache-dit. Use when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or rewriting an existing CUDA or C++ operator into CuTe DSL while preserving correctness and performance expectations.

When should I use Cute Dsl Kernel?

Cute Dsl Kernel fits situations like: optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; rewriting an existing CUDA.

How do I install Cute Dsl Kernel in Claude Code?

Run `npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a claude-code`. Or copy the skill folder (.github/skills/cute-dsl-kernel in vipshop/cache-dit) into .claude/skills/cute-dsl-kernel in your project. Claude Code loads it when a task matches its description.

How do I install Cute Dsl Kernel in Codex?

Run `npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a codex`. Or copy the skill folder (.github/skills/cute-dsl-kernel in vipshop/cache-dit) into .agents/skills/cute-dsl-kernel in your project. Codex loads it when a task matches its description.

Can I use Cute Dsl Kernel in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vipshop/cache-dit --skill cute-dsl-kernel -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/cute-dsl-kernel, .gemini/skills/cute-dsl-kernel, .github/skills/cute-dsl-kernel and .opencode/skills/cute-dsl-kernel in your project.

What does Cute Dsl Kernel need to run?

SKILL.md names no scripts, command-line tools or credentials: Cute Dsl Kernel is instructions for the agent only. Our summary lists: Python 3.

Does Cute Dsl Kernel access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Cute Dsl Kernel safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Cute Dsl Kernel use?

Cute Dsl Kernel is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Cute Dsl Kernel use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Cute Dsl Kernel?

Skills that share tags, products or a category with Cute Dsl Kernel: Paddle Build (PaddlePaddle/Paddle, 24k stars), Ako4all (TongmingLAIC/AKO4ALL, 369 stars), Paddle Op Dev (PaddlePaddle/Paddle, 24k stars) and Cutlass Skill (slowlyC/agent-gpu-skills, 169 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Cute Dsl Kernel?

vipshop (a GitHub organization) maintains it in vipshop/cache-dit, which has 1,289 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: vipshop/cache-dit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.