Search

CUDA · For developers

223 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

sgl-project/sglang37k3 repos~2.1kAutomated safety check: PassApache-2.0today
2

Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

pytorch/pytorch104k1 repo~1.7kAutomated safety check: PassUnknowntoday
3

A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

PaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0yesterday
4

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0yesterday
5

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
6

LeetCUDA 书稿 LaTeX 到 Read the Docs 站点的转换管线维护 skill(站点目录 LeetCUDA/docs/readthedocs/)。当任务涉及:改完书稿后让站点同步、改转换器 convert/、本地构建与预览 build.sh、内容核对 convert.verify、浏览器验收…

xlite-dev/LeetCUDA12k—~902Automated safety check: PassGPL-3.0today
7

Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

ztxz16/fastllm5.1k—~1.8kAutomated safety check: PassApache-2.0yesterday
8

Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks.

NVIDIA/cuda-python3.4k—~1.1kAutomated safety check: PassApache-2.0yesterday
9

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

microsoft/onnxruntime22k—~1.3kAutomated safety check: PassMITtoday
10

Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~1.6kAutomated safety check: PassUnknowntoday
11

LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、…

xlite-dev/LeetCUDA12k—~3.5kAutomated safety check: PassGPL-3.0today
12

Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.

microsoft/onnxruntime22k—~1.4kAutomated safety check: PassMITtoday
13

Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.03 mo ago
14

把 LeetCUDA kernels/interview 教学源码(base/sgemv/sgemm/hgemm/flashattn/ffpaattn/fp8gemm.cuh/fp4gemm.cuh)+ ffpa-attn CuTe sm120 源码(csrc/cuffpa/cute fp8/fp4,commit 861d75e)写成中文 CUDA 技术书(6 Part 38 章 + CuTe…

xlite-dev/LeetCUDA12k—~2.1kAutomated safety check: PassGPL-3.0today
15

Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

stas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.04 days ago
16

Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

greyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.04 days ago
17

Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.

pytorch/executorch5.1k—~2.3kAutomated safety check: NotesUnknownyesterday
18

Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code).

CVCUDA/CV-CUDA2.7k—~1.5kAutomated safety check: PassUnknown24 days ago
19

Read and write real documents on the device - PDF, XLSX, DOCX, PPTX and CSV.

zhongkaifu/TensorSharp568—~4.2kAutomated safety check: PassBSD-3-Clausetoday
20

Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.

microsoft/onnxruntime22k—~1.6kAutomated safety check: PassMITtoday
21

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4652 repos~1.5kAutomated safety check: PassApache-2.019 days ago
22

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT3 days ago
23

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
24

A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

ByteDance-Seed/VeOmni2.2k—~2.8kAutomated safety check: PassApache-2.0today
25
25.Cuda

CUDA kernel development, debugging, and performance optimization for Claude Code.

technillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNo licence9 mo ago
26
26.Cpp

Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics.

crazyguitar/cppcheatsheet290—~1.8kAutomated safety check: PassMITyesterday
27

在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

PaddlePaddle/Paddle24k—~1.4kAutomated safety check: PassApache-2.0yesterday
28

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

huggingface/skills11k3 repos~945Automated safety check: PassApache-2.02 days ago
29

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence6 days ago
30

Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

CVCUDA/CV-CUDA2.7k—~834Automated safety check: PassUnknown24 days ago
31

Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.

meta-pytorch/attention-gym1.3k—~6.4kAutomated safety check: PassBSD-3-Clause3 days ago
32

Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~4.9kAutomated safety check: PassUnknowntoday
33

Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

TongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT26 days ago
34

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

vipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0yesterday
35

Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

microsoft/onnxruntime22k—~2.9kAutomated safety check: PassMITtoday
36

PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

PaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0yesterday
37

A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop…

RightNow-AI/AutoMegaKernel151—~1.8kAutomated safety check: PassMIT23 days ago
38

Diagnose and fix Cosmos3 environment, installation, and runtime errors.

NVIDIA/cosmos-framework560—~1.3kAutomated safety check: NotesUnknownyesterday
39

Review a CV-CUDA operator end-to-end (support / test / bench / docs coverage).

CVCUDA/CV-CUDA2.7k—~481Automated safety check: PassUnknown24 days ago
40

Generate a local phyai environment report for debugging system, Python, CUDA/GPU, dependency, workspace package, git, and PHYAI configuration issues.

mingti-org/phyai131—~590Automated safety check: PassMITtoday
41

Upgrade shared Modal runtime dependencies in kernelbot and verify them end to end.

gpu-mode/kernelbot114—~748Automated safety check: PassUnknown23 days ago
42

Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them.

radixark/miles_diffusion110—~1.6kAutomated safety check: PassApache-2.0today
43

Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…

evo-design/proto-tools135—~2.5kAutomated safety check: NotesMITtoday
44

Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
45

Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

Aseiel/VideoHighlighter166—~839Automated safety check: PassAGPL-3.0today
46

GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI.

spiriMirror/libuipc336—~3.6kAutomated safety check: PassApache-2.08 days ago
47

Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

brevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.03 days ago
48

A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

vipshop/cache-dit1.3k—~2.8kAutomated safety check: PassApache-2.0yesterday