Search
CUDA · For developers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang. | sgl-project/ | 37k | 3 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch. | pytorch/ | 104k | 1 repo | ~1.7k | Automated safety check: Pass | Unknown | today |
| 3 | A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes. | PaddlePaddle/ | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 4 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | LeetCUDA 书稿 LaTeX 到 Read the Docs 站点的转换管线维护 skill(站点目录 LeetCUDA/docs/readthedocs/)。当任务涉及:改完书稿后让站点同步、改转换器 convert/、本地构建与预览 build.sh、内容核对 convert.verify、浏览器验收… | xlite-dev/ | 12k | — | ~902 | Automated safety check: Pass | GPL-3.0 | today |
| 7 | Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm. | ztxz16/ | 5.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 8 | Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks. | NVIDIA/ | 3.4k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 10 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 11 | LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、… | xlite-dev/ | 12k | — | ~3.5k | Automated safety check: Pass | GPL-3.0 | today |
| 12 | Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 13 | 13.Esmfold2 Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. | JimLiu/ | 228 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 14 | 把 LeetCUDA kernels/interview 教学源码(base/sgemv/sgemm/hgemm/flashattn/ffpaattn/fp8gemm.cuh/fp4gemm.cuh)+ ffpa-attn CuTe sm120 源码(csrc/cuffpa/cute fp8/fp4,commit 861d75e)写成中文 CUDA 技术书(6 Part 38 章 + CuTe… | xlite-dev/ | 12k | — | ~2.1k | Automated safety check: Pass | GPL-3.0 | today |
| 15 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 4 days ago |
| 16 | Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI. | greyhaven-ai/ | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 17 | Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks. | pytorch/ | 5.1k | — | ~2.3k | Automated safety check: Notes | Unknown | yesterday |
| 18 | 18.Refactor Op Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code). | CVCUDA/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Unknown | 24 days ago |
| 19 | 19.Documents Read and write real documents on the device - PDF, XLSX, DOCX, PPTX and CSV. | zhongkaifu/ | 568 | — | ~4.2k | Automated safety check: Pass | BSD-3-Clause | today |
| 20 | Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component. | microsoft/ | 22k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 21 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 465 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 19 days ago |
| 22 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | 3 days ago |
| 23 | 23.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 24 | 24.Veomni Debug A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 25 | 25.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 26 | 26.Cpp Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics. | crazyguitar/ | 290 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 27 | 27.Paddle Debug 在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。 | PaddlePaddle/ | 24k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 28 | Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. | huggingface/ | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 29 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 6 days ago |
| 30 | 30.Optimize Op Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary. | CVCUDA/ | 2.7k | — | ~834 | Automated safety check: Pass | Unknown | 24 days ago |
| 31 | Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered. | meta-pytorch/ | 1.3k | — | ~6.4k | Automated safety check: Pass | BSD-3-Clause | 3 days ago |
| 32 | 32.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 33 | 33.Ako4all Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. | TongmingLAIC/ | 369 | — | ~4k | Automated safety check: Pass | MIT | 26 days ago |
| 34 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 35 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 36 | PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML… | PaddlePaddle/ | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 37 | A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop… | RightNow-AI/ | 151 | — | ~1.8k | Automated safety check: Pass | MIT | 23 days ago |
| 38 | Diagnose and fix Cosmos3 environment, installation, and runtime errors. | NVIDIA/ | 560 | — | ~1.3k | Automated safety check: Notes | Unknown | yesterday |
| 39 | 39.Review Op Review a CV-CUDA operator end-to-end (support / test / bench / docs coverage). | CVCUDA/ | 2.7k | — | ~481 | Automated safety check: Pass | Unknown | 24 days ago |
| 40 | Generate a local phyai environment report for debugging system, Python, CUDA/GPU, dependency, workspace package, git, and PHYAI configuration issues. | mingti-org/ | 131 | — | ~590 | Automated safety check: Pass | MIT | today |
| 41 | Upgrade shared Modal runtime dependencies in kernelbot and verify them end to end. | gpu-mode/ | 114 | — | ~748 | Automated safety check: Pass | Unknown | 23 days ago |
| 42 | Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them. | radixark/ | 110 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 43 | 43.Fix Env Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any… | evo-design/ | 135 | — | ~2.5k | Automated safety check: Notes | MIT | today |
| 44 | Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 45 | Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better. | Aseiel/ | 166 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | today |
| 46 | GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 336 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 47 | Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands. | brevdev/ | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 48 | A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or… | vipshop/ | 1.3k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |