Search
Development · CUDA · For developers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch. | pytorch/ | 104k | 1 repo | ~1.7k | Automated safety check: Pass | Unknown | yesterday |
| 2 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks. | NVIDIA/ | 3.4k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 4 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | yesterday |
| 5 | 把 LeetCUDA kernels/interview 教学源码(base/sgemv/sgemm/hgemm/flashattn/ffpaattn/fp8gemm.cuh/fp4gemm.cuh)+ ffpa-attn CuTe sm120 源码(csrc/cuffpa/cute fp8/fp4,commit 861d75e)写成中文 CUDA 技术书(6 Part 38 章 + CuTe… | xlite-dev/ | 12k | — | ~2k | Automated safety check: Pass | GPL-3.0 | 10 days ago |
| 6 | Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | yesterday |
| 7 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 4 days ago |
| 8 | Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks. | pytorch/ | 5.1k | — | ~2.3k | Automated safety check: Notes | Unknown | yesterday |
| 9 | Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code). | CVCUDA/ | 2.7k | — | ~1.5k | Automated safety check: Pass | Unknown | 24 days ago |
| 10 | Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component. | microsoft/ | 22k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 11 | 11.Veomni Debug A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 12 | 12.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 13 | 13.Cpp Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics. | crazyguitar/ | 290 | — | ~1.8k | Automated safety check: Pass | MIT | yesterday |
| 14 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 5 days ago |
| 15 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 16 | Generate a local phyai environment report for debugging system, Python, CUDA/GPU, dependency, workspace package, git, and PHYAI configuration issues. | mingti-org/ | 130 | — | ~590 | Automated safety check: Pass | MIT | 4 days ago |
| 17 | 17.Readable Cpp Readable C/C++/Rust/CUDA code rules inspired by The Art of Readable Code. | crazyguitar/ | 290 | — | ~6.4k | Automated safety check: Pass | MIT | yesterday |
| 18 | Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 19 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM… | vipshop/ | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 20 | A skill your agent uses for performance profiling and optimization. | ByteDance-Seed/ | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 21 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 938 | — | ~2k | Automated safety check: Pass | No licence | 5 days ago |
| 22 | Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging | sgl-project/ | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 23 | 23.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 25 days ago |
| 24 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | yesterday |
| 25 | Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends. | ahrefs/ | 118 | — | ~728 | Automated safety check: Pass | BSD-2-Clause | 4 days ago |
| 26 | 26.Code Review Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. | JakeATX/ | 157 | — | ~5.6k | Automated safety check: Pass | MIT | yesterday |
| 27 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 28 | A skill your agent uses when updating dependencies managed by uv: bumping a package version, upgrading the uv tool itself, updating torch/CUDA stack, switching transformers version, or regenerating… | ByteDance-Seed/ | 2.2k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 29 | 29.Cuda Skill Query current NVIDIA CUDA, PTX ISA, Runtime API, Driver API, Programming Guide, Best Practices, Nsight Compute, and Nsight Systems references. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 30 | 30.Debug Debug crashes and test failures via stack-traces, host/device logging, and DSL buffer inspection. | LuisaGroup/ | 1.1k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 31 | 31.AI Review 使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。 | PaddlePaddle/ | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | 32.Safe Debug Rigor Debug / Rigor Audit skill for deep learning research work. | lllllllama/ | 497 | 1 repo | ~522 | Automated safety check: Pass | MIT | 17 days ago |
| 33 | Measure TileLang kernel and workflow performance, choose benchmark or timeline tools, and compare implementations fairly. | sablin39/ | 145 | — | ~1.6k | Automated safety check: Pass | No licence | 25 days ago |
| 34 | Reference guide for the MPK compilation-to-runtime pipeline. | mirage-project/ | 2.5k | — | ~5.4k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 35 | Optimize MATLAB design files for GPU Coder to generate faster CUDA code. | matlab/ | 1.1k | — | ~4.7k | Automated safety check: Pass | Unknown | 2 days ago |
| 36 | Overview of the main directories and important files in the repository. | spiriMirror/ | 336 | — | ~823 | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 37 | A skill your agent uses when running or debugging interactive Skippy prompts against staged serving, including lab sync, native builds, stage startup, the HTTP prompt REPL, and process lifecycle. | Mesh-LLM/ | 3.5k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 38 | A skill your agent uses when adding, modifying, optimizing, or debugging CuTile autotuning code. | NVIDIA/ | 3.6k | — | ~4.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 39 | Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store… | facebookexperimental/ | 201 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 40 | Generate, verify, refine, and accelerate C/C++ or CUDA code from MATLAB with MATLAB Coder, Embedded Coder, GPU Coder, or MATLAB Test. | matlab/ | 1.1k | — | ~4.2k | Automated safety check: Pass | Unknown | 2 days ago |
| 41 | A skill your agent uses when the operator is authoring, building, loading, or debugging a custom doca-bench plug-in — a versioned shared library with DOCAEXPERIMENTAL-marked C entry points that… | NVIDIA/ | 3.6k | — | ~4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 42 | A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under… | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 43 | How to create a pull request for the intel/torch-xpu-ops repository. | intel/ | 115 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 44 | How HOT-Step's custom flash-attention training ops (GGMLOPFLASHATTNTRAIN/BACK) work, what the AS1.5 DiT trainer campaign proved and disproved, and the exact contract for porting flash mode to the… | scragnog/ | 174 | — | ~5.4k | Automated safety check: Pass | MIT | yesterday |
| 45 | Evidence-gated workflow for MoE performance optimization in Megatron Bridge. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 46 | AI for Science 场景下的昇腾 NPU Profiling 采集与性能分析 Skill,用于在华为 Ascend NPU 上使用 torchnpu.profiler 采集 L0、L1、L2 级性能数据,分析训练或推理中的算子耗时、调用栈、内存与瓶颈,并指导后续调优。 | ascend-ai-coding/ | 174 | — | ~3k | Automated safety check: Pass | No licence | yesterday |
| 47 | 47.Comfy CLI Install, manage, and run ComfyUI instances. An agent skill from sundial-org/awesome-openclaw-skills. | sundial-org/ | 663 | — | ~1.5k | Automated safety check: Pass | No licence | 7 mo ago |
| 48 | CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 252 | — | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |