Search
CUDA · GPU and accelerator computing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 2 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | LeetCUDA 书稿 LaTeX 到 Read the Docs 站点的转换管线维护 skill(站点目录 LeetCUDA/docs/readthedocs/)。当任务涉及:改完书稿后让站点同步、改转换器 convert/、本地构建与预览 build.sh、内容核对 convert.verify、浏览器验收… | xlite-dev/ | 12k | — | ~902 | Automated safety check: Pass | GPL-3.0 | today |
| 4 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 5 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 6 | LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、… | xlite-dev/ | 12k | — | ~3.5k | Automated safety check: Pass | GPL-3.0 | today |
| 7 | 把 LeetCUDA kernels/interview 教学源码(base/sgemv/sgemm/hgemm/flashattn/ffpaattn/fp8gemm.cuh/fp4gemm.cuh)+ ffpa-attn CuTe sm120 源码(csrc/cuffpa/cute fp8/fp4,commit 861d75e)写成中文 CUDA 技术书(6 Part 38 章 + CuTe… | xlite-dev/ | 12k | — | ~2.1k | Automated safety check: Pass | GPL-3.0 | today |
| 8 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 9 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | 10.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 11 | Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. | KernelFlow-ops/ | 214 | — | ~4.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 12 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 6 days ago |
| 13 | 13.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 14 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 16 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. | mit-han-lab/ | 251 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 18 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 19 | GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 336 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 20 | 20.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 188 | — | ~547 | Automated safety check: Pass | No licence | 12 days ago |
| 21 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 144 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 22 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 23 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 938 | — | ~2k | Automated safety check: Pass | No licence | 6 days ago |
| 24 | 24.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 25 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 6 days ago |
| 26 | 26.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 25 days ago |
| 27 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 28 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 29 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 30 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 31 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 32 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 33 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 34 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 35 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 36 | GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 6 days ago |
| 37 | A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer… | mirage-project/ | 2.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 38 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 916 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 39 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.6k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 40 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 41 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 42 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 43 | A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. | NVIDIA/ | 3.6k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 44 | Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store… | facebookexperimental/ | 201 | — | ~1.1k | Automated safety check: Pass | MIT | yesterday |
| 45 | A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation. | NVIDIA/ | 3.6k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 46 | A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the… | NVIDIA/ | 3.6k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 47 | A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests… | NVIDIA/ | 3.6k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 48 | A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under… | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |