Search
CUDA · Deep learning
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch. | pytorch/ | 104k | 1 repo | ~1.7k | Automated safety check: Pass | Unknown | today |
| 2 | A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes. | PaddlePaddle/ | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 4 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 5 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 4 days ago |
| 6 | Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks. | pytorch/ | 5.1k | — | ~2.3k | Automated safety check: Notes | Unknown | yesterday |
| 7 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 8 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | 10.Paddle Debug 在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。 | PaddlePaddle/ | 24k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 11 | Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered. | meta-pytorch/ | 1.3k | — | ~6.4k | Automated safety check: Pass | BSD-3-Clause | 3 days ago |
| 12 | 12.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 13 | 13.Ako4all Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. | TongmingLAIC/ | 369 | — | ~4k | Automated safety check: Pass | MIT | 26 days ago |
| 14 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 15 | PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML… | PaddlePaddle/ | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 16 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop… | RightNow-AI/ | 151 | — | ~1.8k | Automated safety check: Pass | MIT | 23 days ago |
| 18 | 18.Fix Env Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any… | evo-design/ | 135 | — | ~2.5k | Automated safety check: Notes | MIT | today |
| 19 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 20 | Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed… | pytorch/ | 113 | — | ~2k | Automated safety check: Pass | Unknown | today |
| 21 | Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging | sgl-project/ | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 22 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 6 days ago |
| 23 | Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired… | NVIDIA/ | 560 | — | ~2.7k | Automated safety check: Pass | Unknown | yesterday |
| 24 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 25 | Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends. | ahrefs/ | 118 | — | ~728 | Automated safety check: Pass | BSD-2-Clause | 4 days ago |
| 26 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 184 | — | ~3.4k | Automated safety check: Pass | Unknown | yesterday |
| 28 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 29 | 29.AI Review 使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。 | PaddlePaddle/ | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 30 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 31 | 31.Cuda Streams Manage tensor lifetimes across CUDA streams and decide whether recordstream is avoidable. | pytorch/ | 104k | — | ~796 | Automated safety check: Pass | Unknown | today |
| 32 | 32.Safe Debug Rigor Debug / Rigor Audit skill for deep learning research work. | lllllllama/ | 497 | 1 repo | ~522 | Automated safety check: Pass | MIT | 17 days ago |
| 33 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 34 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | yesterday |
| 35 | Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests. | intel/ | 115 | — | ~917 | Automated safety check: Pass | Apache-2.0 | today |
| 36 | 36.GPU Backend pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md. | joselado/ | 145 | — | ~679 | Automated safety check: Pass | GPL-3.0 | 4 days ago |
| 37 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 38 | Update tools/scripts/generatebinarybuildmatrix.py when a PyTorch release goes live. | pytorch/ | 113 | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 39 | A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin… | mirage-project/ | 2.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 40 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 41 | A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels… | xuzhougeng/ | 1k | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 42 | Write, review, and run high-level xTBloom Python GFN2-xTB inference with Calculator, Structure, and BatchCalculator, including single systems, repeated geometry updates, heterogeneous ragged… | jinzhezenggroup/ | 148 | — | ~1.3k | Automated safety check: Pass | LGPL-3.0 | 2 days ago |
| 43 | Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. | amd/ | 182 | — | ~1.9k | Automated safety check: Notes | MIT | 13 days ago |
| 44 | Install or verify the correct PyTorch build for a user's accelerator backend before Quark installation. | amd/ | 182 | — | ~1.6k | Automated safety check: Pass | MIT | 13 days ago |
| 45 | Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. | matlab/ | 1.1k | — | ~2.8k | Automated safety check: Pass | Unknown | 2 days ago |
| 46 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 1.1k | — | ~4.6k | Automated safety check: Pass | Unknown | 2 days ago |
| 47 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 48 | Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | yesterday |