Search

CUDA · Deep learning

59 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

pytorch/pytorch104k1 repo~1.7kAutomated safety check: PassUnknowntoday
2

A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

PaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0yesterday
3

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0yesterday
4

Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~1.6kAutomated safety check: PassUnknowntoday
5

Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

stas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.04 days ago
6

Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.

pytorch/executorch5.1k—~2.3kAutomated safety check: NotesUnknownyesterday
7

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
8

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
9

A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

ByteDance-Seed/VeOmni2.2k—~2.8kAutomated safety check: PassApache-2.0yesterday
10

在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

PaddlePaddle/Paddle24k—~1.4kAutomated safety check: PassApache-2.0yesterday
11

Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.

meta-pytorch/attention-gym1.3k—~6.4kAutomated safety check: PassBSD-3-Clause3 days ago
12

Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~4.9kAutomated safety check: PassUnknowntoday
13

Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

TongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT26 days ago
14

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

vipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0yesterday
15

PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

PaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0yesterday
16

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
17

A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop…

RightNow-AI/AutoMegaKernel151—~1.8kAutomated safety check: PassMIT23 days ago
18

Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…

evo-design/proto-tools135—~2.5kAutomated safety check: NotesMITtoday
19

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
20

Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed…

pytorch/test-infra113—~2kAutomated safety check: PassUnknowntoday
21

Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

sgl-project/sglang37k2 repos~4.9kAutomated safety check: PassApache-2.0today
22

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

lucifer1004/VeloQ128—~1.3kAutomated safety check: PassMIT6 days ago
23

Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired…

NVIDIA/cosmos-framework560—~2.7kAutomated safety check: PassUnknownyesterday
24

Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

microsoft/onnxruntime22k—~6.5kAutomated safety check: PassMITtoday
25

Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends.

ahrefs/ocannl118—~728Automated safety check: PassBSD-2-Clause4 days ago
26
26.At Dispatch V2Official

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0today
27

Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).

matlab/agent-skills-playground184—~3.4kAutomated safety check: PassUnknownyesterday
28

Review 3DGS implementation code for correctness, performance bugs, and best practices.

jaccen/Awesome-Gaussian-Skills161—~2.9kAutomated safety check: PassApache-2.0today
29

使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。

PaddlePaddle/Paddle24k—~303Automated safety check: PassApache-2.0yesterday
30

Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
31

Manage tensor lifetimes across CUDA streams and decide whether recordstream is avoidable.

pytorch/pytorch104k—~796Automated safety check: PassUnknowntoday
32

Rigor Debug / Rigor Audit skill for deep learning research work.

lllllllama/RigorPilot-Skills4971 repo~522Automated safety check: PassMIT17 days ago
33

Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

amd/Quark182—~1.4kAutomated safety check: PassMIT13 days ago
34

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITyesterday
35

Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests.

intel/torch-xpu-ops115—~917Automated safety check: PassApache-2.0today
36

pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md.

joselado/pyqula145—~679Automated safety check: PassGPL-3.04 days ago
37

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0yesterday
38

Update tools/scripts/generatebinarybuildmatrix.py when a PyTorch release goes live.

pytorch/test-infra113—~1.7kAutomated safety check: PassUnknowntoday
39

A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin…

mirage-project/mirage2.5k—~1.7kAutomated safety check: PassApache-2.03 days ago
40

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
41

A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels…

xuzhougeng/wisp-science1k—~3.7kAutomated safety check: PassAGPL-3.0yesterday
42

Write, review, and run high-level xTBloom Python GFN2-xTB inference with Calculator, Structure, and BatchCalculator, including single systems, repeated geometry updates, heterogeneous ragged…

jinzhezenggroup/computational-chemistry-agent-skills148—~1.3kAutomated safety check: PassLGPL-3.02 days ago
43

Diagnose failed Quark installation, PTQ execution, script generation, or export attempts.

amd/Quark182—~1.9kAutomated safety check: NotesMIT13 days ago
44

Install or verify the correct PyTorch build for a user's accelerator backend before Quark installation.

amd/Quark182—~1.6kAutomated safety check: PassMIT13 days ago
45

Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder.

matlab/matlab-agentic-toolkit1.1k—~2.8kAutomated safety check: PassUnknown2 days ago
46

Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).

matlab/matlab-agentic-toolkit1.1k—~4.6kAutomated safety check: PassUnknown2 days ago
47

Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA.

intel/torch-xpu-ops115—~681Automated safety check: PassApache-2.0yesterday
48

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

NVIDIA/skills3.6k—~3.5kAutomated safety check: PassApache-2.0yesterday