Search

CUDA · GPU and accelerator computing

56 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0yesterday
2

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
3

LeetCUDA 书稿 LaTeX 到 Read the Docs 站点的转换管线维护 skill(站点目录 LeetCUDA/docs/readthedocs/)。当任务涉及:改完书稿后让站点同步、改转换器 convert/、本地构建与预览 build.sh、内容核对 convert.verify、浏览器验收…

xlite-dev/LeetCUDA12k—~902Automated safety check: PassGPL-3.0today
4

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

microsoft/onnxruntime22k—~1.3kAutomated safety check: PassMITtoday
5

Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~1.6kAutomated safety check: PassUnknowntoday
6

LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、…

xlite-dev/LeetCUDA12k—~3.5kAutomated safety check: PassGPL-3.0today
7

把 LeetCUDA kernels/interview 教学源码(base/sgemv/sgemm/hgemm/flashattn/ffpaattn/fp8gemm.cuh/fp4gemm.cuh)+ ffpa-attn CuTe sm120 源码(csrc/cuffpa/cute fp8/fp4,commit 861d75e)写成中文 CUDA 技术书(6 Part 38 章 + CuTe…

xlite-dev/LeetCUDA12k—~2.1kAutomated safety check: PassGPL-3.0today
8

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
9

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
10
10.Cuda

CUDA kernel development, debugging, and performance optimization for Claude Code.

technillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNo licence9 mo ago
11

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

KernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT1 mo ago
12

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence6 days ago
13

Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~4.9kAutomated safety check: PassUnknowntoday
14

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

vipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0yesterday
15

Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

microsoft/onnxruntime22k—~2.9kAutomated safety check: PassMITtoday
16

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
17

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

mit-han-lab/ncu-report-skill251—~2kAutomated safety check: PassMIT1 mo ago
18

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
19

GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI.

spiriMirror/libuipc336—~3.6kAutomated safety check: PassApache-2.08 days ago
20

基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

LMIXR/CV_Deployment_skill188—~547Automated safety check: PassNo licence12 days ago
21

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

guqiong96/Lsglang1441 repo~10kAutomated safety check: PassApache-2.06 days ago
22

Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.

huggingface/skills11k2 repos~4.6kAutomated safety check: PassApache-2.02 days ago
23

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2kAutomated safety check: PassNo licence6 days ago
24

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
25

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

lucifer1004/VeloQ128—~1.3kAutomated safety check: PassMIT6 days ago
26
26.Cuda

Draft, debug, and measure CUDA kernels and host launch workflows.

sablin39/tilelang-cuda-skills145—~990Automated safety check: PassNo licence25 days ago
27

Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

microsoft/onnxruntime22k—~6.5kAutomated safety check: PassMITtoday
28
28.At Dispatch V2Official

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0today
29

Review 3DGS implementation code for correctness, performance bugs, and best practices.

jaccen/Awesome-Gaussian-Skills161—~2.9kAutomated safety check: PassApache-2.0today
30

Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
31

Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

Orchestra-Research/AI-Research-SKILLs13k4 repos~3kAutomated safety check: WarnMIT3 mo ago
32

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

facebookexperimental/triton201—~709Automated safety check: PassMITtoday
33

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITyesterday
34

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups

sgl-project/sglang37k—~13kAutomated safety check: PassApache-2.0today
35

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0yesterday
36

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

K-Dense-AI/scientific-agent-skills48k1 repo~3.4kAutomated safety check: PassMIT6 days ago
37

A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer…

mirage-project/mirage2.5k—~1.3kAutomated safety check: PassApache-2.03 days ago
38
38.Hyperpod NcclOfficial

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

awslabs/agent-plugins916—~3.4kAutomated safety check: PassApache-2.0yesterday
39

Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

NVIDIA/skills3.6k1 repo~2.3kAutomated safety check: NotesApache-2.0yesterday
40
40.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.0yesterday
41

5-stage kernel correctness verification protocol for Triton and CUDA kernels.

ZJLi2013/awesome-kernel-skills102—~702Automated safety check: PassNo licence6 mo ago
42

A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

NVIDIA/skills3.6k1 repo~2.4kAutomated safety check: NotesApache-2.0yesterday
43

A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill.

NVIDIA/skills3.6k—~2.1kAutomated safety check: PassApache-2.0yesterday
44

Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…

facebookexperimental/triton201—~1.1kAutomated safety check: PassMITyesterday
45
45.Doca GpiOfficial

A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation.

NVIDIA/skills3.6k—~3.9kAutomated safety check: PassApache-2.0yesterday
46
46.Doca GpunetioOfficial

A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the…

NVIDIA/skills3.6k—~3.7kAutomated safety check: PassApache-2.0yesterday
47

A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests…

NVIDIA/skills3.6k—~4.2kAutomated safety check: PassApache-2.0yesterday
48

A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under…

NVIDIA/skills3.6k—~3.8kAutomated safety check: PassApache-2.0yesterday