Search

GPU and accelerator computing

174 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.

huggingface/skills11k2 repos~4.6kAutomated safety check: PassApache-2.0yesterday
50

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2kAutomated safety check: PassNo licence4 days ago
51

A skill your agent uses for GPU training or inference clusters on SageMaker HyperPod (Slurm or EKS), ParallelCluster, or self-managed EC2/EKS GPU instances.

aws/tools-for-devops-agent103—~5.4kAutomated safety check: PassApache-2.0today
52

Diagnose performance issues on Amazon SageMaker HyperPod clusters — uneven NCCL bandwidth across nodes and poor filesystem throughput.

awslabs/agent-plugins916—~4.1kAutomated safety check: PassApache-2.0today
53

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
54

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

lucifer1004/VeloQ128—~1.3kAutomated safety check: PassMIT5 days ago
55
55.Cuda

Draft, debug, and measure CUDA kernels and host launch workflows.

sablin39/tilelang-cuda-skills145—~990Automated safety check: PassNo licence24 days ago
56

Implements request-to-token-pool copies for Triton kernels on Ascend NPUs, replacing loop-carried vector offsets with fixed-offset block loops that compile reliably.

Krusty84/triton-ascend-agent-dev-kit106—~564Automated safety check: PassApache-2.01 mo ago
57

Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

microsoft/onnxruntime22k—~6.5kAutomated safety check: PassMITtoday
58
58.At Dispatch V2Official

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0today
59

Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.5kAutomated safety check: PassMIT3 mo ago
60

Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.7kAutomated safety check: PassMIT3 mo ago
61

Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
62

Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use.

davila7/claude-code-templates33k10 repos~2.4kAutomated safety check: PassMITtoday
63

Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured.

davila7/claude-code-templates33k11 repos~1.7kAutomated safety check: PassMITtoday
64

Review 3DGS implementation code for correctness, performance bugs, and best practices.

jaccen/Awesome-Gaussian-Skills161—~2.9kAutomated safety check: PassApache-2.0today
65

Writes a Triton kernel pattern for Ascend NPU that gathers several indexed token rows per program into an on-chip buffer and stores one contiguous output tile, for MoE-style token reordering.

Krusty84/triton-ascend-agent-dev-kit106—~597Automated safety check: PassApache-2.01 mo ago
66

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~3.9kAutomated safety check: PassNo licence4 days ago
67

Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
68

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

NVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.0today
69

PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM.

Orchestra-Research/AI-Research-SKILLs13k1 repo~2.8kAutomated safety check: PassMIT3 mo ago
70

Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

Orchestra-Research/AI-Research-SKILLs13k4 repos~3kAutomated safety check: WarnMIT3 mo ago
71

Write optimized Triton GPU kernels for deep learning operations.

vipshop/cache-dit1.3k—~1.1kAutomated safety check: PassApache-2.0today
72

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~7.5kAutomated safety check: PassNo licence4 days ago
73

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

facebookexperimental/triton201—~709Automated safety check: PassMITtoday
74

Builds Triton-Ascend kernels that pack sorted token assignments into fixed expert-capacity tensors, restore top-k outputs and compute router-weight gradients on Ascend NPUs.

Krusty84/triton-ascend-agent-dev-kit106—~662Automated safety check: PassApache-2.01 mo ago
75

Guided step-by-step wizard for producing a new FlyDSL GPU kernel from a requirement: classify the kernel type, pick a skeleton, fill in compute, add control flow / sync / LDS, then test on GPU.

ROCm/FlyDSL291—~4.6kAutomated safety check: NotesUnknowntoday
76

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITtoday
77
77.Ir DebuggingOfficial

Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

facebookexperimental/triton201—~644Automated safety check: PassMITtoday
78

Guides building a FlashAttention-v2-style fused forward attention kernel for Triton-Ascend on Ascend NPU, with online softmax, causal staging and float32 accumulation.

Krusty84/triton-ascend-agent-dev-kit106—~753Automated safety check: PassApache-2.01 mo ago
79

Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
80

Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
81

Fine-tunes and serves Physical Intelligence's pi0, pi0-fast and pi0.5 robot policies with JAX or PyTorch, including checkpoint conversion and policy servers.

Orchestra-Research/AI-Research-SKILLs13k—~3.6kAutomated safety check: PassMIT3 mo ago
82

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups

sgl-project/sglang37k—~13kAutomated safety check: PassApache-2.0today
83

Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train.

mlc-ai/pith-train355—~1.9kAutomated safety check: PassApache-2.0today
84

3DGS compression-to-deployment pipeline: quantization (scalar/VQ/mixed-precision), pruning (coreset/adaptive/variational/merge/Bayesian), progressive streaming & LoD, Web/WebGPU/mobile deployment…

jaccen/Awesome-Gaussian-Skills161—~5.5kAutomated safety check: PassApache-2.0today
85

Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

Krusty84/triton-ascend-agent-dev-kit106—~636Automated safety check: PassApache-2.01 mo ago
86

Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs.

ZJLi2013/awesome-kernel-skills102—~1.1kAutomated safety check: PassNo licence6 mo ago
87

Query live GPU inventory, submit an authenticated Itô fixed-rate RFQ, inspect RFQ or procurement status, revoke device credentials, and run explicitly gated node qualification through the separately…

affaan-m/ECC276k1 repo~1.7kAutomated safety check: PassMITtoday
88

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0today
89

Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming).

yzlnew/infra-skills149—~2.4kAutomated safety check: PassNo licence3 mo ago
90

Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.

K-Dense-AI/scientific-agent-skills48k1 repo~4.5kAutomated safety check: NotesApache-2.04 days ago
91

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

K-Dense-AI/scientific-agent-skills48k1 repo~3.4kAutomated safety check: PassMIT4 days ago
92

Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.

sickn33/agentic-awesome-skills47k2 repos~3.2kAutomated safety check: PassMITyesterday
93

A skill your agent uses when the user asks "why is this B200/Blackwell kernel slow, what should I optimize first, should I fuse or move to Tensor Cores/TMA".

mirage-project/mirage2.5k—~2.2kAutomated safety check: PassApache-2.02 days ago
94

A skill your agent uses when a B200/Blackwell kernel shows wrong results, uncoalesced global memory access, SMEM bank conflicts, a TMA swizzle that mismatches the Tensor Core read, or confused…

mirage-project/mirage2.5k—~1.8kAutomated safety check: PassApache-2.02 days ago
95

A skill your agent uses when you need a NEW or optimized MPK CUDA kernel (a per-task .cuh under include/mirage/persistentkernel/tasks/) that must PROVABLY beat a target WITHOUT the kernel-optimizer…

mirage-project/mirage2.5k—~1.3kAutomated safety check: PassApache-2.02 days ago
96
96.Hyperpod NcclOfficial

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

awslabs/agent-plugins916—~3.4kAutomated safety check: PassApache-2.0today