Topic · AI & LLM Engineering

Best GPU and accelerator computing skills for Claude Code, Codex and other agents.

Skills that provision GPUs and optimise compute-heavy model workloads.
skills
171
official
74

GPU and accelerator computing skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

GPU and accelerator computing skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF.

huggingface/skills11k3 repos~7.2kAutomated safety check: PassApache-2.06 days ago
2

Runs an acceptance-style evaluation of an ML GPU cluster: environment dump, matmul FLOPS, all-reduce, disk benchmarks and a dated markdown report.

stas00/ml-engineering19k—~25kAutomated safety check: NotesCC-BY-SA-4.0yesterday
3

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.07 days ago
4

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
5

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

microsoft/onnxruntime22k—~1.3kAutomated safety check: PassMITtoday
6

Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~1.6kAutomated safety check: PassUnknowntoday
7

Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

huggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.06 days ago
8

Adds Liger Kernel support for a new HuggingFace Transformers model, or modifies existing monkey-patching.

linkedin/Liger-Kernel6.6k—~1.3kAutomated safety check: PassBSD-2-Clausetoday
9

Operates the Inspire ML platform through its local `inspire` CLI: picking account, workspace and resources, launching notebooks, jobs and services, then cleaning up.

realZillionX/InspireSkill548—~1.4kAutomated safety check: PassMIT10 days ago
10

Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.

linkedin/Liger-Kernel6.6k—~799Automated safety check: PassBSD-2-Clausetoday
11

Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA…

fla-org/flash-linear-attention5.8k—~4.2kAutomated safety check: PassMITyesterday
12

Optimizes the performance of existing Liger Kernel Triton kernels.

linkedin/Liger-Kernel6.6k—~1.5kAutomated safety check: PassBSD-2-Clausetoday
13

Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

wshobson/agents40k1 repo~2kAutomated safety check: PassMIT2 days ago
14

Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~2.5kAutomated safety check: PassNo licence2 days ago
15

Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit.

vllm-project/vllm-ascend2.9k—~2.2kAutomated safety check: PassApache-2.0today
16

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.02 mo ago
17

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.09 days ago
18
18.Cuda

CUDA kernel development, debugging, and performance optimization for Claude Code.

technillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNo licence9 mo ago
19

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

KernelFlow-ops/cuda-optimized-skill212—~4.3kAutomated safety check: PassMIT1 mo ago
20

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k1 repo~2kAutomated safety check: PassMIT2 days ago
21

Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno.

inclusionAI/AReno323—~486Automated safety check: PassApache-2.013 days ago
22

Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

huggingface/skills11k1 repo~7.5kAutomated safety check: PassApache-2.06 days ago
23

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~2.8kAutomated safety check: PassNo licence2 days ago
24

Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~4.9kAutomated safety check: PassUnknowntoday
25

Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands.

internet-court/internet-court-skill6.4k1 repo~1.9kAutomated safety check: PassUnknown1 mo ago
26

Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

Orchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT3 mo ago
27

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

vipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.08 days ago
28

Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

microsoft/onnxruntime22k—~2.9kAutomated safety check: PassMITtoday
29

A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/.

ByteDance-Seed/VeOmni2.2k—~3.1kAutomated safety check: PassApache-2.07 days ago
30

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

mit-han-lab/ncu-report-skill244—~2kAutomated safety check: PassMIT1 mo ago
31

Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats.

Orchestra-Research/AI-Research-SKILLs13k9 repos~1.2kAutomated safety check: PassMIT3 mo ago
32

Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

BBuf/AI-Infra-Auto-Driven-SKILLS900—~4.5kAutomated safety check: PassNo licence2 days ago
33

GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI.

spiriMirror/libuipc335—~3.6kAutomated safety check: PassApache-2.04 days ago
34

Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches.

pytorch/pytorch104k—~3.5kAutomated safety check: PassUnknowntoday
35

Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.

NVIDIA/Megatron-LM18k—~1.3kAutomated safety check: PassApache-2.0today
36

Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.1kAutomated safety check: PassMIT3 mo ago
37

GPU-accelerated data curation for LLM training. An agent skill from Orchestra-Research/AI-Research-SKILLs.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.3kAutomated safety check: PassMIT3 mo ago
38

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT3 mo ago
39

Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation.

guqiong96/Lvllm4641 repo~831Automated safety check: PassApache-2.015 days ago
40

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

guqiong96/Lsglang1431 repo~10kAutomated safety check: PassApache-2.02 days ago
41

Shows how to build and run ONNX Runtime WebGPU provider tests on Linux with no GPU, using the Mesa lavapipe software Vulkan adapter, and where that approach falls short.

microsoft/onnxruntime22k—~1.4kAutomated safety check: PassMITtoday
42

Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost.

Orchestra-Research/AI-Research-SKILLs13k4 repos~2.4kAutomated safety check: PassMIT3 mo ago
43

Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

Orchestra-Research/AI-Research-SKILLs13k6 repos~2.1kAutomated safety check: PassMIT3 mo ago
44

基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

LMIXR/CV_Deployment_skill126—~547Automated safety check: PassNo licence8 days ago
45

Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results.

climate-analytics-lab/jax-gcm108—~3.4kAutomated safety check: PassApache-2.0today
46

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
47

dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

dstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.0today
48

Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT3 mo ago

Questions, answered from the data.

What is the best GPU and accelerator computing skill?

Hugging Face LLM Trainer (official) from huggingface/skills ranks first of the 171 GPU and accelerator computing skills listed here, with the highest score: its repository has 11k GitHub stars, 3 other GitHub owners carry a copy, its SKILL.md loads about 7.2k tokens and it passes the automated safety check with no findings. Next come GPU Cluster Evaluation and Paddle Design Compiler.

Which GPU and accelerator computing skills are official?

74 of the 171 GPU and accelerator computing skills are official, published by the vendor's own GitHub organization: Hugging Face LLM Trainer, CUTLASS FMHA Incremental Rebuild, Hugging Face Local Model Evals, Hugging Face Vision Trainer, ONNX Runtime GPU Transformers Tests and 69 more.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.