Topic · AI & LLM Engineering
Best GPU and accelerator computing skills for Claude Code, Codex and other agents.
- skills
- 171
- official
- 74
GPU and accelerator computing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Trains or fine-tunes language and vision models with TRL or Unsloth on Hugging Face Jobs cloud GPUs, then converts the results to GGUF. | huggingface/ | 11k | 3 repos | ~7.2k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 2 | Runs an acceptance-style evaluation of an ML GPU cluster: environment dump, matmul FLOPS, all-reduce, disk benchmarks and a dated markdown report. | stas00/ | 19k | — | ~25k | Automated safety check: Notes | CC-BY-SA-4.0 | yesterday |
| 3 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 4 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 6 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 7 | Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends. | huggingface/ | 11k | 2 repos | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 8 | Adds Liger Kernel support for a new HuggingFace Transformers model, or modifies existing monkey-patching. | linkedin/ | 6.6k | — | ~1.3k | Automated safety check: Pass | BSD-2-Clause | today |
| 9 | Operates the Inspire ML platform through its local `inspire` CLI: picking account, workspace and resources, launching notebooks, jobs and services, then cleaning up. | realZillionX/ | 548 | — | ~1.4k | Automated safety check: Pass | MIT | 10 days ago |
| 10 | Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel. | linkedin/ | 6.6k | — | ~799 | Automated safety check: Pass | BSD-2-Clause | today |
| 11 | Workflow for porting an existing Triton kernel in fla/ops/ to Gluon (triton.experimental.gluon) to gain explicit control over tensor layouts, shared memory, async data movement (cp.async / TMA), MMA… | fla-org/ | 5.8k | — | ~4.2k | Automated safety check: Pass | MIT | yesterday |
| 12 | Optimizes the performance of existing Liger Kernel Triton kernels. | linkedin/ | 6.6k | — | ~1.5k | Automated safety check: Pass | BSD-2-Clause | today |
| 13 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 14 | Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths. | BBuf/ | 900 | — | ~2.5k | Automated safety check: Pass | No licence | 2 days ago |
| 15 | Adapts and debugs Hugging Face or local models to run on vLLM with Ascend NPU, validates them by serving, and delivers the result as one signed commit. | vllm-project/ | 2.9k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 16 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 2 mo ago |
| 17 | 17.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 18 | 18.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 19 | Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass. | KernelFlow-ops/ | 212 | — | ~4.3k | Automated safety check: Pass | MIT | 1 mo ago |
| 20 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | 1 repo | ~2k | Automated safety check: Pass | MIT | 2 days ago |
| 21 | Diagnose failed, hung, slow, OOM, NaN, illegal-memory-access, NCCL, compilation, rollout, or training runs in AReno. | inclusionAI/ | 323 | — | ~486 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 22 | Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub. | huggingface/ | 11k | 1 repo | ~7.5k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 23 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 900 | — | ~2.8k | Automated safety check: Pass | No licence | 2 days ago |
| 24 | 24.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 25 | Guides building on the 0G Compute Network, a decentralized GPU marketplace for AI inference and fine-tuning, with SDK patterns and CLI commands. | internet-court/ | 6.4k | 1 repo | ~1.9k | Automated safety check: Pass | Unknown | 1 mo ago |
| 26 | Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling. | Orchestra-Research/ | 13k | 7 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 28 | Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback. | microsoft/ | 22k | — | ~2.9k | Automated safety check: Pass | MIT | today |
| 29 | A skill your agent uses when adding a new optimized kernel or operator to veomni/ops/. | ByteDance-Seed/ | 2.2k | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 30 | Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill. | mit-han-lab/ | 244 | — | ~2k | Automated safety check: Pass | MIT | 1 mo ago |
| 31 | Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats. | Orchestra-Research/ | 13k | 9 repos | ~1.2k | Automated safety check: Pass | MIT | 3 mo ago |
| 32 | Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks. | BBuf/ | 900 | — | ~4.5k | Automated safety check: Pass | No licence | 2 days ago |
| 33 | GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 335 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 34 | Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches. | pytorch/ | 104k | — | ~3.5k | Automated safety check: Pass | Unknown | today |
| 35 | Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs. | NVIDIA/ | 18k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 36 | Serverless GPU cloud platform for running ML workloads. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 5 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 37 | 37.Nemo Curator GPU-accelerated data curation for LLM training. An agent skill from Orchestra-Research/AI-Research-SKILLs. | Orchestra-Research/ | 13k | 5 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 38 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 39 | Write or review Triton kernels for vLLM, with practical guidance for generated-code inspection, launch grids, indexing, specialization, tuning, and representative performance validation. | guqiong96/ | 464 | 1 repo | ~831 | Automated safety check: Pass | Apache-2.0 | 15 days ago |
| 40 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 143 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 41 | Shows how to build and run ONNX Runtime WebGPU provider tests on Linux with no GPU, using the Mesa lavapipe software Vulkan adapter, and where that approach falls short. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 42 | Runs ML training and batch jobs across clouds with SkyPilot, using spot instances, automatic region selection and managed recovery to cut GPU cost. | Orchestra-Research/ | 13k | 4 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 43 | Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups. | Orchestra-Research/ | 13k | 6 repos | ~2.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 44 | 44.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 126 | — | ~547 | Automated safety check: Pass | No licence | 8 days ago |
| 45 | Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results. | climate-analytics-lab/ | 108 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 46 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 47 | 47.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | today |
| 48 | Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model. | Orchestra-Research/ | 13k | 3 repos | ~2.5k | Automated safety check: Pass | MIT | 3 mo ago |
Questions, answered from the data.
What is the best GPU and accelerator computing skill?
Hugging Face LLM Trainer (official) from huggingface/skills ranks first of the 171 GPU and accelerator computing skills listed here, with the highest score: its repository has 11k GitHub stars, 3 other GitHub owners carry a copy, its SKILL.md loads about 7.2k tokens and it passes the automated safety check with no findings. Next come GPU Cluster Evaluation and Paddle Design Compiler.
Which GPU and accelerator computing skills are official?
74 of the 171 GPU and accelerator computing skills are official, published by the vendor's own GitHub organization: Hugging Face LLM Trainer, CUTLASS FMHA Incremental Rebuild, Hugging Face Local Model Evals, Hugging Face Vision Trainer, ONNX Runtime GPU Transformers Tests and 69 more.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents525
- Deep learning408
- Embeddings381
- LLM inference and serving364
- Retrieval-augmented generation360
- Prompt engineering350
- Fine-tuning313
- LLM evaluation303
- Speech recognition and synthesis272
- Structured output and tool calling271
- LLM cost and token optimization256
- LLM API integration218
- LLM observability217
- Model routing and gateways217
- LLM guardrails208
- Computer vision206
- Model hubs and datasets180
- Diffusion and image models167
- Natural language processing141
- Reinforcement learning67
- AI interpretability23