Search

PyTorch · GPU and accelerator computing

26 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~1.6kAutomated safety check: PassUnknowntoday
2

Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.

linkedin/Liger-Kernel6.7k—~799Automated safety check: PassBSD-2-Clause2 days ago
3

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
4

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
5

Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~4.9kAutomated safety check: PassUnknowntoday
6

Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

Orchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT3 mo ago
7

Debug PyTorch 2 compiler stack failures including Dynamo graph breaks, Inductor codegen errors, AOTAutograd crashes, and accuracy mismatches.

pytorch/pytorch104k—~3.5kAutomated safety check: PassUnknowntoday
8

Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

Orchestra-Research/AI-Research-SKILLs13k5 repos~2.1kAutomated safety check: PassMIT3 mo ago
9

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

lucifer1004/VeloQ128—~1.3kAutomated safety check: PassMIT6 days ago
10
10.At Dispatch V2Official

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0today
11

Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.7kAutomated safety check: PassMIT3 mo ago
12

Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
13

Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use.

davila7/claude-code-templates33k10 repos~2.4kAutomated safety check: PassMITtoday
14

Organizes PyTorch training code into LightningModules, DataModules and Trainers, with multi-GPU strategies, callbacks and logging configured.

davila7/claude-code-templates33k11 repos~1.7kAutomated safety check: PassMITtoday
15

Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
16

PyTorch training reference: architecture choice by data type, scaling rules, a training loop, optimizer and learning-rate choices, and fixes for loss spikes or OOM.

Orchestra-Research/AI-Research-SKILLs13k1 repo~2.8kAutomated safety check: PassMIT3 mo ago
17

Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

Orchestra-Research/AI-Research-SKILLs13k4 repos~3kAutomated safety check: WarnMIT3 mo ago
18

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills408—~2.3kAutomated safety check: PassMITyesterday
19

Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
20

Fine-tunes and serves Physical Intelligence's pi0, pi0-fast and pi0.5 robot policies with JAX or PyTorch, including checkpoint conversion and policy servers.

Orchestra-Research/AI-Research-SKILLs13k—~3.6kAutomated safety check: PassMIT3 mo ago
21

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0yesterday
22
22.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.0yesterday
23

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM…

NVIDIA/skills3.6k—~3.6kAutomated safety check: PassApache-2.0yesterday
24

Run distributed GPU training jobs on CoreWeave with multi-node PyTorch.

jeremylongshore/tons-of-skills-marketplace2.8k—~1.2kAutomated safety check: PassMITtoday
25

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

Mathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT4 days ago
26

Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills252—~1.8kAutomated safety check: PassMIT3 mo ago