Search

Python · GPU and accelerator computing

33 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.08 days ago
2

Trains and fine-tunes object detection, image classification and SAM or SAM2 segmentation models on Hugging Face Jobs cloud GPUs and saves the results to the Hub.

huggingface/skills11k1 repo~7.5kAutomated safety check: PassApache-2.06 days ago
3

Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

Orchestra-Research/AI-Research-SKILLs13k7 repos~2.3kAutomated safety check: PassMIT3 mo ago
4

Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

microsoft/onnxruntime22k—~2.9kAutomated safety check: PassMITtoday
5

Quick-reference help for fine-tuning language models with Axolotl, covering YAML configs, FSDP, context parallelism, compressed saves and dataset formats.

Orchestra-Research/AI-Research-SKILLs13k9 repos~1.2kAutomated safety check: PassMIT3 mo ago
6

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k5 repos~1.3kAutomated safety check: PassMIT3 mo ago
7

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

guqiong96/Lsglang1431 repo~10kAutomated safety check: PassApache-2.03 days ago
8

Adds distributed and mixed-precision training to a PyTorch script with a few Accelerate lines, then launches it on one GPU, many GPUs or DeepSpeed and FSDP setups.

Orchestra-Research/AI-Research-SKILLs13k6 repos~2.1kAutomated safety check: PassMIT3 mo ago
9

dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

dstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.0today
10

Loads large language models in 8-bit or 4-bit with bitsandbytes so they fit smaller GPUs, and sets up QLoRA fine-tuning on a 4-bit base model.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.5kAutomated safety check: PassMIT3 mo ago
11

Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

Orchestra-Research/AI-Research-SKILLs13k3 repos~2.7kAutomated safety check: PassMIT3 mo ago
12

Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.

huggingface/skills11k2 repos~4.6kAutomated safety check: PassApache-2.06 days ago
13

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
14

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
15

Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use.

davila7/claude-code-templates32k10 repos~2.4kAutomated safety check: PassMITtoday
16

Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

Orchestra-Research/AI-Research-SKILLs13k1 repo~3.7kAutomated safety check: PassMIT3 mo ago
17

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins9151 repo~910Automated safety check: PassApache-2.02 days ago
18

Compares SGLang, vLLM, TensorRT-LLM and TokenSpeed on one model and workload, searching server flags to find the best deployment command within a latency SLA.

BBuf/AI-Infra-Auto-Driven-SKILLS911—~7.5kAutomated safety check: PassNo licence2 days ago
19

Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton.

facebookexperimental/triton201—~709Automated safety check: PassMITtoday
20

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills398—~2.3kAutomated safety check: PassMITtoday
21
21.Ir DebuggingOfficial

Debug Triton compilation by dumping IR at each stage (TTIR, TTGIR, LLVM, PTX).

facebookexperimental/triton201—~644Automated safety check: PassMITtoday
22

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups

sgl-project/sglang37k—~13kAutomated safety check: PassApache-2.0today
23

Estimate peak GPU memory for a DualPipeV training run. An agent skill from mlc-ai/pith-train.

mlc-ai/pith-train355—~1.9kAutomated safety check: PassApache-2.03 days ago
24

Builds a fused row-wise softmax kernel for Triton-Ascend that reads and writes each row once, handling padding, strides and masked loads on Ascend NPUs.

Krusty84/triton-ascend-agent-dev-kit106—~636Automated safety check: PassApache-2.01 mo ago
25

Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming).

yzlnew/infra-skills149—~2.4kAutomated safety check: PassNo licence3 mo ago
26

Modal is a serverless cloud platform for running Python on demand, including on-demand GPUs.

K-Dense-AI/scientific-agent-skills48k1 repo~4.5kAutomated safety check: NotesApache-2.03 days ago
27

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

K-Dense-AI/scientific-agent-skills48k1 repo~3.4kAutomated safety check: PassMIT3 days ago
28

A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

NVIDIA/skills3.5k1 repo~2.4kAutomated safety check: NotesApache-2.0today
29

Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures.

NVIDIA/skills3.5k—~3.9kAutomated safety check: PassApache-2.0today
30

A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless…

ericrisco/rsc-harness167—~2.8kAutomated safety check: PassMITtoday
31

Cloud computing platform for running Python on GPUs and serverless infrastructure.

BioTender-max/awesome-bio-agent-skills197—~3.1kAutomated safety check: NotesApache-2.03 mo ago
32

GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT.

majiayu000/claude-skill-registry6661 repo~8.5kAutomated safety check: PassMITtoday
33

Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.

mohitmishra786/low-level-dev-skills253—~1.8kAutomated safety check: PassMIT3 mo ago