Search

CUDA

280 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

sgl-project/sglang37k3 repos~2.1kAutomated safety check: PassApache-2.0today
2

Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

pytorch/pytorch104k1 repo~1.7kAutomated safety check: PassUnknowntoday
3

A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

PaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0today
4

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0today
5

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
6

Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

ztxz16/fastllm5.1k—~1.8kAutomated safety check: PassApache-2.02 days ago
7

Create a CUDA Python pull request from an approved personal or organization-owned fork, including the GitHub CLI GraphQL fallback for renamed organization-owned forks.

NVIDIA/cuda-python3.4k—~1.1kAutomated safety check: PassApache-2.0today
8

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

microsoft/onnxruntime22k—~1.3kAutomated safety check: PassMITtoday
9

Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~1.6kAutomated safety check: PassUnknowntoday
10

Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.

microsoft/onnxruntime22k—~1.4kAutomated safety check: PassMITtoday
11

Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.03 mo ago
12

Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

stas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.03 days ago
13

Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

greyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.02 days ago
14

Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.

pytorch/executorch5.1k—~2.3kAutomated safety check: NotesUnknowntoday
15

Find and safely apply per-operator refactoring / redundancy-reduction opportunities in a CV-CUDA operator (near-duplicate Tensor/VarShape kernels, reinvented shared utilities, dead code).

CVCUDA/CV-CUDA2.7k—~1.5kAutomated safety check: PassUnknown22 days ago
16

Read and write real documents on the device - PDF, XLSX, DOCX, PPTX and CSV.

zhongkaifu/TensorSharp559—~4.2kAutomated safety check: PassBSD-3-Clausetoday
17

Drafts ONNX Runtime release notes from commit history and contributor metadata using named presets for the full runtime or a scoped component.

microsoft/onnxruntime22k—~1.6kAutomated safety check: PassMITtoday
18

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4652 repos~1.5kAutomated safety check: PassApache-2.017 days ago
19

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
20

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMITyesterday
21

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.011 days ago
22

Predicts how small molecules bind to a protein with DiffDock, covering batch docking, pose ranking by confidence and checks on the results; not for binding affinity.

K-Dense-AI/scientific-agent-skills48k1 repo~3kAutomated safety check: NotesMIT4 days ago
23

A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

ByteDance-Seed/VeOmni2.2k—~2.8kAutomated safety check: PassApache-2.0today
24
24.Cuda

CUDA kernel development, debugging, and performance optimization for Claude Code.

technillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNo licence9 mo ago
25

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
26

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

KernelFlow-ops/cuda-optimized-skill213—~4.3kAutomated safety check: PassMIT1 mo ago
27
27.Cpp

Comprehensive C/C++ programming reference covering everything from C11-C23 and C++11-C++23, system programming, CUDA GPU computing, debugging tools, Rust interop, and advanced topics.

crazyguitar/cppcheatsheet290—~1.8kAutomated safety check: PassMIT2 days ago
28

在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

PaddlePaddle/Paddle24k—~1.4kAutomated safety check: PassApache-2.0today
29

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

huggingface/skills11k3 repos~945Automated safety check: PassApache-2.0yesterday
30

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNo licence4 days ago
31

Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

CVCUDA/CV-CUDA2.7k—~834Automated safety check: PassUnknown22 days ago
32

Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.

meta-pytorch/attention-gym1.3k—~6.4kAutomated safety check: PassBSD-3-Clauseyesterday
33

Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~4.9kAutomated safety check: PassUnknowntoday
34

Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

TongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT24 days ago
35

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

vipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.010 days ago
36

Runs the ONNX Runtime transformers Python tests against a GPU wheel and proves the cuDNN flash attention path was used rather than a silent fallback.

microsoft/onnxruntime22k—~2.9kAutomated safety check: PassMITtoday
37

Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

inclusionAI/AReno323—~498Automated safety check: PassApache-2.0today
38

PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

PaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0today
39

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
40

A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop…

RightNow-AI/AutoMegaKernel148—~1.8kAutomated safety check: PassMIT21 days ago
41

Diagnose and fix Cosmos3 environment, installation, and runtime errors.

NVIDIA/cosmos-framework559—~1.3kAutomated safety check: NotesUnknowntoday
42

Review a CV-CUDA operator end-to-end (support / test / bench / docs coverage).

CVCUDA/CV-CUDA2.7k—~481Automated safety check: PassUnknown22 days ago
43

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

mit-han-lab/ncu-report-skill248—~2kAutomated safety check: PassMIT1 mo ago
44

Generate a local phyai environment report for debugging system, Python, CUDA/GPU, dependency, workspace package, git, and PHYAI configuration issues.

mingti-org/phyai130—~590Automated safety check: PassMIT3 days ago
45

Upgrade shared Modal runtime dependencies in kernelbot and verify them end to end.

gpu-mode/kernelbot114—~748Automated safety check: PassUnknown21 days ago
46

Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

intel/auto-round1.6k—~2.3kAutomated safety check: PassApache-2.0today
47

Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them.

radixark/miles_diffusion107—~1.6kAutomated safety check: PassApache-2.0today
48

Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

zhongkaifu/TensorSharp559—~1kAutomated safety check: PassBSD-3-Clausetoday