Search

AI & LLM Engineering · CUDA

209 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang.

sgl-project/sglang37k3 repos~2.1kAutomated safety check: PassApache-2.0today
2

Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

pytorch/pytorch104k1 repo~1.7kAutomated safety check: PassUnknowntoday
3

A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

PaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0yesterday
4

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

PaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0yesterday
5

Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

sgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0today
6

Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm.

ztxz16/fastllm5.1k—~1.8kAutomated safety check: PassApache-2.0yesterday
7

Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

microsoft/onnxruntime22k—~1.3kAutomated safety check: PassMITtoday
8

Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~1.6kAutomated safety check: PassUnknowntoday
9

LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、…

xlite-dev/LeetCUDA12k—~3.5kAutomated safety check: PassGPL-3.0yesterday
10

Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands.

microsoft/onnxruntime22k—~1.4kAutomated safety check: PassMITtoday
11

Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.03 mo ago
12

Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

stas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.04 days ago
13

Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

greyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.04 days ago
14

Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.

pytorch/executorch5.1k—~2.3kAutomated safety check: NotesUnknownyesterday
15

Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL…

guqiong96/Lvllm4652 repos~1.5kAutomated safety check: PassApache-2.019 days ago
16

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
17

Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

Blackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT3 days ago
18

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
19

A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

ByteDance-Seed/VeOmni2.2k—~2.8kAutomated safety check: PassApache-2.0yesterday
20
20.Cuda

CUDA kernel development, debugging, and performance optimization for Claude Code.

technillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNo licence9 mo ago
21

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
22

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

KernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT1 mo ago
23

在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

PaddlePaddle/Paddle24k—~1.4kAutomated safety check: PassApache-2.0yesterday
24

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

huggingface/skills11k3 repos~945Automated safety check: PassApache-2.03 days ago
25

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence6 days ago
26

Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary.

CVCUDA/CV-CUDA2.7k—~834Automated safety check: PassUnknown24 days ago
27

Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered.

meta-pytorch/attention-gym1.3k—~6.4kAutomated safety check: PassBSD-3-Clause3 days ago
28

Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

pytorch/pytorch104k—~4.9kAutomated safety check: PassUnknowntoday
29

Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

TongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT26 days ago
30

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

vipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0yesterday
31

Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

inclusionAI/AReno323—~498Automated safety check: PassApache-2.0yesterday
32

PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

PaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0yesterday
33

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
34

A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop…

RightNow-AI/AutoMegaKernel151—~1.8kAutomated safety check: PassMIT23 days ago
35

Diagnose and fix Cosmos3 environment, installation, and runtime errors.

NVIDIA/cosmos-framework560—~1.3kAutomated safety check: NotesUnknownyesterday
36

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

mit-han-lab/ncu-report-skill251—~2kAutomated safety check: PassMIT1 mo ago
37

Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

intel/auto-round1.6k—~2.3kAutomated safety check: PassApache-2.0yesterday
38

Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

zhongkaifu/TensorSharp568—~1kAutomated safety check: PassBSD-3-Clausetoday
39

Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…

evo-design/proto-tools135—~2.5kAutomated safety check: NotesMITtoday
40

Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
41

Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

Aseiel/VideoHighlighter166—~839Automated safety check: PassAGPL-3.0yesterday
42

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
43

GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI.

spiriMirror/libuipc336—~3.6kAutomated safety check: PassApache-2.08 days ago
44

Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

brevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.03 days ago
45

A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

vipshop/cache-dit1.3k—~2.8kAutomated safety check: PassApache-2.0yesterday
46

基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

LMIXR/CV_Deployment_skill188—~547Automated safety check: PassNo licence12 days ago
47

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

guqiong96/Lsglang1441 repo~10kAutomated safety check: PassApache-2.06 days ago
48

Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage

maxiaosong1124/ncu-cuda-profiling-skill129—~1.6kAutomated safety check: PassMIT4 mo ago