Search
AI & LLM Engineering · CUDA · For developers
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Replay-first debug flow for SGLang serving problems. An agent skill from sgl-project/sglang. | sgl-project/ | 37k | 3 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch. | pytorch/ | 104k | 1 repo | ~1.7k | Automated safety check: Pass | Unknown | today |
| 3 | A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes. | PaddlePaddle/ | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 4 | A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate… | PaddlePaddle/ | 24k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 5 | Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP). | sgl-project/ | 37k | 2 repos | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 6 | Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm. | ztxz16/ | 5.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 7 | Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild. | microsoft/ | 22k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 8 | Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~1.6k | Automated safety check: Pass | Unknown | today |
| 9 | LeetCUDA 中文技术书(584 页,XeLaTeX 源)按需查阅 skill——写、优化、调试或 review CUDA C++/PTX kernel 时的权威参考路由层。当任务涉及:GPU 架构/Roofline/ occupancy、向量化与 coalescing、warp/block reduce、softmax(online/LSE merge)、… | xlite-dev/ | 12k | — | ~3.5k | Automated safety check: Pass | GPL-3.0 | yesterday |
| 10 | Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 11 | 11.Esmfold2 Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. | JimLiu/ | 228 | 4 repos | ~2.5k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 12 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 5 days ago |
| 13 | Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI. | greyhaven-ai/ | 1.3k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 14 | Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks. | pytorch/ | 5.1k | — | ~2.3k | Automated safety check: Notes | Unknown | yesterday |
| 15 | Build, debug, and interpret vLLM GPU kernel microbenchmarks for CUDA, Triton, and CuteDSL, including CUPTI timing, correctness checks, generated-code inspection, multi-GPU measurements, and SOL… | guqiong96/ | 465 | 2 repos | ~1.5k | Automated safety check: Pass | Apache-2.0 | 19 days ago |
| 16 | Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks. | Blackwellboy/ | 135 | — | ~2.1k | Automated safety check: Pass | MIT | 3 days ago |
| 17 | 17.Graphsignal Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 18 | 18.Veomni Debug A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected… | ByteDance-Seed/ | 2.2k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 19 | 19.Cuda CUDA kernel development, debugging, and performance optimization for Claude Code. | technillogue/ | 229 | — | ~2.5k | Automated safety check: Pass | No licence | 9 mo ago |
| 20 | 20.Paddle Debug 在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。 | PaddlePaddle/ | 24k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 21 | Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server. | huggingface/ | 11k | 3 repos | ~945 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 22 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 6 days ago |
| 23 | 23.Optimize Op Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary. | CVCUDA/ | 2.7k | — | ~834 | Automated safety check: Pass | Unknown | 24 days ago |
| 24 | Ensures new Attention Gym eager, Triton, CuTeDSL, and external-library implementations are torch.compile-friendly and correctly registered. | meta-pytorch/ | 1.3k | — | ~6.4k | Automated safety check: Pass | BSD-3-Clause | 3 days ago |
| 25 | 25.Metal Kernel Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch. | pytorch/ | 104k | — | ~4.9k | Automated safety check: Pass | Unknown | today |
| 26 | 26.Ako4all Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. | TongmingLAIC/ | 369 | — | ~4k | Automated safety check: Pass | MIT | 26 days ago |
| 27 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 28 | PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML… | PaddlePaddle/ | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 29 | A skill your agent uses when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose - eval - keep/revert loop… | RightNow-AI/ | 151 | — | ~1.8k | Automated safety check: Pass | MIT | 23 days ago |
| 30 | Diagnose and fix Cosmos3 environment, installation, and runtime errors. | NVIDIA/ | 560 | — | ~1.3k | Automated safety check: Notes | Unknown | yesterday |
| 31 | 31.Fix Env Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any… | evo-design/ | 135 | — | ~2.5k | Automated safety check: Notes | MIT | today |
| 32 | Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 33 | Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better. | Aseiel/ | 166 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 34 | GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 336 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 35 | Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands. | brevdev/ | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 36 | A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or… | vipshop/ | 1.3k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 37 | 37.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 188 | — | ~547 | Automated safety check: Pass | No licence | 12 days ago |
| 38 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 144 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 39 | Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version. | mlc-ai/ | 175 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 40 | 40.New Class Scaffold new Celeritas source and test files with the required copyright header and register them in CMake. | celeritas-project/ | 105 | — | ~496 | Automated safety check: Pass | Unknown | today |
| 41 | Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 42 | Guided workflow for adding a new model architecture to llama.cpp. | JakeATX/ | 166 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 43 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM… | vipshop/ | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 44 | A skill your agent uses for performance profiling and optimization. | ByteDance-Seed/ | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 45 | 45.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 46 | Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed… | pytorch/ | 113 | — | ~2k | Automated safety check: Pass | Unknown | today |
| 47 | A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is… | NVIDIA/ | 138 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 48 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 3 days ago |