Search
AI & LLM Engineering · C++
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes. | PaddlePaddle/ | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Guide for adding Triton-backed CUDA operators to FastLLM. An agent skill from ztxz16/fastllm. | ztxz16/ | 5.1k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 3 | A skill your agent uses when navigating Paddle eager-mode (dynamic graph) source code, tracing forward/backward execution, debugging autograd issues, understanding PyLayer, or investigating… | PaddlePaddle/ | 24k | — | ~562 | Automated safety check: Pass | Apache-2.0 | today |
| 4 | 4.Onnxtxt Read or write ONNX text format ("onnxtxt"). An agent skill from onnx/onnx. | onnx/ | 22k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Builds ONNX Runtime from source with its build scripts, explaining the update, build and test phases, key flags and where the build output lands. | microsoft/ | 22k | — | ~1.4k | Automated safety check: Pass | MIT | today |
| 6 | Measures and shrinks the ExecuTorch runtime binary by building a size test, analyzing it with bloaty and landing each reduction as its own pull request. | pytorch/ | 5.1k | — | ~793 | Automated safety check: Pass | Unknown | today |
| 7 | Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks. | pytorch/ | 5.1k | — | ~2.3k | Automated safety check: Notes | Unknown | today |
| 8 | 8.Ako4all Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. | TongmingLAIC/ | 369 | — | ~4k | Automated safety check: Pass | MIT | 24 days ago |
| 9 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems… | vipshop/ | 1.3k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 10 | PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML… | PaddlePaddle/ | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 11 | 11.Edge Bringup Prepare a macOS or Ubuntu machine for edge-e3 development, diagnose missing Verilator/LLVM/Python dependencies, initialize the public repository, and answer or act on the example prompts in the root… | exeex/ | 110 | — | ~1.7k | Automated safety check: Notes | Apache-2.0 | 13 days ago |
| 12 | Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging. | pytorch/ | 5.1k | — | ~1.8k | Automated safety check: Pass | Unknown | today |
| 13 | Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 14 | GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 336 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 15 | A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or… | vipshop/ | 1.3k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 16 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 143 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 17 | 17.New Class Scaffold new Celeritas source and test files with the required copyright header and register them in CMake. | celeritas-project/ | 105 | — | ~496 | Automated safety check: Pass | Unknown | today |
| 18 | Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~424 | Automated safety check: Pass | Unknown | 22 days ago |
| 19 | 19.Test Runner Runs the narrowest relevant tests to validate changes. An agent skill from noumena-labs/Sipp. | noumena-labs/ | 121 | — | ~854 | Automated safety check: Pass | Apache-2.0 | 19 days ago |
| 20 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM… | vipshop/ | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | 10 days ago |
| 21 | Build vLLM (CPU) end-to-end with the oneDNN + ZenDNN (zen64) backend via direct integration (NO zentorch). | amd/ | 158 | — | ~5.2k | Automated safety check: Pass | Unknown | 2 days ago |
| 22 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 183 | — | ~3.4k | Automated safety check: Pass | Unknown | today |
| 23 | Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks) | sgl-project/ | 37k | 2 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 24 | 将原生 PyTorch 自定义算子库、Torch extension、生态库(TorchCodec/FlashInfer/DeepEP 等)以及 Kernel DSL 生态(Triton/TileLang/TVM FFI 等)以最小修改方式接入 PaddlePaddle。遇到以下场景务必使用:迁移外部算子库到 Paddle;分析 PFCCLab fork 与上游的兼容差异;处理… | PaddlePaddle/ | 24k | — | ~883 | Automated safety check: Pass | Apache-2.0 | today |
| 25 | Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed. | scragnog/ | 173 | — | ~4.9k | Automated safety check: Pass | MIT | 2 days ago |
| 26 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 27 | A checklist for upgrading the pinned ONNX version and opset in ONNX Runtime, covering the files to change, archive hashes, patch rebasing and release-candidate handling. | microsoft/ | 22k | — | ~12k | Automated safety check: Pass | MIT | today |
| 28 | 28.AI Review 使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。 | PaddlePaddle/ | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | today |
| 29 | Review C++ changes for string parameter and call-site efficiency conventions (std::stringview, std::string&&, const std::string&, const char, and TransparentStringMap lookup). | tetherto/ | 683 | — | ~702 | Automated safety check: Pass | Apache-2.0 | today |
| 30 | Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate). | CVCUDA/ | 2.7k | — | ~433 | Automated safety check: Pass | Unknown | 22 days ago |
| 31 | Install or verify the AMD Quark package and its dependencies. | amd/ | 181 | — | ~1.8k | Automated safety check: Notes | MIT | 11 days ago |
| 32 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 33 | 33.Llama Nnc Develop and profile the edge-e3 PyTorch-to-NNC flow, including nnc/compiler.py lowering and generated ABI, example/llama smoke models, cpp/libnn runtime operators, BF16 correctness checks, and… | exeex/ | 110 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 34 | Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix. | CVCUDA/ | 2.7k | — | ~248 | Automated safety check: Pass | Unknown | 22 days ago |
| 35 | Agent skill for sona-learning-optimizer - invoke with $agent-sona-learning-optimizer | ruvnet/ | 74k | 2 repos | ~516 | Automated safety check: Pass | MIT | today |
| 36 | Run, extend, debug, or review the public edge-e3 bare-metal software harness, including encrypted Verilator builds, hello and tensor examples, all example/llama/model smoke cases, PyTorch BF16… | exeex/ | 110 | — | ~1k | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 37 | Build and run the public 64x64 x 128-token tiled circular tensor matmul example on the encrypted Edge Verilator simulator. | exeex/ | 110 | — | ~398 | Automated safety check: Pass | Apache-2.0 | 13 days ago |
| 38 | 推論パスの canonical Python (exportonnx.py / vits/models.py:VitsModel.infer) を変更した PR で、6 ランタイム (Python runtime / Rust / Go / C / C++ / WASM) の inference path が追随しているかを git diff で確認。PR | ayutaz/ | 230 | — | ~972 | Automated safety check: Pass | MIT | yesterday |
| 39 | Team-wide PR dashboard for the DevOps pod, scoped to PRs authored by pod-roster members. | tetherto/ | 683 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 40 | A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the… | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | today |
| 41 | 41.Qv Nx CI Understand and modify the pnpm+Nx consolidated CI — the generic -nx.yml leaves driven by each package's project.json options.ci. | tetherto/ | 683 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 42 | Explains how model files, checkpoints, GGUF quantization, and the Model Manager work in HOT-Step CPP. | scragnog/ | 173 | — | ~6.6k | Automated safety check: Pass | MIT | 2 days ago |
| 43 | GenieAPIService technical documentation retrieval. An agent skill from qualcomm/qai-appbuilder. | qualcomm/ | 247 | — | ~840 | Automated safety check: Pass | Unknown | today |
| 44 | A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance. | NVIDIA/ | 3.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 45 | Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI). | NVIDIA/ | 3.5k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | today |
| 46 | 46.Project Map Maps every HOT-Step CPP feature to its route file, service, UI folder, and engine subsystem, including port topology and the browser-to-engine request path. | scragnog/ | 173 | — | ~5.4k | Automated safety check: Notes | MIT | 2 days ago |
| 47 | Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. | matlab/ | 1.1k | — | ~2.8k | Automated safety check: Pass | Unknown | today |
| 48 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 1.1k | — | ~4.6k | Automated safety check: Pass | Unknown | today |