Search
AI & LLM Engineering · CUDA
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version. | mlc-ai/ | 175 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 50 | 50.New Class Scaffold new Celeritas source and test files with the required copyright header and register them in CMake. | celeritas-project/ | 105 | — | ~496 | Automated safety check: Pass | Unknown | today |
| 51 | Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~424 | Automated safety check: Pass | Unknown | 24 days ago |
| 52 | 52.Setup Guide A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment. | Red-Hat-AI-Innovation-Team/ | 100 | — | ~959 | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 53 | Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 54 | 54.Model Deploy Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). | sohu-mptc/ | 107 | — | ~974 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 55 | Guided workflow for adding a new model architecture to llama.cpp. | JakeATX/ | 166 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 56 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM… | vipshop/ | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 57 | Add, fix, or validate Triton Runner support for an exact Triton version. | toyaix/ | 100 | — | ~1.1k | Automated safety check: Pass | MIT | 24 days ago |
| 58 | A skill your agent uses for performance profiling and optimization. | ByteDance-Seed/ | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 59 | 59.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 60 | Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed… | pytorch/ | 113 | — | ~2k | Automated safety check: Pass | Unknown | today |
| 61 | A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is… | NVIDIA/ | 138 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 62 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 63 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 938 | — | ~2k | Automated safety check: Pass | No licence | 6 days ago |
| 64 | Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks) | sgl-project/ | 37k | 2 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 65 | Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging | sgl-project/ | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 66 | Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves… | Red-Hat-AI-Innovation-Team/ | 100 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 4 days ago |
| 67 | 67.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 68 | 68.Make Op Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done. | CVCUDA/ | 2.7k | — | ~831 | Automated safety check: Pass | Unknown | 24 days ago |
| 69 | A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing… | vipshop/ | 1.3k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 70 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 6 days ago |
| 71 | 71.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 26 days ago |
| 72 | Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired… | NVIDIA/ | 560 | — | ~2.7k | Automated safety check: Pass | Unknown | yesterday |
| 73 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 74 | Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends. | ahrefs/ | 118 | — | ~728 | Automated safety check: Pass | BSD-2-Clause | 4 days ago |
| 75 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 76 | A skill your agent uses when composing the Search(...) call and calling .start(). | NVIDIA/ | 138 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 77 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 184 | — | ~3.4k | Automated safety check: Pass | Unknown | yesterday |
| 78 | 78.Code Review Review llama.cpp changes against project conventions and common reviewer pitfalls before a PR. | JakeATX/ | 166 | — | ~5.6k | Automated safety check: Pass | MIT | yesterday |
| 79 | Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific… | alibaba/ | 168 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 80 | Productionize a vLLM-Omni diffusion model after its Day-0 vertical slice works. | vllm-project/ | 7.1k | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | today |
| 81 | Review 3DGS implementation code for correctness, performance bugs, and best practices. | jaccen/ | 161 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 82 | 82.AI Review 使用 PaddlePaddle 仓库规则评审 Pull Request 和全仓库代码变更,覆盖正确性、兼容性、算子、分布式、数值、性能、安全、测试、构建和 PR 信息。当需要审查 Paddle 的代码、测试、算子 YAML、C++/CUDA/XPU kernel、Python API、分布式逻辑或 CI 配置时使用。 | PaddlePaddle/ | 24k | — | ~303 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 83 | Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 84 | Scaffold a new CV-CUDA operator — a complete, wired, building skeleton — and delegate the implementation to a human or another AI. | CVCUDA/ | 2.7k | — | ~306 | Automated safety check: Pass | Unknown | 24 days ago |
| 85 | Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. | facebookexperimental/ | 201 | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 86 | Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives. | Orchestra-Research/ | 13k | 4 repos | ~3k | Automated safety check: Warn | MIT | 3 mo ago |
| 87 | 87.Cuda Streams Manage tensor lifetimes across CUDA streams and decide whether recordstream is avoidable. | pytorch/ | 104k | — | ~796 | Automated safety check: Pass | Unknown | today |
| 88 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 89 | Workflow for upgrading/integrating cache-dit in SGLang diffusion (multimodalgen): DBCache, DMD calibrator, TaylorSeer, SVDQuant DQ; porting upstream PRs and resolving conflicts against the… | sgl-project/ | 37k | — | ~7.1k | Automated safety check: Pass | Apache-2.0 | today |
| 90 | 90.Safe Debug Rigor Debug / Rigor Audit skill for deep learning research work. | lllllllama/ | 497 | 1 repo | ~522 | Automated safety check: Pass | MIT | 18 days ago |
| 91 | Add or debug an AReno model family, including config conversion, module construction, checkpoint load/save, text or multimodal inference, training backward, tensor parallelism, CUDA graph decode… | inclusionAI/ | 323 | — | ~722 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 92 | Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate). | CVCUDA/ | 2.7k | — | ~433 | Automated safety check: Pass | Unknown | 24 days ago |
| 93 | Recover from GPU-busy / GPU-unavailable failures. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~709 | Automated safety check: Pass | MIT | today |
| 94 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 95 | 95.Debug Test Update .vscode/launch.json to debug a specific CTest test by name. | celeritas-project/ | 105 | — | ~199 | Automated safety check: Pass | Unknown | today |
| 96 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |