Search
CUDA
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 49 | Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them. | radixark/ | 110 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 50 | Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK). | intel/ | 1.6k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 51 | 51.Market Data Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares). | zhongkaifu/ | 568 | — | ~1k | Automated safety check: Pass | BSD-3-Clause | today |
| 52 | 52.Fix Env Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any… | evo-design/ | 135 | — | ~2.5k | Automated safety check: Notes | MIT | today |
| 53 | 53.Chai Structure prediction using Chai-1, a foundation model for molecular structure. | adaptyvbio/ | 164 | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 4 mo ago |
| 54 | Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 55 | Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better. | Aseiel/ | 166 | — | ~839 | Automated safety check: Pass | AGPL-3.0 | today |
| 56 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 57 | GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI. | spiriMirror/ | 336 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | 8 days ago |
| 58 | Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands. | brevdev/ | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 59 | A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or… | vipshop/ | 1.3k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 60 | 60.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 188 | — | ~547 | Automated safety check: Pass | No licence | 12 days ago |
| 61 | 61.Readable Cpp Readable C/C++/Rust/CUDA code rules inspired by The Art of Readable Code. | crazyguitar/ | 290 | — | ~6.4k | Automated safety check: Pass | MIT | yesterday |
| 62 | Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module | guqiong96/ | 144 | 1 repo | ~10k | Automated safety check: Pass | Apache-2.0 | 6 days ago |
| 63 | Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage | maxiaosong1124/ | 129 | — | ~1.6k | Automated safety check: Pass | MIT | 4 mo ago |
| 64 | Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version. | mlc-ai/ | 175 | — | ~5.1k | Automated safety check: Pass | Apache-2.0 | 7 days ago |
| 65 | 65.New Class Scaffold new Celeritas source and test files with the required copyright header and register them in CMake. | celeritas-project/ | 105 | — | ~496 | Automated safety check: Pass | Unknown | today |
| 66 | A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill. | NVIDIA/ | 138 | — | ~1.3k | Automated safety check: Notes | Apache-2.0 | 17 days ago |
| 67 | 67.Research A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages. | zhongkaifu/ | 568 | — | ~2.3k | Automated safety check: Warn | BSD-3-Clause | today |
| 68 | Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~424 | Automated safety check: Pass | Unknown | 24 days ago |
| 69 | 69.Setup Guide A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment. | Red-Hat-AI-Innovation-Team/ | 100 | — | ~959 | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 70 | Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration. | vllm-project/ | 7.1k | — | ~8.7k | Automated safety check: Pass | Apache-2.0 | today |
| 71 | 71.Model Deploy Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling). | sohu-mptc/ | 107 | — | ~974 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 72 | Guided workflow for adding a new model architecture to llama.cpp. | JakeATX/ | 166 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |
| 73 | Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source. | slowlyC/ | 169 | — | ~1.8k | Automated safety check: Pass | MIT | 2 mo ago |
| 74 | A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM… | vipshop/ | 1.3k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 75 | 75.Mindquantum Build, simulate, and analyze quantum circuits with MindQuantum. | mindspore-ai/ | 102 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 19 days ago |
| 76 | Add, fix, or validate Triton Runner support for an exact Triton version. | toyaix/ | 100 | — | ~1.1k | Automated safety check: Pass | MIT | 24 days ago |
| 77 | A skill your agent uses for performance profiling and optimization. | ByteDance-Seed/ | 2.2k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 78 | 78.Cmake CMake build options, custom functions, and backend patterns for LuisaCompute. | LuisaGroup/ | 1.1k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 79 | 79.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 80 | Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed… | pytorch/ | 113 | — | ~2k | Automated safety check: Pass | Unknown | today |
| 81 | A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is… | NVIDIA/ | 138 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 17 days ago |
| 82 | Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits. | huggingface/ | 11k | 2 repos | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 83 | Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. | BBuf/ | 938 | — | ~2k | Automated safety check: Pass | No licence | 6 days ago |
| 84 | Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks) | sgl-project/ | 37k | 2 repos | ~3.4k | Automated safety check: Pass | Apache-2.0 | today |
| 85 | Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging | sgl-project/ | 37k | 2 repos | ~4.9k | Automated safety check: Pass | Apache-2.0 | today |
| 86 | Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves… | Red-Hat-AI-Innovation-Team/ | 100 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 87 | 87.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 88 | Expert guide for Backend.AI distributed computing platform. An agent skill from lablup/backend.ai-webui. | lablup/ | 133 | 1 repo | ~1.8k | Automated safety check: Pass | LGPL-3.0 | today |
| 89 | 89.Make Op Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done. | CVCUDA/ | 2.7k | — | ~831 | Automated safety check: Pass | Unknown | 24 days ago |
| 90 | A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing… | vipshop/ | 1.3k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 91 | Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. | lucifer1004/ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | 6 days ago |
| 92 | 92.Cuda Draft, debug, and measure CUDA kernels and host launch workflows. | sablin39/ | 145 | — | ~990 | Automated safety check: Pass | No licence | 26 days ago |
| 93 | Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired… | NVIDIA/ | 560 | — | ~2.7k | Automated safety check: Pass | Unknown | yesterday |
| 94 | Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing. | microsoft/ | 22k | — | ~6.5k | Automated safety check: Pass | MIT | today |
| 95 | Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends. | ahrefs/ | 118 | — | ~728 | Automated safety check: Pass | BSD-2-Clause | 4 days ago |
| 96 | Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code. | intel/ | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |