Search
AI & LLM Engineering · CUDA
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests. | intel/ | 115 | — | ~917 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 98 | Review a CV-CUDA operator's BENCHMARK coverage — drivers, layout axis, baselines, the basic-tier floor, row counts, and coverage statistics. | CVCUDA/ | 2.7k | — | ~274 | Automated safety check: Pass | Unknown | 24 days ago |
| 99 | 99.GPU Backend pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md. | joselado/ | 145 | — | ~679 | Automated safety check: Pass | GPL-3.0 | 4 days ago |
| 100 | 100.Add Model 给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。 | sohu-mptc/ | 107 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 101 | 101.Add Jit Kernel Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups | sgl-project/ | 37k | — | ~13k | Automated safety check: Pass | Apache-2.0 | today |
| 102 | 102.App Opinionated app components building on top of ./ui primitives | JakeATX/ | 166 | — | ~146 | Automated safety check: Pass | MIT | yesterday |
| 103 | Start, validate, debug, and stop an AReno OpenAI-compatible serving endpoint. | inclusionAI/ | 323 | — | ~409 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 104 | Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server. | vllm-project/ | 102 | — | ~2.5k | Automated safety check: Notes | Apache-2.0 | 6 mo ago |
| 105 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 106 | 106.Optimize For GPU GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 6 days ago |
| 107 | Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining. | NVIDIA/ | 3.6k | — | ~2k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 108 | Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix. | CVCUDA/ | 2.7k | — | ~248 | Automated safety check: Pass | Unknown | 24 days ago |
| 109 | 109.Nvmolkit Usage Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF… | NVIDIA-BioNeMo/ | 479 | — | ~4.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 110 | Update tools/scripts/generatebinarybuildmatrix.py when a PyTorch release goes live. | pytorch/ | 113 | — | ~1.7k | Automated safety check: Pass | Unknown | today |
| 111 | How to build and run GPU targets under Buck in fbcode. An agent skill from facebookexperimental/triton. | facebookexperimental/ | 201 | — | ~998 | Automated safety check: Pass | MIT | today |
| 112 | 112.Vllm Server Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills. | sickn33/ | 47k | 2 repos | ~1.7k | Automated safety check: Pass | MIT | 2 days ago |
| 113 | A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an… | mirage-project/ | 2.5k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 114 | A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin… | mirage-project/ | 2.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 115 | 115.Mixed Precision Use FP16/BF16 mixed precision to accelerate training and reduce memory. | aiming-lab/ | 15k | — | ~275 | Automated safety check: Pass | MIT | 1 mo ago |
| 116 | Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures… | awslabs/ | 916 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 117 | AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU. | sickn33/ | 47k | 1 repo | ~4.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 118 | Install Triton + SageAttention to accelerate ComfyUI (the sageattn attentionmode and inductor torch.compile used by WanVideoWrapper / many video graphs). | artokun/ | 803 | — | ~5k | Automated safety check: Pass | MIT | 6 days ago |
| 119 | Apply Gkeyll naming conventions when creating, editing, or reviewing C, CUDA, or Lua files and code elements. | gkeyllorg/ | 113 | — | ~301 | Automated safety check: Pass | MIT | yesterday |
| 120 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 121 | A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels… | xuzhougeng/ | 1k | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | yesterday |
| 122 | Write, review, and run high-level xTBloom Python GFN2-xTB inference with Calculator, Structure, and BatchCalculator, including single systems, repeated geometry updates, heterogeneous ragged… | jinzhezenggroup/ | 148 | — | ~1.3k | Automated safety check: Pass | LGPL-3.0 | 2 days ago |
| 123 | Install or verify the correct ONNX Runtime build (and the onnx package) for a user's accelerator backend before Quark's ONNX-to-ONNX flow. | amd/ | 182 | — | ~3.2k | Automated safety check: Pass | MIT | 13 days ago |
| 124 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.6k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 125 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 126 | Continue KERMT pretraining on a custom SMILES corpus with a groverbase, cmim, or hybrid checkpoint. | NVIDIA/ | 3.6k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 127 | Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint. | NVIDIA/ | 3.6k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 128 | Finetune a pretrained KERMT encoder on a labeled CSV. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | 1 repo | ~4.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 129 | A skill your agent uses when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches. | NVIDIA/ | 3.6k | 1 repo | ~4.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 130 | Check whether a CV-CUDA operator is READY to optimize (correctness + bench coverage + captured baseline + profiling) per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~255 | Automated safety check: Pass | Unknown | 24 days ago |
| 131 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 132 | 132.Xmake XMake build configuration, options, commands, and patterns for LuisaCompute. | LuisaGroup/ | 1.1k | — | ~13k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 133 | Evaluate whether an existing hot path is a credible NVIDIA Warp candidate. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 134 | A skill your agent uses when training, evaluating, or exporting Workflow policies with online RSL-RL or RLinf, including RL checkpoint and Workflow handoff. | NVIDIA/ | 3.6k | 1 repo | ~3.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 135 | Pretrain a fresh KERMT model from scratch on a user-provided corpus. | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 136 | 136.Heartmula HeartMuLa: Suno-like song generation from lyrics + tags. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 2 repos | ~1.6k | Automated safety check: Pass | MIT | 3 days ago |
| 137 | Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning. | amd/ | 182 | — | ~4.3k | Automated safety check: Pass | MIT | 13 days ago |
| 138 | A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the… | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 139 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 140 | DALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks. | NVIDIA/ | 3.6k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 141 | Mandatory pre-flight compute resource check before running experiments. | OpenLAIR/ | 1.2k | — | ~1.8k | Automated safety check: Pass | MIT | 23 days ago |
| 142 | A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance. | NVIDIA/ | 3.6k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 143 | Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI). | NVIDIA/ | 3.6k | — | ~3.2k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 144 | NVIDIA DeepStream SDK development with Python pyservicemaker API. | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |