Search

AI & LLM Engineering · CUDA

209 skills found, page 3.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Review PyTorch upstream unit-test (UT) PRs that enable Intel GPU (XPU) on existing tests.

intel/torch-xpu-ops115—~917Automated safety check: PassApache-2.0yesterday
98

Review a CV-CUDA operator's BENCHMARK coverage — drivers, layout axis, baselines, the basic-tier floor, row counts, and coverage statistics.

CVCUDA/CV-CUDA2.7k—~274Automated safety check: PassUnknown24 days ago
99

pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md.

joselado/pyqula145—~679Automated safety check: PassGPL-3.04 days ago
100

给 FlashRec 引擎接入一个新模型架构(新的 HF checkpoint / 非 Qwen3 结构)。涵盖模型定义、权重合并加载、FP8 双路径、融合 kernel 接线、CUDA graph 兼容、精度校验、以及压测+trace 验证闭环。当用户要"增加/支持/接入新模型"时使用。

sohu-mptc/FlashRec107—~1.7kAutomated safety check: PassApache-2.0yesterday
101

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang.kernels JIT infrastructure and public operator groups

sgl-project/sglang37k—~13kAutomated safety check: PassApache-2.0today
102
102.App

Opinionated app components building on top of ./ui primitives

JakeATX/llamAmpere166—~146Automated safety check: PassMITyesterday
103

Start, validate, debug, and stop an AReno OpenAI-compatible serving endpoint.

inclusionAI/AReno323—~409Automated safety check: PassApache-2.0yesterday
104

Deploy vLLM using Docker (pre-built images or build-from-source) with NVIDIA GPU support and run the OpenAI-compatible server.

vllm-project/vllm-skills102—~2.5kAutomated safety check: NotesApache-2.06 mo ago
105

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0yesterday
106

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

K-Dense-AI/scientific-agent-skills48k1 repo~3.4kAutomated safety check: PassMIT6 days ago
107

Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining.

NVIDIA/skills3.6k—~2kAutomated safety check: NotesApache-2.02 days ago
108

Review a CV-CUDA operator's input-type, layout, dtype, and channel support matrix.

CVCUDA/CV-CUDA2.7k—~248Automated safety check: PassUnknown24 days ago
109

Write code that calls the installed nvMolKit Python API for GPU-accelerated, batched RDKit-style operations - Morgan fingerprints, Tanimoto/cosine similarity, ETKDG conformer embedding, MMFF/UFF…

NVIDIA-BioNeMo/bionemo-agent-toolkit479—~4.4kAutomated safety check: PassApache-2.02 days ago
110

Update tools/scripts/generatebinarybuildmatrix.py when a PyTorch release goes live.

pytorch/test-infra113—~1.7kAutomated safety check: PassUnknowntoday
111
111.Running With BuckOfficial

How to build and run GPU targets under Buck in fbcode. An agent skill from facebookexperimental/triton.

facebookexperimental/triton201—~998Automated safety check: PassMITtoday
112

Deploy and manage vLLM for high-throughput LLM inference. An agent skill from sickn33/agentic-awesome-skills.

sickn33/agentic-awesome-skills47k2 repos~1.7kAutomated safety check: PassMIT2 days ago
113

A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an…

mirage-project/mirage2.5k—~1.8kAutomated safety check: PassApache-2.03 days ago
114

A skill your agent uses when the user wants to confirm whether an existing CUDA extension/binary can run on B200, configure compute100/sm100 or the architecture-specific sm100a, or check PTX/cubin…

mirage-project/mirage2.5k—~1.7kAutomated safety check: PassApache-2.03 days ago
115

Use FP16/BF16 mixed precision to accelerate training and reduce memory.

aiming-lab/AutoResearchClaw15k—~275Automated safety check: PassMIT1 mo ago
116
116.Hyperpod NcclOfficial

Diagnose NCCL failures and adjacent training-pod failures on HyperPod GPU clusters (EKS or Slurm) — training hangs, AllReduce / collective-op timeouts, EFA or libfabric errors, rendezvous failures…

awslabs/agent-plugins916—~3.4kAutomated safety check: PassApache-2.0yesterday
117

AI demos and GPU compute with Gradio Spaces and Hugging Face Spaces ZeroGPU.

sickn33/agentic-awesome-skills47k1 repo~4.6kAutomated safety check: PassApache-2.02 days ago
118

Install Triton + SageAttention to accelerate ComfyUI (the sageattn attentionmode and inductor torch.compile used by WanVideoWrapper / many video graphs).

artokun/comfyui-mcp803—~5kAutomated safety check: PassMIT6 days ago
119

Apply Gkeyll naming conventions when creating, editing, or reviewing C, CUDA, or Lua files and code elements.

gkeyllorg/gkeyll113—~301Automated safety check: PassMITyesterday
120

Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
121

A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels…

xuzhougeng/wisp-science1k—~3.7kAutomated safety check: PassAGPL-3.0yesterday
122

Write, review, and run high-level xTBloom Python GFN2-xTB inference with Calculator, Structure, and BatchCalculator, including single systems, repeated geometry updates, heterogeneous ragged…

jinzhezenggroup/computational-chemistry-agent-skills148—~1.3kAutomated safety check: PassLGPL-3.02 days ago
123

Install or verify the correct ONNX Runtime build (and the onnx package) for a user's accelerator backend before Quark's ONNX-to-ONNX flow.

amd/Quark182—~3.2kAutomated safety check: PassMIT13 days ago
124

Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

NVIDIA/skills3.6k1 repo~2.3kAutomated safety check: NotesApache-2.02 days ago
125
125.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.02 days ago
126

Continue KERMT pretraining on a custom SMILES corpus with a groverbase, cmim, or hybrid checkpoint.

NVIDIA/skills3.6k1 repo~4.1kAutomated safety check: PassApache-2.02 days ago
127
127.Kermt EmbedOfficial

Extract per-molecule embeddings from any encoder-bearing KERMT checkpoint.

NVIDIA/skills3.6k1 repo~1.9kAutomated safety check: PassApache-2.02 days ago
128
128.Kermt FinetuneOfficial

Finetune a pretrained KERMT encoder on a labeled CSV. An agent skill from NVIDIA/skills.

NVIDIA/skills3.6k1 repo~4.1kAutomated safety check: PassApache-2.02 days ago
129
129.Nvmolkit UsageOfficial

A skill your agent uses when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches.

NVIDIA/skills3.6k1 repo~4.8kAutomated safety check: PassApache-2.02 days ago
130

Check whether a CV-CUDA operator is READY to optimize (correctness + bench coverage + captured baseline + profiling) per .agents/guidance/OPTIMIZATIONGUIDELINES.md.

CVCUDA/CV-CUDA2.7k—~255Automated safety check: PassUnknown24 days ago
131

5-stage kernel correctness verification protocol for Triton and CUDA kernels.

ZJLi2013/awesome-kernel-skills102—~702Automated safety check: PassNo licence6 mo ago
132
132.Xmake

XMake build configuration, options, commands, and patterns for LuisaCompute.

LuisaGroup/LuisaCompute1.1k—~13kAutomated safety check: PassApache-2.0yesterday
133
133.Warp EvalOfficial

Evaluate whether an existing hot path is a credible NVIDIA Warp candidate.

NVIDIA/skills3.6k—~4.9kAutomated safety check: PassApache-2.02 days ago
134

A skill your agent uses when training, evaluating, or exporting Workflow policies with online RSL-RL or RLinf, including RL checkpoint and Workflow handoff.

NVIDIA/skills3.6k1 repo~3.8kAutomated safety check: PassApache-2.02 days ago
135

Pretrain a fresh KERMT model from scratch on a user-provided corpus.

NVIDIA/skills3.6k1 repo~2.4kAutomated safety check: PassApache-2.02 days ago
136

HeartMuLa: Suno-like song generation from lyrics + tags. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1712 repos~1.6kAutomated safety check: PassMIT3 days ago
137

Inspect a target ONNX model and prepare metadata for Quark ONNX PTQ planning.

amd/Quark182—~4.3kAutomated safety check: PassMIT13 days ago
138

A skill your agent uses when compile time or startup time is the problem in code that uses Warp: a request to improve, optimize, or cut compile times; an app that is slow to start or stalls at the…

NVIDIA/skills3.6k—~3.5kAutomated safety check: NotesApache-2.02 days ago
139
139.Jetson Video SetupOfficial

A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

NVIDIA/skills3.6k1 repo~2.4kAutomated safety check: NotesApache-2.02 days ago
140
140.Dali Dynamic ModeOfficial

DALI imperative dynamic mode (nvidia.dali.experimental.dynamic, ndd): use when working on ndd code or migrating pipelines; skip pipeline-only tasks.

NVIDIA/skills3.6k—~3.7kAutomated safety check: PassApache-2.02 days ago
141

Mandatory pre-flight compute resource check before running experiments.

OpenLAIR/dr-claw1.2k—~1.8kAutomated safety check: PassMIT23 days ago
142
142.Cudaq GuideOfficial

A skill your agent uses for CUDA-Q setup, simulation targets, QPU access, and @cudaq.kernel authoring guidance.

NVIDIA/skills3.6k—~1.3kAutomated safety check: PassApache-2.02 days ago
143
143.Cuopt DeveloperOfficial

Modify, build, test, debug, and contribute to NVIDIA cuOpt (C++/CUDA, Python, server, CI).

NVIDIA/skills3.6k—~3.2kAutomated safety check: NotesApache-2.02 days ago
144
144.Deepstream DevOfficial

NVIDIA DeepStream SDK development with Python pyservicemaker API.

NVIDIA/skills3.6k—~3.3kAutomated safety check: PassApache-2.02 days ago