Search

CUDA

282 skills found, page 2.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
49

Fallback installer for milesdiffusion on a bare CUDA 12.9 Linux GPU box, reproducing the official radixark/milesdiffusion image's package versions and verifying them.

radixark/miles_diffusion110—~1.6kAutomated safety check: PassApache-2.0today
50

Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

intel/auto-round1.6k—~2.3kAutomated safety check: PassApache-2.0yesterday
51

Use only for current stock/share prices, ticker quotes, and financial market movers (gainers, losers, most-traded shares).

zhongkaifu/TensorSharp568—~1kAutomated safety check: PassBSD-3-Clausetoday
52

Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…

evo-design/proto-tools135—~2.5kAutomated safety check: NotesMITtoday
53
53.Chai

Structure prediction using Chai-1, a foundation model for molecular structure.

adaptyvbio/protein-design-skills1643 repos~1.5kAutomated safety check: PassMIT4 mo ago
54

Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
55

Train a custom VideoHighlighter action or object model from a few videos the user provides — cut into samples, sort with CLIP, review contact sheets, build, train, install only if better.

Aseiel/VideoHighlighter166—~839Automated safety check: PassAGPL-3.0today
56

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
57

GPU optimization workflow using uipc.profile, uipc.profile.nsight, and Nsight Compute CLI.

spiriMirror/libuipc336—~3.6kAutomated safety check: PassApache-2.08 days ago
58

Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

brevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.03 days ago
59

A skill your agent uses when writing, modifying, porting, or optimizing CuTe DSL GPU kernels in Python; reading CuTe DSL API reference material; integrating a CuTe DSL kernel into a project; or…

vipshop/cache-dit1.3k—~2.8kAutomated safety check: PassApache-2.0yesterday
60

基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

LMIXR/CV_Deployment_skill188—~547Automated safety check: PassNo licence12 days ago
61

Readable C/C++/Rust/CUDA code rules inspired by The Art of Readable Code.

crazyguitar/cppcheatsheet290—~6.4kAutomated safety check: PassMITyesterday
62

Step-by-step tutorial for adding a new lightweight JIT CUDA kernel to sglang's jitkernel module

guqiong96/Lsglang1441 repo~10kAutomated safety check: PassApache-2.06 days ago
63

Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage

maxiaosong1124/ncu-cuda-profiling-skill129—~1.6kAutomated safety check: PassMIT4 mo ago
64

Register, extend, or audit instructions in the table-driven T.ptx dialect (python/tvm/backend/cuda/ptx/table.py), and move the table to a newer PTX ISA version.

mlc-ai/relax175—~5.1kAutomated safety check: PassApache-2.07 days ago
65

Scaffold new Celeritas source and test files with the required copyright header and register them in CMake.

celeritas-project/celeritas105—~496Automated safety check: PassUnknowntoday
66

A skill your agent uses when starting a fresh CompileIQ project, hitting a socket timeout, or before running any other compileiq- skill.

NVIDIA/CompileIQ138—~1.3kAutomated safety check: NotesApache-2.017 days ago
67

A skill your agent uses for web searches and current information lookups, finding sources, fact-checking, researching questions, comparing sources, or summarising web pages.

zhongkaifu/TensorSharp568—~2.3kAutomated safety check: WarnBSD-3-Clausetoday
68

Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md.

CVCUDA/CV-CUDA2.7k—~424Automated safety check: PassUnknown24 days ago
69

A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment.

Red-Hat-AI-Innovation-Team/training_hub100—~959Automated safety check: PassApache-2.03 days ago
70

Integrate a new text-to-speech model into vLLM-Omni from HuggingFace reference implementation through production-ready serving with streaming and CUDA graph acceleration.

vllm-project/vllm-omni7.1k—~8.7kAutomated safety check: PassApache-2.0today
71

Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).

sohu-mptc/FlashRec107—~974Automated safety check: PassApache-2.0yesterday
72

Guided workflow for adding a new model architecture to llama.cpp.

JakeATX/llamAmpere166—~4.1kAutomated safety check: PassMITyesterday
73

Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.

slowlyC/agent-gpu-skills169—~1.8kAutomated safety check: PassMIT2 mo ago
74

A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUTLASS or CuTe C++ kernels and templates; navigating CUTLASS examples, collectives, epilogues, pipelines, GEMM…

vipshop/cache-dit1.3k—~2.2kAutomated safety check: PassApache-2.0yesterday
75

Build, simulate, and analyze quantum circuits with MindQuantum.

mindspore-ai/mindquantum102—~2.3kAutomated safety check: PassApache-2.019 days ago
76

Add, fix, or validate Triton Runner support for an exact Triton version.

toyaix/triton-runner100—~1.1kAutomated safety check: PassMIT24 days ago
77

A skill your agent uses for performance profiling and optimization.

ByteDance-Seed/VeOmni2.2k—~1.7kAutomated safety check: PassApache-2.0yesterday
78

CMake build options, custom functions, and backend patterns for LuisaCompute.

LuisaGroup/LuisaCompute1.1k—~2.4kAutomated safety check: PassApache-2.0yesterday
79

Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware.

Orchestra-Research/AI-Research-SKILLs13k3 repos~1.5kAutomated safety check: PassMIT3 mo ago
80

Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed…

pytorch/test-infra113—~2kAutomated safety check: PassUnknowntoday
81
81.Compileiq DebugOfficial

A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is…

NVIDIA/CompileIQ138—~2.9kAutomated safety check: NotesApache-2.017 days ago
82

Covers the rules for writing Gradio Spaces on ZeroGPU hardware: the @spaces.GPU decorator, duration and quota tuning, process isolation and CUDA build limits.

huggingface/skills11k2 repos~4.6kAutomated safety check: PassApache-2.02 days ago
83

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2kAutomated safety check: PassNo licence6 days ago
84

Step-by-step tutorial for adding a heavyweight AOT CUDA/C++ kernel to sgl-kernel (including tests & benchmarks)

sgl-project/sglang37k2 repos~3.4kAutomated safety check: PassApache-2.0today
85

Call this skill when you need to debug CUDA crashes in SGLang using kernel API logging

sgl-project/sglang37k2 repos~4.9kAutomated safety check: PassApache-2.0today
86

Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves…

Red-Hat-AI-Innovation-Team/training_hub100—~2.8kAutomated safety check: PassApache-2.03 days ago
87

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
88

Expert guide for Backend.AI distributed computing platform. An agent skill from lablup/backend.ai-webui.

lablup/backend.ai-webui1331 repo~1.8kAutomated safety check: PassLGPL-3.0today
89

Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done.

CVCUDA/CV-CUDA2.7k—~831Automated safety check: PassUnknown24 days ago
90

A skill your agent uses when doing operator migration or kernel migration for CUDA, Triton, or custom ops in cache-dit; porting kernels from nunchaku, deepcompressor, or other repos; designing…

vipshop/cache-dit1.3k—~3.8kAutomated safety check: PassApache-2.0yesterday
91

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

lucifer1004/VeloQ128—~1.3kAutomated safety check: PassMIT6 days ago
92
92.Cuda

Draft, debug, and measure CUDA kernels and host launch workflows.

sablin39/tilelang-cuda-skills145—~990Automated safety check: PassNo licence26 days ago
93

Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired…

NVIDIA/cosmos-framework560—~2.7kAutomated safety check: PassUnknownyesterday
94

Patterns and pitfalls for the ONNX-domain Attention operator's CUDA implementation in ONNX Runtime: dispatch cascade, eligibility limits, mask and bias kernels, and test routing.

microsoft/onnxruntime22k—~6.5kAutomated safety check: PassMITtoday
95

Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends.

ahrefs/ocannl118—~728Automated safety check: PassBSD-2-Clause4 days ago
96
96.At Dispatch V2Official

Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0today