Search

AI & LLM Engineering · CUDA · For data scientists

47 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

JimLiu/science-skills2284 repos~2.5kAutomated safety check: PassApache-2.03 mo ago
2

Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

stas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.05 days ago
3

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
4

A skill your agent uses when running, debugging, interpreting, or documenting mesh-llm benchmark tune model-serving throughput trials, including choosing…

Mesh-LLM/mesh-llm3.5k—~1.6kAutomated safety check: PassApache-2.0today
5

Iteratively optimize a CUDA/CUTLASS/Triton kernel only when strict on-device compilation, correctness, timing, and NCU evidence gates pass.

KernelFlow-ops/cuda-optimized-skill214—~4.3kAutomated safety check: PassMIT1 mo ago
6

Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

huggingface/skills11k3 repos~945Automated safety check: PassApache-2.03 days ago
7

Develop, optimize, debug, and validate an AReno CUDA, Triton, fused, attention, convolution, routing, or MoE operator.

inclusionAI/AReno323—~498Automated safety check: PassApache-2.0yesterday
8

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
9

Profile CUDA kernels with Nsight Compute on B200 / sm100. An agent skill from mit-han-lab/ncu-report-skill.

mit-han-lab/ncu-report-skill251—~2kAutomated safety check: PassMIT1 mo ago
10

Add a new hardware inference backend to AutoRound for deploying quantized models (e.g., CUDA/Marlin, Triton, CPU, HPU, ARK).

intel/auto-round1.6k—~2.3kAutomated safety check: PassApache-2.0yesterday
11

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k—~2kAutomated safety check: PassMIT6 days ago
12

Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

brevdev/workshop-build-an-agent146—~5.2kAutomated safety check: PassApache-2.03 days ago
13

Automated NCU (Nsight Compute) profiling workflow with full metrics collection and persistent storage

maxiaosong1124/ncu-cuda-profiling-skill129—~1.6kAutomated safety check: PassMIT4 mo ago
14

Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md.

CVCUDA/CV-CUDA2.7k—~424Automated safety check: PassUnknown24 days ago
15

A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment.

Red-Hat-AI-Innovation-Team/training_hub100—~959Automated safety check: PassApache-2.04 days ago
16

Deploy and serve GenRec checkpoints with FlashRec (install, serve.sh, SID trie, wide-beam knobs, health/curl, FP8, profiling).

sohu-mptc/FlashRec107—~974Automated safety check: PassApache-2.0yesterday
17

Add, fix, or validate Triton Runner support for an exact Triton version.

toyaix/triton-runner100—~1.1kAutomated safety check: PassMIT24 days ago
18

Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves…

Red-Hat-AI-Innovation-Team/training_hub100—~2.8kAutomated safety check: PassApache-2.04 days ago
19

Touch-lists for common OCANNL extension tasks: adding a primitive operation, adding or extending a backend, extending shape inference, and diagnosing output differences between backends.

ahrefs/ocannl118—~728Automated safety check: PassBSD-2-Clause4 days ago
20

Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…

alibaba/atrex-kernel-agent168—~1.3kAutomated safety check: PassApache-2.0yesterday
21

Guide to using Mamba selective state-space models for linear-time sequence modeling, from the Mamba block and pretrained checkpoints to Mamba-2 and speed comparisons.

Orchestra-Research/AI-Research-SKILLs13k2 repos~1.8kAutomated safety check: PassMIT3 mo ago
22

Guide to renting GPUs on Lambda Labs for ML training and inference: on-demand instances, 1-Click Clusters, SSH access, persistent filesystems and alternatives.

Orchestra-Research/AI-Research-SKILLs13k4 repos~3kAutomated safety check: WarnMIT3 mo ago
23

Manage tensor lifetimes across CUDA streams and decide whether recordstream is avoidable.

pytorch/pytorch104k—~796Automated safety check: PassUnknowntoday
24

Add or debug an AReno model family, including config conversion, module construction, checkpoint load/save, text or multimodal inference, training backward, tensor parallelism, CUDA graph decode…

inclusionAI/AReno323—~722Automated safety check: PassApache-2.0yesterday
25

Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).

CVCUDA/CV-CUDA2.7k—~433Automated safety check: PassUnknown24 days ago
26

Review a CV-CUDA operator's BENCHMARK coverage — drivers, layout axis, baselines, the basic-tier floor, row counts, and coverage statistics.

CVCUDA/CV-CUDA2.7k—~274Automated safety check: PassUnknown24 days ago
27

pyqula's CPU/GPU switch (src/pyqula/gpu.py), how a routine is routed onto the device, per-call precision, and the tiered porting plan in documentation/gpuportingplan.md.

joselado/pyqula145—~679Automated safety check: PassGPL-3.04 days ago
28

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

K-Dense-AI/scientific-agent-skills48k1 repo~3.4kAutomated safety check: PassMIT6 days ago
29

Turns a parquet of image file paths into a parquet of embeddings with CLIP, SigLIP or a TAO checkpoint, using the TAO Data Services container, ahead of neighbor mining.

NVIDIA/skills3.6k—~2kAutomated safety check: NotesApache-2.02 days ago
30

How to build and run GPU targets under Buck in fbcode. An agent skill from facebookexperimental/triton.

facebookexperimental/triton201—~998Automated safety check: PassMITtoday
31

A skill your agent uses when the user needs to choose the tile, dtype, ctagroup::1/2, SMEM operand layout, or TMEM accumulator mapping for a tcgen05 MMA on B200/Blackwell, or to implement an…

mirage-project/mirage2.5k—~1.8kAutomated safety check: PassApache-2.03 days ago
32

Use FP16/BF16 mixed precision to accelerate training and reduce memory.

aiming-lab/AutoResearchClaw15k—~275Automated safety check: PassMIT1 mo ago
33

5-stage kernel correctness verification protocol for Triton and CUDA kernels.

ZJLi2013/awesome-kernel-skills102—~702Automated safety check: PassNo licence6 mo ago
34

A skill your agent uses when training, evaluating, or exporting Workflow policies with online RSL-RL or RLinf, including RL checkpoint and Workflow handoff.

NVIDIA/skills3.6k1 repo~3.8kAutomated safety check: PassApache-2.02 days ago
35

Build deterministic forecast scripts with Earth2Studio (model, data source, IO, inference).

NVIDIA/skills3.6k—~1.4kAutomated safety check: PassApache-2.02 days ago
36

NV-Tesseract AD Diffusion — diffusion-based anomaly detection and fine-tuning for multivariate time series.

NVIDIA/skills3.6k—~2.9kAutomated safety check: NotesApache-2.02 days ago
37

NV-Tesseract Forecasting — transformer-based multivariate time series forecasting with DARR (context-enhanced kNN retrieval), interpretability, and fine-tuning.

NVIDIA/skills3.6k—~3.2kAutomated safety check: NotesApache-2.02 days ago
38

Build a Quark ONNX PTQ quantization plan from modelanalysis.json and user intent.

amd/Quark182—~4.8kAutomated safety check: PassMIT13 days ago
39

Guide installing Earth2Studio via uv or pip, selecting model extras, and configuring the environment.

NVIDIA/skills3.6k—~1.6kAutomated safety check: NotesApache-2.02 days ago
40
40.Nv Segment CtmrOfficial

Used for running NV-Segment-CTMR on CT or MRI NIfTI volumes and recording label-map evidence.

NVIDIA/skills3.6k—~2.3kAutomated safety check: NotesApache-2.02 days ago
41

Generate inorganic material structures using MatterGen, a diffusion-based generative model.

learningmatter-mit/AtomisticSkills176—~1.8kAutomated safety check: PassMIT3 days ago
42

GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

Mathews-Tom/armory329—~3.5kAutomated safety check: NotesMIT5 days ago
43

Use this repo skill for Make-It-3D single-image 3D creation, including CUDA asset setup, alpha-image validation, coarse NeRF optimization, refinement, rendering, export, and troubleshooting.

VectorSpaceLab/AREX-Skill331—~1.4kAutomated safety check: PassApache-2.01 mo ago
44

Use this repo skill for Microsoft Swin-Transformer image-classification model, config, data, checkpoint, SimMIM, Swin-MoE, and optional CUDA acceleration workflows.

VectorSpaceLab/AREX-Skill331—~1.2kAutomated safety check: PassMIT1 mo ago
45

Generate novel crystal structures and molecules using ADiT (All-atom Diffusion Transformer), a unified latent diffusion model.

learningmatter-mit/AtomisticSkills176—~1.4kAutomated safety check: PassMIT3 days ago
46

Boltz2 蛋白质结构预测模型的昇腾 NPU 迁移与复现 Skill,适用于在华为 Ascend 910、910B、910C 上准备权重、适配 Lightning 和 CUDA only kernel、完成 Boltz2 端到端结构预测推理,并沉淀可复现的环境与验证命令。

ascend-ai-coding/awesome-ascend-skills174—~2.8kAutomated safety check: PassNo licenceyesterday
47
47.Uma

Run structure relaxation and phonon calculations using Meta's UMA (Universal Materials Accelerator) via fairchem

lamm-mit/scienceclaw246—~3.5kAutomated safety check: PassApache-2.01 mo ago