Search

NVIDIA AI Platform · GPU and accelerator computing

72 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0yesterday
2

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

BBuf/AI-Infra-Auto-Driven-SKILLS938—~2.8kAutomated safety check: PassNo licence5 days ago
3

Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs.

NVIDIA/Megatron-LM18k—~1.3kAutomated safety check: PassApache-2.0today
4

Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

NVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0today
5

Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

wshobson/agents40k—~2kAutomated safety check: PassMIT5 days ago
6

Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision.

wshobson/agents40k—~2kAutomated safety check: PassMIT5 days ago
7

基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。

LMIXR/CV_Deployment_skill188—~547Automated safety check: PassNo licence11 days ago
8

Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck.

fla-org/flash-linear-attention5.8k—~1.6kAutomated safety check: PassMITtoday
9

Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command.

Orchestra-Research/AI-Research-SKILLs13k4 repos~1.3kAutomated safety check: PassMIT3 mo ago
10

dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.

dstackai/dstack2.3k—~6.2kAutomated safety check: WarnMPL-2.0yesterday
11

Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source.

slowlyC/agent-gpu-skills169—~1.3kAutomated safety check: PassMIT2 mo ago
12

Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count.

Orchestra-Research/AI-Research-SKILLs13k2 repos~2.4kAutomated safety check: PassMIT3 mo ago
13

Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use.

davila7/claude-code-templates33k10 repos~2.4kAutomated safety check: PassMITtoday
14

Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

NVIDIA/skills3.6k—~4.7kAutomated safety check: NotesApache-2.0yesterday
15

Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
16

Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups.

Orchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT3 mo ago
17

Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs.

ZJLi2013/awesome-kernel-skills102—~1.1kAutomated safety check: PassNo licence6 mo ago
18

Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

awslabs/agent-plugins916—~910Automated safety check: PassApache-2.0today
19

Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming).

yzlnew/infra-skills149—~2.4kAutomated safety check: PassNo licence3 mo ago
20

GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

K-Dense-AI/scientific-agent-skills48k1 repo~3.4kAutomated safety check: PassMIT5 days ago
21

Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.

sickn33/agentic-awesome-skills47k2 repos~3.2kAutomated safety check: PassMITyesterday
22

Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink.

NVIDIA/skills3.6k1 repo~1.6kAutomated safety check: PassApache-2.0yesterday
23

Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes.

NVIDIA/skills3.6k1 repo~2.7kAutomated safety check: PassApache-2.0yesterday
24

Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

NVIDIA/skills3.6k1 repo~2.3kAutomated safety check: NotesApache-2.0yesterday
25
25.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.6k1 repo~1.8kAutomated safety check: PassApache-2.0yesterday
26

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.6k1 repo~2.9kAutomated safety check: PassApache-2.0yesterday
27

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.6k1 repo~3.1kAutomated safety check: PassApache-2.0yesterday
28

A skill your agent uses when you need to print Jetson device info (module model, L4T version, kernel, OS version, current power mode) from a running Jetson target.

NVIDIA/skills3.6k1 repo~1.3kAutomated safety check: PassApache-2.0yesterday
29

Per-pin SFIO / direction / initial-state configurator for a Jetson Orin or Thor custom carrier from the pinmux XLSM.

NVIDIA/skills3.6k—~2.4kAutomated safety check: PassApache-2.0yesterday
30

Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both.

NVIDIA/skills3.6k—~3.5kAutomated safety check: NotesApache-2.0yesterday
31

A skill your agent uses when measuring Jetson Video Codec SDK or PyNvVideoCodec encode/decode throughput, comparing presets or surfaces, testing codec-worker capacity with authenticated samples and…

NVIDIA/skills3.6k1 repo~1.6kAutomated safety check: PassApache-2.0yesterday
32

A skill your agent uses when Jetson codec, profile, chroma, bit-depth, dimension, engine-count, or operational support must be reconciled from live APIs, authenticated NVIDIA samples, and NVIDIA…

NVIDIA/skills3.6k1 repo~2kAutomated safety check: PassApache-2.0yesterday
33

A skill your agent uses when planning, executing, and independently validating Jetson Video Codec SDK or PyNvVideoCodec encode/decode, transcode, segmentation, container decode, AV1, or concise…

NVIDIA/skills3.6k1 repo~2.2kAutomated safety check: PassApache-2.0yesterday
34

A skill your agent uses when turning a Jetson encoder use case into one surface-neutral recipe with native Video Codec SDK and PyNvVideoCodec projections.

NVIDIA/skills3.6k1 repo~1.3kAutomated safety check: PassApache-2.0yesterday
35

A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

NVIDIA/skills3.6k1 repo~2.4kAutomated safety check: NotesApache-2.0yesterday
36

Plan and apply safe Jetson headless-mode changes to reclaim GUI and daemon memory.

NVIDIA/skills3.6k1 repo~1.9kAutomated safety check: PassApache-2.0yesterday
37
37.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.6k1 repo~3kAutomated safety check: NotesApache-2.0yesterday
38

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

NVIDIA/skills3.6k1 repo~1.2kAutomated safety check: PassApache-2.0yesterday
39

A skill your agent uses when you need to rebuild the BSP overlay — DT, OOT modules, or kernel — from changes under bspsources/.

NVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.0yesterday
40

A skill your agent uses to lock/cap Jetson CPU/GPU/EMC clocks, toggle EMC/CPU DVFS, or change cpufreq governors by editing BPMP DTB and nvpower.sh pre-flash.

NVIDIA/skills3.6k—~4.8kAutomated safety check: NotesApache-2.0yesterday
41

A skill your agent uses to flash a promoted BSP image to a Jetson DUT in RCM mode via flash.sh or l4tinitrdflash.sh.

NVIDIA/skills3.6k—~4.9kAutomated safety check: NotesApache-2.0yesterday
42

A skill your agent uses to promote overlay files and built artifacts into the staged BSP image.

NVIDIA/skills3.6k—~5kAutomated safety check: NotesApache-2.0yesterday
43

A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill.

NVIDIA/skills3.6k—~2.1kAutomated safety check: PassApache-2.0yesterday
44

Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…

jeremylongshore/tons-of-skills-marketplace2.8k—~3.2kAutomated safety check: PassMITtoday
45
45.Doca GpiOfficial

A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation.

NVIDIA/skills3.6k—~3.9kAutomated safety check: PassApache-2.0yesterday
46
46.Doca GpunetioOfficial

A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the…

NVIDIA/skills3.6k—~3.7kAutomated safety check: PassApache-2.0yesterday
47

A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests…

NVIDIA/skills3.6k—~4.2kAutomated safety check: PassApache-2.0yesterday
48

A skill your agent uses when you need to add, remove, edit, list, or change the boot default of an nvfancontrol fan profile on a Jetson/Tegra (Orin, Thor) target.

NVIDIA/skills3.6k—~4.3kAutomated safety check: NotesApache-2.0yesterday