Topic · AI & LLM Engineering

Best GPU and accelerator computing skills, page 3

Skills #97–144 of 176, ranked by score.

GPU and accelerator computing skills, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

GPU and accelerator computing skills, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
97

Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing.

awslabs/agent-plugins915—~5.2kAutomated safety check: PassApache-2.02 days ago
98

Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink.

NVIDIA/skills3.5k1 repo~1.6kAutomated safety check: PassApache-2.0today
99

Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes.

NVIDIA/skills3.5k1 repo~2.7kAutomated safety check: PassApache-2.0today
100

Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

NVIDIA/skills3.5k1 repo~2.3kAutomated safety check: NotesApache-2.0today
101
101.Jetson PackageOfficial

Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

NVIDIA/skills3.5k1 repo~1.8kAutomated safety check: PassApache-2.0today
102
102.Tao Run On SlurmOfficial

Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results.

NVIDIA/skills3.5k1 repo~4.8kAutomated safety check: NotesApache-2.0today
103

Run GPU workloads on Modal — training, fine-tuning, inference, batch processing.

AI4Scientist/nano-scientist1284 repos~3.1kAutomated safety check: NotesNo licence4 mo ago
104

5-stage kernel correctness verification protocol for Triton and CUDA kernels.

ZJLi2013/awesome-kernel-skills102—~702Automated safety check: PassNo licence6 mo ago
105
105.Jetson LLM ServeOfficial

Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin.

NVIDIA/skills3.5k2 repos~3kAutomated safety check: NotesApache-2.0today
106

Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

NVIDIA/skills3.5k1 repo~2.9kAutomated safety check: PassApache-2.0today
107

Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output.

NVIDIA/skills3.5k1 repo~3.1kAutomated safety check: PassApache-2.0today
108

A skill your agent uses when you need to print Jetson device info (module model, L4T version, kernel, OS version, current power mode) from a running Jetson target.

NVIDIA/skills3.5k1 repo~1.3kAutomated safety check: PassApache-2.0today
109
109.Qzcli

Manage GPU compute jobs on the Qizhi (启智) platform using qzcli — a kubectl-style CLI tool.

AI4Scientist/nano-scientist1283 repos~1.9kAutomated safety check: NotesNo licence4 mo ago
110
110.Local AI AgentsOfficial

Bumuo ng mga local-first AI agents na tumatakbo nang buong-buo sa isang developer workstation gamit ang Microsoft Foundry Local at Qwen function-calling models.

microsoft/ai-agents-for-beginners77k—~1.7kAutomated safety check: PassMIT19 days ago
111
111.Webgpu Threejs TslOfficial

Comprehensive guide for developing WebGPU-enabled Three.js applications using TSL (Three.js Shading Language).

JetBrains/skills3642 repos~833Automated safety check: PassNo licence3 mo ago
112

Add a new target-specific Mma / Copy Op type to a FlyDSL backend dialect (lib/Dialect/Fly<TARGET/<SUBTARGET/ + include/flydsl/Dialect/Fly<TARGET/IR/).

ROCm/FlyDSL288—~6.9kAutomated safety check: NotesUnknowntoday
113

Per-pin SFIO / direction / initial-state configurator for a Jetson Orin or Thor custom carrier from the pinmux XLSM.

NVIDIA/skills3.5k—~2.4kAutomated safety check: PassApache-2.0today
114

Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both.

NVIDIA/skills3.5k—~3.5kAutomated safety check: NotesApache-2.0today
115

A skill your agent uses when measuring Jetson Video Codec SDK or PyNvVideoCodec encode/decode throughput, comparing presets or surfaces, testing codec-worker capacity with authenticated samples and…

NVIDIA/skills3.5k1 repo~1.6kAutomated safety check: PassApache-2.0today
116

A skill your agent uses when Jetson codec, profile, chroma, bit-depth, dimension, engine-count, or operational support must be reconciled from live APIs, authenticated NVIDIA samples, and NVIDIA…

NVIDIA/skills3.5k1 repo~2kAutomated safety check: PassApache-2.0today
117

A skill your agent uses when planning, executing, and independently validating Jetson Video Codec SDK or PyNvVideoCodec encode/decode, transcode, segmentation, container decode, AV1, or concise…

NVIDIA/skills3.5k1 repo~2.2kAutomated safety check: PassApache-2.0today
118

A skill your agent uses when turning a Jetson encoder use case into one surface-neutral recipe with native Video Codec SDK and PyNvVideoCodec projections.

NVIDIA/skills3.5k1 repo~1.3kAutomated safety check: PassApache-2.0today
119
119.Jetson Video SetupOfficial

A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame…

NVIDIA/skills3.5k1 repo~2.4kAutomated safety check: NotesApache-2.0today
120
120.Modal

Serverless GPU cloud for ML jobs and model APIs. An agent skill from Luciole-Studio/Misaka-Agent.

Luciole-Studio/Misaka-Agent1391 repo~2.3kAutomated safety check: PassMITtoday
121

Plan and apply safe Jetson headless-mode changes to reclaim GUI and daemon memory.

NVIDIA/skills3.5k1 repo~1.9kAutomated safety check: PassApache-2.0today
122
122.Jetson Init ImageOfficial

Extract Jetson Linux + sample-rootfs tarballs and run applybinaries.sh for the active target, then record bspimage in the profile.

NVIDIA/skills3.5k1 repo~1.8kAutomated safety check: NotesApache-2.0today
123

Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck.

NVIDIA/skills3.5k1 repo~1.2kAutomated safety check: PassApache-2.0today
124

Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint.

open-thoughts/OpenThoughts-Agent301—~3.8kAutomated safety check: PassApache-2.09 days ago
125
125.Extract Asm TritonOfficial

Extract GPU ISA from Triton kernels on XPU. An agent skill from intel/torch-xpu-ops.

intel/torch-xpu-ops115—~806Automated safety check: PassApache-2.0today
126

A skill your agent uses when you need to rebuild the BSP overlay — DT, OOT modules, or kernel — from changes under bspsources/.

NVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0today
127

A skill your agent uses to lock/cap Jetson CPU/GPU/EMC clocks, toggle EMC/CPU DVFS, or change cpufreq governors by editing BPMP DTB and nvpower.sh pre-flash.

NVIDIA/skills3.5k—~4.8kAutomated safety check: NotesApache-2.0today
128
128.Jetson Flash ImageOfficial

A skill your agent uses to flash a promoted BSP image to a Jetson DUT in RCM mode via flash.sh or l4tinitrdflash.sh.

NVIDIA/skills3.5k—~4.9kAutomated safety check: NotesApache-2.0today
129

A skill your agent uses to promote overlay files and built artifacts into the staged BSP image.

NVIDIA/skills3.5k—~5kAutomated safety check: NotesApache-2.0today
130

A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill.

NVIDIA/skills3.5k—~2.1kAutomated safety check: PassApache-2.0today
131

Diagnose the most expensive silent failure on a CoreWeave multi-node GPU job: GPUDirect RDMA falling back from InfiniBand to TCP.

jeremylongshore/tons-of-skills-marketplace2.8k—~3.4kAutomated safety check: NotesMITtoday
132

Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a…

jeremylongshore/tons-of-skills-marketplace2.8k—~3.5kAutomated safety check: PassMITtoday
133

Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently…

jeremylongshore/tons-of-skills-marketplace2.8k—~3.2kAutomated safety check: PassMITtoday
134
134.Autows AuthoringOfficial

Author Triton kernels with automatic warp specialization (AutoWS).

facebookexperimental/triton201—~7.1kAutomated safety check: PassMITtoday
135

Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store…

facebookexperimental/triton201—~1.1kAutomated safety check: PassMITtoday
136
136.Doca GpiOfficial

A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation.

NVIDIA/skills3.5k—~3.9kAutomated safety check: PassApache-2.0today
137
137.Doca GpunetioOfficial

A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the…

NVIDIA/skills3.5k—~3.7kAutomated safety check: PassApache-2.0today
138

A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests…

NVIDIA/skills3.5k—~4.2kAutomated safety check: PassApache-2.0today
139

A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under…

NVIDIA/skills3.5k—~3.8kAutomated safety check: PassApache-2.0today
140

A skill your agent uses when you need to add, remove, edit, list, or change the boot default of an nvfancontrol fan profile on a Jetson/Tegra (Orin, Thor) target.

NVIDIA/skills3.5k—~4.3kAutomated safety check: NotesApache-2.0today
141

Enable Jetson Thor 25G/10G/1G MGBE QSFP via kernel-DT overlay.

NVIDIA/skills3.5k—~2.1kAutomated safety check: PassApache-2.0today
142

A skill your agent uses when you need to add, remove, edit, list, or change the boot default of an nvpmodel power mode on a Jetson/Tegra (Orin, Thor) target.

NVIDIA/skills3.5k—~4.4kAutomated safety check: NotesApache-2.0today
143

Download NVIDIA Jetson Linux BSP artifacts (BSP tarball, sample rootfs, publicsources, x-tools, guides) for the active target.

NVIDIA/skills3.5k—~3kAutomated safety check: PassApache-2.0today
144

Reclaim DRAM by disabling unused subsystems across MB1 BCT, MB2 BCT, kernel reserved-memory, and SWIOTLB.

NVIDIA/skills3.5k—~2kAutomated safety check: NotesApache-2.0today