Topic · AI & LLM Engineering
Best GPU and accelerator computing skills, page 3
GPU and accelerator computing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 97 | Diagnose and remediate per-node issues on a HyperPod cluster (EKS or Slurm) — a specific node is unhealthy, unresponsive, stuck, or needs replacing. | awslabs/ | 915 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 98 | Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. | NVIDIA/ | 3.5k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 99 | Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes. | NVIDIA/ | 3.5k | 1 repo | ~2.7k | Automated safety check: Pass | Apache-2.0 | today |
| 100 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.5k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | today |
| 101 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.5k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 102 | Remote SLURM GPU cluster execution over SSH with sbatch/srun, Pyxis/Enroot containers, and Lustre-backed results. | NVIDIA/ | 3.5k | 1 repo | ~4.8k | Automated safety check: Notes | Apache-2.0 | today |
| 103 | 103.Serverless Modal Run GPU workloads on Modal — training, fine-tuning, inference, batch processing. | AI4Scientist/ | 128 | 4 repos | ~3.1k | Automated safety check: Notes | No licence | 4 mo ago |
| 104 | 5-stage kernel correctness verification protocol for Triton and CUDA kernels. | ZJLi2013/ | 102 | — | ~702 | Automated safety check: Pass | No licence | 6 mo ago |
| 105 | Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin. | NVIDIA/ | 3.5k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | today |
| 106 | Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson. | NVIDIA/ | 3.5k | 1 repo | ~2.9k | Automated safety check: Pass | Apache-2.0 | today |
| 107 | Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output. | NVIDIA/ | 3.5k | 1 repo | ~3.1k | Automated safety check: Pass | Apache-2.0 | today |
| 108 | A skill your agent uses when you need to print Jetson device info (module model, L4T version, kernel, OS version, current power mode) from a running Jetson target. | NVIDIA/ | 3.5k | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 109 | 109.Qzcli Manage GPU compute jobs on the Qizhi (启智) platform using qzcli — a kubectl-style CLI tool. | AI4Scientist/ | 128 | 3 repos | ~1.9k | Automated safety check: Notes | No licence | 4 mo ago |
| 110 | Bumuo ng mga local-first AI agents na tumatakbo nang buong-buo sa isang developer workstation gamit ang Microsoft Foundry Local at Qwen function-calling models. | microsoft/ | 77k | — | ~1.7k | Automated safety check: Pass | MIT | 19 days ago |
| 111 | Comprehensive guide for developing WebGPU-enabled Three.js applications using TSL (Three.js Shading Language). | JetBrains/ | 364 | 2 repos | ~833 | Automated safety check: Pass | No licence | 3 mo ago |
| 112 | Add a new target-specific Mma / Copy Op type to a FlyDSL backend dialect (lib/Dialect/Fly<TARGET/<SUBTARGET/ + include/flydsl/Dialect/Fly<TARGET/IR/). | ROCm/ | 288 | — | ~6.9k | Automated safety check: Notes | Unknown | today |
| 113 | Per-pin SFIO / direction / initial-state configurator for a Jetson Orin or Thor custom carrier from the pinmux XLSM. | NVIDIA/ | 3.5k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 114 | Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. | NVIDIA/ | 3.5k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | today |
| 115 | A skill your agent uses when measuring Jetson Video Codec SDK or PyNvVideoCodec encode/decode throughput, comparing presets or surfaces, testing codec-worker capacity with authenticated samples and… | NVIDIA/ | 3.5k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 116 | A skill your agent uses when Jetson codec, profile, chroma, bit-depth, dimension, engine-count, or operational support must be reconciled from live APIs, authenticated NVIDIA samples, and NVIDIA… | NVIDIA/ | 3.5k | 1 repo | ~2k | Automated safety check: Pass | Apache-2.0 | today |
| 117 | A skill your agent uses when planning, executing, and independently validating Jetson Video Codec SDK or PyNvVideoCodec encode/decode, transcode, segmentation, container decode, AV1, or concise… | NVIDIA/ | 3.5k | 1 repo | ~2.2k | Automated safety check: Pass | Apache-2.0 | today |
| 118 | A skill your agent uses when turning a Jetson encoder use case into one surface-neutral recipe with native Video Codec SDK and PyNvVideoCodec projections. | NVIDIA/ | 3.5k | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 119 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.5k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | today |
| 120 | 120.Modal Serverless GPU cloud for ML jobs and model APIs. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | today |
| 121 | Plan and apply safe Jetson headless-mode changes to reclaim GUI and daemon memory. | NVIDIA/ | 3.5k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | today |
| 122 | Extract Jetson Linux + sample-rootfs tarballs and run applybinaries.sh for the active target, then record bspimage in the profile. | NVIDIA/ | 3.5k | 1 repo | ~1.8k | Automated safety check: Notes | Apache-2.0 | today |
| 123 | Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck. | NVIDIA/ | 3.5k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | today |
| 124 | Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint. | open-thoughts/ | 301 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 125 | Extract GPU ISA from Triton kernels on XPU. An agent skill from intel/torch-xpu-ops. | intel/ | 115 | — | ~806 | Automated safety check: Pass | Apache-2.0 | today |
| 126 | A skill your agent uses when you need to rebuild the BSP overlay — DT, OOT modules, or kernel — from changes under bspsources/. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | today |
| 127 | A skill your agent uses to lock/cap Jetson CPU/GPU/EMC clocks, toggle EMC/CPU DVFS, or change cpufreq governors by editing BPMP DTB and nvpower.sh pre-flash. | NVIDIA/ | 3.5k | — | ~4.8k | Automated safety check: Notes | Apache-2.0 | today |
| 128 | A skill your agent uses to flash a promoted BSP image to a Jetson DUT in RCM mode via flash.sh or l4tinitrdflash.sh. | NVIDIA/ | 3.5k | — | ~4.9k | Automated safety check: Notes | Apache-2.0 | today |
| 129 | A skill your agent uses to promote overlay files and built artifacts into the staged BSP image. | NVIDIA/ | 3.5k | — | ~5k | Automated safety check: Notes | Apache-2.0 | today |
| 130 | A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. | NVIDIA/ | 3.5k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 131 | Diagnose the most expensive silent failure on a CoreWeave multi-node GPU job: GPUDirect RDMA falling back from InfiniBand to TCP. | jeremylongshore/ | 2.8k | — | ~3.4k | Automated safety check: Notes | MIT | today |
| 132 | Hunt down CoreWeave GPU cost leaks — idle reserved capacity, wrong-GPU-type right-sizing waste, allocated-but-idle instances, and on-demand spend that should be committed — then produce a… | jeremylongshore/ | 2.8k | — | ~3.5k | Automated safety check: Pass | MIT | today |
| 133 | Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently… | jeremylongshore/ | 2.8k | — | ~3.2k | Automated safety check: Pass | MIT | today |
| 134 | Author Triton kernels with automatic warp specialization (AutoWS). | facebookexperimental/ | 201 | — | ~7.1k | Automated safety check: Pass | MIT | today |
| 135 | Diagnose CUDA "illegal instruction" / kernel crashes on Triton kernels that reference to TMA loads or stores (maketensordescriptor, TensorDescriptor, descriptor.load, descriptor.store… | facebookexperimental/ | 201 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 136 | A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation. | NVIDIA/ | 3.5k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | today |
| 137 | A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the… | NVIDIA/ | 3.5k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | today |
| 138 | A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests… | NVIDIA/ | 3.5k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 139 | A skill your agent uses when the user is measuring GPU-kernel-initiated RDMA WRITE latency through doca-gpunetio — building and running the gpunetioibwritelat client + server pair under… | NVIDIA/ | 3.5k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 140 | A skill your agent uses when you need to add, remove, edit, list, or change the boot default of an nvfancontrol fan profile on a Jetson/Tegra (Orin, Thor) target. | NVIDIA/ | 3.5k | — | ~4.3k | Automated safety check: Notes | Apache-2.0 | today |
| 141 | Enable Jetson Thor 25G/10G/1G MGBE QSFP via kernel-DT overlay. | NVIDIA/ | 3.5k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 142 | A skill your agent uses when you need to add, remove, edit, list, or change the boot default of an nvpmodel power mode on a Jetson/Tegra (Orin, Thor) target. | NVIDIA/ | 3.5k | — | ~4.4k | Automated safety check: Notes | Apache-2.0 | today |
| 143 | Download NVIDIA Jetson Linux BSP artifacts (BSP tarball, sample rootfs, publicsources, x-tools, guides) for the active target. | NVIDIA/ | 3.5k | — | ~3k | Automated safety check: Pass | Apache-2.0 | today |
| 144 | Reclaim DRAM by disabling unused subsystems across MB1 BCT, MB2 BCT, kernel reserved-memory, and SWIOTLB. | NVIDIA/ | 3.5k | — | ~2k | Automated safety check: Notes | Apache-2.0 | today |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23