Topic · AI & LLM Engineering
Best GPU and accelerator computing skills, page 4
GPU and accelerator computing skills, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | A skill your agent uses when you need to print Jetson BSP info (L4T version, board configs, rootfs state) from a LinuxforTegra root on the host PC. | NVIDIA/ | 3.5k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 146 | 146.Lambda Labs On-demand GPU cloud instances for ML training. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 139 | 2 repos | ~3k | Automated safety check: Warn | MIT | today |
| 147 | Enable MIPI/GMSL camera sensors on a Jetson Thor or Orin custom carrier by rendering a kernel-DT overlay from the in-tree sensor DTSI. | NVIDIA/ | 3.5k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 148 | Per-controller PCIe enable / disable / lanes / link-speed for a Jetson Thor or Orin custom carrier via ODMDATA + kernel-DT overlay. | NVIDIA/ | 3.5k | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | today |
| 149 | Configure Jetson UPHY lane allocation (uphy0/uphy1-config) on Orin/Thor custom carriers. | NVIDIA/ | 3.5k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | today |
| 150 | Enable/disable Jetson USB2/USB3 SS ports via kernel-DT overlay. | NVIDIA/ | 3.5k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | today |
| 151 | Set up the BSP source workspace: LinuxforTegra overlay tracker, bspsources, Crosstool-NG toolchain. | NVIDIA/ | 3.5k | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | today |
| 152 | Author a new Jetson target-platform profile (referencedevkit + optional customcarrier) and update the active pointer. | NVIDIA/ | 3.5k | — | ~4.8k | Automated safety check: Pass | Apache-2.0 | today |
| 153 | Convert single-node scripts to multi-node Slurm sbatch jobs and debug common multi-node failures. | NVIDIA/ | 3.5k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | today |
| 154 | 154.Runpod A skill your agent uses when running GPU compute on RunPod and deciding between Pods (hourly, always-on) and Serverless (per-second, autoscaling) for training, fine-tuning or inference — serverless… | ericrisco/ | 167 | — | ~2.8k | Automated safety check: Pass | MIT | today |
| 155 | Compare Triton, TLX, and inductor kernel numerics across A/B compiler builds with per-config isolation tests and grouped impact reporting. | facebookexperimental/ | 201 | — | ~792 | Automated safety check: Pass | MIT | today |
| 156 | 156.Modal Cloud computing platform for running Python on GPUs and serverless infrastructure. | BioTender-max/ | 197 | — | ~3.1k | Automated safety check: Notes | Apache-2.0 | 3 mo ago |
| 157 | Bootstrap a custom carrier board by forking carrier files and scaffolding a DT overlay from the reference devkit. | NVIDIA/ | 3.5k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | today |
| 158 | Build a per-target knowledge-base markdown next to the active profile by walking the BSP root and source tree. | NVIDIA/ | 3.5k | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | today |
| 159 | Entry skill for Jetson / IGX BSP customization. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.5k | — | ~4.7k | Automated safety check: Pass | Apache-2.0 | today |
| 160 | Switch the active Jetson target-platform pointer to an existing profile YAML. | NVIDIA/ | 3.5k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | today |
| 161 | Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM… | NVIDIA/ | 3.5k | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | today |
| 162 | Diagnose and fix CoreWeave GPU scheduling, pod, and networking errors. | jeremylongshore/ | 2.8k | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 163 | Run distributed GPU training jobs on CoreWeave with multi-node PyTorch. | jeremylongshore/ | 2.8k | — | ~1.2k | Automated safety check: Pass | MIT | today |
| 164 | Optimize CoreWeave GPU cloud costs with right-sizing and scheduling. | jeremylongshore/ | 2.8k | — | ~1.3k | Automated safety check: Pass | MIT | today |
| 165 | 165.GPU Optimizer GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile. | Mathews-Tom/ | 328 | — | ~3.5k | Automated safety check: Notes | MIT | 2 days ago |
| 166 | 166.Optimize For GPU GPU-accelerate Python code using CuPy, Numba CUDA, Warp, cuDF, cuML, cuGraph, KvikIO, cuCIM, cuxfilter, cuVS, cuSpatial, and RAFT. | majiayu000/ | 666 | 1 repo | ~8.5k | Automated safety check: Pass | MIT | today |
| 167 | 167.Cuda CUDA C/C++ skill for NVIDIA GPU kernel programming. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 168 | 168.Cuda Debugging CUDA debugging skill for GPU program correctness. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 169 | 169.Cuda Profiling CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |
| 170 | 170.Hip Rocm HIP and ROCm skill for AMD GPU programming. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.6k | Automated safety check: Notes | MIT | 3 mo ago |
| 171 | 171.Triton Lang Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills. | mohitmishra786/ | 253 | — | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 172 | 172.Cublas Cudnn Expert integration with NVIDIA GPU-accelerated math libraries. | majiayu000/ | 666 | 1 repo | ~2.4k | Automated safety check: Notes | MIT | today |
| 173 | 173.Cutlass Triton High-performance kernel template libraries and DSLs. An agent skill from majiayu000/claude-skill-registry. | majiayu000/ | 666 | 1 repo | ~2.6k | Automated safety check: Notes | MIT | today |
| 174 | NVIDIA Collective Communications Library integration for multi-GPU operations. | majiayu000/ | 666 | 1 repo | ~1.9k | Automated safety check: Notes | MIT | today |
| 175 | Bind pre-downloaded Jetson reference docs (developer guide, design guide, pinmux, schematics) into the active profile documents block. | NVIDIA/ | 3.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | today |
| 176 | Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar. | NVIDIA/ | 3.5k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | today |
Explore related skills
Category
More topics in AI & LLM Engineering
- Building AI agents563
- Deep learning415
- Embeddings386
- LLM inference and serving372
- Prompt engineering360
- Retrieval-augmented generation358
- Fine-tuning309
- LLM evaluation308
- Speech recognition and synthesis308
- Structured output and tool calling276
- LLM cost and token optimization259
- LLM API integration255
- Model routing and gateways255
- LLM observability240
- LLM guardrails221
- Computer vision203
- Model hubs and datasets180
- Diffusion and image models166
- Natural language processing131
- Reinforcement learning66
- AI interpretability23