Search
AI & LLM Engineering · NVIDIA AI Platform · For devops and sre engineers
- AI & LLM Engineering (remove filter)
- NVIDIA AI Platform (remove filter)
- For devops and sre engineers (remove filter)
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up. | NVIDIA/ | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 3 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 461 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 4 | This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a. | brevdev/ | 146 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 5 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 6 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | 7.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | 2 days ago |
| 8 | Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL. | areal-project/ | 5.8k | — | ~6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. | NVIDIA/ | 3.6k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 10 | A skill your agent uses for MolMIM, NVIDIA's BioNeMo NIM microservice for small-molecule latent-space generation and optimization. | NVIDIA/ | 3.6k | 1 repo | ~1.9k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 11 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | sickn33/ | 47k | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 12 | 12.Slime RL post-training for LLMs with Megatron and SGLang. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | 3 days ago |
| 13 | How to swap the DeepStream CV detection model in the VSS Alerts Blueprint verification (2dcv) mode - covers ONNX export, custom bbox parsers, compose mount gotchas, nvinfer config, runtime TRT… | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 14 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 15 | Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar. | NVIDIA/ | 3.6k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | yesterday |