Search
NVIDIA AI Platform · GPU and accelerator computing
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 2 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 5 days ago |
| 3 | Creates one-node GitHub merge-request variants of existing two-node GB200 functional tests in Megatron-LM, adjusting parallelism settings to fit four GPUs. | NVIDIA/ | 18k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 5 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 6 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 5 days ago |
| 7 | 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 188 | — | ~547 | Automated safety check: Pass | No licence | 11 days ago |
| 8 | Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck. | fla-org/ | 5.8k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 9 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 10 | 10.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | yesterday |
| 11 | 11.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 12 | Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 13 | Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use. | davila7/ | 33k | 10 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 14 | Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 15 | Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 16 | Fine-tunes and evaluates OpenVLA-OFT and OFT+ robot policies with LoRA and continuous action heads on LIBERO simulation and ALOHA real-robot setups. | Orchestra-Research/ | 13k | — | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 17 | Optimize dense matrix multiplication (GEMM) kernels in Triton for NVIDIA and AMD GPUs. | ZJLi2013/ | 102 | — | ~1.1k | Automated safety check: Pass | No licence | 6 mo ago |
| 18 | Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)… | awslabs/ | 916 | — | ~910 | Automated safety check: Pass | Apache-2.0 | today |
| 19 | Write, optimize, and debug high-performance AI compute kernels using TileLang (a Python DSL for GPU programming). | yzlnew/ | 149 | — | ~2.4k | Automated safety check: Pass | No licence | 3 mo ago |
| 20 | GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster. | K-Dense-AI/ | 48k | 1 repo | ~3.4k | Automated safety check: Pass | MIT | 5 days ago |
| 21 | Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls. | sickn33/ | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | yesterday |
| 22 | Validate that a Dynamo deployment's NIXL/UCX/NCCL interconnect is ready for disaggregated serving over RDMA/NVLink. | NVIDIA/ | 3.6k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 23 | Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes. | NVIDIA/ | 3.6k | 1 repo | ~2.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 24 | Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data. | NVIDIA/ | 3.6k | 1 repo | ~2.3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 25 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 26 | Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson. | NVIDIA/ | 3.6k | 1 repo | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 27 | Benchmark Jetson LLM/VLM serving performance across vLLM, llama.cpp, and Ollama with structured JSON output. | NVIDIA/ | 3.6k | 1 repo | ~3.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 28 | A skill your agent uses when you need to print Jetson device info (module model, L4T version, kernel, OS version, current power mode) from a running Jetson target. | NVIDIA/ | 3.6k | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 29 | Per-pin SFIO / direction / initial-state configurator for a Jetson Orin or Thor custom carrier from the pinmux XLSM. | NVIDIA/ | 3.6k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 30 | Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 31 | A skill your agent uses when measuring Jetson Video Codec SDK or PyNvVideoCodec encode/decode throughput, comparing presets or surfaces, testing codec-worker capacity with authenticated samples and… | NVIDIA/ | 3.6k | 1 repo | ~1.6k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 32 | A skill your agent uses when Jetson codec, profile, chroma, bit-depth, dimension, engine-count, or operational support must be reconciled from live APIs, authenticated NVIDIA samples, and NVIDIA… | NVIDIA/ | 3.6k | 1 repo | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 33 | A skill your agent uses when planning, executing, and independently validating Jetson Video Codec SDK or PyNvVideoCodec encode/decode, transcode, segmentation, container decode, AV1, or concise… | NVIDIA/ | 3.6k | 1 repo | ~2.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 34 | A skill your agent uses when turning a Jetson encoder use case into one surface-neutral recipe with native Video Codec SDK and PyNvVideoCodec projections. | NVIDIA/ | 3.6k | 1 repo | ~1.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 35 | A skill your agent uses when installing, repairing, reusing, inspecting, or verifying readiness of the native NVIDIA Video Codec SDK or PyNvVideoCodec on Jetson, including the one-frame… | NVIDIA/ | 3.6k | 1 repo | ~2.4k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 36 | Plan and apply safe Jetson headless-mode changes to reclaim GUI and daemon memory. | NVIDIA/ | 3.6k | 1 repo | ~1.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 37 | Stand up vLLM or SGLang serving on Jetson, using upstream vLLM on Thor and Orin JetPack 7.2+, and NVIDIA-AI-IOT vLLM on older Orin. | NVIDIA/ | 3.6k | 1 repo | ~3k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 38 | Add EAGLE-3 or draft-model speculative decoding to a Jetson vLLM server when TPOT is the bottleneck. | NVIDIA/ | 3.6k | 1 repo | ~1.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 39 | A skill your agent uses when you need to rebuild the BSP overlay — DT, OOT modules, or kernel — from changes under bspsources/. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 40 | A skill your agent uses to lock/cap Jetson CPU/GPU/EMC clocks, toggle EMC/CPU DVFS, or change cpufreq governors by editing BPMP DTB and nvpower.sh pre-flash. | NVIDIA/ | 3.6k | — | ~4.8k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 41 | A skill your agent uses to flash a promoted BSP image to a Jetson DUT in RCM mode via flash.sh or l4tinitrdflash.sh. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 42 | A skill your agent uses to promote overlay files and built artifacts into the staged BSP image. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 43 | A skill your agent uses for NVIDIA-related requests where an NVIDIA skill might help, even if the user did not ask for a skill. | NVIDIA/ | 3.6k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 44 | Triage a dead or degraded GPU on a CoreWeave node fast — decide reschedule vs GPU-reset vs node-reboot vs RMA from an Xid code or a pasted dmesg / nvidia-smi blob, so a bad card does not silently… | jeremylongshore/ | 2.8k | — | ~3.2k | Automated safety check: Pass | MIT | today |
| 45 | A skill your agent uses for hands-on DOCA GPI programming — wiring a GPU-Packet-Initiator context so a CUDA kernel drives RDMA queues directly from GPU memory without host CPU mediation. | NVIDIA/ | 3.6k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 46 | A skill your agent uses when the user is doing hands-on DOCA GPUNetIO programming — wiring a CUDA kernel on an NVIDIA GPU to a doca-eth queue via docagpuethrxq / docagpuethtxq, standing up the… | NVIDIA/ | 3.6k | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 47 | A skill your agent uses when the user is building, running, or interpreting the doca/tools/gpunetioibwritebw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests… | NVIDIA/ | 3.6k | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 48 | A skill your agent uses when you need to add, remove, edit, list, or change the boot default of an nvfancontrol fan profile on a Jetson/Tegra (Orin, Thor) target. | NVIDIA/ | 3.6k | — | ~4.3k | Automated safety check: Notes | Apache-2.0 | yesterday |