Search
AI & LLM Engineering · NVIDIA AI Platform
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search. | decolua/ | 31k | — | ~604 | Automated safety check: Pass | MIT | 3 days ago |
| 3 | Moves Megatron-LM CI to a newer NVIDIA PyTorch base image, updating both the GitHub and GitLab pins together and handling the CI follow-up. | NVIDIA/ | 18k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | today |
| 4 | Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others. | decolua/ | 31k | — | ~914 | Automated safety check: Pass | MIT | 3 days ago |
| 5 | Reviews recent code changes and checks if documentation needs updates. | NVlabs/ | 1.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 23 days ago |
| 6 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 7 | Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 8 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 9 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 6 days ago |
| 10 | 10.Optimize Op Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary. | CVCUDA/ | 2.7k | — | ~834 | Automated safety check: Pass | Unknown | 24 days ago |
| 11 | Add a new step under src/nemotron/steps/<category/<stepid/ — manifest (step.toml), runner glue, configs, and per-step README.md. | NVIDIA-NeMo/ | 2.1k | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 12 | Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner). | tryonlabs/ | 551 | — | ~1.1k | Automated safety check: Pass | Unknown | yesterday |
| 13 | Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis. | NVIDIA/ | 18k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | today |
| 14 | Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable. | EGalahad/ | 146 | — | ~1.1k | Automated safety check: Pass | No licence | 13 days ago |
| 15 | Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 16 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 461 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 17 | This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a. | brevdev/ | 146 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 18 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 19 | Prepare, validate, build, and use Nemotron Customizer airgap image bundles for offline clusters. | NVIDIA-NeMo/ | 2.1k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 20 | YOLO 2026 — state-of-the-art real-time object detection. An agent skill from SharpAI/DeepCamera. | SharpAI/ | 3.1k | — | ~1.5k | Automated safety check: Pass | MIT | 24 days ago |
| 21 | A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion. | sgl-project/ | 37k | 2 repos | ~5k | Automated safety check: Pass | Apache-2.0 | today |
| 22 | Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 23 | Preflight checks and diagnosis for ten known failure modes of ML training on NVIDIA DGX Spark's GB10, spanning launch errors, memory, thermals, bandwidth and precision. | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 24 | Shows how to call NEAR AI Cloud through an OpenAI-compatible API and verify that inference ran in a TEE, using attestation checks and signed chat responses. | internet-court/ | 6.6k | 1 repo | ~1.3k | Automated safety check: Pass | Unknown | 1 mo ago |
| 25 | Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands. | brevdev/ | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 26 | 26.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 188 | — | ~547 | Automated safety check: Pass | No licence | 12 days ago |
| 27 | Reference desk for Nemotron 3 Nano / Llama-Nemotron Nano 3 — architecture, training data, recipes, evaluation, quantization, deployment. | NVIDIA-NeMo/ | 2.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 28 | Verify a CV-CUDA optimization campaign's deterministic definition-of-done and concise versioned MR summary per .agents/guidance/OPTIMIZATIONGUIDELINES.md. | CVCUDA/ | 2.7k | — | ~424 | Automated safety check: Pass | Unknown | 24 days ago |
| 29 | Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck. | fla-org/ | 5.8k | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 30 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 31 | 31.Keirouter Entry point for KeiRouter — local/remote AI gateway with OpenAI-compatible REST for chat, image, TTS, embeddings, web search, web fetch. | mydisha/ | 147 | — | ~995 | Automated safety check: Pass | MIT | 1 mo ago |
| 32 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 408 | — | ~4k | Automated safety check: Notes | MIT | yesterday |
| 33 | Reference desk for NVIDIA Nemotron 3 Super — architecture, training data, recipes (pretrain/SFT/RL/eval/quantization), and deployment notes. | NVIDIA-NeMo/ | 2.1k | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 34 | 34.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | 2 days ago |
| 35 | GGUF format and llama.cpp quantization for efficient CPU/GPU inference. | Orchestra-Research/ | 13k | 3 repos | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 36 | 36.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 37 | A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is… | NVIDIA/ | 138 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 38 | 38.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 39 | 39.Make Op Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done. | CVCUDA/ | 2.7k | — | ~831 | Automated safety check: Pass | Unknown | 24 days ago |
| 40 | Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed. | scragnog/ | 174 | — | ~4.9k | Automated safety check: Pass | MIT | 2 days ago |
| 41 | Estimate GPU memory usage for Megatron-based MoE (Mixture of Experts) and dense models. | yzlnew/ | 149 | — | ~2.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 42 | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. | Orchestra-Research/ | 13k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 43 | Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 44 | Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 45 | Generate vector embeddings via KeiRouter /v1/embeddings using OpenAI / Gemini / Mistral / Voyage / Nvidia embedding models for RAG, semantic search, similarity. | mydisha/ | 147 | — | ~577 | Automated safety check: Pass | MIT | 1 mo ago |
| 46 | Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling. | NVIDIA/ | 3.6k | — | ~5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 47 | Run the Nemotron-3.5 Lightning Text2SQL LoRA fine-tuning tutorial (NeMo Megatron-Bridge) end-to-end for the user on a single node: data prep, checkpoint conversion, LoRA fine-tuning of the 30B-A3B… | NVIDIA-NeMo/ | 2.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 48 | Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download. | NVIDIA/ | 3.6k | — | ~4.7k | Automated safety check: Notes | Apache-2.0 | 2 days ago |