Search
AI & LLM Engineering · NVIDIA AI Platform · For developers
- AI & LLM Engineering (remove filter)
- NVIDIA AI Platform (remove filter)
- For developers (remove filter)
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed. | sgl-project/ | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | today |
| 2 | Generates vector embeddings through the 9Router /v1/embeddings endpoint, using models from providers such as OpenAI, Gemini, Mistral and Voyage for RAG and semantic search. | decolua/ | 31k | — | ~604 | Automated safety check: Pass | MIT | 3 days ago |
| 3 | Transcribes audio files into text or subtitles through 9Router's Whisper-compatible endpoint, using models from OpenAI, Groq, Gemini, Deepgram and others. | decolua/ | 31k | — | ~914 | Automated safety check: Pass | MIT | 3 days ago |
| 4 | Reviews recent code changes and checks if documentation needs updates. | NVlabs/ | 1.3k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 23 days ago |
| 5 | Use with the dstack skill for model-serving work when the image, serving command, resources, backend/fleet choice, or service behavior is not proven. | dstackai/ | 2.3k | — | ~1.6k | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 6 | Guides moving Megatron Core GPTModel checkpoints, configs, training commands and launch scripts to HybridModel, following the repository's migration document. | NVIDIA/ | 18k | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | today |
| 7 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 8 | Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. | BBuf/ | 938 | — | ~2.8k | Automated safety check: Pass | No licence | 6 days ago |
| 9 | Drive a single-operator optimization campaign per .agents/guidance/OPTIMIZATIONGUIDELINES.md, with a deterministically enforced definition-of-done and versioned MR summary. | CVCUDA/ | 2.7k | — | ~834 | Automated safety check: Pass | Unknown | 24 days ago |
| 10 | Integrates a hosted API or local/open-weight model end-to-end across OpenTryOn (adapter, CLI registry, MCP, docs) and TryOn Studio (catalog, Connect keys, planner). | tryonlabs/ | 551 | — | ~1.1k | Automated safety check: Pass | Unknown | 2 days ago |
| 11 | Install, convert, debug, and benchmark sim2real ONNX GPU and TensorRT inference backends on onboard JetPack 5 Orin hosts such as g1-cable. | EGalahad/ | 146 | — | ~1.1k | Automated safety check: Pass | No licence | 13 days ago |
| 12 | Builds a NeMo Gym (NVIDIA) variant of an RL environment. An agent skill from adithya-s-k/FineEnvs. | adithya-s-k/ | 461 | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 13 | This skill should be used when the user wants to set up, install, deploy, bootstrap, or "spin up" the Build-an-Agent workshop (a.k.a. | brevdev/ | 146 | — | ~2.3k | Automated safety check: Notes | Apache-2.0 | 3 days ago |
| 14 | Create and manage dstack presets: a toolkit that streamlines model inference optimization with agents, and a portable preset format. | dstackai/ | 2.3k | — | ~403 | Automated safety check: Pass | MPL-2.0 | 2 days ago |
| 15 | Prepare, validate, build, and use Nemotron Customizer airgap image bundles for offline clusters. | NVIDIA-NeMo/ | 2.1k | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 16 | A skill your agent uses when quantizing a diffusion DiT with NVIDIA ModelOpt and making the resulting FP8 or NVFP4 checkpoint loadable, verifiable, and benchmarkable in SGLang Diffusion. | sgl-project/ | 37k | 2 repos | ~5k | Automated safety check: Pass | Apache-2.0 | today |
| 17 | Write, debug, and optimize CUTLASS, CuTe, and CuTeDSL GPU kernels from local upstream source, examples, and headers. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 18 | Shows how to call NEAR AI Cloud through an OpenAI-compatible API and verify that inference ran in a TEE, using attestation checks and signed chat responses. | internet-court/ | 6.6k | 1 repo | ~1.3k | Automated safety check: Pass | Unknown | 1 mo ago |
| 19 | Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands. | brevdev/ | 146 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | 3 days ago |
| 20 | 20.Cv Deploy 基于 helpfile 工程经验,协助 agent 配置 CV 主机和边缘设备环境、编译视觉与推理依赖、接入摄像头视频并打包部署服务。适用于 Ubuntu、CentOS、Windows、macOS、Jetson、树莓派和 RK3399 的 CV 工程实施与故障排查,以及相关移动端配套工具;模型训练和纯算法设计不属于本技能主线。 | LMIXR/ | 188 | — | ~547 | Automated safety check: Pass | No licence | 12 days ago |
| 21 | Port an FLA Triton kernel to Gluon when explicit layouts, asynchronous transfers, or scheduling can address a measured bottleneck. | fla-org/ | 5.8k | — | ~1.6k | Automated safety check: Pass | MIT | yesterday |
| 22 | Optimizes and serves LLMs on NVIDIA GPUs with TensorRT-LLM, covering quantization, in-flight batching, multi-GPU parallelism and the trtllm-serve command. | Orchestra-Research/ | 13k | 4 repos | ~1.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 23 | 23.Keirouter Entry point for KeiRouter — local/remote AI gateway with OpenAI-compatible REST for chat, image, TTS, embeddings, web search, web fetch. | mydisha/ | 147 | — | ~995 | Automated safety check: Pass | MIT | 1 mo ago |
| 24 | Serves AI models on AMD Instinct GPU hardware using vLLM. An agent skill from amd/skills. | amd/ | 408 | — | ~4k | Automated safety check: Notes | MIT | 2 days ago |
| 25 | 25.Dstack dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters. | dstackai/ | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | 2 days ago |
| 26 | GGUF format and llama.cpp quantization for efficient CPU/GPU inference. | Orchestra-Research/ | 13k | 3 repos | ~2.6k | Automated safety check: Pass | MIT | 3 mo ago |
| 27 | 27.Llama Cpp Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 28 | A skill your agent uses when something is wrong: Search() hangs, all evaluations return INVALIDSCORE, scores aren't improving, every config returns the same number, ptxas errors fill the log, CV% is… | NVIDIA/ | 138 | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | 18 days ago |
| 29 | 29.Triton Skill Write, debug, and optimize Triton and Gluon GPU kernels from local upstream tutorials, production kernels, language definitions, and compiler source. | slowlyC/ | 169 | — | ~1.3k | Automated safety check: Pass | MIT | 2 mo ago |
| 30 | 30.Make Op Add a new CV-CUDA operator end-to-end per .agents/guidance/MAKEOPGUIDELINES.md, with a deterministically-enforced definition-of-done. | CVCUDA/ | 2.7k | — | ~831 | Automated safety check: Pass | Unknown | 24 days ago |
| 31 | Explains where HOT-Step generation time goes (LM/DiT/VAE), how the TensorRT paths activate, how to benchmark from logs, and which knobs trade quality for speed. | scragnog/ | 174 | — | ~4.9k | Automated safety check: Pass | MIT | 2 days ago |
| 32 | Estimate GPU memory usage for Megatron-based MoE (Mixture of Experts) and dense models. | yzlnew/ | 149 | — | ~2.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 33 | Evaluates LLMs across 100+ benchmarks from 18+ harnesses (MMLU, HumanEval, GSM8K, safety, VLM) with multi-backend execution. | Orchestra-Research/ | 13k | 2 repos | ~3.1k | Automated safety check: Pass | MIT | 3 mo ago |
| 34 | Covers serving LLMs with SGLang, whose RadixAttention reuses cached prefixes, and constraining output to JSON, regex or grammar for agent and tool-calling workloads. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 35 | Sets up large-scale LLM training with NVIDIA Megatron-Core, choosing tensor, pipeline, data, context and expert parallelism for a given model size and GPU count. | Orchestra-Research/ | 13k | 2 repos | ~2.4k | Automated safety check: Pass | MIT | 3 mo ago |
| 36 | Generate vector embeddings via KeiRouter /v1/embeddings using OpenAI / Gemini / Mistral / Voyage / Nvidia embedding models for RAG, semantic search, similarity. | mydisha/ | 147 | — | ~577 | Automated safety check: Pass | MIT | 1 mo ago |
| 37 | Run the Nemotron-3.5 Lightning Text2SQL LoRA fine-tuning tutorial (NeMo Megatron-Bridge) end-to-end for the user on a single node: data prep, checkpoint conversion, LoRA fine-tuning of the 30B-A3B… | NVIDIA-NeMo/ | 2.1k | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 5 days ago |
| 38 | Scaffold a new CV-CUDA operator — a complete, wired, building skeleton — and delegate the implementation to a human or another AI. | CVCUDA/ | 2.7k | — | ~306 | Automated safety check: Pass | Unknown | 24 days ago |
| 39 | Calculate training costs for Tinker fine-tuning jobs. An agent skill from sundial-org/skills. | sundial-org/ | 153 | — | ~1.2k | Automated safety check: Pass | No licence | 2 mo ago |
| 40 | Run NVIDIA compute-sanitizer (memcheck, racecheck, initcheck, synccheck) against a Triton/TLX kernel to find runtime memory and synchronization bugs. | facebookexperimental/ | 201 | — | ~1.6k | Automated safety check: Pass | MIT | today |
| 41 | Run, monitor, stop and report Primus convergence tests -- training a model on a real corpus and checking that the loss curve is healthy -- from a plain-language request such as "run convergence test… | AMD-AGI/ | 131 | — | ~2.1k | Automated safety check: Pass | Unknown | today |
| 42 | Quick install and deploy vLLM, start serving with a simple LLM, and test OpenAI API. | vllm-project/ | 102 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 6 mo ago |
| 43 | A skill your agent uses when deploying or operating standalone RTVI-CV-3D / MV3DT multi-camera 3D tracking for calibrated MP4/file inputs and live RTSP streams: missing-calibration handoff to AMC… | NVIDIA-AI-Blueprints/ | 1.9k | — | ~5.1k | Automated safety check: Notes | Apache-2.0 | yesterday |
| 44 | Uses Meta's LlamaGuard moderation model to screen prompts and model replies against six safety categories, with vLLM, FastAPI and NeMo Guardrails setups. | Orchestra-Research/ | 13k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | 3 mo ago |
| 45 | Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate). | CVCUDA/ | 2.7k | — | ~433 | Automated safety check: Pass | Unknown | 24 days ago |
| 46 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 47 | A skill your agent uses when an SGLang, vLLM, TensorRT-LLM, or TokenSpeed serving/model optimization task needs prior model-family PR evidence. | BBuf/ | 938 | — | ~1.5k | Automated safety check: Pass | No licence | 6 days ago |
| 48 | 48.Upgrade Deps Upgrade focused runtime dependencies in AReaL. An agent skill from areal-project/AReaL. | areal-project/ | 5.8k | — | ~6k | Automated safety check: Pass | Apache-2.0 | yesterday |