Search
PyTorch · LLM inference and serving
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 2 | Runs LLM inference on CPU, Apple Silicon, and consumer GPUs without NVIDIA hardware. | Orchestra-Research/ | 13k | 3 repos | ~1.5k | Automated safety check: Pass | MIT | 3 mo ago |
| 3 | QAI ModelBuilder. An agent skill from qualcomm/qai-appbuilder. | qualcomm/ | 247 | — | ~4.1k | Automated safety check: Pass | BSD-3-Clause | yesterday |
| 4 | Run, resume, monitor, diagnose, and report Quark Quant-Perf workflows for PyTorch and HuggingFace transformers models. | amd/ | 182 | — | ~3k | Automated safety check: Pass | MIT | 13 days ago |
| 5 | Half-Quadratic Quantization for LLMs without calibration data. | Orchestra-Research/ | 13k | 2 repos | ~2.9k | Automated safety check: Pass | MIT | 3 mo ago |
| 6 | Explains RWKV, a hybrid that trains in parallel like a GPT and runs inference like an RNN with constant memory per token, plus usage, fine-tuning and troubleshooting. | Orchestra-Research/ | 13k | 2 repos | ~1.8k | Automated safety check: Pass | MIT | 3 mo ago |
| 7 | Author a new ShapeShifter graph-transformation pass for AMD Quark (ONNX or PyTorch) so it conforms to the pass framework's conventions and auto-registers. | amd/ | 182 | — | ~2.9k | Automated safety check: Pass | MIT | 13 days ago |
| 8 | Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning. | amd/ | 182 | — | ~1.4k | Automated safety check: Pass | MIT | 13 days ago |
| 9 | Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks. | davila7/ | 33k | 9 repos | ~2.3k | Automated safety check: Pass | MIT | today |
| 10 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 408 | — | ~2.3k | Automated safety check: Pass | MIT | 2 days ago |
| 11 | Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13). | wshobson/ | 40k | — | ~2k | Automated safety check: Pass | MIT | 6 days ago |
| 12 | Root-cause a vLLM torch-nightly CI regression report. An agent skill from pytorch/test-infra. | pytorch/ | 113 | — | ~2.4k | Automated safety check: Pass | Unknown | yesterday |
| 13 | Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices. | NVIDIA/ | 3.6k | 1 repo | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 14 | Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer. | pytorch/ | 113 | — | ~1.3k | Automated safety check: Pass | Unknown | yesterday |
| 15 | Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills. | databricks/ | 345 | — | ~4.6k | Automated safety check: Pass | Unknown | yesterday |
| 16 | CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment. | NVIDIA/ | 3.6k | — | ~4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 17 | InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 18 | PyTorch-based TAO image classification. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 19 | Apply existing ShapeShifter graph passes to an .onnx model via the quark-cli shapeshifter CLI or a ShapeShifter YAML. | amd/ | 182 | — | ~1.9k | Automated safety check: Pass | MIT | 13 days ago |
| 20 | Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. | amd/ | 182 | — | ~1.9k | Automated safety check: Notes | MIT | 13 days ago |
| 21 | L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate. | amd/ | 182 | — | ~2.6k | Automated safety check: Pass | MIT | 13 days ago |
| 22 | Route Quark user goals to the correct atomic skill or workflow. | amd/ | 182 | — | ~1.9k | Automated safety check: Pass | MIT | 13 days ago |
| 23 | Route Quark user goals to the correct atomic skill or workflow. | amd/ | 182 | — | ~1.8k | Automated safety check: Pass | MIT | 13 days ago |
| 24 | Installs or verifies AMD Quark and ensures the selected Python environment has an accelerator-matched PyTorch. | amd/ | 182 | — | ~3.5k | Automated safety check: Notes | MIT | 13 days ago |
| 25 | Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request… | amd/ | 182 | — | ~2.3k | Automated safety check: Pass | MIT | 13 days ago |
| 26 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 27 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | 3 days ago |
| 28 | Use this sub-skill for Torch-TensorRT model compilation, dynamic input planning, torch.export workflows, save/load formats, raw TensorRT engines, and compile-time troubleshooting. | VectorSpaceLab/ | 331 | — | ~1.3k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 29 | A skill your agent uses for Torch-TensorRT tasks: compiling PyTorch models with TensorRT, dynamic-shape/export workflows, runtime optimization, Triton/C++/distributed deployment, debugging… | VectorSpaceLab/ | 331 | — | ~1.5k | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |
| 30 | A skill your agent uses when planning or auditing TorchVision reference training/evaluation workflows for classification, quantization, detection, segmentation, video classification, optical flow… | VectorSpaceLab/ | 331 | — | ~684 | Automated safety check: Pass | BSD-3-Clause | 1 mo ago |