Search
AI & LLM Engineering · PyTorch
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 145 | A skill your agent uses when writing or debugging nvMolKit Python code for GPU-accelerated RDKit fingerprints, similarity, conformers, clustering, and molecular searches. | NVIDIA/ | 3.6k | 1 repo | ~4.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 146 | 146.Model Scaffold A skill your agent uses when you need a runnable PyTorch training repo for a medical-imaging task (segmentation, classification, detection, synthesis, self-supervised, or fine-tuning a pretrained… | Aperivue/ | 333 | — | ~3.1k | Automated safety check: Pass | MIT | 6 days ago |
| 147 | 147.Runtime Skills Universal Runtime best practices for PyTorch inference, Transformers models, and FastAPI serving. | llama-farm/ | 836 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 4 mo ago |
| 148 | Review vLLM-routed torch-nightly root causes against existing upstream vLLM issues using read-only search results, and emit a validated upstream-checks artifact for the filer. | pytorch/ | 113 | — | ~1.3k | Automated safety check: Pass | Unknown | yesterday |
| 149 | Build and debug masked one-dimensional elementwise kernels for Triton-Ascend, including launch wrappers and PyTorch/NPU correctness checks. | Krusty84/ | 106 | — | ~584 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 150 | 150.Accelerate Run PyTorch training across GPUs with minimal changes. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~2.3k | Automated safety check: Pass | MIT | 3 days ago |
| 151 | 151.Flash Attention Speed up long-sequence transformer training and inference. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~2.7k | Automated safety check: Pass | MIT | 3 days ago |
| 152 | 152.Pytorch Fsdp Fully sharded data-parallel training for large models. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~619 | Automated safety check: Pass | MIT | 3 days ago |
| 153 | 153.Torchtitan Pretrain LLMs at scale with PyTorch 4D parallelism. An agent skill from Luciole-Studio/Misaka-Agent. | Luciole-Studio/ | 171 | 1 repo | ~2.6k | Automated safety check: Pass | MIT | 3 days ago |
| 154 | Convert existing Hugging Face Transformers Trainer or TRL SFTTrainer training code into an NVFLARE federated job using flare.patch(trainer), local validation, and job export; use when the user names… | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 155 | 155.CI Metrics Fallback loader for the canonical PyTorch ci-metrics skill. An agent skill from pytorch/test-infra. | pytorch/ | 113 | — | ~452 | Automated safety check: Pass | Unknown | yesterday |
| 156 | Train ML models on Databricks. An agent skill from databricks/databricks-agent-skills. | databricks/ | 345 | — | ~4.6k | Automated safety check: Pass | Unknown | yesterday |
| 157 | Official NVIDIA-authored guidance for navigating PhysicsNeMo — pick the model, datapipe, or example for a SciML/AI4Science task (surrogates, forecasting, downscaling, physics-informed, inverse… | NVIDIA/ | 3.6k | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 158 | Official NVIDIA-authored guidance for PhysicsNeMo ShardTensor domain parallelism — integrate domain parallelism into training/inference scripts (new or existing) with DDP or FSDP2, write and… | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 159 | CLIP vision-language model for image-text retrieval, zero-shot classification, embedding extraction, ONNX export, and TensorRT deployment. | NVIDIA/ | 3.6k | — | ~4k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 160 | Fine-tune any HuggingFace CV / VLM / LLM model on local NVIDIA GPUs inside an NGC PyTorch container when no dedicated TAO model skill matches. | NVIDIA/ | 3.6k | — | ~4.9k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 161 | InternVideo2-CLIP L14 (TAO videoclip) for video-text retrieval, zero-shot classification, embedding extraction, LoRA fine-tuning, ONNX export, and TensorRT deployment. | NVIDIA/ | 3.6k | — | ~3.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 162 | PyTorch-based TAO image classification. An agent skill from NVIDIA/skills. | NVIDIA/ | 3.6k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 163 | Apply existing ShapeShifter graph passes to an .onnx model via the quark-cli shapeshifter CLI or a ShapeShifter YAML. | amd/ | 182 | — | ~1.9k | Automated safety check: Pass | MIT | 13 days ago |
| 164 | Diagnose failed Quark installation, PTQ execution, script generation, or export attempts. | amd/ | 182 | — | ~1.9k | Automated safety check: Notes | MIT | 13 days ago |
| 165 | Install or verify the correct PyTorch build for a user's accelerator backend before Quark installation. | amd/ | 182 | — | ~1.6k | Automated safety check: Pass | MIT | 13 days ago |
| 166 | L3 recipe that runs a Torch LLM PTQ end-to-end for AMD Quark — for PyTorch / HuggingFace transformers models (safetensors input): quantize → validate → evaluate. | amd/ | 182 | — | ~2.6k | Automated safety check: Pass | MIT | 13 days ago |
| 167 | Route Quark user goals to the correct atomic skill or workflow. | amd/ | 182 | — | ~1.9k | Automated safety check: Pass | MIT | 13 days ago |
| 168 | Route Quark user goals to the correct atomic skill or workflow. | amd/ | 182 | — | ~1.8k | Automated safety check: Pass | MIT | 13 days ago |
| 169 | Generate C/C++ or CUDA code from an AI model (PyTorch, LiteRT) using MATLAB Coder or GPU Coder. | matlab/ | 1.1k | — | ~2.8k | Automated safety check: Pass | Unknown | 3 days ago |
| 170 | Creates MATLAB interfaces to Python image processing and computer vision models from GitHub repositories or pip-installable packages using MPyReq. | matlab/ | 1.1k | — | ~3.7k | Automated safety check: Pass | Unknown | 3 days ago |
| 171 | Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder). | matlab/ | 1.1k | — | ~4.6k | Automated safety check: Pass | Unknown | 3 days ago |
| 172 | Validate Triton-Ascend kernel outputs against PyTorch references with dtype-aware tolerances, exact integer checks, bfloat16 promotion, and boolean handling. | Krusty84/ | 106 | — | ~649 | Automated safety check: Pass | Apache-2.0 | 1 mo ago |
| 173 | This skill provides guidance for implementing PyTorch pipeline parallelism for distributed training of large language models. | lazyFrogLOL/ | 128 | — | ~2.1k | Automated safety check: Pass | No licence | 4 mo ago |
| 174 | This skill provides guidance for implementing tensor parallelism in PyTorch, specifically column-parallel and row-parallel linear layers. | lazyFrogLOL/ | 128 | — | ~2.8k | Automated safety check: Pass | No licence | 4 mo ago |
| 175 | Convert existing PyTorch Lightning training code into an NVFLARE federated job using the Lightning Client API patch, local validation, and job export; use only when the request names… | NVIDIA/ | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 176 | Convert existing plain or manual PyTorch training code into an NVFLARE federated job using Client API model exchange, local validation, and job export; use when the user names plain PyTorch or… | NVIDIA/ | 3.6k | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 177 | Extract Intel GPU ISA (assembly) from any XPU kernel. An agent skill from intel/torch-xpu-ops. | intel/ | 115 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 178 | A skill your agent uses when asked to implement a fix for a root-caused failure, apply a proposed patch, or produce a staged code change from triage output. | intel/ | 115 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 179 | A skill your agent uses when asked to verify a fix works, confirm a staged patch resolves a failure, or produce a before/after summary of a fix. | intel/ | 115 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 180 | A skill your agent uses when setting up a new torch-xpu-ops release branch corresponding to a PyTorch release. | intel/ | 115 | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 181 | Build PyTorch from source with Intel XPU (GPU) support. An agent skill from intel/torch-xpu-ops. | intel/ | 115 | — | ~939 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 182 | 182.Quark Install Installs or verifies AMD Quark and ensures the selected Python environment has an accelerator-matched PyTorch. | amd/ | 182 | — | ~3.5k | Automated safety check: Notes | MIT | 13 days ago |
| 183 | 183.Quark Torch Ptq Runs an end-to-end AMD Quark post-training quantization workflow for PyTorch / Hugging Face LLMs: inspect a Hub or local model, choose a quantization plan, create reproducible artifacts, request… | amd/ | 182 | — | ~2.3k | Automated safety check: Pass | MIT | 13 days ago |
| 184 | Fine-tune a DPA3 model in DeePMD-kit using the PyTorch backend. | jinzhezenggroup/ | 148 | — | ~3.1k | Automated safety check: Pass | LGPL-3.0-or-later | 2 days ago |
| 185 | Integrate a HuggingFace Computer Vision model into the NVIDIA TAO Toolkit ecosystem (tao-core config, tao-pytorch trainer, tao-deploy TensorRT pipeline). | NVIDIA/ | 3.6k | — | ~4.5k | Automated safety check: Notes | Apache-2.0 | 2 days ago |
| 186 | Generate and analyze T1/T2/R roofline reports for PyTorch OOB workloads comparing Intel XPU and NVIDIA CUDA. | intel/ | 115 | — | ~681 | Automated safety check: Pass | Apache-2.0 | yesterday |
| 187 | Find upstream PyTorch behavior or fixes that may require XPU parity work, validate them on XPU, and produce independently reviewed evidence. | intel/ | 115 | — | ~2k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 188 | 188.Scholar Compute Design and execute computational social science analyses across 11 modules: text-as-data/NLP (STM, BERTopic, Wordfish, BERT, conText embedding regression, LLM annotation + DSL bias correction… | joshzyj/ | 168 | — | ~15k | Automated safety check: Pass | Unknown | 23 days ago |
| 189 | PyTorch Geometric (PyG) for graph neural networks: node/graph classification, link prediction with GCN, GAT, GraphSAGE, GIN. | jaechang-hits/ | 374 | 1 repo | ~5.1k | Automated safety check: Pass | MIT | 12 days ago |
| 190 | 通过 PyTorch torch.distributed 接口测试昇腾 NPU 通信算子性能。支持指定任意 tensor shape、dtype,使用 torchrun 启动,贴近真实训练场景的通信算子测试与性能分析。Use for testing collective communication operators (AllReduce, AllGather, ReduceScatter… | ascend-ai-coding/ | 174 | — | ~2.2k | Automated safety check: Pass | No licence | yesterday |
| 191 | A skill your agent uses when asked to reproduce a bug, verify a nightly CI failure, or confirm a failure still exists on latest source. | intel/ | 115 | — | ~7.7k | Automated safety check: Pass | Apache-2.0 | yesterday |
| 192 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | 3 days ago |