Search
PyTorch · Performance optimization
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Measures and shrinks the ExecuTorch runtime binary by building a size test, analyzing it with bloaty and landing each reduction as its own pull request. | pytorch/ | 5.1k | — | ~793 | Automated safety check: Pass | Unknown | today |
| 2 | Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness. | stas00/ | 1.7k | — | ~6.1k | Automated safety check: Notes | CC-BY-SA-4.0 | 2 days ago |
| 3 | Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged. | open-infra-skills/ | 141 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 3 mo ago |
| 4 | Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint. | graphsignal/ | 257 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 11 days ago |
| 5 | Inspect a PyTorch baseline with torch.compile graphs/code and torch.profiler before drafting a custom kernel, then compare TileLang and PyTorch timelines, forward/backward regions, and allocations. | sablin39/ | 145 | — | ~1.3k | Automated safety check: Pass | No licence | 24 days ago |
| 6 | Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use. | davila7/ | 32k | 10 repos | ~2.4k | Automated safety check: Pass | MIT | today |
| 7 | 7.Ascendc End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project. | ascend-ai-coding/ | 174 | — | ~3.5k | Automated safety check: Pass | No licence | today |
| 8 | Profile ExecuTorch model execution. Use when measuring performance, analyzing operator timing, or debugging slow models. | pytorch/ | 5.1k | — | ~165 | Automated safety check: Pass | Unknown | today |
| 9 | Benchmarks LLM inference and drives GPU kernel optimization with Magpie. | amd/ | 406 | — | ~2.3k | Automated safety check: Pass | MIT | today |
| 10 | Optimize or review eager CPU-only Albucore PyTorch runtime paths with benchmark-backed decisions. | albumentations-team/ | 123 | — | ~895 | Automated safety check: Pass | MIT | 4 days ago |
| 11 | Orchestrates modular PyTorch profiler trace analysis with TraceLens: generates perf reports, prepares category data, runs system-level and compute-kernel subagents in parallel, validates outputs… | amd/ | 406 | — | ~760 | Automated safety check: Pass | MIT | today |
| 12 | 优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供… | majiayu000/ | 287 | — | ~1.1k | Automated safety check: Pass | MIT | today |
| 13 | 面向 Ascend PyTorch Profiler / msprof DB(如 ascendpytorchprofiler.db、msprof.db)的 SQL 分析技能。将自然语言问题(算子耗时、通信、下发、调度、schema/table 查询)转为安全可执行 SQL,并按需从官方文档提取表结构详情。 | ascend-ai-coding/ | 174 | — | ~1.4k | Automated safety check: Pass | No licence | today |