Search

PyTorch · Performance optimization

13 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Measures and shrinks the ExecuTorch runtime binary by building a size test, analyzing it with bloaty and landing each reduction as its own pull request.

pytorch/executorch5.1k—~793Automated safety check: PassUnknowntoday
2

Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

stas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.02 days ago
3

Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

open-infra-skills/infra-skills141—~1.7kAutomated safety check: PassApache-2.03 mo ago
4

Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

graphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.011 days ago
5

Inspect a PyTorch baseline with torch.compile graphs/code and torch.profiler before drafting a custom kernel, then compare TileLang and PyTorch timelines, forward/backward regions, and allocations.

sablin39/tilelang-cuda-skills145—~1.3kAutomated safety check: PassNo licence24 days ago
6

Detects CPU, GPU, memory and disk resources before heavy scientific tasks and writes a JSON file with advice on parallelism, out-of-core work and GPU use.

davila7/claude-code-templates32k10 repos~2.4kAutomated safety check: PassMITtoday
7

End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.

ascend-ai-coding/awesome-ascend-skills174—~3.5kAutomated safety check: PassNo licencetoday
8

Profile ExecuTorch model execution. Use when measuring performance, analyzing operator timing, or debugging slow models.

pytorch/executorch5.1k—~165Automated safety check: PassUnknowntoday
9

Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

amd/skills406—~2.3kAutomated safety check: PassMITtoday
10

Optimize or review eager CPU-only Albucore PyTorch runtime paths with benchmark-backed decisions.

albumentations-team/albucore123—~895Automated safety check: PassMIT4 days ago
11

Orchestrates modular PyTorch profiler trace analysis with TraceLens: generates perf reports, prepares category data, runs system-level and compute-kernel subagents in parallel, validates outputs…

amd/skills406—~760Automated safety check: PassMITtoday
12

优化实际模型推理链路,将正确性对齐、分段 profiling、显存与数据搬运、TensorRT/ONNX/PyTorch 后端、attention/kernel、FP8/compile、缓存与少步采样、质量回归、GPU 成本和服务验收串成同一实验闭环。当用户要求推理提速、降低显存或 GPU 成本、复现模型效果、定位 GPU 利用率低、优化图像/视频/扩散模型或自托管 LLM 时使用,提供…

majiayu000/spellbook287—~1.1kAutomated safety check: PassMITtoday
13

面向 Ascend PyTorch Profiler / msprof DB(如 ascendpytorchprofiler.db、msprof.db)的 SQL 分析技能。将自然语言问题(算子耗时、通信、下发、调度、schema/table 查询)转为安全可执行 SQL,并按需从官方文档提取表结构详情。

ascend-ai-coding/awesome-ascend-skills174—~1.4kAutomated safety check: PassNo licencetoday