Cuda Index Width
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.
$ npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install lucifer1004/VeloQ pytorch-profile-analysis --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/lucifer1004/VeloQ.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/veloq/skills/pytorch-profile-analysis .claude/skills/pytorch-profile-analysis && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pytorch-profile-analysis" agent skill from https://github.com/lucifer1004/VeloQ/tree/main/plugins/veloq/skills/pytorch-profile-analysis into .claude/skills/pytorch-profile-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytorch-profile-analysis", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/lucifer1004/VeloQ/tree/main/plugins/veloq/skills/pytorch-profile-analysisType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install lucifer1004/VeloQ pytorch-profile-analysis --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/lucifer1004/VeloQ.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/veloq/skills/pytorch-profile-analysis .agents/skills/pytorch-profile-analysis && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pytorch-profile-analysis" agent skill from https://github.com/lucifer1004/VeloQ/tree/main/plugins/veloq/skills/pytorch-profile-analysis into .agents/skills/pytorch-profile-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytorch-profile-analysis", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install lucifer1004/VeloQ pytorch-profile-analysis --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/lucifer1004/VeloQ.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/veloq/skills/pytorch-profile-analysis .cursor/skills/pytorch-profile-analysis && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pytorch-profile-analysis" agent skill from https://github.com/lucifer1004/VeloQ/tree/main/plugins/veloq/skills/pytorch-profile-analysis into .cursor/skills/pytorch-profile-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytorch-profile-analysis", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/lucifer1004/VeloQ.git --path plugins/veloq/skills/pytorch-profile-analysis--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install lucifer1004/VeloQ pytorch-profile-analysis --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/lucifer1004/VeloQ.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/veloq/skills/pytorch-profile-analysis .gemini/skills/pytorch-profile-analysis && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pytorch-profile-analysis" agent skill from https://github.com/lucifer1004/VeloQ/tree/main/plugins/veloq/skills/pytorch-profile-analysis into .gemini/skills/pytorch-profile-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytorch-profile-analysis", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install lucifer1004/VeloQ pytorch-profile-analysisInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/lucifer1004/VeloQ.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/veloq/skills/pytorch-profile-analysis .github/skills/pytorch-profile-analysis && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pytorch-profile-analysis" agent skill from https://github.com/lucifer1004/VeloQ/tree/main/plugins/veloq/skills/pytorch-profile-analysis into .github/skills/pytorch-profile-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytorch-profile-analysis", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install lucifer1004/VeloQ pytorch-profile-analysis --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/lucifer1004/VeloQ.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/veloq/skills/pytorch-profile-analysis .opencode/skills/pytorch-profile-analysis && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pytorch-profile-analysis" agent skill from https://github.com/lucifer1004/VeloQ/tree/main/plugins/veloq/skills/pytorch-profile-analysis into .opencode/skills/pytorch-profile-analysis/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pytorch-profile-analysis", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pytorch-profile-analysisAnalyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.
Pytorch Profile Analysis is an agent skill from lucifer1004/VeloQ. Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. Use for CPU/CUDA/kernel correlation, ProfilerStep/annotation slicing, memory/shape grouping, and single-trace NCCL evidence.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Deep learning and GPU and accelerator computing. It works with PyTorch and CUDA. The repository describes itself as: Agent-friendly GPU profile-query CLI. The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d69af56. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pytorch Profile Analysis loads about 1.3k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 513 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from lucifer1004/VeloQ at commit d69af56, republished under its MIT licence (© lucifer1004). 513 words, ~1,312 tokens.
.claude/skills/pytorch-profile-analysis/SKILL.md (or your agent's skills folder).Use veloq pytorch for PyTorch/Kineto Chrome traces:
veloq pytorch summary T
veloq pytorch search T --type kernel --name-regex 'nccl|gemm' --limit 20
veloq pytorch inspect T kernel:91
veloq pytorch correlate T kernel:91
veloq pytorch slices T --aggregate --group-by step
veloq pytorch collectives TThis skill requires the VeloQ CLI on PATH. If veloq is missing,
install it before analysis.
Use veloq pytorch verbs as the analysis interface. Do not query
<input>.veloq/pytorch/ sidecars, generated Parquet files, or raw
Kineto trace tables directly with DuckDB, PyArrow, pandas, or ad hoc
SQL unless the user explicitly asks for raw-trace exploration or you
are developing VeloQ itself.
veloq pytorch prep T only builds/checks sidecars. After prep, continue
with summary, search, inspect, stats, correlate, timeline,
slices, or collectives.
veloq pytorch commands accept one Chrome trace named .json
or .json.gz..pt.trace.json and
.pt.trace.json.gz; explicitly select pytorch for other JSON filenames.PyTorch row ids use <kind>:<stable_index>, where the stable index is
derived from the original traceEvents order after non-event flow markers
are skipped. Do not use Kineto Ev Idx as a stable key.
Use veloq pytorch schema <target> for the authoritative response field
inventory; do not infer the public contract from raw Kineto fields.
Common prefixes:
| Type | Row id prefix |
|---|---|
| CPU op | cpu_op:N |
| Annotation | annotation:N |
| Step | step:N |
| Runtime | runtime:N |
| Driver | driver:N |
| Kernel | kernel:N |
| Memcpy | memcpy:N |
| Memset | memset:N |
| Memory | memory:N |
| Python | python:N |
| Comm | comm:N |
Inventory first:
veloq pytorch summary TRead data.auxiliary.capabilities before choosing a path.
Find events:
veloq pytorch search T --type cpu-op --name '*aten::*' --limit 20
veloq pytorch search T --type kernel --is-comm --limit 20Drill into one event:
veloq pytorch inspect T ROW_IDInspect returns raw args, typed args, parent/children, enclosing step, and link metadata.
Answer launch-cause questions:
veloq pytorch correlate T kernel:91Read data.rows[0].events[] for the CPU op, annotation/step,
runtime/driver, and GPU activity chain.
Attribute CPU overhead to Python context when captured:
Traces exported from torch.profiler.profile(..., with_stack=True)
include python_function events. inspect returns
python_context / python_stack, and stats can group CPU work by
python-context or python-path:
veloq pytorch stats T --type cpu-op --group-by python-path,name --limit 20
veloq pytorch inspect T cpu_op:42Slice ProfilerStep/user annotation ranges:
slices --from/--to selects ranges that overlap the time window and
clips attributed GPU/comm time to that window. Slice row start_ns and
duration_ns remain the original trace range so the row id still points
to the inspectable event.
For communication questions, stay within one trace file:
veloq pytorch stats T --type comm --group-by comm-kind,rank
veloq pytorch search T --type kernel --is-comm --limit 20collectives groups single-trace communication evidence and reports
linked CPU/NCCL row ids. If a trace file contains multiple rank values,
rank-scoped commands (search, stats, timeline, slices, and
collectives) require --rank <n> or --all-ranks. inspect and
correlate operate on explicit row ids and are not rank-scope gated.
Device ids are rank-local and stream ids are device-local: filter a
stream with --rank <n> --device <id> --stream <id>, or compare lanes
with --group-by rank,device,stream. VeloQ does not compute
cross-rank skew in PyTorch v0:
veloq pytorch collectives T--type accepts cpu-op, annotation, step, runtime, driver,
kernel, memcpy, memset, memory, python, comm, or all.
comm is a communication-related set. Use --type kernel --is-comm to
focus on NCCL kernels.
PyTorch support is experimental (source.version = "v0"). Classification
is based on Kineto category/name/arg conventions and may need extension
for profiler variants not yet represented by tests. Treat documented
fields, schema targets, row ids/keys, command ids, and output modes as the
versioned source contract even while the source remains v0.
© lucifer1004, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/veloq/skills/pytorch-profile-analysis of lucifer1004/VeloQ.
Open the folder on GitHubat commit d69af56
Pytorch Profile Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pytorch Profile Analysis this skilllucifer1004/VeloQ | 128 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Graphsignalgraphsignal/graphsignal | 257 | — | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| Metal Kernelpytorch/pytorch | 104k | — | ~4.9k | Automated safety check: Pass | Custom licence | |
| At Dispatch V2intel/torch-xpu-ops | 115 | 3 repos | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Magpie Kernel Evaluatoramd/skills | 408 | — | ~2.3k | Automated safety check: Pass | MIT |
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
graphsignal/graphsignal
Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.
pytorch/pytorch
Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.
intel/torch-xpu-ops
Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.
amd/skills
Benchmarks LLM inference and drives GPU kernel optimization with Magpie.
awslabs/agent-plugins
Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…
Categories
Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. Pytorch Profile Analysis is an agent skill from lucifer1004/VeloQ.gz) files using the VeloQ CLI.
Pytorch Profile Analysis fits situations like: CPU/CUDA/kernel correlation; profilerStep/annotation slicing; memory/shape grouping; single-trace NCCL evidence.
Run `npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a claude-code`. Or copy the skill folder (plugins/veloq/skills/pytorch-profile-analysis in lucifer1004/VeloQ) into .claude/skills/pytorch-profile-analysis in your project. Claude Code loads it when a task matches its description.
Run `npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a codex`. Or copy the skill folder (plugins/veloq/skills/pytorch-profile-analysis in lucifer1004/VeloQ) into .agents/skills/pytorch-profile-analysis in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytorch-profile-analysis, .gemini/skills/pytorch-profile-analysis, .github/skills/pytorch-profile-analysis and .opencode/skills/pytorch-profile-analysis in your project.
SKILL.md names no scripts, command-line tools or credentials: Pytorch Profile Analysis is instructions for the agent only. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Pytorch Profile Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Pytorch Profile Analysis: Cuda Index Width (pytorch/pytorch, 104k stars), Graphsignal (graphsignal/graphsignal, 257 stars), Metal Kernel (pytorch/pytorch, 104k stars) and At Dispatch V2 (intel/torch-xpu-ops, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
lucifer1004 (a GitHub user) maintains it in lucifer1004/VeloQ, which has 128 GitHub stars. The repository was last updated on October 4, 2026.
Source: lucifer1004/VeloQ on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.