Agent skill

Pytorch Profile Analysis

by lucifer1004 in lucifer1004/VeloQ

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

MITAuto-check passedAI & LLM Engineering

Install Pytorch Profile Analysis

skills CLI
$ npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lucifer1004/VeloQ pytorch-profile-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lucifer1004/VeloQ.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/veloq/skills/pytorch-profile-analysis .claude/skills/pytorch-profile-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pytorch-profile-analysis
GitHub stars
128
Token cost
~1.3k tokens
SKILL.md length
513 words
Files
1
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI.

  • Works in 7 steps: Inventory first → Find events → Drill into one event → …
  • CPU/CUDA/kernel correlation
  • SKILL.md covers Tool Boundary, Inputs, Row IDs and Workflow, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Pytorch Profile Analysis is an agent skill from lucifer1004/VeloQ. Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. Use for CPU/CUDA/kernel correlation, ProfilerStep/annotation slicing, memory/shape grouping, and single-trace NCCL evidence.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning and GPU and accelerator computing. It works with PyTorch and CUDA. The repository describes itself as: Agent-friendly GPU profile-query CLI. The licence is MIT.

When your agent uses it

  • CPU/CUDA/kernel correlation
  • ProfilerStep/annotation slicing
  • Memory/shape grouping
  • Single-trace NCCL evidence

Example prompts

  • “/pytorch-profile-analysis”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Inventory first
  2. Find events
  3. Drill into one event
  4. Answer launch-cause questions
  5. Attribute CPU overhead to Python context when captured
  6. Slice ProfilerStep/user annotation ranges
  7. For communication questions, stay within one trace file

What it can do on your machine

Read from SKILL.md and the folder at commit d69af56. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pytorch Profile Analysis loads about 1.3k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 513 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from lucifer1004/VeloQ at commit d69af56, republished under its MIT licence (© lucifer1004). 513 words, ~1,312 tokens.

Download SKILL.mdSave it as .claude/skills/pytorch-profile-analysis/SKILL.md (or your agent's skills folder).
name
pytorch-profile-analysis
description
Analyze single-file PyTorch/Kineto Chrome trace `.json(.gz)` files using the VeloQ CLI. Use for CPU/CUDA/kernel correlation, ProfilerStep/annotation slicing, memory/shape grouping, and single-trace NCCL evidence.

PyTorch Profile Analysis

Use veloq pytorch for PyTorch/Kineto Chrome traces:

bash
veloq pytorch summary T
veloq pytorch search T --type kernel --name-regex 'nccl|gemm' --limit 20
veloq pytorch inspect T kernel:91
veloq pytorch correlate T kernel:91
veloq pytorch slices T --aggregate --group-by step
veloq pytorch collectives T

This skill requires the VeloQ CLI on PATH. If veloq is missing, install it before analysis.

Tool Boundary

Use veloq pytorch verbs as the analysis interface. Do not query <input>.veloq/pytorch/ sidecars, generated Parquet files, or raw Kineto trace tables directly with DuckDB, PyArrow, pandas, or ad hoc SQL unless the user explicitly asks for raw-trace exploration or you are developing VeloQ itself.

veloq pytorch prep T only builds/checks sidecars. After prep, continue with summary, search, inspect, stats, correlate, timeline, slices, or collectives.

Inputs

  • Explicit veloq pytorch commands accept one Chrome trace named .json or .json.gz.
  • Automatic source detection only claims .pt.trace.json and .pt.trace.json.gz; explicitly select pytorch for other JSON filenames.
  • Directory inputs are not supported in PyTorch v0. Ask the user to choose one trace file if they point at a directory.

Row IDs

PyTorch row ids use <kind>:<stable_index>, where the stable index is derived from the original traceEvents order after non-event flow markers are skipped. Do not use Kineto Ev Idx as a stable key. Use veloq pytorch schema <target> for the authoritative response field inventory; do not infer the public contract from raw Kineto fields.

Common prefixes:

TypeRow id prefix
CPU opcpu_op:N
Annotationannotation:N
Stepstep:N
Runtimeruntime:N
Driverdriver:N
Kernelkernel:N
Memcpymemcpy:N
Memsetmemset:N
Memorymemory:N
Pythonpython:N
Commcomm:N

Workflow

  1. Inventory first:

    bash
    veloq pytorch summary T

    Read data.auxiliary.capabilities before choosing a path.

  2. Find events:

    bash
    veloq pytorch search T --type cpu-op --name '*aten::*' --limit 20
    veloq pytorch search T --type kernel --is-comm --limit 20
  3. Drill into one event:

    bash
    veloq pytorch inspect T ROW_ID

    Inspect returns raw args, typed args, parent/children, enclosing step, and link metadata.

  4. Answer launch-cause questions:

    bash
    veloq pytorch correlate T kernel:91

    Read data.rows[0].events[] for the CPU op, annotation/step, runtime/driver, and GPU activity chain.

  5. Attribute CPU overhead to Python context when captured:

    Traces exported from torch.profiler.profile(..., with_stack=True) include python_function events. inspect returns python_context / python_stack, and stats can group CPU work by python-context or python-path:

    bash
    veloq pytorch stats T --type cpu-op --group-by python-path,name --limit 20
    veloq pytorch inspect T cpu_op:42
  6. Slice ProfilerStep/user annotation ranges:

    slices --from/--to selects ranges that overlap the time window and clips attributed GPU/comm time to that window. Slice row start_ns and duration_ns remain the original trace range so the row id still points to the inspectable event.

  7. For communication questions, stay within one trace file:

    bash
    veloq pytorch stats T --type comm --group-by comm-kind,rank
    veloq pytorch search T --type kernel --is-comm --limit 20

    collectives groups single-trace communication evidence and reports linked CPU/NCCL row ids. If a trace file contains multiple rank values, rank-scoped commands (search, stats, timeline, slices, and collectives) require --rank <n> or --all-ranks. inspect and correlate operate on explicit row ids and are not rank-scope gated. Device ids are rank-local and stream ids are device-local: filter a stream with --rank <n> --device <id> --stream <id>, or compare lanes with --group-by rank,device,stream. VeloQ does not compute cross-rank skew in PyTorch v0:

    bash
    veloq pytorch collectives T
Show full SKILL.md (80 more words)Show less

Event Types

--type accepts cpu-op, annotation, step, runtime, driver, kernel, memcpy, memset, memory, python, comm, or all.

comm is a communication-related set. Use --type kernel --is-comm to focus on NCCL kernels.

Limits

PyTorch support is experimental (source.version = "v0"). Classification is based on Kineto category/name/arg conventions and may need extension for profiler variants not yet represented by tests. Treat documented fields, schema targets, row ids/keys, command ids, and output modes as the versioned source contract even while the source remains v0.

© lucifer1004, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/veloq/skills/pytorch-profile-analysis of lucifer1004/VeloQ.

Open the folder on GitHubat commit d69af56

Compare with similar skills

Pytorch Profile Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pytorch Profile Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pytorch Profile Analysis this skilllucifer1004/VeloQ128—~1.3kAutomated safety check: PassMIT
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence
Graphsignalgraphsignal/graphsignal257—~6.3kAutomated safety check: PassApache-2.0
Metal Kernelpytorch/pytorch104k—~4.9kAutomated safety check: PassCustom licence
At Dispatch V2intel/torch-xpu-ops1153 repos~2.2kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills408—~2.3kAutomated safety check: PassMIT

Similar skills

  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Metal Kernel

    pytorch/pytorch

    Write Metal/MPS kernels for PyTorch operators. An agent skill from pytorch/pytorch.

    104k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • At Dispatch V2

    intel/torch-xpu-ops

    Official

    Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

    115 GitHub starsUsed in 3 repos~2.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    408 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    916 GitHub stars~910 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

Works with

Questions about Pytorch Profile Analysis

What does Pytorch Profile Analysis do?

Analyze single-file PyTorch/Kineto Chrome trace .json(.gz) files using the VeloQ CLI. Pytorch Profile Analysis is an agent skill from lucifer1004/VeloQ.gz) files using the VeloQ CLI.

When should I use Pytorch Profile Analysis?

Pytorch Profile Analysis fits situations like: CPU/CUDA/kernel correlation; profilerStep/annotation slicing; memory/shape grouping; single-trace NCCL evidence.

How do I install Pytorch Profile Analysis in Claude Code?

Run `npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a claude-code`. Or copy the skill folder (plugins/veloq/skills/pytorch-profile-analysis in lucifer1004/VeloQ) into .claude/skills/pytorch-profile-analysis in your project. Claude Code loads it when a task matches its description.

How do I install Pytorch Profile Analysis in Codex?

Run `npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a codex`. Or copy the skill folder (plugins/veloq/skills/pytorch-profile-analysis in lucifer1004/VeloQ) into .agents/skills/pytorch-profile-analysis in your project. Codex loads it when a task matches its description.

Can I use Pytorch Profile Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lucifer1004/VeloQ --skill pytorch-profile-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pytorch-profile-analysis, .gemini/skills/pytorch-profile-analysis, .github/skills/pytorch-profile-analysis and .opencode/skills/pytorch-profile-analysis in your project.

What does Pytorch Profile Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: Pytorch Profile Analysis is instructions for the agent only. Our summary lists: Python 3.

Does Pytorch Profile Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Pytorch Profile Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pytorch Profile Analysis use?

Pytorch Profile Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pytorch Profile Analysis use?

About 1.3k tokens (SKILL.md is roughly 5.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Pytorch Profile Analysis?

Skills that share tags, products or a category with Pytorch Profile Analysis: Cuda Index Width (pytorch/pytorch, 104k stars), Graphsignal (graphsignal/graphsignal, 257 stars), Metal Kernel (pytorch/pytorch, 104k stars) and At Dispatch V2 (intel/torch-xpu-ops, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pytorch Profile Analysis?

lucifer1004 (a GitHub user) maintains it in lucifer1004/VeloQ, which has 128 GitHub stars. The repository was last updated on October 4, 2026.

Source: lucifer1004/VeloQ on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.