Agent skill

Torch Profiler Layer Track

by BBuf in BBuf/AI-Infra-Auto-Driven-SKILLS

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

No licenceAuto-check passedDevelopment

Install Torch Profiler Layer Track

skills CLI
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill torch-profiler-layer-track -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS torch-profiler-layer-track --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/torch-profiler-layer-track .claude/skills/torch-profiler-layer-track && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
torch-profiler-layer-track
GitHub stars
911
Token cost
~2k tokens
SKILL.md length
1,002 words
Files
6 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

  • Works in 4 steps: Check the report's PID, layer count,… → Verify that reversing the recorded… → Inspect the guide beside the GPU… → …
  • Navigating a Torch Profiler trace by transformer layer
  • SKILL.md covers Establish what the labels mean, Add the auxiliary track, Open and inspect and Verify and deliver
  • Runs Python scripts from its folder; calls python3

What it does

The skill writes a normal Chrome JSON trace with layer guides labeled L0, L1 and so on beside at most ten synthetic GPU activity lanes per device, and it keeps your source file. Compaction only moves events: original kernel names, timestamps, durations and `args.stream` stay intact and CPU events are untouched. The output is labeled a compact view, not a runtime stream optimization.

Layer labels must be earned from evidence. You use your existing local trace, take the layer count from the matching model config, pick the rank, phase and one complete forward pass, and choose a GPU kernel with a verified per-layer cadence, using `--anchor-stride` and `--extra-anchors-per-pass` where needed. The helper never guesses L0 from counts alone, and if evidence is too thin it reports the ambiguity. Three scripts cover adding the layer track, compacting GPU tracks and serving a trace locally.

When your agent uses it

  • Navigating a Torch Profiler trace by transformer layer
  • Reducing hundreds of CUDA Graph stream rows to a compact view
  • Adding verified L0 and L1 layer guides to an existing Chrome JSON trace

Example prompts

  • “Add layer guides to my TP-0 profiler trace and compact the GPU stream rows.”
  • “Make a compact view of this CUDA Graph trace with at most ten lanes per device.”
  • “Label the layers in the trace using a verified per-layer kernel anchor.”

Requirements

  • Python 3.10 or later with the standard library only
  • An existing Torch Profiler Chrome JSON trace

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Check the report's PID, layer count, anchor offset, phase and source hash
  2. Verify that reversing the recorded placement changes and restoring retired
  3. Inspect the guide beside the GPU kernels. If requested, capture the real
  4. Return the compact trace, report, and screenshot if requested. State the

What it can do on your machine

Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Torch Profiler Layer Track loads about 2k tokens when it runs, and up to ~4.7k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 1,002 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,002 words (~2,002 tokens).

“Create a normal Chrome JSON trace with L0, L1, ... guides beside at most ten synthetic GPU activity lanes per device. Keep the source file. Compaction changes event placement, not execution: original kernel names, timestamps, durations and args.stream remain intact…”

— opening of SKILL.md by BBuf
name
torch-profiler-layer-track

Read the full SKILL.md on GitHub

Files

SKILL.md and 5 other files (scripts, references) in skills/torch-profiler-layer-track of BBuf/AI-Infra-Auto-Driven-SKILLS.

  • SKILL.md
  • agents/openai.yaml
  • references/navigation.md
  • scripts/add_layer_track.py
  • scripts/compact_gpu_tracks.py
  • scripts/serve_local_trace.py

Open the folder on GitHubat commit 6dc9c66

Compare with similar skills

Torch Profiler Layer Track next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Torch Profiler Layer Track compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Torch Profiler Layer Track this skillBBuf/AI-Infra-Auto-Driven-SKILLS911—~2kAutomated safety check: PassNone
Cudatechnillogue/ptx-isa-markdown229—~2.5kAutomated safety check: PassNone
Cuda Profilingmohitmishra786/low-level-dev-skills253—~1.6kAutomated safety check: NotesMIT
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills398—~2.3kAutomated safety check: PassMIT
Debug Distributed Hangsgl-project/sglang37k2 repos~2.4kAutomated safety check: PassApache-2.0

Similar skills

  • Cuda

    technillogue/ptx-isa-markdown

    CUDA kernel development, debugging, and performance optimization for Claude Code.

    229 GitHub stars~2.5k tokensUpdated 9 mo ago
    DevelopmentAuto-check passed
  • Cuda Profiling

    mohitmishra786/low-level-dev-skills

    CUDA profiling skill for NVIDIA GPU performance analysis. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check: notes
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    398 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Debug Distributed Hang

    sgl-project/sglang

    Debug hanging issues in SGLang distributed inference (TP/PP/DP/EP).

    37k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed
  • Official

    Explains why editing CUTLASS fused-MHA headers in ONNX Runtime can leave stale CUDA kernels after an incremental build, and how to force and verify a real rebuild.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed

More from BBuf/AI-Infra-Auto-Driven-SKILLS

All 10 skills in this repo
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    911 GitHub stars~2.3k tokensUpdated 3 days ago
    Auto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    911 GitHub stars~2.5k tokensUpdated 3 days ago
    Auto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    911 GitHub stars~2.8k tokensUpdated 3 days ago
    Auto-check passed
  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    911 GitHub stars~1.2k tokensUpdated 3 days ago
    Auto-check passed
  • Model Compute Simulator

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

    911 GitHub stars~4.5k tokensUpdated 3 days ago
    Auto-check passed
  • SGLang Maintainer-Style Review

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.

    911 GitHub stars~4.6k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Torch Profiler Layer Track

What does Torch Profiler Layer Track do?

Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran. The skill writes a normal Chrome JSON trace with layer guides labeled L0, L1 and so on beside at most ten synthetic GPU activity lanes per device, and it keeps your source file.stream` stay intact and CPU events are untouched.

When should I use Torch Profiler Layer Track?

Torch Profiler Layer Track fits situations like: navigating a Torch Profiler trace by transformer layer; reducing hundreds of CUDA Graph stream rows to a compact view; adding verified L0 and L1 layer guides to an existing Chrome JSON trace.

How do I install Torch Profiler Layer Track in Claude Code?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill torch-profiler-layer-track -a claude-code`. Or copy the skill folder (skills/torch-profiler-layer-track in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/torch-profiler-layer-track in your project. Claude Code loads it when a task matches its description.

How do I install Torch Profiler Layer Track in Codex?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill torch-profiler-layer-track -a codex`. Or copy the skill folder (skills/torch-profiler-layer-track in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/torch-profiler-layer-track in your project. Codex loads it when a task matches its description.

Can I use Torch Profiler Layer Track in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill torch-profiler-layer-track -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/torch-profiler-layer-track, .gemini/skills/torch-profiler-layer-track, .github/skills/torch-profiler-layer-track and .opencode/skills/torch-profiler-layer-track in your project.

What does Torch Profiler Layer Track need to run?

Going by SKILL.md and its folder, Torch Profiler Layer Track needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.10 or later with the standard library only; An existing Torch Profiler Chrome JSON trace.

Does Torch Profiler Layer Track access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Torch Profiler Layer Track safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Torch Profiler Layer Track use?

No licence was found for Torch Profiler Layer Track or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Torch Profiler Layer Track use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to Torch Profiler Layer Track?

Skills that share tags, products or a category with Torch Profiler Layer Track: Cuda (technillogue/ptx-isa-markdown, 229 stars), Cuda Profiling (mohitmishra786/low-level-dev-skills, 253 stars), Graphsignal (graphsignal/graphsignal, 257 stars) and Magpie Kernel Evaluator (amd/skills, 398 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Torch Profiler Layer Track?

BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 911 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.

Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.