Agent skill

LLM Torch Profiler Trace Analysis

by BBuf in BBuf/AI-Infra-Auto-Driven-SKILLS

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

No licenceAuto-check passedAI & LLM Engineering

Install LLM Torch Profiler Trace Analysis

skills CLI
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-torch-profiler-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-torch-profiler-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-torch-profiler-analysis .claude/skills/llm-torch-profiler-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-torch-profiler-analysis
GitHub stars
925
Token cost
~2.8k tokens
SKILL.md length
1,231 words
Files
15 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

  • Works in 6 steps: Confirm nonzero GPU events, a complete… → Read all relevant GPU streams. A CPU… → Separate sum of kernel durations, union… → …
  • Finding which kernels dominate an LLM decode or prefill trace
  • SKILL.md covers Choose the evidence, Read a trace, Capture against a running server and Interpret before optimizing, plus 1 more section
  • Runs Python and Shell scripts from its folder; calls python3

What it does

This skill turns a PyTorch profiler trace from an LLM serving framework into three tables: kernel and source attribution, overlap opportunities and fusion patterns. It supports SGLang, vLLM, TensorRT-LLM and TokenSpeed, and recognizes SGLang Omni traces. It is used for questions about model dispatch, CUDA Graph gaps, stream contention and exposed kernel tails.

There are three evidence modes. An existing Chrome JSON or gzip trace with GPU events needs no GPU or framework install, and the Python analyzers use only the standard library. A single live capture uses a shared output directory and a supported HTTP profiler, recording the launch arguments and package revisions. A mapping-plus-formal run takes call-site context from an eager trace and timing from a warmed graph trace, and eager timing is never carried over.

Guidance warns against collapsing distributed runs to rank zero when studying rank skew, and against treating source checks or historical captures as fresh GPU validation. Reference files hold heuristics, fusion and overlap catalogs and a source map, and helper scripts start profiling on SGLang or vLLM hosts.

When your agent uses it

  • Finding which kernels dominate an LLM decode or prefill trace
  • Looking for CUDA Graph gaps or exposed kernel tails
  • Checking a trace for stream contention and overlap opportunities
  • Spotting operator fusion candidates in a serving framework's trace

Example prompts

  • “Analyze this SGLang decode trace and give me the kernel attribution table.”
  • “Capture a profile from the running vLLM server and list the best fusion opportunities.”
  • “Why is there a gap between graph replays in this trace? Look for stream contention.”

Requirements

  • Python 3, standard library only for the analyzers
  • A Torch Profiler Chrome trace, or a running SGLang or vLLM server to capture from

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Confirm nonzero GPU events, a complete warmed forward, actual dispatch and
  2. Read all relevant GPU streams. A CPU cudaGraphLaunch span is host time;
  3. Separate sum of kernel durations, union of GPU busy intervals, and end-to-end
  4. Identify the producer/consumer join and exposed tail. PDL consumers may
  5. Consult source dispatch before proposing fusion. Existing routed-expert
  6. Time changes without profiling; run kernel correctness and actual-model

What it can do on your machine

Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Torch Profiler Trace Analysis loads about 2.8k tokens when it runs, and up to ~451k if it reads all its reference files. Until then it costs about 83 tokens; SKILL.md has 1,231 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~451k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,231 words (~2,828 tokens).

“Produce three tables: kernel/source attribution, overlap opportunities, and fusion patterns. Start from the user's trace or running server. Inspect the actual model, framework revision, phase, device/rank and parallelism before choosing a fast path. Existing traces need no GPU or framework…”

— opening of SKILL.md by BBuf
name
llm-torch-profiler-analysis

Read the full SKILL.md on GitHub

Files

SKILL.md and 14 other files (scripts, references) in skills/llm-torch-profiler-analysis of BBuf/AI-Infra-Auto-Driven-SKILLS.

  • SKILL.md
  • references/fuse-overlap-catalog.md
  • references/heuristics.md
  • references/overlap-catalog.md
  • references/source-map.md
  • references/vllm-torch-compile-fusions.md
  • scripts/analyze_llm_torch_profile.py
  • scripts/analyze_sglang_torch_profile.py
  • scripts/probe_llm_server.py
  • scripts/profile_common.py
  • scripts/render_triage_markdown_bundle.py
  • scripts/run_sglang_torch_profile_host.sh
  • scripts/run_vllm_torch_profile_host.sh
  • scripts/triage_kernel_helpers.py
  • scripts/triage_overlap_helpers.py

Open the folder on GitHubat commit 6dc9c66

Compare with similar skills

LLM Torch Profiler Trace Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Torch Profiler Trace Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Torch Profiler Trace Analysis this skillBBuf/AI-Infra-Auto-Driven-SKILLS925—~2.8kAutomated safety check: PassNone
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills406—~2.3kAutomated safety check: PassMIT
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0
Jetson Memory AuditNVIDIA/skills3.5k1 repos~2.3kAutomated safety check: NotesApache-2.0
Jetson PackageNVIDIA/skills3.5k1 repos~1.8kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    406 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Jetson Memory Audit

    NVIDIA/skills

    Official

    Measure Jetson DRAM/NvMap usage and verify before/after memory reclamation with live audit data.

    3.5k GitHub starsUsed in 1 repo~2.3k tokens
    AI & LLM EngineeringAuto-check: notes
  • Jetson Package

    NVIDIA/skills

    Official

    Pick Jetson-compatible containers, vLLM runtime images, and Jetson AI Lab PyPI indexes; maps Orin SM 8.7 vs Thor SM 11.0 and JetPack-specific package choices.

    3.5k GitHub starsUsed in 1 repo~1.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Pick the serving stack and per-runtime memory flags (vLLM, SGLang, llama.cpp, TensorRT Edge-LLM) for an LLM/VLM workload on any NVIDIA Jetson.

    3.5k GitHub starsUsed in 1 repo~2.9k tokens
    AI & LLM EngineeringAuto-check passed

More from BBuf/AI-Infra-Auto-Driven-SKILLS

All 10 skills in this repo
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    925 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    925 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    925 GitHub stars~1.2k tokensUpdated 4 days ago
    Auto-check passed
  • Model Compute Simulator

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

    925 GitHub stars~4.5k tokensUpdated 4 days ago
    Auto-check passed
  • SGLang Maintainer-Style Review

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.

    925 GitHub stars~4.6k tokensUpdated 4 days ago
    Auto-check passed
  • Torch Profiler Layer Track

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Adds verified layer guides such as L0 and L1 and compact GPU lanes to an existing Torch Profiler Chrome trace, changing how it looks but not how it ran.

    925 GitHub stars~2k tokensUpdated 4 days ago
    Auto-check passed

Questions about LLM Torch Profiler Trace Analysis

What does LLM Torch Profiler Trace Analysis do?

Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables. This skill turns a PyTorch profiler trace from an LLM serving framework into three tables: kernel and source attribution, overlap opportunities and fusion patterns. It supports SGLang, vLLM, TensorRT-LLM and TokenSpeed, and recognizes SGLang Omni traces.

When should I use LLM Torch Profiler Trace Analysis?

LLM Torch Profiler Trace Analysis fits situations like: finding which kernels dominate an LLM decode or prefill trace; looking for CUDA Graph gaps or exposed kernel tails; checking a trace for stream contention and overlap opportunities; spotting operator fusion candidates in a serving framework's trace.

How do I install LLM Torch Profiler Trace Analysis in Claude Code?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-torch-profiler-analysis -a claude-code`. Or copy the skill folder (skills/llm-torch-profiler-analysis in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/llm-torch-profiler-analysis in your project. Claude Code loads it when a task matches its description.

How do I install LLM Torch Profiler Trace Analysis in Codex?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-torch-profiler-analysis -a codex`. Or copy the skill folder (skills/llm-torch-profiler-analysis in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/llm-torch-profiler-analysis in your project. Codex loads it when a task matches its description.

Can I use LLM Torch Profiler Trace Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-torch-profiler-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-torch-profiler-analysis, .gemini/skills/llm-torch-profiler-analysis, .github/skills/llm-torch-profiler-analysis and .opencode/skills/llm-torch-profiler-analysis in your project.

What does LLM Torch Profiler Trace Analysis need to run?

Going by SKILL.md and its folder, LLM Torch Profiler Trace Analysis needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3, standard library only for the analyzers; A Torch Profiler Chrome trace, or a running SGLang or vLLM server to capture from.

Does LLM Torch Profiler Trace Analysis access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is LLM Torch Profiler Trace Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does LLM Torch Profiler Trace Analysis use?

No licence was found for LLM Torch Profiler Trace Analysis or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does LLM Torch Profiler Trace Analysis use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 448k tokens, read only when the agent opens those files.

What are the alternatives to LLM Torch Profiler Trace Analysis?

Skills that share tags, products or a category with LLM Torch Profiler Trace Analysis: Graphsignal (graphsignal/graphsignal, 257 stars), Magpie Kernel Evaluator (amd/skills, 406 stars), LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars) and Jetson Memory Audit (NVIDIA/skills, 3.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Torch Profiler Trace Analysis?

BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 925 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.

Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.