Agent skill

LLM Pipeline Profiler Analysis

by BBuf in BBuf/AI-Infra-Auto-Driven-SKILLS

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

No licenceAuto-check passedAI & LLM Engineering

Install LLM Pipeline Profiler Analysis

skills CLI
$ npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-pipeline-analysis -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BBuf/AI-Infra-Auto-Driven-SKILLS llm-pipeline-analysis --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BBuf/AI-Infra-Auto-Driven-SKILLS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/llm-pipeline-analysis .claude/skills/llm-pipeline-analysis && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
llm-pipeline-analysis
GitHub stars
925
Token cost
~3.9k tokens
SKILL.md length
1,511 words
Files
5 (incl. scripts)
Skills in repo
10
Repo updated
First seen
Licence
None found

At a glance

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect.

  • Works in 8 steps: layer_timeline_analyzer.py — Per-layer… → layer_kernel_breakdown.py — Per-layer… → perfetto_time_mapper.py — Perfetto UI… → …
  • Finding which layers dominate a forward pass
  • SKILL.md covers Overview, When To Use It, Verify inputs from artifacts and Model Profiles, plus 7 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Four bundled Python scripts read a Chrome-trace JSON file from a profiler run, locate the anchor kernels that mark layer boundaries, and group kernels into forward passes and layers. The output is a set of timing tables that show which layers cost the most and give time ranges to jump to in the Perfetto UI.

Model profiles tell the scripts how to find boundaries and classify kernels. They are inferred from config.json or chosen with --profile, and the listed ones cover DeepSeek, GLM, Kimi, Qwen, Nemotron-H and LongCat-Flash families, among others. The skill also asks you to confirm model config, phase, rank and parallelism from the run manifest, server arguments and trace rather than assuming defaults.

When your agent uses it

  • Finding which layers dominate a forward pass
  • Comparing cold-start and steady-state forward passes
  • Locating a specific layer's time range in Perfetto
  • Picking representative layers for a deep dive

Example prompts

  • “Show per-layer timings for this torch profiler trace of DeepSeek-V4.”
  • “Map the slowest layer in this trace to a Perfetto time range.”
  • “Compare the first forward pass with the steady-state ones in this trace.”

Requirements

  • Python to run the bundled scripts
  • A torch profiler trace saved as Chrome-trace JSON

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. layer_timeline_analyzer.py — Per-layer timeline and cluster stats
  2. layer_kernel_breakdown.py — Per-layer kernel detail and compute flow
  3. perfetto_time_mapper.py — Perfetto UI time navigation
  4. Identify steady-state forward pass
  5. Per-layer breakdown on steady-state pass
  6. Compute flow for representative layer(s)
  7. Compare layer types (optional)
  8. Navigate in Perfetto UI (optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 6dc9c66. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 4 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

LLM Pipeline Profiler Analysis loads about 3.9k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 1,511 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 1,511 words (~3,870 tokens).

“Use this when a whole-trace profiler summary is too coarse. The scripts read a Chrome-trace JSON file, find layer-boundary anchor kernels, group kernels into forward passes and layers, and print timing tables you can use for Perfetto navigation or detailed…”

— opening of SKILL.md by BBuf
name
llm-pipeline-analysis

Read the full SKILL.md on GitHub

Files

SKILL.md and 4 other files (scripts) in skills/llm-pipeline-analysis of BBuf/AI-Infra-Auto-Driven-SKILLS.

  • SKILL.md
  • scripts/layer_kernel_breakdown.py
  • scripts/layer_timeline_analyzer.py
  • scripts/model_profiles.py
  • scripts/perfetto_time_mapper.py

Open the folder on GitHubat commit 6dc9c66

Compare with similar skills

LLM Pipeline Profiler Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

LLM Pipeline Profiler Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
LLM Pipeline Profiler Analysis this skillBBuf/AI-Infra-Auto-Driven-SKILLS925—~3.9kAutomated safety check: PassNone
Graphsignalgraphsignal/graphsignal257—~6.2kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills406—~2.3kAutomated safety check: PassMIT
Vllm Daily PR Issue Trackerascend-ai-coding/awesome-ascend-skills174—~731Automated safety check: PassNone
External Gitcode Ascend Vllm Ascend Deployascend-ai-coding/awesome-ascend-skills174—~1.2kAutomated safety check: PassNone
LLM Torch Profiler Analysissgl-project/sglang37k2 repos~6.4kAutomated safety check: PassApache-2.0

Similar skills

  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    406 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Vllm Daily PR Issue Tracker

    ascend-ai-coding/awesome-ascend-skills

    Track daily PRs and Issues from vllm-project/vllm and vllm-project/vllm-ascend, filter by model (DeepSeek/Qwen/GLM/MiniMax/Kimi) and tech topics (PD disaggregation, MTP, quantization, graph mode…

    174 GitHub stars~731 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • External Gitcode Ascend Vllm Ascend Deploy

    ascend-ai-coding/awesome-ascend-skills

    昇腾 NPU 平台 vLLM 大模型推理服务一键部署。触发:用户说'部署 模型名'、'NPU 部署模型'、'vllm serve'。流程:SSH检查 → NPU检查 → 配置发现(必须验证) → 用户确认 → 部署 → cron监控 → 验证。约束:(1) 配置必须从官方文档验证,禁止猜测;(2) 后台启动必须用cron监控,禁止手动轮询。支持…

    174 GitHub stars~1.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • LLM Torch Profiler Analysis

    sgl-project/sglang

    Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.

    37k GitHub starsUsed in 2 repos~6.4k tokens
    DevelopmentAuto-check passed
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from BBuf/AI-Infra-Auto-Driven-SKILLS

All 10 skills in this repo
  • SGLang Model Day-0 Support

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Plans and audits Day-0 SGLang support for a new model release: scope, architecture gaps, PR order, validation gates and sanitized public evidence.

    925 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check passed
  • LLM Serving Capacity Planner

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reads SGLang or vLLM startup logs to show where GPU memory went and estimates how many concurrent requests fit at common token lengths.

    925 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • LLM Torch Profiler Trace Analysis

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.

    925 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check passed
  • Model Architecture Diagram Finder

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Looks up public original architecture diagrams for named LLM, vision-language, MoE, diffusion and OCR models and returns the image with its source attribution.

    925 GitHub stars~1.2k tokensUpdated 4 days ago
    Auto-check passed
  • Model Compute Simulator

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Builds an operator-level compute template for an LLM and estimates FLOPs and MFU for a serving shape, with tensor shapes and parallelism what-if checks.

    925 GitHub stars~4.5k tokensUpdated 4 days ago
    Auto-check passed
  • SGLang Maintainer-Style Review

    BBuf/AI-Infra-Auto-Driven-SKILLS

    Reviews SGLang changes the way its maintainers do, drawing on a bundled corpus of public PR review threads and a flowchart of how the diff runs.

    925 GitHub stars~4.6k tokensUpdated 4 days ago
    Auto-check passed

Questions about LLM Pipeline Profiler Analysis

What does LLM Pipeline Profiler Analysis do?

Breaks LLM torch profiler traces down by forward pass, layer and kernel, with timing tables and Perfetto time ranges for the layers you want to inspect. Four bundled Python scripts read a Chrome-trace JSON file from a profiler run, locate the anchor kernels that mark layer boundaries, and group kernels into forward passes and layers. The output is a set of timing tables that show which layers cost the most and give time ranges to jump to in the Perfetto UI.

When should I use LLM Pipeline Profiler Analysis?

LLM Pipeline Profiler Analysis fits situations like: finding which layers dominate a forward pass; comparing cold-start and steady-state forward passes; locating a specific layer's time range in Perfetto; picking representative layers for a deep dive.

How do I install LLM Pipeline Profiler Analysis in Claude Code?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-pipeline-analysis -a claude-code`. Or copy the skill folder (skills/llm-pipeline-analysis in BBuf/AI-Infra-Auto-Driven-SKILLS) into .claude/skills/llm-pipeline-analysis in your project. Claude Code loads it when a task matches its description.

How do I install LLM Pipeline Profiler Analysis in Codex?

Run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-pipeline-analysis -a codex`. Or copy the skill folder (skills/llm-pipeline-analysis in BBuf/AI-Infra-Auto-Driven-SKILLS) into .agents/skills/llm-pipeline-analysis in your project. Codex loads it when a task matches its description.

Can I use LLM Pipeline Profiler Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BBuf/AI-Infra-Auto-Driven-SKILLS --skill llm-pipeline-analysis -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/llm-pipeline-analysis, .gemini/skills/llm-pipeline-analysis, .github/skills/llm-pipeline-analysis and .opencode/skills/llm-pipeline-analysis in your project.

What does LLM Pipeline Profiler Analysis need to run?

Going by SKILL.md and its folder, LLM Pipeline Profiler Analysis needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python to run the bundled scripts; A torch profiler trace saved as Chrome-trace JSON.

Does LLM Pipeline Profiler Analysis access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is LLM Pipeline Profiler Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does LLM Pipeline Profiler Analysis use?

No licence was found for LLM Pipeline Profiler Analysis or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does LLM Pipeline Profiler Analysis use?

About 3.9k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to LLM Pipeline Profiler Analysis?

Skills that share tags, products or a category with LLM Pipeline Profiler Analysis: Graphsignal (graphsignal/graphsignal, 257 stars), Magpie Kernel Evaluator (amd/skills, 406 stars), Vllm Daily PR Issue Tracker (ascend-ai-coding/awesome-ascend-skills, 174 stars) and External Gitcode Ascend Vllm Ascend Deploy (ascend-ai-coding/awesome-ascend-skills, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains LLM Pipeline Profiler Analysis?

BBuf (a GitHub user) maintains it in BBuf/AI-Infra-Auto-Driven-SKILLS, which has 925 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 5, 2026.

Source: BBuf/AI-Infra-Auto-Driven-SKILLS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.