LLM Torch Profiler Analysis
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.
$ npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install alibaba/atrex-kernel-agent ppu-acu-joint-profile --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ppu-acu-joint-profile .claude/skills/ppu-acu-joint-profile && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ppu-acu-joint-profile" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/ppu-acu-joint-profile into .claude/skills/ppu-acu-joint-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ppu-acu-joint-profile", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/ppu-acu-joint-profileType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install alibaba/atrex-kernel-agent ppu-acu-joint-profile --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ppu-acu-joint-profile .agents/skills/ppu-acu-joint-profile && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ppu-acu-joint-profile" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/ppu-acu-joint-profile into .agents/skills/ppu-acu-joint-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ppu-acu-joint-profile", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install alibaba/atrex-kernel-agent ppu-acu-joint-profile --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ppu-acu-joint-profile .cursor/skills/ppu-acu-joint-profile && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ppu-acu-joint-profile" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/ppu-acu-joint-profile into .cursor/skills/ppu-acu-joint-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ppu-acu-joint-profile", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/alibaba/atrex-kernel-agent.git --path skills/ppu-acu-joint-profile--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install alibaba/atrex-kernel-agent ppu-acu-joint-profile --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ppu-acu-joint-profile .gemini/skills/ppu-acu-joint-profile && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ppu-acu-joint-profile" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/ppu-acu-joint-profile into .gemini/skills/ppu-acu-joint-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ppu-acu-joint-profile", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install alibaba/atrex-kernel-agent ppu-acu-joint-profileInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ppu-acu-joint-profile .github/skills/ppu-acu-joint-profile && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ppu-acu-joint-profile" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/ppu-acu-joint-profile into .github/skills/ppu-acu-joint-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ppu-acu-joint-profile", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install alibaba/atrex-kernel-agent ppu-acu-joint-profile --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ppu-acu-joint-profile .opencode/skills/ppu-acu-joint-profile && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ppu-acu-joint-profile" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/ppu-acu-joint-profile into .opencode/skills/ppu-acu-joint-profile/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ppu-acu-joint-profile", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ppu-acu-joint-profileChoose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.
Ppu Acu Joint Profile is an agent skill from alibaba/atrex-kernel-agent. Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel. Use for device-level bottleneck diagnosis, kernel-internal critical-path questions, or evidence that genuinely needs both; do not require all three modes.
Its SKILL.md is about 5.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including scripts and reference files (for example `backends/ppu_backend/adapter.py`, `references/acu_collection.md` and `references/ppu0015_bottlenecks.md`).
It sits in Development, covering Performance optimization. It works with NVIDIA AI Platform. The repository describes itself as: An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a… The licence is Apache-2.0.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 602bc38. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ppu Acu Joint Profile loads about 5.2k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 2,153 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from alibaba/atrex-kernel-agent at commit 602bc38, republished under its Apache-2.0 licence (© alibaba). 2,153 words, ~5,189 tokens.
.claude/skills/ppu-acu-joint-profile/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.Use this only after a correct runnable kernel and representative workload exist. Follow the NVIDIA profiling pattern: start from the missing fact, choose the least intrusive evidence route that can answer it, and escalate only when the current route leaves a concrete ambiguity. ACU, timeline, and joint analysis are three selectable modes, not mandatory stages of one pipeline.
Start and finish every optimization iteration with the probe-free kernel. Before invoking any profiler, name the unresolved performance question and how its answer could change the next edit. Skip profiling when source inspection, compiler output, the clean benchmark, or still-valid evidence from an earlier iteration already separates the plausible bottlenecks.
Do not repeat ACU or timeline merely because an earlier iteration used it. Reuse accepted evidence while the relevant kernel specialization, launch topology, workload, device, and control or pipeline structure remain comparable. Collect new evidence only when a change invalidates the evidence needed for the current decision, or when clean results expose a new ambiguity. Timeline is an escalation for a kernel-internal timing question, never a required per-round validation step.
Across long-horizon episodes, read reusable evidence from
memory/vN.json.profile_evidence.accepted_ppu_diagnostics; the raw episode archive is not a
prerequisite. Compare every recorded identity and invalidation_conditions entry with the current clean
kernel before reuse. A canonical evidence reference identifies the prior conclusion, but does not
make it valid after specialization, workload, device, topology, or pipeline changes.
| Route | Choose it when | What it can establish |
|---|---|---|
| ACU only | The bottleneck or expensive kernel is not yet localized at device level. This is the default first profiler for an unknown bottleneck. | Kernel duration, launch resources, occupancy, and device-wide compute, memory, cache, and tail behavior. |
| Timeline only | A specific kernel-internal ordering or dependency question already exists, whether or not ACU was run first. | Ordering and intervals within explicitly selected writers, such as issue, wait, consume, MMA, and epilogue boundaries. |
| Optional joint | Both accepted evidence sets exist and the remaining question depends on their relationship. | Possible and guaranteed overlap under a bounded owner-origin uncertainty; never direct ownership of a device-global metric. |
If the probe-free benchmark already answers the question, do not profile. After each selected route,
stop when its evidence answers the question. Do not collect timeline merely because ACU ran, collect
ACU merely because timeline ran, or invoke merge.py merely because both artifacts exist.
Keep mode-specific attempts separate, for example under <PROFILE_DIR>/acu/attempt-N,
<PROFILE_DIR>/timeline/attempt-N, and <PROFILE_DIR>/joint/attempt-N. Never combine events from
different launches into one apparent execution.
In a long-horizon episode, add accepted_ppu_diagnostics to the terminal journal outcome only for
ACU, timeline, joint, comparison, or envelope conclusions that still apply to the terminal probe-free kernel. Omit an
invalidated intermediate capture. Each row records the question and finding, the exact comparison
identity, how it affected the optimization decision, and the conditions that require collection of
new evidence:
{
"accepted_ppu_diagnostics": [
{
"route": "timeline",
"question": "Does the tensor wait serialize the steady-state load pipeline?",
"kernel_specialization": "target kernel specialization and compile-time parameters",
"workload_identity": "representative shape, dtype, layout, and cache policy",
"device_identity": "physical device and PPU runtime architecture",
"launch_topology": "grid, block, selected writer roles, and relevant occupancy facts",
"control_pipeline_identity": "mainloop stages, waits, barriers, and epilogue structure",
"finding": "owner-local ranges show the wait on the measured critical path",
"decision_impact": "next edit targets the load/tensor handoff instead of the epilogue",
"evidence": {
"artifact": "profiles/episode_N/timeline/attempt-N/fine.timeline.receipt.json",
"sha256": "lowercase SHA-256 of that exact JSON artifact",
"schema": "ppu-fixed-slot-receipt/v5",
"evidence_id": "evidence_id read from the artifact"
},
"invalidation_conditions": [
"a change to the measured specialization or workload",
"a change to launch topology or mainloop synchronization"
]
}
]
}Allowed terminal schemas:
| Route | Accepted schemas |
|---|---|
acu | ppu-acu-extraction/v4 |
timeline | ppu-fixed-slot-receipt/v5, ppu-critical-path-report/v3 |
joint | ppu-joint-profile/v4 |
comparison | ppu-acu-comparison/v1 |
envelope | ppu-envelope-measurement/v1 |
The supervisor resolves each workspace-relative artifact, recomputes its outer and transitive hashes,
and runs profile_report.py validate before accepting an accepted, decision-grade artifact whose
schema, authoritative kernel, binding payload, and evidence id match the row. It writes stable
source_memory_version, source_episode, memory_ref, and hash-bound evidence_ref fields into
canonical memory. Diagnostic-grade or warning artifacts may guide the current investigation but
must not enter terminal-reusable memory. An empty or omitted list is valid when profiling was
skipped or all collected evidence was invalidated.
For any route, set PPU_PROFILE_SKILL to the directory containing this SKILL.md.
When a ppu0015 question depends on Tensor Cell, AIU, shared-memory, occupancy, or timer semantics,
read references/ppu0015_bottlenecks.md. It is an interpretation
aid, not a reason to profile an otherwise understood kernel.
Read references/acu_collection.md and collect the smallest useful metric set on one exact probe-free target launch. Export the raw page and PM windows without creating any timeline manifest or instrumented source. Do not default to FP8: select Tensor metrics only when they match the kernel's actual dtype, and use no Tensor metric when it is irrelevant. Treat the reference's verified metric list as selectable examples rather than one required bundle:
Use the collection reference for the exact export commands and the single-device GPM retry.
Read the compact validation and metric summaries in profile.extract.json, then inspect the relevant
original rows in profile.raw.csv and profile.samples.csv; the summary never replaces those raw
artifacts. accepted means the extractor found no data-quality issue, not that ACU proves a root
cause. A warning does not trigger a retry or timeline automatically. Use the ACU report
independently to classify device-level compute, memory, cache, occupancy, or tail evidence, and stop
when it selects or rejects the optimization hypothesis. Do not infer which source interval owns a
device-global metric.
Use only accepted decision-grade ACU receipts with matching workload, physical device, runtime, cache, and clock identities. The comparison recomputes launch and PM deltas from both hash-bound artifact graphs:
python "$PPU_PROFILE_SKILL/scripts/profile_report.py" compare-acu \
--incumbent incumbent.extract.json --candidate candidate.extract.json \
--output candidate-vs-incumbent.jsonBuild each hardware bound from a measured calibration rather than a datasheet claim. First seal a
ppu-calibration-spec/v1 containing the device/runtime/cache/clock identity, named measurements, and
at least one raw JSON benchmark artifact. Each measurement declares kind, unit, direction,
source_artifact, and json_pointer; the sealing command reads the value from that hash-bound source
rather than trusting a copied scalar:
python "$PPU_PROFILE_SKILL/scripts/profile_report.py" seal-calibration \
--spec calibration.spec.json --output calibration.receipt.json
python "$PPU_PROFILE_SKILL/scripts/profile_report.py" envelope \
--current candidate.extract.json --calibration calibration.receipt.json \
--kind compute \
--current-pointer '/metric_summaries/packet_0:cu__inst_executed.avg.pct_of_peak_sustained_elapsed/time_weighted_mean' \
--bound-name compute_peak --output compute.envelope.jsonThe envelope output uses ppu-envelope-measurement/v1, binds the current kernel and both source
artifacts, and reports measured headroom. Create separate bounds for compute, bandwidth,
launch/merge, and random gather when applicable; a missing direction remains unknown rather than
being treated as zero headroom.
Choose this route only for a falsifiable kernel-internal timing question. ACU is not a prerequisite when the question is already precise. The agent owns the hypothesis, coarse/fine transition, selected blocks, writer threads or roles, owner count, sites, density, and stop condition.
One owner is one declared writer in one launch. Construct one recorder for that owner and reuse it; do not let several threads race for the same owner id. Every owner has an independent local origin, so compare events within an owner and never order starts from different owners.
Read references/recorder.md, then choose the smallest topology that can separate the current alternatives. Examples are choices, not defaults:
Set capture_mode: "coarse", record the selection in sampling_rationale, and enumerate the exact
(owner, block, thread) writers. Compile the temporary snapshot with PPU_TIMELINE_ENABLED:
#include "ppu_timeline.cuh"
namespace ptl = ppu_acu_profile::timeline;
const bool selected_writer = /* exact block/thread/role predicate */;
const unsigned owner = /* dense id for that selected writer */;
ptl::Recorder trace(params.timeline_buffer, owner, selected_writer);
trace.range_begin(10); // one semantic coarse phase
// Existing kernel work; control flow and synchronization stay unchanged.
trace.range_end(10);When an immediate record write would perturb the region being measured, capture only the timer at the real boundaries and flush the pair later, in chronological order, after the sensitive work:
const auto phase_begin = trace.timestamp();
// Existing sensitive work.
const auto phase_end = trace.timestamp();
// Flush outside the sensitive region; do not insert synchronization to move this flush.
trace.range_at(10, phase_begin, phase_end);Use immediate writes when their cost does not change the conclusion. Use deferred writes when a density check shows local interval distortion or when the hypothesis concerns a short critical region. A deferred timestamp reduces the boundary operation to the timer read; it does not make the capture probe-free, and records must still be flushed in owner-local timestamp order.
Do not add barriers, waits, atomics, predicates, or control-flow changes to simplify the trace. Map each timestamp to a real semantic boundary. An asynchronous issue marker is not completion; observe the original wait, barrier, or first dependent consume when completion matters.
Use the target project's real PPU compiler, runtime, architecture flags, and launch path; do not replace them with a standalone CUDA-SDK executable or a synthetic launch route.
For a remote attempt, read references/remote-capture.md. Upload this skill as an explicit sandbox input, keep clean/instrumented snapshots under one attempt directory, decode before the remote job exits, and synchronize only that attempt's evidence.
Read references/recorder.md for the timer contract,
optional sanity experiment with a declared error bound, and correctness artifact requirements.
The decoder rejects timer_tick_ns, timer_calibration, invalid source/unit declarations, and
correctness evidence that does not pass for the exact kernel, workload, and device.
Initialize the ABI buffer on the host, copy it to the device, launch once, synchronize, and copy the entire allocation back. Emit manifest v5 and the event dictionary from actual launch and source facts, then decode:
python "$PPU_PROFILE_SKILL/scripts/timeline.py" decode \
--raw coarse.timeline.bin \
--manifest coarse.timeline.manifest.json \
--event-dictionary coarse.timeline.events.json \
--output-prefix coarse.timelineUse only an accepted receipt. A diagnostic receipt is intentionally incomplete and may answer a local exploratory question; a joint merge or terminal-reusable conclusion requires a decision-grade receipt with source, compiled-binary, and workload-input bindings. If the coarse trace answers the question, stop without fine probes, ACU, or merge. Read references/timeline_contract.md when interpreting decoder outputs or preparing a fine capture for optional joint analysis.
Create a new reversible fine snapshot only when coarse evidence leaves a narrower question. Retain
the minimum context and add only the issue/wait/consume or subphase boundaries needed inside that
region. Set capture_mode: "fine". A fine timeline remains valid standalone evidence and does not
require ACU or merge.py.
Declare analysis.owner, analysis.window_site_id, and analysis.site_ids only when preparing a
fine capture for optional joint analysis or when those labels help the standalone interpretation.
The decoder rejects declared analysis ranges outside the window. Analyze another owner in another
attempt instead of pretending unsynchronized owner clocks share an axis.
Preserve A (clean), B (minimal useful probes), and, only when density sensitivity matters, C (denser
nearby probes). Each timing command must warm up, iterate, synchronize, validate the representative
output, and emit one __PPU_TIMELINE_SAMPLE__=... JSON line:
python "$PPU_PROFILE_SKILL/scripts/timeline.py" measure \
--baseline-command '["python","run_a.py"]' \
--instrumented-command '["python","run_b.py"]' \
--workload-identity case-id --warmup 10 --iterations 100 \
--max-relative-change 0.03 \
--output fine.perturbation-a-b.json
python "$PPU_PROFILE_SKILL/scripts/timeline.py" measure \
--baseline-command '["python","run_b.py"]' \
--instrumented-command '["python","run_c.py"]' \
--workload-identity case-id --warmup 10 --iterations 100 \
--max-relative-change 0.03 \
--output fine.perturbation-b-c.jsonEvery emitted sample includes positive finite latency_ms, correctness: "passed",
synchronized: true, the exact workload/warmup/iteration, physical-device and runtime identity,
kernel_sha256, cache policy, clock configuration, and a stable allocation_identity such as the
authorized Pod UID plus allocation/job id. Measurement v2 also records each command argv hash. The
helper rejects any kernel, runtime, cache, clock, device, allocation, or command drift across the
interleaved schedule.
A/B measures end-to-end probe overhead; B/C measures the end-to-end effect of probe density.
Declare --max-relative-change before the experiment (0.03 above is an example, not a hardware
constant). The helper records the bound and rejects an absolute median latency change above it;
joint validation recomputes that gate. Compare common owner-local intervals separately when the
claim depends on their stability, using a critical-path plan with stability.material_relative_spread.
Reduce or move sites when either declared gate fails.
If the question is whether selected subranges explain an enclosing range, or whether the slowest owner is stable across captures, declare the semantic parent and components in a plan and analyze the accepted canonical captures:
python "$PPU_PROFILE_SKILL/scripts/critical_path.py" \
--plan critical-path.plan.json \
--capture attempt-1/fine.timeline.canonical.json attempt-1/fine.timeline.receipt.json \
--capture attempt-2/fine.timeline.canonical.json attempt-2/fine.timeline.receipt.json \
--output critical-path.report.jsonThe plan, not the tool, chooses phases, owner-topology comparison, and any material-spread threshold. The analyzer verifies each canonical hash against its receipt and rejects grid, block, capture-mode, identity, or selected-site semantic drift. It computes component interval union rather than double-counting overlap and leaves uncovered time unattributed. Escalate representative owner to more warps or blocks only when the observed spread or remaining topology ambiguity can change the optimization decision. See references/timeline_contract.md for the plan contract.
Choose this only after ACU and fine timeline were each collected and interpreted independently, and the unresolved question requires their relationship. The fine capture must declare one analysis owner, one enclosing window, and a non-empty site list. Verify kernel, grid, block, workload, device, and duration agreement before interpreting the output (warning above 3%, rejection above 5%; warning evidence cannot enter reusable memory).
python "$PPU_PROFILE_SKILL/scripts/merge.py" \
--timeline fine.timeline.perfetto.json \
--timeline-receipt fine.timeline.receipt.json \
--pm-csv profile.samples.csv \
--acu-raw-csv profile.raw.csv \
--acu-metadata profile.extract.json \
--perturbation fine.perturbation-a-b.json \
--density-sensitivity fine.perturbation-b-c.json \
--output-prefix fine.jointThe perturbation and density-sensitivity inputs are optional; pass those flags only when the corresponding validated artifacts exist. Joint merge is fail-closed: it requires decision-grade timeline and ACU bindings, compares kernel specialization, workload, physical device, runtime, cache policy, clock configuration, grid, block, and duration, and has no duration-mismatch override.
Interpret the result as four separate claims:
The merger never aligns raw clocks or rescales either run. The exact optional full-block survival gates are in references/timeline_contract.md.
Name the selected route in the result and keep its evidence scope explicit. A valid ACU-only or timeline-only conclusion is complete without a joint artifact. Restore the probe-free kernel and use the normal evaluator for final correctness and performance; profile evidence does not replace the probe-free result.
© alibaba, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (scripts, references) in skills/ppu-acu-joint-profile of alibaba/atrex-kernel-agent.
Open the folder on GitHubat commit 602bc38
Ppu Acu Joint Profile next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ppu Acu Joint Profile this skillalibaba/atrex-kernel-agent | 168 | — | ~5.2k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Analysissgl-project/sglang | 37k | 2 repos | ~6.4k | Automated safety check: Pass | Apache-2.0 | |
| Tilelang SkillslowlyC/agent-gpu-skills | 169 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Kernel ProfilingZJLi2013/awesome-kernel-skills | 102 | — | ~696 | Automated safety check: Pass | None | |
| Nemo Mbridge Perf Moe Optimization WorkflowNVIDIA/skills | 3.6k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| LLM Torch Profiler Trace AnalysisBBuf/AI-Infra-Auto-Driven-SKILLS | 938 | — | ~2.8k | Automated safety check: Pass | None |
sgl-project/sglang
Unified LLM torch-profiler triage skill for sglang, vllm, TensorRT-LLM, and TokenSpeed.
slowlyC/agent-gpu-skills
Write, debug, and optimize TileLang kernels from local upstream language, JIT, autotuning, profiling, compiler, test, and example source.
ZJLi2013/awesome-kernel-skills
Profile GPU kernels using NCU (NVIDIA) or rocprof (AMD) to collect performance metrics.
NVIDIA/skills
Evidence-gated workflow for MoE performance optimization in Megatron Bridge.
BBuf/AI-Infra-Auto-Driven-SKILLS
Analyzes Torch Profiler traces from SGLang, vLLM and TensorRT-LLM servers into kernel attribution, overlap and fusion tables.
AI4Scientist/nano-scientist
Profile a target (script, process, GPU, memory, interconnect) using external tools and code instrumentation.
alibaba/atrex-kernel-agent
Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU…
alibaba/atrex-kernel-agent
Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki.
alibaba/atrex-kernel-agent
Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…
alibaba/atrex-kernel-agent
Generate a structured implementation plan from an evidence draft.
alibaba/atrex-kernel-agent
Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel.
alibaba/atrex-kernel-agent
Run the evidence loop of one long-horizon GPU kernel optimization episode.
Works with
Categories
Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel. Ppu Acu Joint Profile is an agent skill from alibaba/atrex-kernel-agent. Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.
Ppu Acu Joint Profile fits situations like: device-level bottleneck diagnosis; kernel-internal critical-path questions; evidence that genuinely needs both; do not require all three modes.
Run `npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a claude-code`. Or copy the skill folder (skills/ppu-acu-joint-profile in alibaba/atrex-kernel-agent) into .claude/skills/ppu-acu-joint-profile in your project. Claude Code loads it when a task matches its description.
Run `npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a codex`. Or copy the skill folder (skills/ppu-acu-joint-profile in alibaba/atrex-kernel-agent) into .agents/skills/ppu-acu-joint-profile in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alibaba/atrex-kernel-agent --skill ppu-acu-joint-profile -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ppu-acu-joint-profile, .gemini/skills/ppu-acu-joint-profile, .github/skills/ppu-acu-joint-profile and .opencode/skills/ppu-acu-joint-profile in your project.
Going by SKILL.md and its folder, Ppu Acu Joint Profile needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Ppu Acu Joint Profile is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.2k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 10k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ppu Acu Joint Profile: LLM Torch Profiler Analysis (sgl-project/sglang, 37k stars), Tilelang Skill (slowlyC/agent-gpu-skills, 169 stars), Kernel Profiling (ZJLi2013/awesome-kernel-skills, 102 stars) and Nemo Mbridge Perf Moe Optimization Workflow (NVIDIA/skills, 3.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
alibaba (a GitHub organization) maintains it in alibaba/atrex-kernel-agent, which has 168 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 10, 2026.
Source: alibaba/atrex-kernel-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.