Agent skill

Analyze Nsys Profile

by mlc-ai in mlc-ai/pith-train

Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

Apache-2.0Auto-check passedDatabases

Install Analyze Nsys Profile

skills CLI
$ npx skills add mlc-ai/pith-train --skill analyze-nsys-profile -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mlc-ai/pith-train analyze-nsys-profile --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/analyze-nsys-profile .claude/skills/analyze-nsys-profile && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyze-nsys-profile
GitHub stars
355
Token cost
~1.9k tokens
SKILL.md length
931 words
Files
8 (incl. scripts, references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

  • Works in 3 steps: Export the trace to SQLite → Shared preparation (always run first) → Measure overlap per DualPipeV stage
  • The user asks to analyze an nsys profile
  • SKILL.md covers Prerequisites, Step 1 — Export the trace to…, Step 2 — Shared preparation… and Step 3 — Measure overlap per…, plus 6 more sections
  • Runs Python scripts from its folder; calls python

What it does

Analyze Nsys Profile is an agent skill from mlc-ai/pith-train. Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior. Use when the user asks to "analyze an nsys profile", "check overlap quality", "find exposed comm", "which stage is the bottleneck", or any question that starts from an existing .nsys-rep file. Assumes the trace was already captured (see capture-nsys-profile); provides query primitives the agent composes for the specific question being asked.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/conventions.md`, `references/examples.md` and `scripts/classify_streams.py`).

It sits in Databases. It works with SQLite. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.

When your agent uses it

  • The user asks to analyze an nsys profile
  • Check overlap quality
  • Find exposed comm
  • Which stage is the bottleneck

Example prompts

  • “analyze an nsys profile”
  • “check overlap quality”
  • “find exposed comm”
  • “/analyze-nsys-profile”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Export the trace to SQLite
  2. Shared preparation (always run first)
  3. Measure overlap per DualPipeV stage

What it can do on your machine

Read from SKILL.md and the folder at commit 87208d9. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyze Nsys Profile loads about 1.9k tokens when it runs, and up to ~4.2k if it reads all its reference files. Until then it costs about 134 tokens; SKILL.md has 931 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~134
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from mlc-ai/pith-train at commit 87208d9, republished under its Apache-2.0 licence (© mlc-ai). 931 words, ~1,927 tokens.

Download SKILL.mdSave it as .claude/skills/analyze-nsys-profile/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
analyze-nsys-profile
description
Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior. Use when the user asks to "analyze an nsys profile", "check overlap quality", "find exposed comm", "which stage is the bottleneck", or any question that starts from an existing `.nsys-rep` file. Assumes the trace was already captured (see capture-nsys-profile); provides query primitives the agent composes for the specific question being asked.

Analyze Nsys Profile

A passive query toolkit for PithTrain nsys traces. The agent asks a specific question; the skill provides primitives that answer it fast and correctly. The skill does not produce an unsolicited full report. It expects the agent to compose the right query for the question being asked.

Prerequisites

  • A captured .nsys-rep exists (default location: workspace/capture-nsys-profile/pithtrain_node*.nsys-rep).
  • The repo venv is active: source .venv/bin/activate.
  • nsys CLI on PATH (for the one-time SQLite export).

Step 1 — Export the trace to SQLite

bash
nsys export --type=sqlite --force-overwrite=true --output=workspace/capture-nsys-profile/pithtrain_node0.sqlite workspace/capture-nsys-profile/pithtrain_node0.nsys-rep

All subsequent queries hit the SQLite, not the raw .nsys-rep.

Step 2 — Shared preparation (always run first)

Three primitives establish who, when, and what — every downstream analysis depends on the data they surface. Run (or at least understand the output of) all three before reaching for the analysis scripts below.

QuestionPrimitive
What ranks are in this trace? What's the per-rank setup?show_setup.py
What's the steady-state analysis window for each rank?find_window.py
Which streams are compute / comm, and what's each comm stream's purpose?classify_streams.py

Pipeline: show_setup → find_window → classify_streams. show_setup gives you the mapping pid ↔ rank ↔ mesh coordinates; find_window picks the median DualPipeV chunk per rank (deterministic across re-runs, so before/after comparisons are valid); classify_streams identifies which CUDA streams in that window are compute vs comm, and labels the comm streams' purpose (ep_a2a, cp_ring, pp_p2p).

Step 3 — Measure overlap per DualPipeV stage

bash
python .agents/skills/analyze-nsys-profile/scripts/compute_overlap.py workspace/capture-nsys-profile/pithtrain_node0.sqlite

Emits one row per (rank, stage) with columns: pid | stage | exposed_ns | overlap | overlap_min | overlap_max. The overlap column is the time-weighted hidden fraction across the stage's comm kernels; overlap_min / overlap_max are the extremes of the per-kernel overlap percentage and surface whether the stage is uniformly bad or bimodal.

See references/examples.md for recipes that compose this with the Step 2 primitives.

Critical conventions

Before composing a custom SQL query, read references/conventions.md. Highlights:

  • pid (Linux PID) is the per-rank join key, extracted exactly as the nsys docs prescribe: globalPid / 0x1000000 % 0x1000000 (kernel rows) == globalTid / 0x1000000 % 0x1000000 (NVTX rows). Single SQLite per node → PIDs unique within a trace.
  • Always filter start >= 0 — pre-cudaProfilerStart NCCL init ranges have negative timestamps.
  • Per-rank setup label is the first in-window NVTX event per rank: rank=N; pp=R/S dp=R/S cp=R/S ep=R/S; mbs=M seq=Q.
  • Chunk anchor for steady state: the median-indexed forward chunk X (phaseY) backward chunk Z (phaseW) NVTX range emitted by DualPipeV. Match with LIKE 'forward chunk%backward chunk%' to disambiguate from per-stage forward markers.
  • Compute-vs-comm classification: a kernel is communication if its short name starts with nccl, otherwise compute. A stream is a comm stream iff every one of its kernels is NCCL.
  • Comm-stream purpose is discovered from the PithTrain stage-NVTX (layer*.stageN_*) enclosing each kernel at its CPU-side launch time — every kernel must agree on the label (unanimity), otherwise mixed.

Non-fragile classification rules

Avoid these heuristics — they break across configs:

  • "Stream with > N kernels of type X is comm" (N depends on layer count, chunks, seq length).
  • "Kernel duration > T µs means data movement" (long duration can also be a straggler wait).
  • "Stream with avg µs < threshold is EP" (depends on token volume per rank).

Use these instead:

  • One-sided purity check for compute-vs-comm streams.
  • NVTX-context labeling for stream purpose: look up the innermost PithTrain stage range enclosing each kernel and require unanimous agreement. Implemented in classify_streams.py.

Worked examples

See references/examples.md for recipe-style answers to:

  • How well is each EP phase overlapped with compute?
  • Which EP phase has the worst overlap?
  • Are the PP stages balanced?
  • Which (rank, stage) carries the most exposed comm?
Show full SKILL.md (365 more words)Show less

Output guidance

  • Scripts emit a fixed-width table to stdout. One column per record field; agents read it directly.
  • Cross-script joins are by pid — every script's table includes pid as the per-rank identifier; downstream rows compose against show_setup's mapping pid ↔ rank ↔ setup.
  • When reporting to the human user, summarize as plain prose with a small table extracted from the relevant columns.
  • Always cite the analysis window. An overlap percentage with no window is meaningless.

Gotchas (surfaced by prior agent runs)

  • classify_streams.py only reports streams active in the analysis window, not every stream that exists in the trace. A rank typically has 6-8 streams overall but only 2-3 inside a single steady-state chunk. This is intentional — analyzing a small window does not need the inactive streams.
  • PP P2P kernels rarely appear in a single chunk window — they fire between chunks. Widen the window (--start NS --end NS on the analysis script) if you specifically want to see the PP P2P comm stream.
  • compute_overlap.py's percent cells include a trailing % (58.4%, not 0.584). Sort/compare numerically by stripping the % first. Absolute time columns (exposed_ns) are bare integer nanoseconds.
  • CPU launch time vs GPU execution time — for any NVTX-context lookup on a kernel, use kernel["launch_start"] (CPU-side cudaLaunchKernel time) rather than kernel["start"] (GPU-side execution time). All scripts already do this; if you write an ad-hoc query, call common.innermost_nvtx on launch_start values.
  • Comm-stream purpose uses unanimity, not majority — every kernel on the stream must agree on its enclosing-NVTX category, otherwise the label is mixed. A single mis-categorized kernel surfaces as mixed instead of silently being out-voted.

Common Issues

no such table: NVTX_EVENTS

The .nsys-rep has not been exported yet. Run the nsys export command from Step 1.

PP P2P comm is missing from the overlap output

By design. compute_overlap.py buckets kernels by their enclosing PithTrain stage NVTX (stage1_* through stage5_*); PP P2P kernels live inside the pipeline send/recv wrapper, which is not a stage marker, so they are filtered out. Widen the window with --start NS --end NS if you need to investigate them — they typically fire between chunks, not inside.

Negative timestamps on NVTX events

NCCL init opened these ranges before cudaProfilerStart. Filter WHERE start >= 0 to scope to the profiled window.

© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in .agents/skills/analyze-nsys-profile of mlc-ai/pith-train.

  • SKILL.md
  • references/conventions.md
  • references/examples.md
  • scripts/classify_streams.py
  • scripts/common.py
  • scripts/compute_overlap.py
  • scripts/find_window.py
  • scripts/show_setup.py

Open the folder on GitHubat commit 87208d9

Compare with similar skills

Analyze Nsys Profile next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyze Nsys Profile compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyze Nsys Profile this skillmlc-ai/pith-train355—~1.9kAutomated safety check: PassApache-2.0
Iptvnator Sqlite DB Worker4gray/iptvnator7.3k—~824Automated safety check: PassMIT
Reactive Sqlite UIfastrepl/anarlog9.5k—~699Automated safety check: PassMIT
Composer Forensicsdxos/dxos526—~3.1kAutomated safety check: PassCustom licence
Sqlite Schema Designfastrepl/anarlog9.5k—~1.9kAutomated safety check: PassMIT
MakemigrationsdeusXmachina-dev/memorylane121—~973Automated safety check: PassGPL-3.0

Similar skills

  • A skill your agent uses when changing Electron SQLite IPC, database-worker operations, request-scoped progress or cancellation, worker packaging, or runtime verification of non-EPG database work.

    7.3k GitHub stars~824 tokensUpdated today
    DatabasesAuto-check passed
  • Reactive Sqlite UI

    fastrepl/anarlog

    Build SQLite-backed reactive UI in apps/desktop using stable patterns for reads, selection, forms, writes, and loading states.

    9.5k GitHub stars~699 tokensUpdated today
    DatabasesAuto-check passed
  • Forensically inspect and repair Composer browser profiles — offline (Chrome OPFS / SQLite extract) or live via /recovery.html debug port.

    526 GitHub stars~3.1k tokensUpdated today
    DatabasesAuto-check passed
  • Sqlite Schema Design

    fastrepl/anarlog

    Design or review schemas for crates/cloudsync using SQLite Sync constraints, not generic SQLite advice.

    9.5k GitHub stars~1.9k tokensUpdated today
    DatabasesAuto-check passed
  • Makemigrations

    deusXmachina-dev/memorylane

    Create SQLite migrations for MemoryLane storage schema changes.

    121 GitHub stars~973 tokensUpdated yesterday
    DatabasesAuto-check passed
  • Toon

    butttons/dora

    Token-Oriented Object Notation is a compact, human-readable encoding of the JSON data model that minimizes tokens and makes structure easy for models to follow.

    109 GitHub stars~9.4k tokensUpdated 7 mo ago
    DatabasesAuto-check passed

More from mlc-ai/pith-train

All 10 skills in this repo
  • Capture Nsys Profile

    mlc-ai/pith-train

    Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.

    355 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Validate Correctness

    mlc-ai/pith-train

    Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

    355 GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Validate Performance

    mlc-ai/pith-train

    Measures the throughput difference between two branches with force-balanced routing.

    355 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Setup Benchmark Inputs

    mlc-ai/pith-train

    Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.

    355 GitHub stars~399 tokensUpdated today
    Auto-check passed
  • Add New Model

    mlc-ai/pith-train

    Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

    355 GitHub stars~4.6k tokensUpdated today
    Auto-check passed
  • Wandb Tracking

    mlc-ai/pith-train

    Read, analyze, and manage Weights & Biases (wandb) experiment data for PithTrain runs.

    355 GitHub stars~1.1k tokensUpdated today
    Auto-check: warnings

Works with

Categories

Questions about Analyze Nsys Profile

What does Analyze Nsys Profile do?

Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior. Analyze Nsys Profile is an agent skill from mlc-ai/pith-train. Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

When should I use Analyze Nsys Profile?

Analyze Nsys Profile fits situations like: the user asks to analyze an nsys profile; check overlap quality; find exposed comm; which stage is the bottleneck.

How do I install Analyze Nsys Profile in Claude Code?

Run `npx skills add mlc-ai/pith-train --skill analyze-nsys-profile -a claude-code`. Or copy the skill folder (.agents/skills/analyze-nsys-profile in mlc-ai/pith-train) into .claude/skills/analyze-nsys-profile in your project. Claude Code loads it when a task matches its description.

How do I install Analyze Nsys Profile in Codex?

Run `npx skills add mlc-ai/pith-train --skill analyze-nsys-profile -a codex`. Or copy the skill folder (.agents/skills/analyze-nsys-profile in mlc-ai/pith-train) into .agents/skills/analyze-nsys-profile in your project. Codex loads it when a task matches its description.

Can I use Analyze Nsys Profile in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill analyze-nsys-profile -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyze-nsys-profile, .gemini/skills/analyze-nsys-profile, .github/skills/analyze-nsys-profile and .opencode/skills/analyze-nsys-profile in your project.

What does Analyze Nsys Profile need to run?

Going by SKILL.md and its folder, Analyze Nsys Profile needs Python for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Analyze Nsys Profile access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Analyze Nsys Profile safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Analyze Nsys Profile use?

Analyze Nsys Profile is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyze Nsys Profile use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.3k tokens, read only when the agent opens those files.

What are the alternatives to Analyze Nsys Profile?

Skills that share tags, products or a category with Analyze Nsys Profile: Iptvnator Sqlite DB Worker (4gray/iptvnator, 7.3k stars), Reactive Sqlite UI (fastrepl/anarlog, 9.5k stars), Composer Forensics (dxos/dxos, 526 stars) and Sqlite Schema Design (fastrepl/anarlog, 9.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyze Nsys Profile?

mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 9, 2026.

Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.