Agent skill

Profile Training

by marin-community in marin-community/marin

Profile a named JAX, Levanter, or Marin run, or investigate a measured startup, compilation, initialization, or throughput bottleneck.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Profile Training

skills CLI
$ npx skills add marin-community/marin --skill profile-training -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install marin-community/marin profile-training --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/marin-community/marin.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/profile-training .claude/skills/profile-training && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
profile-training
GitHub stars
3.9k
Token cost
~2.5k tokens
SKILL.md length
782 words
Files
1
Skills in repo
41
Repo updated
First seen
Licence
Apache-2.0

At a glance

Profile a named JAX, Levanter, or Marin run, or investigate a measured startup, compilation, initialization, or throughput bottleneck.

  • Works in 8 steps: Measure: generate before.json. → Change: apply one bounded patch/config… → Re-measure: generate after.json. → …
  • Tasks that involve Deep learning
  • SKILL.md covers Scope, Capture Profiles, Ingest to Structured Summary and Query the summary, plus 1 more section
  • Calls uv

What it does

Profile Training is an agent skill from marin-community/marin. Profile a named JAX, Levanter, or Marin run, or investigate a measured startup, compilation, initialization, or throughput bottleneck.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning, Performance optimization and gRPC and Protobuf. It works with gRPC. The repository describes itself as: Open-source framework for the research and development of foundation models. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Deep learning
  • Tasks that involve Performance optimization
  • Tasks that involve gRPC and Protobuf

Example prompts

  • “/profile-training”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Measure: generate before.json.
  2. Change: apply one bounded patch/config tweak.
  3. Re-measure: generate after.json.
  4. Compare
  5. Track (thresholded pass/warn/fail + history)
  6. History summary (regression trend tracking)
  7. One-shot compare bundle
  8. Publish summary/report back to W&B

What it can do on your machine

Read from SKILL.md and the folder at commit 61bb85c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.jax.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Profile Training loads about 2.5k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 782 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from marin-community/marin at commit 61bb85c, republished under its Apache-2.0 licence (© marin-community). 782 words, ~2,497 tokens.

Download SKILL.mdSave it as .claude/skills/profile-training/SKILL.md (or your agent's skills folder).
name
profile-training
description
Profile a named JAX, Levanter, or Marin run, or investigate a measured startup, compilation, initialization, or throughput bottleneck.

Profile JAX training

Scope

Ingestion sources:

  • XPlane protobufs inside Levanter profile directories (source of truth):
    • plugins/profile/<timestamp>/*.xplane.pb
    • explicit local *.xplane.pb files via --xplane-file
  • xprof aggregate tables exported from the same XPlane protobuf when the optional xprof package is available: step overview timing, kernel stats, collective breakdowns, xprof bottleneck statements.
  • Perfetto trace JSON as an explicit/fallback source for older profiles:
    • plugins/profile/<timestamp>/perfetto_trace.json.gz
    • plugins/profile/<timestamp>/*.trace.json.gz

Prefer XPlane protobuf for new work. Perfetto trace JSON commonly hits the trace event cap; XPlane contains the uncapped timeline events needed for named-scope regions, pre-op gaps, gap context, process/thread metadata, and xprof aggregate tables. Use --trace-file only for a specific Perfetto JSON trace or an older profile with no XPlane protobuf.

Capture Profiles

Use Levanter profiler flags so profiles land under <trainer.log_dir>/<run_id>/profiler. Remote Marin runs also upload to MARIN_PREFIX TTL storage and print an XProf link:

bash
uv run ... \
  --trainer.profiler.enabled true \
  --trainer.profiler.start_step 5 \
  --trainer.profiler.num_steps 10 \
  --trainer.profiler.upload.ttl_days 30

For profiles where xprof/HLO protobuf tables matter, enable JAX profile options through the Levanter profiler config:

bash
uv run ... \
  --trainer.profiler.enabled true \
  --trainer.profiler.start_step 5 \
  --trainer.profiler.num_steps 5 \
  --trainer.profiler.profile_options.host_tracer_level 1 \
  --trainer.profiler.profile_options.python_tracer_level 0 \
  --trainer.profiler.profile_options.device_tracer_level 0 \
  --trainer.profiler.profile_options.enable_hlo_proto true

HLO metadata increases artifact size, so keep these profile windows short. The XProf profile: link appears after upload. Set --trainer.profiler.upload.enabled false for local-only capture. Do not copy profiles to another GCS region for inspection.

Known-good TensorBoard scope recipe from CoreWeave Grug MoE profiling: trainer.profiler.enabled=true, trainer.profiler.start_step=3, trainer.profiler.num_steps=2, trainer.profiler.perfetto_link=false, trainer.profiler.profile_options.host_tracer_level=1, trainer.profiler.profile_options.python_tracer_level=0, and trainer.profiler.profile_options.enable_hlo_proto=true preserved useful jax.named_scope / named_call regions in TensorBoard for GM2560-MAY-120S4096-W2048-B8-R1-E8M1-FA4PROFILE-S3B-N1-cw-20260617-2353. Leave device_tracer_level unset unless device timelines are specifically needed; this profile retained useful hierarchical host/XLA metadata without it.

On GPU, command buffers can collapse or suppress the visible name stack in TensorBoard/Perfetto. For profile-readability runs, disable command buffers:

bash
export XLA_FLAGS="${XLA_FLAGS:-} --xla_gpu_enable_command_buffer=''"

This hurts performance, so use it only when the goal is semantic trace attribution; leave it out of throughput comparisons unless command-buffer behavior is the axis being tested.

For GPU throughput runs, keep profile-readability flags separate from XLA code generation and scheduling flags. Start from JAX's GPU performance guide, especially the code generation flags section: https://docs.jax.dev/en/latest/gpu_performance_tips.html#code-generation-flags. The exact set of useful XLA flags is jaxlib-version dependent, so record the full XLA_FLAGS value with each profile or W&B run.

For better profile readability, use haliax.jax_utils.named_call and jax.named_scope liberally in model code; these names flow into trace annotations and make region-level summaries far more actionable.

Reference:

Ingest to Structured Summary

Use /tmp for ephemeral downloads. Use scratch/ only when the working tree must retain an uncommitted analysis artifact.

bash
# /tmp (ephemeral)
uv run python lib/marin/tools/profile_summary.py summarize \
  --run-target marin-community/marin/<run_id> \
  --download-root /tmp/marin-profiles \
  --breakdown-mode exclusive_global \
  --output /tmp/profile_summary.json
Option A: From a W&B artifact reference
bash
uv run python lib/marin/tools/profile_summary.py summarize \
  --artifact marin-community/marin/run-grug-125m-profile-apples-pallas_tpu-20260217-225239-055ab2-profiler:v0 \
  --download-root /tmp/marin-profiles \
  --output /tmp/profile_summary.json

--run-target accepts: a bare run id (requires --entity and --project), entity/project/run_id, or a full W&B run URL. The profiler directory is resolved from trainer.log_dir in the run config.

Option B: From a local artifact directory
bash
uv run python lib/marin/tools/profile_summary.py summarize \
  --profile-dir /path/to/profiler_dir \
  --output /tmp/profile_summary.json

If the directory contains *.xplane.pb, --profile-dir uses the XPlane path automatically. When both *.xplane.pb and Perfetto trace JSON are present, --profile-dir reads the XPlane protobuf by default (Perfetto exports are often capped). Use --trace-file to force a specific Perfetto JSON file.

Show full SKILL.md (315 more words)Show less
Option C: From a specific trace file
bash
uv run python lib/marin/tools/profile_summary.py summarize \
  --trace-file /path/to/perfetto_trace.json.gz \
  --output /tmp/profile_summary.json
Option D: From a specific XPlane protobuf

Direct XPlane timeline parsing uses protobuf and does not require TensorFlow-generated xplane_pb2 modules. If xprof is installed, ingestion also exports compact xprof table JSON and augments the timeline summary with aggregate step, kernel, collective, and bottleneck evidence.

bash
uv run --with xprof --with protobuf python lib/marin/tools/profile_summary.py summarize \
  --xplane-file /path/to/profile.xplane.pb \
  --xplane-output-dir /tmp/profile_xprof_tables \
  --xplane-count-trace-events \
  --output /tmp/profile_summary.json

Without --xplane-output-dir the command still parses XPlane timeline events directly. Add --with xprof for xprof aggregate table augmentation; add --xplane-output-dir to preserve the exported table JSON (this flag requires the optional xprof package).

XPlane summaries expose hierarchical named-scope regions, pre-op gaps, gap region context, process/thread/timeline event metadata, step timing (when step markers or xprof overview rows exist), xprof bottleneck statements, kernel stats, collective breakdowns, and optimization candidates.

Summary version tag: profile_summary.v1

Generate a deterministic markdown root-cause report:

bash
uv run python lib/marin/tools/profile_summary.py report \
  --summary /tmp/profile_summary.json \
  --output /tmp/profile_report.md

Trace quality checks are surfaced in trace_overview:

  • suspected_truncation: true when event counts match a known export cap.
  • quality_warnings: warnings to treat hotspot/gap attribution with caution.

Query the summary

bash
uv run python lib/marin/tools/profile_summary.py query \
  --summary /tmp/profile_summary.json \
  --question "<top ops, compute vs communication, gap, region, or op context>"

Query top exclusive-time ops, compute/communication balance and collectives, specific pre-op gaps, hierarchical regions, noisy-op context, and suggested optimizations.

Useful query forms include:

  • What are the top 10 ops by exclusive time?
  • Is comm or compute dominating? Which collective is worst?
  • gap before _linear_softmax_cross_entropy_loss_bwd_pallas_mosaic_tpu_combined.1
  • show hierarchical regions
  • show context for op copy.564
  • What should we try next?

Pre-op gap attribution is marker-aware:

  • gap_before_ops[].payload_op: op where useful work starts after the idle period.
  • gap_before_ops[].marker_op: first op observed after the gap (often lightweight setup like iota.*).

Optimization Workflow

Use a strict workflow:

  1. Measure: generate before.json.
  2. Change: apply one bounded patch/config tweak.
  3. Re-measure: generate after.json.
  4. Compare:
bash
uv run python lib/marin/tools/profile_summary.py compare \
  --before /tmp/profile_before.json \
  --after /tmp/profile_after.json \
  --strict-provenance
  1. Track (thresholded pass/warn/fail + history):
bash
uv run python lib/marin/tools/profile_summary.py track \
  --before /tmp/profile_before.json \
  --after /tmp/profile_after.json \
  --label "pallas-kernel-attempt-3" \
  --history /tmp/profile_regression_history.jsonl
  1. History summary (regression trend tracking):
bash
uv run python lib/marin/tools/profile_summary.py history \
  --history /tmp/profile_regression_history.jsonl
  1. One-shot compare bundle:
bash
uv run python lib/marin/tools/profile_summary.py bundle \
  --before-run-target marin-community/marin/<baseline_run_id> \
  --after-run-target marin-community/marin/<candidate_run_id> \
  --output-dir /tmp/profile_bundle \
  --history /tmp/profile_regression_history.jsonl
  1. Publish summary/report back to W&B:
bash
uv run python lib/marin/tools/profile_summary.py publish \
  --summary /tmp/profile_summary.json \
  --report /tmp/profile_report.md \
  --alias latest

The comparison reports: steady-state step-time delta, step class deltas (light/heavy when detected), compute/comm/host/stall share deltas, semantic family deltas with workload-normalized metrics, provenance checks (trace hash/run identity), and regressed/improved ops by exclusive duration.

© marin-community, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/profile-training of marin-community/marin.

Open the folder on GitHubat commit 61bb85c

Compare with similar skills

Profile Training next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Profile Training compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Profile Training this skillmarin-community/marin3.9k—~2.5kAutomated safety check: PassApache-2.0
Projectsamchon/typia5.9k—~1.8kAutomated safety check: PassMIT
Performance Optimizationalbumentations-team/AlbumentationsX567—~1.7kAutomated safety check: PassAGPL-3.0
Torch Performance Optimizationalbumentations-team/albucore123—~895Automated safety check: PassMIT
GcloudKilo-Org/kilo-marketplace190—~2.9kAutomated safety check: PassApache-2.0
Golang Proantoniopaya22/go-rest-template1723 repos~1.2kAutomated safety check: PassMIT

Similar skills

  • Project

    samchon/typia

    Defines the typia product contract, workspace layout, package boundaries, and canonical commands.

    5.9k GitHub stars~1.8k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Performance Optimization

    albumentations-team/AlbumentationsX

    Systematic performance audit for AlbumentationsX runtime code.

    567 GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Torch Performance Optimization

    albumentations-team/albucore

    Optimize or review eager CPU-only Albucore PyTorch runtime paths with benchmark-backed decisions.

    123 GitHub stars~895 tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Gcloud

    Kilo-Org/kilo-marketplace

    Interacts with Google Cloud services using the gcloud CLI safely and efficiently.

    190 GitHub stars~2.9k tokensUpdated 11 days ago
    Backend & APIsAuto-check passed
  • Golang Pro

    antoniopaya22/go-rest-template

    Implements concurrent Go patterns using goroutines and channels, designs and builds microservices with gRPC or REST, optimizes Go application performance with pprof, and enforces idiomatic Go with…

    172 GitHub starsUsed in 3 repos~1.2k tokens
    Backend & APIsAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes

More from marin-community/marin

All 41 skills in this repo
  • Noslop

    marin-community/marin

    Deslop, simplify, or review low-value tests and prose only when explicitly requested for a branch or diff.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Use Iris

    marin-community/marin

    Use Iris to submit, inspect, debug, monitor, or recover jobs and tasks; diagnose scheduling and federation; deploy controllers; or reserve dev GPUs and TPUs.

    3.9k GitHub stars~745 tokensUpdated today
    Auto-check passed
  • Launch Rl

    marin-community/marin

    Define, validate, submit, or restart a Marin SkyRL experiment through its artifact main.

    3.9k GitHub stars~894 tokensUpdated today
    Auto-check passed
  • Marina Applet

    marin-community/marin

    Build, validate, publish, update, inspect, query, roll back, or archive a dynamic Marina applet.

    3.9k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Query Finelog

    marin-community/marin

    Query Finelog logs and telemetry for Iris tasks, workers, profiles, training, vLLM, and cross-cluster forwarding.

    3.9k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Trace Pulumi Diff

    marin-community/marin

    Run a read-only preview for a specified Marin infra/pulumi stack and trace each pending resource change to merged pull requests since its latest successful update when that update records a clean…

    3.9k GitHub stars~663 tokensUpdated today
    Auto-check passed

Works with

Questions about Profile Training

What does Profile Training do?

Profile a named JAX, Levanter, or Marin run, or investigate a measured startup, compilation, initialization, or throughput bottleneck. Profile Training is an agent skill from marin-community/marin. Profile a named JAX, Levanter, or Marin run, or investigate a measured startup, compilation, initialization, or throughput bottleneck.

When should I use Profile Training?

Profile Training fits situations like: tasks that involve Deep learning; tasks that involve Performance optimization; tasks that involve gRPC and Protobuf.

How do I install Profile Training in Claude Code?

Run `npx skills add marin-community/marin --skill profile-training -a claude-code`. Or copy the skill folder (.agents/skills/profile-training in marin-community/marin) into .claude/skills/profile-training in your project. Claude Code loads it when a task matches its description.

How do I install Profile Training in Codex?

Run `npx skills add marin-community/marin --skill profile-training -a codex`. Or copy the skill folder (.agents/skills/profile-training in marin-community/marin) into .agents/skills/profile-training in your project. Codex loads it when a task matches its description.

Can I use Profile Training in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add marin-community/marin --skill profile-training -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/profile-training, .gemini/skills/profile-training, .github/skills/profile-training and .opencode/skills/profile-training in your project.

What does Profile Training need to run?

Going by SKILL.md and its folder, Profile Training needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Profile Training access the network?

SKILL.md names 1 domain. As links in the text: docs.jax.dev. This is read from the text; nothing was executed.

Is Profile Training safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Profile Training use?

Profile Training is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Profile Training use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Profile Training?

Skills that share tags, products or a category with Profile Training: Project (samchon/typia, 5.9k stars), Performance Optimization (albumentations-team/AlbumentationsX, 567 stars), Torch Performance Optimization (albumentations-team/albucore, 123 stars) and Gcloud (Kilo-Org/kilo-marketplace, 190 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Profile Training?

marin-community (a GitHub organization) maintains it in marin-community/marin, which has 3,920 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on October 9, 2026.

Source: marin-community/marin on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.