Agent skill

Add Memory Prints

by mlc-ai in mlc-ai/pith-train

Add detailed memory profiling prints throughout the training framework.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Add Memory Prints

skills CLI
$ npx skills add mlc-ai/pith-train --skill add-memory-prints -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mlc-ai/pith-train add-memory-prints --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mlc-ai/pith-train.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/add-memory-prints .claude/skills/add-memory-prints && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
add-memory-prints
GitHub stars
355
Token cost
~6.2k tokens
SKILL.md length
1,190 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Add detailed memory profiling prints throughout the training framework.

  • Works in 3 steps: model (required): Model name matching a… → ranks (optional, default: {0}):… → detail-layers (optional, default: auto):…
  • User asks to add memory prints
  • SKILL.md covers Arguments, Correctness Guarantee, Before Starting and Group 1: Helpers, plus 8 more sections
  • Calls ruff

What it does

Add Memory Prints is an agent skill from mlc-ai/pith-train. Add detailed memory profiling prints throughout the training framework. Instruments distributed setup, model creation, checkpoint loading, pipeline scheduling, per-layer activations, saved tensor profiling, expert MLP internals, and memory snapshot dumps. Use when user asks to "add memory prints", "instrument memory", "profile memory", "memory breakdown", or "debug memory".

Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering. The repository describes itself as: Compact and Agent-Native MoE Training System. The licence is Apache-2.0.

When your agent uses it

  • User asks to add memory prints
  • Instrument memory
  • Memory breakdown

Example prompts

  • “add memory prints”
  • “instrument memory”
  • “profile memory”
  • “/add-memory-prints”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. model (required): Model name matching a file under pithtrain/models/ (e.g., qwen3-moe maps to pithtrain/models/qwen3_moe.py).
  2. ranks (optional, default: {0}): Comma-separated GPU global ranks to print on. E.g., --ranks 3,14 becomes the Python set {3, 14}. If not…
  3. detail-layers (optional, default: auto): Comma-separated layer indices for fine-grained per-layer profiling. If omitted, auto-select by…

What it can do on your machine

Read from SKILL.md and the folder at commit c7c8b1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ruff

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Add Memory Prints loads about 6.2k tokens when it runs. Until then it costs about 99 tokens; SKILL.md has 1,190 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~99
When it runs · the whole SKILL.md, loaded when a task matches
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mlc-ai/pith-train at commit c7c8b1d, republished under its Apache-2.0 licence (© mlc-ai). 1,190 words, ~6,163 tokens.

Download SKILL.mdSave it as .claude/skills/add-memory-prints/SKILL.md (or your agent's skills folder).
name
add-memory-prints
description
Add detailed memory profiling prints throughout the training framework. Instruments distributed setup, model creation, checkpoint loading, pipeline scheduling, per-layer activations, saved tensor profiling, expert MLP internals, and memory snapshot dumps. Use when user asks to "add memory prints", "instrument memory", "profile memory", "memory breakdown", or "debug memory".
argument-hint
<model-name> [--ranks <rank-list>] [--detail-layers <layer-indices>]

Memory Prints Instrumentation

Add comprehensive memory profiling instrumentation to the pithtrain training framework. This skill adds 7 groups of memory prints across 6 files, covering every phase from distributed init through the pipeline loop.

Arguments

Parse the following from $ARGUMENTS:

  1. model (required): Model name matching a file under pithtrain/models/ (e.g., qwen3-moe maps to pithtrain/models/qwen3_moe.py).
  2. --ranks (optional, default: {0}): Comma-separated GPU global ranks to print on. E.g., --ranks 3,14 becomes the Python set {3, 14}. If not provided, defaults to {0}.
  3. --detail-layers (optional, default: auto): Comma-separated layer indices for fine-grained per-layer profiling. If omitted, auto-select by reading the model file to find the MoE-vs-dense layer condition, then pick the first MoE layer and the next one (2 layers total).

Throughout this document, RANKS means the parsed rank set (e.g., {3, 14}), and DETAIL_LAYERS means the parsed or auto-selected layer indices tuple (e.g., (5, 6)).

Correctness Guarantee

All edits are observation-only:

  • _lmem() and _mem_gb() call torch.cuda.synchronize() + read memory_allocated(). No tensor modification.
  • _SavedTensorsProfiler uses saved_tensors_hooks with a pack function that returns the tensor unchanged.
  • Expert detail prints refactor return self.down_proj(silu_mul(g, u), ...) into gu = silu_mul(g, u); out = self.down_proj(gu, ...); return out — functionally identical.
  • Pipeline prints add synchronize/print between existing operations. Adds latency, does not change computation order.

Before Starting

  1. Check for existing prints: Search for _mem_gb, _lmem, _layer_mem_profile, _setup_mem, or memory_profiling in the codebase. If found, ask the user whether to update ranks/layers or skip.
  2. Read all target files (listed in Execution Order below) to find exact insertion points.
  3. Parse $ARGUMENTS for model, ranks, and detail layers.

Group 1: Helpers

1A: Pipeline helpers in pithtrain/pipeline/dualpipev.py

Add right before class DualPipeV:

python
def _mem_gb() -> float:
    """Return current CUDA memory allocated in GiB."""
    return torch.cuda.memory_allocated() / 1024**3


def _mem_detail() -> str:
    """Return allocated, cached-pool, and non-pytorch memory in GiB."""
    free, total = torch.cuda.mem_get_info()
    allocated = torch.cuda.memory_allocated()
    reserved = torch.cuda.memory_reserved()
    G = 1024**3
    cached = reserved - allocated
    non_pytorch = total - free - reserved
    return f"alloc={allocated / G:.2f} cached={cached / G:.2f} non-pt={non_pytorch / G:.2f}"

In DualPipeV.__init__, add after self.comm_stream = ...:

python
self.memory_profiling = True  # Set to True to enable per-step memory logging
1B: Layer-level helpers in pithtrain/pipeline/execution.py

Add after imports, before any function definitions:

python
# -- Per-layer activation profiling --
_layer_mem_profile = False
_layer_mem_ranks = RANKS


class _SavedTensorsProfiler:
    """Context manager that logs tensors saved by autograd, distinguishing weights from activations."""

    def __init__(self, layer_idx: int, stage_name: str, weight_data_ptrs: set):
        self._layer_idx = layer_idx
        self._stage_name = stage_name
        self._weight_data_ptrs = weight_data_ptrs
        self._log: list[str] = []
        self._act_bytes = 0
        self._wt_bytes = 0

    def _pack(self, t: torch.Tensor) -> torch.Tensor:
        nbytes = t.nelement() * t.element_size()
        is_wt = t.data_ptr() in self._weight_data_ptrs
        tag = "weight" if is_wt else "activ"
        self._log.append(
            f"    saved ({tag}): {tuple(t.shape)} {t.dtype} ({nbytes / 1024**2:.1f} MB)"
        )
        if is_wt:
            self._wt_bytes += nbytes
        else:
            self._act_bytes += nbytes
        return t

    def __enter__(self):
        self._ctx = torch.autograd.graph.saved_tensors_hooks(self._pack, lambda t: t)
        self._ctx.__enter__()
        return self

    def __exit__(self, *args):
        self._ctx.__exit__(*args)

    def print_summary(self):
        rank = torch.distributed.get_rank() if torch.distributed.is_initialized() else 0
        if rank not in _layer_mem_ranks:
            return
        hdr = f"layer{self._layer_idx} {self._stage_name} saved tensors"
        print(f"[rank={rank}] {hdr}:", flush=True)
        for line in self._log:
            print(f"[rank={rank}] {line}", flush=True)
        print(
            f"[rank={rank}]   activ={self._act_bytes / 1024**2:.1f} MB, "
            f"weight={self._wt_bytes / 1024**2:.1f} MB, "
            f"total={len(self._log)} tensors",
            flush=True,
        )


def _lmem(label: str) -> None:
    """Print memory at a layer-internal checkpoint."""
    if not _layer_mem_profile:
        return
    rank = torch.distributed.get_rank() if torch.distributed.is_initialized() else 0
    if rank not in _layer_mem_ranks:
        return
    torch.cuda.synchronize()
    alloc = torch.cuda.memory_allocated() / 1024**3
    print(f"[rank={rank}]   {label}: alloc={alloc:.2f}", flush=True)
1C: Setup helper in pithtrain/modules/training.py

Add before setup_model:

python
def _setup_mem(label: str) -> None:
    """Print CUDA memory at a setup checkpoint."""
    if torch.distributed.get_rank() in RANKS:
        torch.cuda.synchronize()
        G = 1024**3
        alloc = torch.cuda.memory_allocated()
        reserved = torch.cuda.memory_reserved()
        free, total = torch.cuda.mem_get_info()
        cached = reserved - alloc
        non_pytorch = total - free - reserved
        print(
            f"[rank={torch.distributed.get_rank()}] setup_model | {label}: "
            f"alloc={alloc / G:.2f} cached={cached / G:.2f} non-pt={non_pytorch / G:.2f}",
            flush=True,
        )

Group 2: Distributed Setup Prints (pithtrain/modules/distributed.py)

In setup_default_process_group

Important: this runs BEFORE init_process_group, so torch.distributed.get_rank() is not available, and distributed.rank is not published until setup_device_mesh. Use the device_id parameter for the rank guard.

torch.cuda.mem_get_info reads the current device, which nothing has set this early, so the probe sets it itself; setup_device_mesh sets it again later, harmlessly. Add a memory probe before and after init_process_group:

python
# Probe: isolate CUDA context cost from NCCL communicator init
torch.cuda.set_device(device_id)
torch.cuda.synchronize()
_free0, _total = torch.cuda.mem_get_info()
_non_pt0 = _total - _free0
G = 1024**3

# ... existing init_process_group code ...

torch.cuda.synchronize()
_free1, _ = torch.cuda.mem_get_info()
_non_pt1 = _total - _free1
if device_id in RANKS:
    print(
        f"[rank={torch.distributed.get_rank()}] init_process_group | "
        f"cuda_ctx={_non_pt0 / G:.2f} "
        f"after_nccl_world={_non_pt1 / G:.2f} "
        f"nccl_world_cost={(_non_pt1 - _non_pt0) / G:.2f}",
        flush=True,
    )
In setup_device_mesh

Before and after init_device_mesh:

python
torch.cuda.synchronize()
_free_before, _total = torch.cuda.mem_get_info()
_non_pt_before = _total - _free_before

# ... existing init_device_mesh code ...

torch.cuda.synchronize()
_free_after, _ = torch.cuda.mem_get_info()
_non_pt_after = _total - _free_after
G = 1024**3
if device_id in RANKS:
    print(
        f"[rank={distributed.rank}] init_device_mesh | "
        f"non-pt={_non_pt_after / G:.2f} "
        f"mesh_cost={(_non_pt_after - _non_pt_before) / G:.2f}",
        flush=True,
    )

Group 3: Model Setup Prints (pithtrain/modules/training.py)

Add _setup_mem(...) calls at these 5 points in setup_model:

  1. Before modules = [] — _setup_mem("before model creation")
  2. After creating module[0] — _setup_mem("after module[0] creation")
  3. After creating module[1] — _setup_mem("after module[1] creation")
  4. After init_weights loop — _setup_mem("after init_weights")
  5. After apply_fsdp(...) — _setup_mem("after apply_fsdp")

Group 4: Pipeline-Level Prints (pithtrain/pipeline/dualpipev.py)

This is the most complex group. All insertions go into DualPipeV.step(). The code below shows every insertion with its exact anchor point.

4A: Profiling flag setup — insert after the first-rank micro-batch setup (self.objective = objective), before # Step 1:

python
_profiling = self.memory_profiling and distributed.rank in RANKS
if _profiling:
    torch.cuda.synchronize()
    _m0 = _mem_gb()
    print(
        f"[rank={distributed.rank} pp={pp_rank}] Before pipeline: {_m0:.2f} GiB | {_mem_detail()}",
        flush=True,
    )

4B: Step 1 (nF0) — modify the existing loop body. The original loop is:

python
for i in range(step_1):
    self._forward_chunk(0)

Replace with:

python
for i in range(step_1):
    if _profiling and i == 1:
        import pithtrain.pipeline.execution as _mod
        _mod._layer_mem_profile = True
    self._forward_chunk(0)
    if _profiling and i == 1:
        _mod._layer_mem_profile = False
    if _profiling:
        torch.cuda.synchronize()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step1 F0 i={i}: {_mem_gb():.2f} GiB (+{_mem_gb() - _m0:.2f}) | {_mem_detail()}",
            flush=True,
        )

After the loop, before Step 2:

python
if _profiling:
    torch.cuda.synchronize()
    _m1 = _mem_gb()
    print(
        f"[rank={distributed.rank} pp={pp_rank}] After Step1 ({step_1} F0): {_m1:.2f} GiB (+{_m1 - _m0:.2f}) | {_mem_detail()}",
        flush=True,
    )

4C: Step 2 (nF0F1) — modify the loop body. Original:

python
for i in range(step_2):
    self._forward_chunk(0, recv=False, send=False)
    self._recv_forward(0)
    self._forward_chunk(1, send=(not self.is_last_pp_rank) or (i < step_2 - 1))
    self._send_forward(0)

Replace with:

python
for i in range(step_2):
    self._forward_chunk(0, recv=False, send=False)
    if _profiling:
        torch.cuda.synchronize()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step2 i={i} forward_chunk(0): {_mem_gb():.2f} GiB (+{_mem_gb() - _m0:.2f}) | {_mem_detail()}",
            flush=True,
        )
    self._recv_forward(0)
    if _profiling and i == 0:
        import pithtrain.pipeline.execution as _mod
        _mod._layer_mem_profile = True
    self._forward_chunk(1, send=(not self.is_last_pp_rank) or (i < step_2 - 1))
    if _profiling and i == 0:
        _mod._layer_mem_profile = False
    if _profiling:
        torch.cuda.synchronize()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step2 i={i} forward_chunk(1): {_mem_gb():.2f} GiB (+{_mem_gb() - _m0:.2f}) | {_mem_detail()}",
            flush=True,
        )
    self._send_forward(0)

After the loop:

python
if _profiling:
    torch.cuda.synchronize()
    _m2 = _mem_gb()
    print(
        f"[rank={distributed.rank} pp={pp_rank}] After Step2 ({step_2} F0F1): {_m2:.2f} GiB (+{_m2 - _m0:.2f}) | {_mem_detail()}",
        flush=True,
    )

4D: Step 3 (nB1W1F1) — modify the loop body. Original:

python
for i in range(step_3):
    self._backward_chunk(1, enable_zb=True)
    self._recv_forward(1)
    self._weight_chunk()
    self._forward_chunk(1, recv=False)

Replace with:

python
for i in range(step_3):
    if _profiling:
        torch.cuda.synchronize()
        _ms3 = _mem_gb()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step3 i={i} before B1: {_ms3:.2f} GiB | {_mem_detail()}",
            flush=True,
        )
    self._backward_chunk(1, enable_zb=True)
    if _profiling:
        torch.cuda.synchronize()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step3 i={i} after B1: {_mem_gb():.2f} GiB (delta={_mem_gb() - _ms3:+.2f}) | {_mem_detail()}",
            flush=True,
        )
    self._recv_forward(1)
    self._weight_chunk()
    if _profiling:
        torch.cuda.synchronize()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step3 i={i} after W1: {_mem_gb():.2f} GiB (delta={_mem_gb() - _ms3:+.2f}) | {_mem_detail()}",
            flush=True,
        )
    self._forward_chunk(1, recv=False)
    if _profiling:
        torch.cuda.synchronize()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step3 i={i} after F1: {_mem_gb():.2f} GiB (delta={_mem_gb() - _ms3:+.2f}) | {_mem_detail()}",
            flush=True,
        )

After the loop:

python
if _profiling:
    torch.cuda.synchronize()
    _m3 = _mem_gb()
    print(
        f"[rank={distributed.rank} pp={pp_rank}] After Step3 ({step_3} B1W1F1): {_m3:.2f} GiB (+{_m3 - _m0:.2f}) | {_mem_detail()}",
        flush=True,
    )

4E: Step 4 (nF0B1F1B0) — after the last self._forward_backward_chunk(1, 0) inside the loop (there is one at the end of every iteration), add:

python
    if _profiling:
        torch.cuda.synchronize()
        print(
            f"[rank={distributed.rank} pp={pp_rank}] Step4 i={i}: {_mem_gb():.2f} GiB | {_mem_detail()}",
            flush=True,
        )

After the loop:

python
if _profiling:
    torch.cuda.synchronize()
    _m4 = _mem_gb()
    print(
        f"[rank={distributed.rank} pp={pp_rank}] After Step4 ({step_4} F0B1F1B0): {_m4:.2f} GiB | {_mem_detail()}",
        flush=True,
    )

4F: Steps 5-7 — no prints needed.

4G: After Step 8 — after assert WeightGradStore.funcs_queue.empty(), before self._commit_and_wait_comm():

python
if _profiling:
    torch.cuda.synchronize()
    _m8 = _mem_gb()
    print(
        f"[rank={distributed.rank} pp={pp_rank}] After Step8 (end of pipeline): {_m8:.2f} GiB | {_mem_detail()}",
        flush=True,
    )

Group 5: Per-Layer Prints (pithtrain/pipeline/execution.py + model file)

5A: In layer_forward() (execution.py)

layer_forward(layer, hidden_states, rotary_posemb, layer_record, cu_seqlens=None) runs the five stages in order, each under its own # Stage N. comment block. At the top of the function, before # Stage 1., add:

python
_do_lmem = _layer_mem_profile and layer.idx in DETAIL_LAYERS
_do_saved = _layer_mem_profile and layer.idx in DETAIL_LAYERS

_wt_ptrs: set = set()
if _do_saved:
    _wt_ptrs = {p.data_ptr() for p in layer.parameters()}

Then instrument each stage:

Stage 1 — wrap the layer.forward_stage1(...) call:

python
if _do_saved:
    _prof = _SavedTensorsProfiler(layer.idx, "stage1", _wt_ptrs)
    with _prof:
        dispatch_tokens, residual, routing = layer.forward_stage1(next_hidden_states, rotary_posemb, cu_seqlens)
    _prof.print_summary()
else:
    dispatch_tokens, residual, routing = layer.forward_stage1(next_hidden_states, rotary_posemb, cu_seqlens)

After the Stage 1 record is populated: if _do_lmem: _lmem(f"layer{layer.idx} after stage1 (forward_stage1: attn + gate + dispatch_prep)")

Stage 2 — after the nvtx.range_pop() closing Stage 2: if _do_lmem: _lmem(f"layer{layer.idx} after stage2 (dispatch a2a)")

Stage 3 — same wrapping pattern as Stage 1, using _SavedTensorsProfiler(layer.idx, "stage3", _wt_ptrs) around the layer.forward_stage3(...) call:

python
if _do_saved:
    _prof = _SavedTensorsProfiler(layer.idx, "stage3", _wt_ptrs)
    with _prof:
        moe_outs = layer.forward_stage3(gathered_tokens, routing.expert_idxs if has_experts else None, routing.expand_idx if has_experts else None)
    _prof.print_summary()
else:
    moe_outs = layer.forward_stage3(gathered_tokens, routing.expert_idxs if has_experts else None, routing.expand_idx if has_experts else None)

After: if _do_lmem: _lmem(f"layer{layer.idx} after stage3 (forward_stage3)")

Stage 4 — after the nvtx.range_pop() closing Stage 4: if _do_lmem: _lmem(f"layer{layer.idx} after stage4 (combine a2a)")

Stage 5 — same wrapping pattern around the layer.forward_stage5(...) call with _SavedTensorsProfiler(layer.idx, "stage5", _wt_ptrs):

python
if _do_saved:
    _prof = _SavedTensorsProfiler(layer.idx, "stage5", _wt_ptrs)
    with _prof:
        hidden_states = layer.forward_stage5(moe_outs, moe_local_idxs, topk_weight, residual)
    _prof.print_summary()
else:
    hidden_states = layer.forward_stage5(moe_outs, moe_local_idxs, topk_weight, residual)

After: if _do_lmem: _lmem(f"layer{layer.idx} after stage5 (forward_stage5)")

5B: In model decoder layer forward_stage1 (model file)

forward_stage1 runs the compiled forward_stage1_compute helper (LN + attention + LN + router gate) and then prepare_dispatch. Add at the top:

python
from pithtrain.pipeline.execution import _layer_mem_profile, _lmem
_do = _layer_mem_profile and self.idx in DETAIL_LAYERS

Add _lmem prints:

  • Before the forward_stage1_compute(...) call: if _do: _lmem(f" layer{self.idx} stage1: before forward_stage1_compute")
  • After the forward_stage1_compute(...) call: if _do: _lmem(f" layer{self.idx} stage1: after forward_stage1_compute (LN+Attn+LN+gate)")
  • After prepare_dispatch(...): if _do: _lmem(f" layer{self.idx} stage1: after dispatch_prep dispatch={tuple(dispatch_tokens.shape)}")

The router gate runs inside the compiled forward_stage1_compute, so its allocation is captured by the "after forward_stage1_compute" probe rather than a separate print.

Show full SKILL.md (446 more words)Show less
5C: In model decoder layer forward_stage3 (model file)

Add at the top:

python
from pithtrain.pipeline.execution import _layer_mem_profile, _lmem
_do = _layer_mem_profile and self.idx in DETAIL_LAYERS

Add _lmem prints:

  • Before the padded_index_gather(gathered_tokens, expand_idx) gather (ep_size > 1 branch): if _do: _lmem(f" layer{self.idx} stage3: before expand gather gathered={tuple(gathered_tokens.shape)}")
  • After the expand gather: if _do: _lmem(f" layer{self.idx} stage3: after expand gather expanded={tuple(gathered_tokens.shape)}")
  • After scatter_for_grouped_gemm: if _do: _lmem(f" layer{self.idx} stage3: after scatter output_tokens={tuple(output_tokens.shape)}")
  • After self.mlp.experts(...): if _do: _lmem(f" layer{self.idx} stage3: after experts outs={tuple(outs.shape)}")
  • After the reverse-shuffle padded_index_gather(outs, reverse_shuffle_idxs): if _do: _lmem(f" layer{self.idx} stage3: after unshuffle outs={tuple(outs.shape)}")

forward_stage3 returns padded_index_gather(outs, reverse_shuffle_idxs) directly; assign it to outs first so the unshuffle print has a value to report before the return.

Critical: Also pass _do_mem=_do to the experts forward call. Change:

python
outs = self.mlp.experts(output_tokens, grouped_mm_offs, ks=ks, ks_tensor=ks_tensor)

to:

python
outs = self.mlp.experts(output_tokens, grouped_mm_offs, ks=ks, ks_tensor=ks_tensor, _do_mem=_do)
5D: In model experts class forward (model file)
  1. Add _do_mem: bool = False parameter to the forward() signature.
  2. Add from pithtrain.pipeline.execution import _lmem at the top.
  3. Add prints guarded by if _do_mem::
    • Before gate_proj: shape of input x
    • After gate_proj: shape of g
    • After up_proj: shape of u
    • After silu_mul(g, u): shape of gu
    • After down_proj: shape of out
  4. Refactor the return: Change return self.down_proj(silu_mul(g, u), **kwargs) to:
    python
    gu = silu_mul(g, u)
    if _do_mem:
        _lmem(f"    experts: after silu_mul    gu={tuple(gu.shape)}")
    out = self.down_proj(gu, **kwargs)
    if _do_mem:
        _lmem(f"    experts: after down_proj   out={tuple(out.shape)}")
    return out
5E: In model forward_prolog / forward_epilog (model file)

Prolog (embed) and epilog (norm + lm_head) compute live in the model's forward_prolog and forward_epilog methods, which the pipeline invokes on the first and last stage respectively. Add the import in each method that uses it:

python
from pithtrain.pipeline.execution import _layer_mem_profile, _lmem

Then:

  • In forward_prolog, after embedding the tokens (assign to a local, print, then return): if _layer_mem_profile: _lmem("after prolog (embed_tokens)")
  • In forward_epilog, at the top before self.norm(...): if _layer_mem_profile: _lmem("before epilog (norm + lm_head)")
  • In forward_epilog, after self.norm(...): if _layer_mem_profile: _lmem("after norm, before lm_head")
  • In forward_epilog, after self.lm_head(...): if _layer_mem_profile: _lmem("after lm_head")

Group 6: Checkpoint Load Prints (pithtrain/tasks/pretrain_lm.py)

In load_checkpoint, after dcp.load(...):

python
rank = torch.distributed.get_rank()
if rank in RANKS:
    torch.cuda.synchronize()
    optim_mem = sum(
        (s._local_tensor if isinstance(s, DTensor) else s).nelement()
        * (s._local_tensor if isinstance(s, DTensor) else s).element_size()
        for state in optimizer.state.values()
        for s in state.values()
        if isinstance(s, torch.Tensor)
    )
    G = 1024**3
    alloc = torch.cuda.memory_allocated()
    reserved = torch.cuda.memory_reserved()
    free, total = torch.cuda.mem_get_info()
    cached = reserved - alloc
    non_pytorch = total - free - reserved
    print(
        f"[rank={rank}] load_ckpt | optimizer state: {optim_mem / G:.2f} GiB "
        f"({len(optimizer.state)} param entries) | "
        f"alloc={alloc / G:.2f} cached={cached / G:.2f} non-pt={non_pytorch / G:.2f}",
        flush=True,
    )

Ensure DTensor is imported: from torch.distributed.tensor import DTensor. Check if it's already imported before adding.


Group 7: Memory Snapshot (pithtrain/tasks/pretrain_lm.py)

In train_step, instrument the first training step with a full memory timeline snapshot.

Before the forward/backward call (after model.train()):

python
_mem_profile = step == 0
_mem_snapshot = _mem_profile and torch.distributed.get_rank() in RANKS
if _mem_profile:
    model.memory_profiling = True
if _mem_snapshot:
    torch.cuda.memory._record_memory_history(max_entries=1048576)

Important: if the file already has torch.cuda.memory._record_memory_history for the configurable profiler (memory_profile_start), do NOT conflict with it. The snapshot code above is separate — it runs on step 0 unconditionally, while the configurable profiler starts at a user-specified step.

Wrap model.step(...) in try/except:

python
try:
    objective_outputs = model.step(microbatches, objective)
except torch.OutOfMemoryError:
    if _mem_snapshot:
        snapshot = torch.cuda.memory._snapshot()
        from pickle import dump
        snapshot_path = f"/tmp/memory_snapshot_rank{torch.distributed.get_rank()}.pickle"
        with open(snapshot_path, "wb") as f:
            dump(snapshot, f)
        print(f"[rank={torch.distributed.get_rank()}] Memory snapshot saved to {snapshot_path}")
        torch.cuda.memory._record_memory_history(enabled=None)
    raise

After the model.step call (before optimizer step):

python
if _mem_snapshot:
    snapshot = torch.cuda.memory._snapshot()
    from pickle import dump
    snapshot_path = f"/tmp/memory_snapshot_rank{torch.distributed.get_rank()}_step0.pickle"
    with open(snapshot_path, "wb") as f:
        dump(snapshot, f)
    print(f"[rank={torch.distributed.get_rank()}] Memory snapshot saved to {snapshot_path}")
    torch.cuda.memory._record_memory_history(enabled=None)

if _mem_profile:
    if hasattr(model, "memory_profiling"):
        model.memory_profiling = False

Execution Order

  1. pithtrain/pipeline/execution.py — Groups 1B, 5A
  2. pithtrain/pipeline/dualpipev.py — Groups 1A, 4
  3. pithtrain/models/<model>.py — Groups 5B, 5C, 5D, 5E
  4. pithtrain/modules/distributed.py — Group 2
  5. pithtrain/modules/training.py — Groups 1C, 3
  6. pithtrain/tasks/pretrain_lm.py — Groups 6, 7

After Completing

Run ruff check --fix and ruff format on all modified files to ensure code style compliance.

© mlc-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/add-memory-prints of mlc-ai/pith-train.

Open the folder on GitHubat commit c7c8b1d

Compare with similar skills

Add Memory Prints next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Add Memory Prints compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Add Memory Prints this skillmlc-ai/pith-train355—~6.2kAutomated safety check: PassApache-2.0
Agent BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
LLM Benchmarking with lm-evaluation-harnessOrchestra-Research/AI-Research-SKILLs13k8 repos~3kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.9k14 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.9k GitHub starsUsed in 14 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Planning With Files

    jarrodwatts/claude-code-config

    Transforms workflow to use Manus-style persistent markdown files for planning, progress tracking, and knowledge storage.

    1.1k GitHub starsUsed in 5 repos~967 tokens
    AI & LLM EngineeringAuto-check passed

More from mlc-ai/pith-train

All 10 skills in this repo
  • Analyze Nsys Profile

    mlc-ai/pith-train

    Query a captured PithTrain Nsight Systems profile to measure compute/communication overlap, locate exposed comm by DualPipeV stage, and inspect per-rank stream behavior.

    355 GitHub stars~1.9k tokensUpdated 4 days ago
    Auto-check passed
  • Capture Nsys Profile

    mlc-ai/pith-train

    Capture a Nsight Systems (.nsys-rep) profile of a short PithTrain run for performance analysis.

    355 GitHub stars~1k tokensUpdated 4 days ago
    Auto-check passed
  • Validate Correctness

    mlc-ai/pith-train

    Validates that code changes do not break training correctness by comparing loss deltas against a base-vs-base run-to-run envelope.

    355 GitHub stars~2.3k tokensUpdated 4 days ago
    Auto-check passed
  • Validate Performance

    mlc-ai/pith-train

    Measures the throughput difference between two branches with force-balanced routing.

    355 GitHub stars~1.4k tokensUpdated 4 days ago
    Auto-check passed
  • Setup Benchmark Inputs

    mlc-ai/pith-train

    Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain.

    355 GitHub stars~399 tokensUpdated 4 days ago
    Auto-check passed
  • Add New Model

    mlc-ai/pith-train

    Adds support for a new MoE language model to PithTrain. An agent skill from mlc-ai/pith-train.

    355 GitHub stars~4.6k tokensUpdated 4 days ago
    Auto-check passed

Questions about Add Memory Prints

What does Add Memory Prints do?

Add detailed memory profiling prints throughout the training framework. Add Memory Prints is an agent skill from mlc-ai/pith-train. Add detailed memory profiling prints throughout the training framework.

When should I use Add Memory Prints?

Add Memory Prints fits situations like: user asks to add memory prints; instrument memory; memory breakdown.

How do I install Add Memory Prints in Claude Code?

Run `npx skills add mlc-ai/pith-train --skill add-memory-prints -a claude-code`. Or copy the skill folder (.agents/skills/add-memory-prints in mlc-ai/pith-train) into .claude/skills/add-memory-prints in your project. Claude Code loads it when a task matches its description.

How do I install Add Memory Prints in Codex?

Run `npx skills add mlc-ai/pith-train --skill add-memory-prints -a codex`. Or copy the skill folder (.agents/skills/add-memory-prints in mlc-ai/pith-train) into .agents/skills/add-memory-prints in your project. Codex loads it when a task matches its description.

Can I use Add Memory Prints in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mlc-ai/pith-train --skill add-memory-prints -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/add-memory-prints, .gemini/skills/add-memory-prints, .github/skills/add-memory-prints and .opencode/skills/add-memory-prints in your project.

What does Add Memory Prints need to run?

Going by SKILL.md and its folder, Add Memory Prints needs the command-line tools its instructions call (ruff). Our summary lists: Python 3.

Does Add Memory Prints access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Add Memory Prints safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Add Memory Prints use?

Add Memory Prints is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Add Memory Prints use?

About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Add Memory Prints?

Skills that share tags, products or a category with Add Memory Prints: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), LLM Benchmarking with lm-evaluation-harness (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Add Memory Prints?

mlc-ai (a GitHub organization) maintains it in mlc-ai/pith-train, which has 355 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 4, 2026.

Source: mlc-ai/pith-train on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.