Agent skill

Debug Distributed

by areal-project in areal-project/AReaL

Guide for debugging distributed training issues in AReaL. An agent skill from areal-project/AReaL.

Apache-2.0Auto-check passedDevelopment

Install Debug Distributed

skills CLI
$ npx skills add areal-project/AReaL --skill debug-distributed -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install areal-project/AReaL debug-distributed --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/areal-project/AReaL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/debug-distributed .claude/skills/debug-distributed && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
debug-distributed
GitHub stars
5.8k
Used in
1 other repo
Token cost
~1.6k tokens
SKILL.md length
275 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

Guide for debugging distributed training issues in AReaL. An agent skill from areal-project/AReaL.

  • Works in 4 steps: Hang Debugging (Deadlocks,… → Wrong Results (Gradient, Reduction Issues) → OOM Issues (Memory, Sharding) → …
  • User encounters hangs
  • SKILL.md covers When to Use, Debugging Principles, Step-by-Step Debugging Guide and Debugging Tools, plus 1 more section
  • Calls pip

What it does

Debug Distributed is an agent skill from areal-project/AReaL. Guide for debugging distributed training issues in AReaL. Use when user encounters hangs, wrong results, OOM, or communication errors.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Deep learning and Debugging. The repository describes itself as: The RL Bridge for LLM-based Agent Applications. Made Simple & Flexible. The licence is Apache-2.0.

When your agent uses it

  • User encounters hangs
  • Communication errors

Example prompts

  • “/debug-distributed”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Hang Debugging (Deadlocks, Synchronization)
  2. Wrong Results (Gradient, Reduction Issues)
  3. OOM Issues (Memory, Sharding)
  4. Communication Errors

What it can do on your machine

Read from SKILL.md and the folder at commit 2fad2d0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Debug Distributed loads about 1.6k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 275 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from areal-project/AReaL at commit 2fad2d0, republished under its Apache-2.0 licence (© areal-project). 275 words, ~1,591 tokens.

Download SKILL.mdSave it as .claude/skills/debug-distributed/SKILL.md (or your agent's skills folder).
name
debug-distributed
description
Guide for debugging distributed training issues in AReaL. Use when user encounters hangs, wrong results, OOM, or communication errors.

Debug Distributed Training

Debugging guide for distributed training issues in AReaL (FSDP2, TP, CP, EP).

When to Use

This skill is triggered when:

  • Training hangs or deadlocks
  • Results differ across ranks or are numerically wrong
  • OOM errors in distributed settings
  • NCCL/communication errors or device mesh issues

Debugging Principles

Minimal Reproduction

Always follow the minimal demo principle: Reproduce with the least amount of code to narrow down the issue faster.

python
# Bad: Debug in full training loop
# Good: Create minimal script
import torch
import torch.distributed as dist

dist.init_process_group("nccl")
rank = dist.get_rank()

# Reproduce the exact operation that fails
tensor = torch.ones(10).cuda()
dist.all_reduce(tensor)  # <-- Isolate the failing op
print(f"Rank {rank}: {tensor}")

Reduction strategy:

  1. Remove unrelated model components
  2. Use small tensor sizes
  3. Reduce world_size to minimum (e.g., 2 GPUs)
  4. Remove torch.compile if possible
  5. Disable activation checkpointing

Step-by-Step Debugging Guide

1. Hang Debugging (Deadlocks, Synchronization)

Environment Variables for Debugging:

bash
# Full debug logging
export TORCH_DISTRIBUTED_DEBUG=DETAIL
export NCCL_DEBUG=INFO
export NCCL_DEBUG_SUBSYS=ALL

# torch.compile debugging
export TORCH_LOGS="+dynamo,recompiles"
export TORCHDYNAMO_VERBOSE=1

Dump Call Stack with py-spy (for hung processes):

bash
# Find process IDs
ps aux | grep python

# Dump call stack of specific rank
py-spy dump --pid <PID>

# Record flame graph for performance analysis
py-spy record -o profile.svg --pid <PID> --duration 30

Common Causes:

  1. Mismatched Collectives: One rank calls all_reduce, another doesn't.
  2. Wrong Process Group: Using wrong group for collective.
  3. Tensor Shape Mismatch: Different shapes across ranks.

Debug Steps:

python
# Verify group membership
mesh = parallel_dims.get_mesh("dp_shard_cp")
group = mesh.get_group()
print(f"Rank {dist.get_rank()}: group size = {dist.get_world_size(group)}")

# Print shapes on all ranks
print(f"Rank {dist.get_rank()}: tensor.shape = {tensor.shape}")
dist.barrier()

Timeout Adjustment (for debugging only):

python
from areal.engine.core.distributed import patch_dist_group_timeout
from datetime import timedelta
patch_dist_group_timeout(timedelta(minutes=30))
2. Wrong Results (Gradient, Reduction Issues)

Check DTensor Placements:

python
from torch.distributed.tensor import DTensor
if isinstance(param, DTensor):
    print(f"Param {name}: placements={param.placements}, mesh={param.device_mesh}")

Verify Gradient Reduction:

python
for name, param in model.named_parameters():
    if param.grad is not None:
        print(f"Rank {dist.get_rank()}: {name} grad_sum = {param.grad.sum().item()}")
3. OOM Issues (Memory, Sharding)

Check Memory Usage:

python
print(f"Rank {dist.get_rank()}: "
      f"allocated={torch.cuda.memory_allocated()/1e9:.2f}GB, "
      f"reserved={torch.cuda.memory_reserved()/1e9:.2f}GB")

Check FSDP Coverage:

python
for name, param in model.named_parameters():
    is_dtensor = isinstance(param, DTensor)
    print(f"{name}: is_dtensor={is_dtensor}, shape={param.shape}")
4. Communication Errors
ErrorCauseSolution
NCCL WARN Cuda failureGPU communicationCheck NCCL version, GPU topology
RuntimeError: Timed outRank synchronizationIncrease timeout, check code paths
Invalid device meshMesh configurationVerify world_size = dp * tp * cp

Debugging Tools

Environment Variables Reference
VariablePurpose
TORCH_DISTRIBUTED_DEBUG=DETAILDetailed distributed logging
NCCL_DEBUG=INFONCCL communication logging
NCCL_DEBUG_SUBSYS=ALLAll NCCL subsystems
TORCH_LOGS="+dynamo,recompiles"torch.compile logging
TORCHDYNAMO_VERBOSE=1Dynamo verbose output
CUDA_LAUNCH_BLOCKING=1Synchronous CUDA (slow, for debugging)
py-spy for Call Stack Analysis
bash
# Install
pip install py-spy

# Dump call stack of hung process
py-spy dump --pid <PID>

# Dump all Python processes
pgrep -f python | xargs -I {} py-spy dump --pid {}

# Record flame graph
py-spy record -o profile.svg --pid <PID> --duration 30
Rank-Conditional Printing
python
def print_all_ranks(msg):
    for r in range(dist.get_world_size()):
        if dist.get_rank() == r:
            print(f"[Rank {r}] {msg}")
        dist.barrier()
Check Device Mesh
python
def debug_mesh(parallel_dims):
    mesh = parallel_dims.world_mesh
    for dim_name in mesh.mesh_dim_names:
        submesh = parallel_dims.get_mesh(dim_name)
        if submesh:
            print(f"Rank {dist.get_rank()}: {dim_name} size={submesh.size()}")
Validate Tensor Consistency
python
def check_tensor_consistency(tensor, name, group=None):
    local_sum = tensor.sum().item()
    tensor_sums = [None] * dist.get_world_size(group)
    dist.all_gather_object(tensor_sums, local_sum, group=group)
    if dist.get_rank() == 0 and len(set(tensor_sums)) > 1:
        print(f"WARNING: {name} inconsistent: {tensor_sums}")

Key Files Reference

ComponentFile
Parallel Dimsareal/experimental/models/archon/parallel_dims.py
Expert Parallelareal/experimental/models/archon/expert_parallel.py
Ulysses (CP)areal/experimental/models/archon/ulysses.py
FSDP/TP Applyareal/experimental/models/archon/qwen2/infra/parallelize.py

© areal-project, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/debug-distributed of areal-project/AReaL.

Open the folder on GitHubat commit 2fad2d0

Used in 1 other repository

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in areal-project/AReaL, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Debug Distributed next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Debug Distributed compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Debug Distributed this skillareal-project/AReaL5.8k1 repos~1.6kAutomated safety check: PassApache-2.0
Aoti Debugpytorch/pytorch104k1 repos~1.7kAutomated safety check: PassCustom licence
The Art of Debuggingstas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.0
Veomni DebugByteDance-Seed/VeOmni2.2k—~2.8kAutomated safety check: PassApache-2.0
Ascendcascend-ai-coding/awesome-ascend-skills174—~3.5kAutomated safety check: PassNone
Cuda Cpp Kernelvipshop/cache-dit1.3k—~2.3kAutomated safety check: PassApache-2.0

Similar skills

  • Aoti Debug

    pytorch/pytorch

    Debug AOTInductor (AOTI) errors and crashes. An agent skill from pytorch/pytorch.

    104k GitHub starsUsed in 1 repo~1.7k tokens
    DevelopmentAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Veomni Debug

    ByteDance-Seed/VeOmni

    A skill your agent uses for ANY bug, error, crash, wrong output, loss divergence, gradient explosion, test failure, CUDA error, distributed training hang, checkpoint load failure, or unexpected…

    2.2k GitHub stars~2.8k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Ascendc

    ascend-ai-coding/awesome-ascend-skills

    End-to-end AscendC custom operator development for Ascend NPU in an ascend-kernel (csrc/ops + build.sh + torchnpu PyTorch custom op) project.

    174 GitHub stars~3.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Cuda Cpp Kernel

    vipshop/cache-dit

    A skill your agent uses when writing, debugging, porting, reviewing, or optimizing CUDA C++ or PTX kernels; investigating CUDA Runtime or Driver API behavior; profiling kernels with Nsight Systems…

    1.3k GitHub stars~2.3k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Helps build, test and extend the Qualcomm AI Engine Direct (QNN) backend in ExecuTorch, with routes for new ops, model export, Buck-vs-CMake parity fixes and per-layer accuracy debugging.

    5.1k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from areal-project/AReaL

All 11 skills in this repo
  • Add Dataset

    areal-project/AReaL

    Guide for adding a new dataset loader to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Add Reward

    areal-project/AReaL

    Guide for adding a new reward function to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub starsUsed in 1 repo~1.2k tokens
    Auto-check passed
  • Add Archon Model

    areal-project/AReaL

    Guide for adding a new model to the Archon engine. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~4.9k tokensUpdated 4 days ago
    Auto-check passed
  • Add Unit Tests

    areal-project/AReaL

    Guide for adding unit tests to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~1.8k tokensUpdated 4 days ago
    Auto-check passed
  • Add Workflow

    areal-project/AReaL

    Guide for adding a new RolloutWorkflow to AReaL. An agent skill from areal-project/AReaL.

    5.8k GitHub stars~1.1k tokensUpdated 4 days ago
    Auto-check passed
  • Review PR

    areal-project/AReaL

    Read-only pull request review workflow with risk analysis, targeted checklists, and Codex subagent consultation.

    5.8k GitHub stars~704 tokensUpdated 4 days ago
    Auto-check passed

Questions about Debug Distributed

What does Debug Distributed do?

Guide for debugging distributed training issues in AReaL. An agent skill from areal-project/AReaL. Debug Distributed is an agent skill from areal-project/AReaL. Guide for debugging distributed training issues in AReaL.

When should I use Debug Distributed?

Debug Distributed fits situations like: user encounters hangs; communication errors.

How do I install Debug Distributed in Claude Code?

Run `npx skills add areal-project/AReaL --skill debug-distributed -a claude-code`. Or copy the skill folder (.agents/skills/debug-distributed in areal-project/AReaL) into .claude/skills/debug-distributed in your project. Claude Code loads it when a task matches its description.

How do I install Debug Distributed in Codex?

Run `npx skills add areal-project/AReaL --skill debug-distributed -a codex`. Or copy the skill folder (.agents/skills/debug-distributed in areal-project/AReaL) into .agents/skills/debug-distributed in your project. Codex loads it when a task matches its description.

Can I use Debug Distributed in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add areal-project/AReaL --skill debug-distributed -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/debug-distributed, .gemini/skills/debug-distributed, .github/skills/debug-distributed and .opencode/skills/debug-distributed in your project.

What does Debug Distributed need to run?

Going by SKILL.md and its folder, Debug Distributed needs the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Debug Distributed access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Debug Distributed safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Debug Distributed use?

Debug Distributed is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Debug Distributed use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Debug Distributed?

Skills that share tags, products or a category with Debug Distributed: Aoti Debug (pytorch/pytorch, 104k stars), The Art of Debugging (stas00/the-art-of-debugging, 1.7k stars), Veomni Debug (ByteDance-Seed/VeOmni, 2.2k stars) and Ascendc (ascend-ai-coding/awesome-ascend-skills, 174 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Debug Distributed?

areal-project (a GitHub organization) maintains it in areal-project/AReaL, which has 5,814 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 3, 2026.

Source: areal-project/AReaL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.