Docstring
pytorch/pytorch
Write docstrings for PyTorch functions and methods following PyTorch conventions.
Multi-GPU setup for PerforatedAI with DataParallel or DistributedDataParallel (DDP).
$ npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install PerforatedAI/PerforatedAI perforatedai-distributed --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/PerforatedAI/PerforatedAI.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/perforatedai-distributed .claude/skills/perforatedai-distributed && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "perforatedai-distributed" agent skill from https://github.com/PerforatedAI/PerforatedAI/tree/main/skills/perforatedai-distributed into .claude/skills/perforatedai-distributed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "perforatedai-distributed", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/PerforatedAI/PerforatedAI/tree/main/skills/perforatedai-distributedType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install PerforatedAI/PerforatedAI perforatedai-distributed --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PerforatedAI/PerforatedAI.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/perforatedai-distributed .agents/skills/perforatedai-distributed && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "perforatedai-distributed" agent skill from https://github.com/PerforatedAI/PerforatedAI/tree/main/skills/perforatedai-distributed into .agents/skills/perforatedai-distributed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "perforatedai-distributed", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install PerforatedAI/PerforatedAI perforatedai-distributed --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PerforatedAI/PerforatedAI.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/perforatedai-distributed .cursor/skills/perforatedai-distributed && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "perforatedai-distributed" agent skill from https://github.com/PerforatedAI/PerforatedAI/tree/main/skills/perforatedai-distributed into .cursor/skills/perforatedai-distributed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "perforatedai-distributed", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/PerforatedAI/PerforatedAI.git --path skills/perforatedai-distributed--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install PerforatedAI/PerforatedAI perforatedai-distributed --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PerforatedAI/PerforatedAI.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/perforatedai-distributed .gemini/skills/perforatedai-distributed && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "perforatedai-distributed" agent skill from https://github.com/PerforatedAI/PerforatedAI/tree/main/skills/perforatedai-distributed into .gemini/skills/perforatedai-distributed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "perforatedai-distributed", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install PerforatedAI/PerforatedAI perforatedai-distributedInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/PerforatedAI/PerforatedAI.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/perforatedai-distributed .github/skills/perforatedai-distributed && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "perforatedai-distributed" agent skill from https://github.com/PerforatedAI/PerforatedAI/tree/main/skills/perforatedai-distributed into .github/skills/perforatedai-distributed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "perforatedai-distributed", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install PerforatedAI/PerforatedAI perforatedai-distributed --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/PerforatedAI/PerforatedAI.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/perforatedai-distributed .opencode/skills/perforatedai-distributed && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "perforatedai-distributed" agent skill from https://github.com/PerforatedAI/PerforatedAI/tree/main/skills/perforatedai-distributed into .opencode/skills/perforatedai-distributed/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "perforatedai-distributed", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
perforatedai-distributedMulti-GPU setup for PerforatedAI with DataParallel or DistributedDataParallel (DDP).
Perforatedai Distributed is an agent skill from PerforatedAI/PerforatedAI. Multi-GPU setup for PerforatedAI with DataParallel or DistributedDataParallel (DDP). Invoked automatically by the perforatedai skill when multi-GPU training is detected. Handles initialization workflow, checkpoint loading, rank 0 handling, and shell script generation for DDP.
Its SKILL.md is about 5.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Deep learning and Shell scripting. It works with Bash and PyTorch. The repository describes itself as: Add Dendrites to your PyTorch Project. The licence is Apache-2.0.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 9d317e6. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Perforatedai Distributed loads about 5.5k tokens when it runs. Until then it costs about 75 tokens; SKILL.md has 1,411 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from PerforatedAI/PerforatedAI at commit 9d317e6, republished under its Apache-2.0 licence (© PerforatedAI). 1,411 words, ~5,489 tokens.
.claude/skills/perforatedai-distributed/SKILL.md (or your agent's skills folder).This skill handles DataParallel and DistributedDataParallel (DDP) setup for PerforatedAI. It is automatically invoked by the main perforatedai skill when multi-GPU training is detected.
Prerequisites:
UPA.perforate_model() has been addedUse this section when user is using torch.nn.DataParallel.
Educate them about the additional complexity:
Tell them:
Important: DataParallel Setup with PerforatedAI
DataParallel with PerforatedAI requires a simple two-step initialization process:
- First run: Initialize multi-GPU settings (exits after one batch)
- Second run: Actual training with multi-GPU support
This is a one-time setup. After initialization, you just run your training normally.
Steps:
Find their argparse section (or add one if they don't have it) and add:
parser.add_argument('--perforate_model_parallel', action='store_true',
help='Initialize PAI settings for multi-GPU (run once on single GPU)')Find their DataParallel line:
model = torch.nn.DataParallel(model)Replace it with:
# PAI multi-GPU setup - initialize on single GPU, then use DataParallel
if not args.perforate_model_parallel:
# Normal training mode - load settings and wrap with DataParallel
GPA.pai_tracker.initialize_tracker_settings()
model = torch.nn.DataParallel(model)
# else: initialization mode - stay on single GPU, will save settings after first batch2.5. Setup Optimizer and Scheduler:
Find where their optimizer and scheduler are currently defined in their script.
🚨 CRITICAL RULE: PRESERVE USER'S EXACT OPTIMIZER AND SCHEDULER TYPES AND ARGUMENTS
Replace their optimizer/scheduler code with the PAI pattern while keeping the EXACT SAME types and arguments.
Example: Original:
optimizer = torch.optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-4)
scheduler = torch.optim.lr_scheduler.StepLR(optimizer, step_size=30, gamma=0.1)Correct PAI conversion (PRESERVES their choices):
GPA.pai_tracker.set_optimizer(torch.optim.Adam) # SAME optimizer type
GPA.pai_tracker.set_scheduler(torch.optim.lr_scheduler.StepLR) # SAME scheduler type
optimArgs = {'params': model.parameters(), 'lr': 0.001, 'weight_decay': 1e-4} # SAME args
schedArgs = {'step_size': 30, 'gamma': 0.1} # SAME scheduler args
optimizer, scheduler = GPA.pai_tracker.setup_optimizer(model, optimArgs, schedArgs)⚠️ IMPORTANT: This must be added INSIDE the batch/iteration loop, immediately after loss.backward(), so it exits after processing just ONE batch (not after a full epoch).
Find their batch loop (usually for batch in train_loader: or similar) and add this code immediately after loss.backward():
# Inside the batch loop:
for batch in train_loader: # Their loop
# ... their training code ...
loss.backward()
# PAI initialization mode - save settings and exit after FIRST batch
if args.perforate_model_parallel:
GPA.pai_tracker.save_tracker_settings()
print(f"PAI multi-GPU settings saved to {{save_name}}/")
print("Initialization complete. Now run without --perforate_model_parallel flag for training.")
exit(0)
optimizer.step() # Their code continues
# ...Optional: Add gradient pre-initialization (if encountering gradient-related errors):
DataParallel typically doesn't require this, but if they encounter errors about undefined gradients, add this AFTER optimizer.zero_grad() and BEFORE loss.backward():
optimizer.zero_grad()
# Pre-initialize gradients for multi-GPU compatibility (optional)
for param in model.parameters():
if param.requires_grad and param.grad is None:
param.grad = torch.zeros_like(param)
loss.backward()
optimizer.step()Tell them:
"I've set up your script for DataParallel. To train:
- First run:
python train.py --perforate_model_parallel(runs on single GPU, exits after ONE batch)- Second run:
python train.py(uses DataParallel for multi-GPU training)- If you change any PAI configuration settings later, re-run step 1"
Use this section when user is using torch.nn.parallel.DistributedDataParallel.
Educate them about the additional complexity:
Tell them:
Important: DistributedDataParallel Setup with PerforatedAI
DistributedDataParallel with PerforatedAI requires a special workflow because when dendrites are added, the process needs to exit and restart. I'll create a shell script that automates this entire process:
- The script handles initialization and continuous training automatically
- When training is interrupted (e.g., dendrites added), it automatically restarts and resumes
- You just run one command and let it handle everything
Steps:
Find their argparse section (or add one if they don't have it) and add:
parser.add_argument('--perforate_model_parallel', action='store_true',
help='Initialize PAI settings for DDP (run once)')
parser.add_argument('--pai_load_folder', type=str, default=None,
help='Folder to load PAI state from (for automatic resumption)')⚠️ CRITICAL: This must be added RIGHT AFTER model creation and perforate_model(), BEFORE creating the optimizer.
Find where they create their model and call perforate_model:
model = YourModel(...)
model = UPA.perforate_model(model, save_name="...", maximizing_score=...)
model = model.to(device)Add this loading logic immediately after (still BEFORE any optimizer creation):
# Load checkpoint if resuming from dendrite restructure (DDP mode)
if args.pai_load_folder is not None:
# Find the highest numbered switch_x.pt file
import glob
switch_files = glob.glob(f"{args.pai_load_folder}/switch_*.pt")
if switch_files:
# Extract switch numbers and find the maximum
switch_numbers = []
for f in switch_files:
try:
num = int(f.split('switch_')[1].split('.pt')[0])
switch_numbers.append(num)
except:
pass
if switch_numbers:
max_switch = max(switch_numbers)
model = UPA.load_system(model, args.pai_load_folder, f'switch_{max_switch}', True)
print(f"Loaded PAI state from {args.pai_load_folder}/switch_{max_switch}.pt")
else:
print(f"Starting from beginning (no valid switch_x.pt found in {args.pai_load_folder})")
else:
print(f"Starting from beginning (no switch_x.pt files found in {args.pai_load_folder})")
# NEXT: Add optimizer setup HERE (see substep 2.5 below)
# THEN: Add DDP wrapper (see substep 3 below)Why this order matters: When dendrites are added, the model structure changes. load_system loads the new structure. The optimizer needs to be created with the correct model parameters, so loading must happen BEFORE optimizer creation.
⚠️ CRITICAL ORDERING FOR YOUR SCRIPT:
Substep 2.5: Setup Optimizer and Scheduler (ADD BETWEEN CHECKPOINT LOADING AND DDP WRAPPER)
Find where their optimizer and scheduler are currently defined in their script.
🚨 CRITICAL RULE: PRESERVE USER'S EXACT OPTIMIZER AND SCHEDULER TYPES AND ARGUMENTS
If their setup is clean (2-5 lines in one place):
Replace their optimizer/scheduler code with the PAI pattern while keeping the EXACT SAME types and arguments.
Example - User has Adam optimizer with StepLR scheduler: Original:
optimizer = torch.optim.Adam(model.parameters(), lr=0.001, weight_decay=1e-4)
scheduler = torch.optim.lr_scheduler.StepLR(optimizer, step_size=30, gamma=0.1)Correct PAI conversion (PRESERVES their choices):
GPA.pai_tracker.set_optimizer(torch.optim.Adam) # SAME optimizer type
GPA.pai_tracker.set_scheduler(torch.optim.lr_scheduler.StepLR) # SAME scheduler type
optimArgs = {'params': model.parameters(), 'lr': 0.001, 'weight_decay': 1e-4} # SAME args
schedArgs = {'step_size': 30, 'gamma': 0.1} # SAME scheduler args
optimizer, scheduler = GPA.pai_tracker.setup_optimizer(model, optimArgs, schedArgs)⚠️ CRITICAL: Place this code in their script RIGHT AFTER the checkpoint loading block from substep 2, BEFORE the DDP wrapper from substep 3.
If their setup is complex (scattered across many lines or conditional):
Ask the user: "Your optimizer/scheduler setup is complex. Can you point me to where you want PAI's optimizer setup to go, and confirm which optimizer and scheduler types you want to use?"
Find their DDP initialization:
model = torch.nn.parallel.DistributedDataParallel(model, ...)Replace it with:
# PAI DDP setup - initialize on single GPU, then use DDP
if not args.perforate_model_parallel:
# Normal training mode - initialize settings and wrap with DDP
GPA.pai_tracker.initialize_tracker_settings()
# Wrap with DDP - IMPORTANT: find_unused_parameters=True is required for PerforatedAI
model = torch.nn.parallel.DistributedDataParallel(
model,
device_ids=[...], # Their original device_ids if they had them
find_unused_parameters=True # Required for PerforatedAI dendrite growth
)
# else: initialization mode - stay on single GPU, will save settings after first batchIMPORTANT: If their original DDP call had other arguments (like device_ids, output_device, broadcast_buffers, etc.), preserve those arguments and add find_unused_parameters=True to them.
⚠️ CRITICAL: This fixes a DDP crash caused by PerforatedAI's selective training.
Problem: PerforatedAI's selective training (Cascade Correlation) means some parameters intentionally don't receive gradients during backward(). DDP requires ALL parameters to have gradients during allreduce, or it crashes with: RuntimeError: Encountered gradient which is undefined, but still allreduced by DDP reducer.
Solution: Pre-initialize all parameter gradients to zeros AFTER optimizer.zero_grad() and BEFORE loss.backward().
🔴 EXACT PLACEMENT REQUIREMENT:
Find their training loop and locate this sequence:
optimizer.zero_grad()
loss.backward()
optimizer.step()Insert the gradient pre-initialization code BETWEEN optimizer.zero_grad() and loss.backward():
optimizer.zero_grad()
# Pre-initialize gradients for DDP compatibility
# PerforatedAI's selective training means some params won't get gradients from backward()
# Initialize them to zero so DDP's allreduce doesn't encounter None
if args.distributed:
for param in model.parameters():
if param.requires_grad and param.grad is None:
param.grad = torch.zeros_like(param)
loss.backward()
optimizer.step()Why this placement:
optimizer.zero_grad() clears old gradients (prevents memory leak)loss.backward() accumulates real gradients onto the zeros (correct behavior)⚠️ If they use AMP (Automatic Mixed Precision):
They might have two code paths (with/without scaler). Add the same fix to BOTH:
optimizer.zero_grad()
# Pre-initialize gradients for DDP compatibility
if args.distributed:
for param in model.parameters():
if param.requires_grad and param.grad is None:
param.grad = torch.zeros_like(param)
if scaler is not None:
scaler.scale(loss).backward()
scaler.step(optimizer)
scaler.update()
else:
loss.backward()
optimizer.step()⚠️ IMPORTANT: This must be added INSIDE the batch/iteration loop, immediately after loss.backward(), so it exits after processing just ONE batch (not after a full epoch).
Find their batch loop (usually for batch in train_loader: or similar) and add this code immediately after loss.backward():
# Inside the batch loop:
for batch in train_loader: # Their loop
# ... their training code ...
loss.backward()
# PAI initialization mode - save settings and exit after FIRST batch
if args.perforate_model_parallel:
GPA.pai_tracker.save_tracker_settings()
print(f"PAI DDP settings saved to {{save_name}}/")
print("Initialization complete. Now run without --perforate_model_parallel flag for training.")
# Barrier ensures all ranks finish before any exit
torch.distributed.barrier()
exit(0)
optimizer.step() # Their code continues
# ...⚠️ CRITICAL: DistributedDataParallel requires special handling for PAI tracker functions.
In the section where they call add_validation_score (and any add_extra_score calls), replace it with this DDP-aware pattern:
# After validation completes and you have val_score/val_acc/val_loss
# Support both DDP and single GPU modes
if args.distributed:
# DDP mode - only rank 0 calls PAI tracker functions
if torch.distributed.get_rank() == 0:
# Unwrap model from DDP wrapper
model_unwrapped = model.module
model_unwrapped, restructured, training_complete = GPA.pai_tracker.add_validation_score(val_acc, model_unwrapped)
# If adding extra scores, do those here too:
# model_unwrapped = GPA.pai_tracker.add_extra_score(extra_score, model_unwrapped, score_name="extra_metric")
# Re-wrap the updated model
model.module = model_unwrapped
model.module = model.module.to(device)
else:
# Other ranks skip PAI calls
restructured = False
training_complete = False
# Broadcast restructured and training_complete from rank 0 to all ranks
restructured_tensor = torch.tensor([1 if restructured else 0], dtype=torch.int, device=device)
training_complete_tensor = torch.tensor([1 if training_complete else 0], dtype=torch.int, device=device)
torch.distributed.broadcast(restructured_tensor, src=0)
torch.distributed.broadcast(training_complete_tensor, src=0)
restructured = bool(restructured_tensor.item())
training_complete = bool(training_complete_tensor.item())
else:
# Single GPU mode
model, restructured, training_complete = GPA.pai_tracker.add_validation_score(val_acc, model)
model = model.to(device)
# Handle training completion
if training_complete:
print("PAI training complete!")
if args.distributed:
# Create completion marker for shell script (rank 0 only, but all ranks will exit)
if torch.distributed.get_rank() == 0:
os.makedirs(save_name, exist_ok=True)
with open(f"{save_name}/.training_complete", "w") as f:
f.write("complete")
# Barrier ensures file write completes before process group destruction
torch.distributed.barrier()
# Clean up distributed process group (all ranks)
torch.distributed.destroy_process_group()
sys.exit(0)
# Handle model restructuring (dendrite added)
elif restructured:
print("Model restructured! Exiting for restart...")
if args.distributed:
# Barrier ensures all ranks are ready before process group destruction
torch.distributed.barrier()
# Clean up distributed process group (all ranks)
torch.distributed.destroy_process_group()
exit(0) # Shell script will restart training automaticallyIMPORTANT NOTES:
args.distributed flagval_acc with their actual validation metric variable namesave_name with their actual save folder variableadd_validation_score() and add_extra_score() calls.module) before PAI callstorch.distributed.barrier() calls are mandatory before destroy_process_group() to prevent file corruption - barriers ensure all ranks complete their I/O operations before any process exitsdestroy_process_group() on all ranks when exitingBefore creating the shell script, ask:
Required:
Optional (only ask if relevant):
torchrun to launch DDP, or a different launcher like python -m torch.distributed.launch?"Wait for their answers.
Based on their answers, create a file train_distributed.sh in the same directory.
If they're using torchrun (most common):
#!/bin/bash
# PerforatedAI DistributedDataParallel Training Script
# This script handles automatic restarting when dendrites are added
SAVE_NAME="[their_save_name]" # Use the save_name from UPA.perforate_model
PYTHON_SCRIPT="[their_script_name]" # Their actual script filename
NUM_GPUS=[their_num_gpus] # Number of GPUs they specified
echo "Step 1: Initializing PAI DDP settings..."
# Run initialization on single GPU (no DDP launcher)
python $PYTHON_SCRIPT --perforate_model_parallel
echo ""
echo "Initialization complete. Starting continuous DDP training loop..."
echo "Press Ctrl+C to stop training"
echo ""
# Continuous training loop with DDP launcher
while true; do
# Check for completion at loop start (in case of restart after completion)
if [ -f "${SAVE_NAME}/.training_complete" ]; then
echo "Training already completed!"
break
fi
# Check if any switch_*.pt checkpoint files exist
if ls "${SAVE_NAME}"/switch_*.pt 1> /dev/null 2>&1; then
echo "Resuming training from checkpoint..."
torchrun --nproc_per_node=$NUM_GPUS $PYTHON_SCRIPT --pai_load_folder $SAVE_NAME
else
echo "Starting training from beginning..."
torchrun --nproc_per_node=$NUM_GPUS $PYTHON_SCRIPT
fi
# Check exit code
EXIT_CODE=$?
# Check if training completed successfully
if [ -f "${SAVE_NAME}/.training_complete" ]; then
echo "Training completed successfully!"
break
fi
if [ $EXIT_CODE -eq 0 ]; then
echo "Model restructured. Restarting in 2 seconds..."
sleep 2
else
echo "Error occurred. Exiting..."
exit $EXIT_CODE
fi
doneIf they specified GPU IDs:
Add this environment variable at the top of the script:
export CUDA_VISIBLE_DEVICES=[their_gpu_ids] # e.g., "0,1,2"If they're using a different launcher:
Replace the torchrun commands with their launcher. For example, if using python -m torch.distributed.launch:
python -m torch.distributed.launch --nproc_per_node=$NUM_GPUS $PYTHON_SCRIPT --pai_load_folder $SAVE_NAMEFill in the actual values:
[their_save_name] with the save_name from UPA.perforate_model() call[their_script_name] with their Python script filename[their_num_gpus] with the number they specified[their_gpu_ids] if they specified specific IDs (e.g., "0,1")Instruct the user to make it executable:
chmod +x train_distributed.shTell them:
"I've set up your script for DistributedDataParallel and created
train_distributed.shwith your configuration ([NUM_GPUS] GPUs). To run it, execute./train_distributed.shin your terminal. The script will:
- Initialize DDP settings (single GPU)
- Start continuous training with [NUM_GPUS] GPUs
- Automatically restart and resume when dendrites are added (model restructured)
- Press Ctrl+C to stop training"
Replace [NUM_GPUS] with the actual number they specified.
After completing either DataParallel or DDP setup:
✅ Multi-GPU setup is COMPLETE. You have:
📍 RETURN TO MAIN SKILL:
© PerforatedAI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/perforatedai-distributed of PerforatedAI/PerforatedAI.
Open the folder on GitHubat commit 9d317e6
Perforatedai Distributed next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Perforatedai Distributed this skillPerforatedAI/PerforatedAI | 237 | — | ~5.5k | Automated safety check: Pass | Apache-2.0 | |
| Docstringpytorch/pytorch | 104k | 2 repos | ~2.6k | Automated safety check: Pass | Custom licence | |
| CI Metricspytorch/pytorch | 104k | — | ~1.1k | Automated safety check: Pass | Custom licence | |
| Cuda Index Widthpytorch/pytorch | 104k | — | ~1.6k | Automated safety check: Pass | Custom licence | |
| Benchmark Pyreflyfacebook/pyrefly | 7.1k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Document Public APIspytorch/pytorch | 104k | — | ~4.2k | Automated safety check: Pass | Custom licence |
pytorch/pytorch
Write docstrings for PyTorch functions and methods following PyTorch conventions.
pytorch/pytorch
Query PyTorch CI, GitHub Actions, HUD, Grafana, and infrastructure metrics.
pytorch/pytorch
Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.
facebook/pyrefly
Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.
pytorch/pytorch
Document undocumented public APIs in PyTorch by removing functions from coverageignorefunctions and coverageignoreclasses in docs/source/conf.py, running Sphinx coverage, and adding the appropriate…
linkedin/Liger-Kernel
Develops production-ready Triton kernels for Liger Kernel. An agent skill from linkedin/Liger-Kernel.
PerforatedAI/PerforatedAI
Solutions for non-trivial PerforatedAI integration scenarios.
PerforatedAI/PerforatedAI
HuggingFace Transformers integration for PerforatedAI. An agent skill from PerforatedAI/PerforatedAI.
PerforatedAI/PerforatedAI
Render a single-panel PAI figure of score versus parameter count from sweep CSVs, PAI run folders, or hand-supplied numbers.
PerforatedAI/PerforatedAI
Expert in PerforatedAI library for adding artificial dendrites to PyTorch neural networks.
PerforatedAI/PerforatedAI
Analyze PerforatedAI training results and provide optimization recommendations.
PerforatedAI/PerforatedAI
WandB-specific PerforatedAI integration guardrail skill. An agent skill from PerforatedAI/PerforatedAI.
Categories
Multi-GPU setup for PerforatedAI with DataParallel or DistributedDataParallel (DDP). Perforatedai Distributed is an agent skill from PerforatedAI/PerforatedAI. Multi-GPU setup for PerforatedAI with DataParallel or DistributedDataParallel (DDP).
Perforatedai Distributed fits situations like: tasks that involve Deep learning; tasks that involve Shell scripting.
Run `npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a claude-code`. Or copy the skill folder (skills/perforatedai-distributed in PerforatedAI/PerforatedAI) into .claude/skills/perforatedai-distributed in your project. Claude Code loads it when a task matches its description.
Run `npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a codex`. Or copy the skill folder (skills/perforatedai-distributed in PerforatedAI/PerforatedAI) into .agents/skills/perforatedai-distributed in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PerforatedAI/PerforatedAI --skill perforatedai-distributed -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/perforatedai-distributed, .gemini/skills/perforatedai-distributed, .github/skills/perforatedai-distributed and .opencode/skills/perforatedai-distributed in your project.
Going by SKILL.md and its folder, Perforatedai Distributed needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Perforatedai Distributed is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.5k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Perforatedai Distributed: Docstring (pytorch/pytorch, 104k stars), CI Metrics (pytorch/pytorch, 104k stars), Cuda Index Width (pytorch/pytorch, 104k stars) and Benchmark Pyrefly (facebook/pyrefly, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
PerforatedAI (a GitHub organization) maintains it in PerforatedAI/PerforatedAI, which has 237 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 8, 2026.
Source: PerforatedAI/PerforatedAI on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.