Agent skill

Paddle Design Distributed

by PaddlePaddle in PaddlePaddle/Paddle

A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Paddle Design Distributed

skills CLI
$ npx skills add PaddlePaddle/Paddle --skill paddle-design-distributed -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PaddlePaddle/Paddle paddle-design-distributed --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PaddlePaddle/Paddle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/paddle-design-distributed .claude/skills/paddle-design-distributed && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paddle-design-distributed
GitHub stars
24k
Token cost
~660 tokens
SKILL.md length
152 words
Files
3 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

  • Works in 3 steps: 推导 Einsum 表示:matmul(X[M,K], Y[K,N]) ->… → 合并输入 dims_mapping:同一 Einsum 轴的切分必须一致,冲突时… → 推导输出 dims_mapping:根据合并后的轴切分信息映射到输出
  • Working with Paddles distributed training system: understanding parallelism strategies (DP
  • SKILL.md covers 分布式范式速查, 三种编程范式, SPMD 推导规则 and Pipeline 调度, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Paddle Design Distributed is an agent skill from PaddlePaddle/Paddle. Use when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor, SPMD inference rules, pipeline scheduling, or autoparallel Engine.

Its SKILL.md is about 660 tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/distributed-primer.md` and `references/spmd-rules.md`).

It sits in AI & LLM Engineering, covering Deep learning. It works with Python. The repository describes itself as: PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署). The licence is Apache-2.0.

When your agent uses it

  • Working with Paddles distributed training system: understanding parallelism strategies (DP
  • Semi-automatic parallel with ProcessMesh + shardtensor
  • SPMD inference rules
  • Pipeline scheduling

Example prompts

  • “/paddle-design-distributed”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. 推导 Einsum 表示:matmul(X[M,K], Y[K,N]) -> Z[M,N] → mk,kn->mn
  2. 合并输入 dims_mapping:同一 Einsum 轴的切分必须一致,冲突时 reshard
  3. 推导输出 dims_mapping:根据合并后的轴切分信息映射到输出

What it can do on your machine

Read from SKILL.md and the folder at commit 5434c21. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paddle Design Distributed loads about 660 tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 152 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~660
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PaddlePaddle/Paddle at commit 5434c21, republished under its Apache-2.0 licence (© PaddlePaddle). 152 words, ~660 tokens.

Download SKILL.mdSave it as .claude/skills/paddle-design-distributed/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
paddle-design-distributed
description
Use when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shard_tensor, SPMD inference rules, pipeline scheduling, or auto_parallel Engine.

Paddle 分布式训练

分布式范式速查

范式核心思想通信原语
Data Parallel复制模型,切分数据,AllReduce 梯度AllReduce
Group Sharded (ZeRO)Stage1 切 optimizer / Stage2 + 切 grad / Stage3 + 切 weightBroadcast, ReduceScatter, AllGather
Model Parallel (Tensor)Column Parallel 切权重列 / Row Parallel 切权重行AllReduce / AllGather
Pipeline ParallelF-then-B / 1F1B 交错前反向Send / Recv (P2P)
Sequence Parallel沿 sequence 维度切分 LayerNorm/DropoutAllGather / ReduceScatter

三种编程范式

范式入口适用场景
手动并行fleet.meta_parallel灵活但代码量大,适合深度定制
半自动动态图ProcessMesh + shard_tensor用户标注切分方式,框架自动推导通信,兼具易用性和灵活性
半自动静态图auto_parallel.Engine基于 PIR 做全局优化,追求极致性能
半自动动态图示例
python
import paddle.distributed as dist

mesh = dist.ProcessMesh([0, 1, 2, 3], dim_names=["x"])
x = dist.shard_tensor(x, mesh, [dist.Shard(0)])  # 沿 dim 0 切分

SPMD 推导规则

半自动并行的核心:用户只标注部分 Tensor 的切分方式(Placement),框架自动推导每个算子的输入输出应如何切分,必要时插入通信算子。

基于 Einsum 的计算类规则
  1. 推导 Einsum 表示:matmul(X[M,K], Y[K,N]) -> Z[M,N] → mk,kn->mn
  2. 合并输入 dims_mapping:同一 Einsum 轴的切分必须一致,冲突时 reshard
  3. 推导输出 dims_mapping:根据合并后的轴切分信息映射到输出
形状变换类规则

使用 DimTrans 系统描述维度映射:InputDim / Flatten / Split / Singleton。

Pipeline 调度

调度策略特点
F-then-B所有前向完毕再反向,简单但显存占用大
1F1Bwarm-up → steady(前反交替)→ cool-down,显存减少约 37.5%

什么场景看什么文件

场景参考文档
分布式策略原理(DP/Sharded/MP/PP/SP)references/distributed-primer.md
SPMD 推导规则与 Pipeline 调度references/spmd-rules.md

源码入口

模块路径
分布式总目录python/paddle/distributed/
ProcessMesh / shard_tensor APIpython/paddle/distributed/auto_parallel/
auto_parallel Enginepython/paddle/distributed/auto_parallel/high_level_api.py
fully_shard(ZeRO-like)python/paddle/distributed/auto_parallel/fully_shard.py
SPMD 推导规则 C++paddle/phi/infermeta/spmd_rules/
SPMD 规则注册paddle/phi/infermeta/spmd_rules/rules.h
Pipeline 调度 Passpython/paddle/distributed/passes/pipeline_scheduler_pass/
分布式策略配置python/paddle/distributed/auto_parallel/strategy.py
SPMD 规则测试test/auto_parallel/spmd_rules/

© PaddlePaddle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/paddle-design-distributed of PaddlePaddle/Paddle.

  • SKILL.md
  • references/distributed-primer.md
  • references/spmd-rules.md

Open the folder on GitHubat commit 5434c21

Compare with similar skills

Paddle Design Distributed next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paddle Design Distributed compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paddle Design Distributed this skillPaddlePaddle/Paddle24k—~660Automated safety check: PassApache-2.0
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k8 repos~3.3kAutomated safety check: PassMIT
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k7 repos~1.7kAutomated safety check: PassMIT
Onnxtxtonnx/onnx22k—~1.3kAutomated safety check: PassApache-2.0
Scaffold Examplecomet-ml/comet-examples175—~1kAutomated safety check: PassNone
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Onnxtxt

    onnx/onnx

    Read or write ONNX text format ("onnxtxt"). An agent skill from onnx/onnx.

    22k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Scaffold Example

    comet-ml/comet-examples

    Scaffold a brand-new Comet example in this repo from the canonical template under templates/integration-example/.

    175 GitHub stars~1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from PaddlePaddle/Paddle

All 11 skills in this repo
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated today
    Auto-check passed
  • Paddle Eager Graph

    PaddlePaddle/Paddle

    A skill your agent uses when navigating Paddle eager-mode (dynamic graph) source code, tracing forward/backward execution, debugging autograd issues, understanding PyLayer, or investigating…

    24k GitHub stars~562 tokensUpdated today
    Auto-check passed
  • Paddle Phi Kernel

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's PHI kernel system: registering new kernels, debugging kernel selection/dispatch, understanding code auto-generation from YAML, or implementing…

    24k GitHub stars~656 tokensUpdated today
    Auto-check passed
  • Paddle Debug

    PaddlePaddle/Paddle

    在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

    24k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Paddle Design Distributed

What does Paddle Design Distributed do?

A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…. Paddle Design Distributed is an agent skill from PaddlePaddle/Paddle. Use when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor, SPMD inference rules, pipeline scheduling, or autoparallel Engine.

When should I use Paddle Design Distributed?

Paddle Design Distributed fits situations like: working with Paddles distributed training system: understanding parallelism strategies (DP; semi-automatic parallel with ProcessMesh + shardtensor; SPMD inference rules; pipeline scheduling.

How do I install Paddle Design Distributed in Claude Code?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-design-distributed -a claude-code`. Or copy the skill folder (.agents/skills/paddle-design-distributed in PaddlePaddle/Paddle) into .claude/skills/paddle-design-distributed in your project. Claude Code loads it when a task matches its description.

How do I install Paddle Design Distributed in Codex?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-design-distributed -a codex`. Or copy the skill folder (.agents/skills/paddle-design-distributed in PaddlePaddle/Paddle) into .agents/skills/paddle-design-distributed in your project. Codex loads it when a task matches its description.

Can I use Paddle Design Distributed in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PaddlePaddle/Paddle --skill paddle-design-distributed -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paddle-design-distributed, .gemini/skills/paddle-design-distributed, .github/skills/paddle-design-distributed and .opencode/skills/paddle-design-distributed in your project.

What does Paddle Design Distributed need to run?

SKILL.md names no scripts, command-line tools or credentials: Paddle Design Distributed is instructions for the agent only. Our summary lists: Python 3.

Does Paddle Design Distributed access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Paddle Design Distributed safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paddle Design Distributed use?

Paddle Design Distributed is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paddle Design Distributed use?

About 660 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.7k tokens, read only when the agent opens those files.

What are the alternatives to Paddle Design Distributed?

Skills that share tags, products or a category with Paddle Design Distributed: Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Onnxtxt (onnx/onnx, 22k stars) and Scaffold Example (comet-ml/comet-examples, 175 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paddle Design Distributed?

PaddlePaddle (a GitHub organization) maintains it in PaddlePaddle/Paddle, which has 24,121 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 9, 2026.

Source: PaddlePaddle/Paddle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.