Agent skill

Paddle Debug

by PaddlePaddle in PaddlePaddle/Paddle

在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

Apache-2.0Auto-check passedAI & LLM Engineering

Install Paddle Debug

skills CLI
$ npx skills add PaddlePaddle/Paddle --skill paddle-debug -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PaddlePaddle/Paddle paddle-debug --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PaddlePaddle/Paddle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/paddle-debug .claude/skills/paddle-debug && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paddle-debug
GitHub stars
24k
Token cost
~1.4k tokens
SKILL.md length
364 words
Files
3 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

  • Works in 4 steps: 描述问题并构造最小复现 → 代码定位与多假设验证 → 先写问题分析报告,再做最小修复 → …
  • Tasks that involve Deep learning
  • SKILL.md covers 调试流程概览, 步骤 1:描述问题并构造最小复现, 步骤 2:代码定位与多假设验证 and 步骤 3:先写问题分析报告,再做最小修复, plus 4 more sections
  • Calls python and git

What it does

Paddle Debug is an agent skill from PaddlePaddle/Paddle. 在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/case-studies.md` and `references/cuda-debug.md`).

It sits in AI & LLM Engineering, covering Deep learning. It works with CUDA, Git and Python. The repository describes itself as: PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署). The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Deep learning

Example prompts

  • “/paddle-debug”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. 描述问题并构造最小复现
  2. 代码定位与多假设验证
  3. 先写问题分析报告,再做最小修复
  4. 利用 Git / CI 收束和巩固结论

What it can do on your machine

Read from SKILL.md and the folder at commit 4793e33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paddle Debug loads about 1.4k tokens when it runs, and up to ~8.6k if it reads all its reference files. Until then it costs about 39 tokens; SKILL.md has 364 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PaddlePaddle/Paddle at commit 4793e33, republished under its Apache-2.0 licence (© PaddlePaddle). 364 words, ~1,398 tokens.

Download SKILL.mdSave it as .claude/skills/paddle-debug/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
paddle-debug
description
在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

Paddle 仓库调试

调试流程概览

调试遵循以下步骤:

  1. 描述问题并构造最小复现
  2. 代码定位与多假设验证
  3. 先写问题分析报告,再做最小修复
  4. 利用 Git / CI 收束和巩固结论

步骤 1:描述问题并构造最小复现

用简洁的自然语言说明:

  • 触发步骤(命令、脚本、关键配置)
  • 期望行为 vs 实际行为
  • 是否只在特定环境 / 机器 / 设备 / 数据子集上出现

先确认 bug 能被稳定复现。若无法复现:

  • 检查命令是否抄错、参数是否缺失
  • 比对并对齐环境(Paddle / Python / CUDA / CUDNN / 驱动 / 显卡型号等)
  • 确认与最初出问题的环境一致后再继续

抽取独立的 Python 脚本承载问题:

  • 固定随机种子(numpy / random / paddle.seed 等)
  • 使用固定、可序列化的小数据
  • 去掉与问题无关的逻辑

目标:一条命令即可复现 python reproduce_xxx.py。

步骤 2:代码定位与多假设验证

使用工具定位代码
  • ast-grep:用于结构化代码搜索,快速定位特定代码模式
带观测点的复现

阅读报错栈和相关代码时,先列出多个可能原因假设(数据异常、shape 错误、数值不稳定、环境不一致、算子实现问题等),不要立刻改代码。

围绕假设在关键路径上加入观测点:

观测方式用途
打印与断言在关键算子调用前后,打印 Tensor 的 shape、dtype、device、数值范围(min/max/mean)
对比法对同一逻辑分别在 CPU / GPU 上运行,比较中间结果差异
版本与环境信息记录 paddle.__version__、CUDA/CUDNN 版本、驱动信息等

每完成一次带观测点的复现:

  • 基于运行时数据排除不成立的假设
  • 在更窄的范围内继续加观测点,逐步缩小问题所在的模块 / 算子 / 配置

将调试日志保存到 .paddle-agent/debug-logs/ 目录。

步骤 3:先写问题分析报告,再做最小修复

基于已有观测和对比结果,先完成问题分析报告:

markdown
# [问题标题]

## 复现方式
- 命令:
- 环境:
- 最小脚本路径:

## 现象描述
[错误信息或异常行为]

## 根因分析
[配置 / 数据 / 框架 / 算子 / 环境中的哪一处有问题]

## 关键证据
[日志片段、对比结果、重要观测点输出]

报告存放在 .paddle-agent/debug-analysis/ 目录,没有该目录请创建。

归因时考虑以下维度:

  • 接口 / 形状 / dtype:哪个 Tensor 的 shape / dtype 与预期不符
  • NaN / Inf / 数值发散:哪一层首次出现异常数值
  • 性能与显存:瓶颈在 CPU、IO 还是 GPU kernel

只有在分析结论较为充分时,才进入最小修复阶段:

  • 设计改动面尽量小的修改来验证根因
  • 先用最小复现脚本验证修复
  • 再用完整训练 / 推理脚本验证关键业务路径

步骤 4:利用 Git / CI 收束和巩固结论,最后总结保存为文件

当判断问题可能由近期提交引入时:

  • 使用 git bisect 对可疑提交范围做二分定位

对已定位的问题:

  • 补充覆盖最小复现脚本逻辑的单测
  • 留意 CI 中相关用例是否出现新增失败
  • 将最终结论沉淀到 .paddle-agent/debug-analysis/

CUDA / GPU 调试

详细的 CUDA 调试技巧:见 references/cuda-debug.md

快速参考:

bash
# 启用错误检查环境变量复现问题
export PYTHONPATH=$(pwd)/Paddle/build/python
FLAGS_check_cuda_error=1 FLAGS_use_system_allocator=1 python reproduce.py

关键点:

  • CUDA 错误通常是异步的,使用 FLAGS_check_cuda_error=1 让错误立即暴露
  • GPU kernel 调用前必须检查 numel/shape 是否为空
  • 空 Tensor(numel=0)会导致 grid size=0,触发 CUDA error(9)
  • CUDA API 返回值必须全部检查,忽略返回值会导致 sticky error 残留
  • PADDLE_ENFORCE_GPU_SUCCESS 不会调用 cudaGetLastError(),在错误路径上需手动清除
  • CUDA context 在 fork 后不可用:父进程初始化了 CUDA 后 fork 子进程(如 DataLoader worker),子进程中所有 CUDA API 调用都会返回 cudaErrorInitializationError(3)——详见 references/cuda-debug.md 中的 CUDA Fork Safety 章节

注意事项

  • 调试的第一目标是稳定复现并缩小范围,不要一开始就尝试大规模重构
  • 任何「只在某些机器上出现」的问题,优先从环境差异入手
  • 在 Paddle 仓库遇到 bug 时,优先按本 skill 流程执行,再考虑具体修复实现
算子修复注意事项
  • 前向和反向 kernel 要一并检查:反向 kernel 往往复用相同的计算逻辑,同样存在边界问题
  • 检查所有入口函数:底层公共函数可能被多个入口调用,确保边界检查在正确的层级
  • 头文件修改需完整重编:修改 .h 后需重新编译所有引用它的 .cu,并重新链接 .so
Show full SKILL.md (162 more words)Show less
CUDA API 与 Sticky Error 注意事项
  • 所有 CUDA API 返回值必须检查:包括 cudaEventSynchronize、cudaStreamSynchronize 等,忽略返回值不仅丢失错误信息,还会导致 CUDA runtime 中残留 sticky error
  • 错误路径必须清除 last error:在 PADDLE_ENFORCE_GPU_SUCCESS 抛出异常之前,手动调用 cudaGetLastError() 清除残留错误,否则 Python try/except 捕获异常后 CUDA 状态仍被污染
  • 跨测试状态污染:unittest 中一个测试的 CUDA sticky error 会影响后续所有测试,排查时需关注测试执行顺序
  • 定位 sticky error 污染源:通过逐步删减测试来二分定位产生残留错误的源头测试
CUDA Fork Safety 注意事项
  • CUDA context 在 fork 后不可用:如果父进程已初始化 CUDA(创建了 GPU tensor、调用过 CUDA API),fork 出的子进程中所有 CUDA 调用都会返回 cudaErrorInitializationError(3)
  • 典型触发场景:主进程中运行了 GPU 测试/训练后,DataLoader 使用 num_workers > 0 fork 子进程;子进程继承了父进程中 GPU tensor 的引用,GC 回收时触发 cudaFree
  • 修复模式:在 CUDA API 调用前检测 context 是否可用,对 fork 后不可用的场景做 graceful skip
  • 判断依据:cudaGetDevice() 返回 cudaErrorInitializationError(3)、cudaErrorNoDevice(100)、cudaErrorInsufficientDriver(35) 均表示 CUDA 不可用
  • 安全性:跳过 cudaFree 是安全的,因为 fork 后子进程中的 GPU 内存不属于该进程,进程退出时由 OS/driver 回收
Paddle 编译验证流程

修改 kernel 头文件后的增量编译:

bash
cd build
# 编译修改的 kernel
ninja paddle/phi/CMakeFiles/phi_gpu.dir/kernels/gpu/<kernel_name>.cu.o -j512
# 重新链接 phi_gpu
ninja phi_gpu -j512
# 重新链接 libpaddle.so
ninja paddle/fluid/pybind/libpaddle.so -j512
# 如果 Python 库未自动更新,手动复制
cp paddle/fluid/pybind/libpaddle.so python/paddle/base/libpaddle.so
.so 部署验证(关键踩坑点)

Paddle 构建产物存在两套路径,增量编译后 Python 加载的可能仍是旧版本:

构建产物路径Python 加载路径说明
build/paddle/phi/libphi_core.sobuild/python/paddle/libs/libphi_core.sophi core 库
build/paddle/phi/libphi_gpu.sobuild/python/paddle/libs/libphi_gpu.sophi GPU 库
build/paddle/fluid/pybind/libpaddle.sobuild/python/paddle/base/libpaddle.so主绑定库

增量编译后务必检查:

bash
# 确认 Python 实际加载了哪个 .so
python -c "import paddle; import os; print(os.path.realpath(paddle.__file__))"

# 比较构建时间戳
stat build/paddle/phi/libphi_core.so
stat build/python/paddle/libs/libphi_core.so

# 如果时间戳不一致,手动同步
cp build/paddle/phi/libphi_core.so build/python/paddle/libs/libphi_core.so
cp build/paddle/phi/libphi_gpu.so build/python/paddle/libs/libphi_gpu.so

典型症状:修改了源码并重新编译,但运行时错误信息中的行号不变——这说明 Python 加载的仍是旧 .so。

多路径调用链分析方法

当崩溃发生在公共底层函数(如 GetCurrentDeviceId、cudaFree)时,需穷举所有调用路径来定位真正的入口:

  1. 从崩溃点出发,向上追溯:用 Grep 搜索崩溃函数的所有调用者,逐层向上展开
  2. 结合分配器类型缩小范围:根据 FLAGS(如 FLAGS_use_system_allocator)确定实际使用的分配器链路
  3. Tensor 生命周期追踪:DenseTensor::~DenseTensor -> shared_ptr<Allocation> -> AllocationDeleter -> 具体分配器的 FreeImpl
  4. 在崩溃点添加 backtrace 日志:临时加入 backtrace_symbols_fd 打印调用栈,确认实际触发路径
  5. 注意虚函数/宏展开:PADDLE_ENFORCE_GPU_SUCCESS 是宏,行号由 __LINE__ 决定;FreeImpl 是虚函数,实际调用取决于运行时类型

调试案例

见 references/case-studies.md 了解实际调试案例。

© PaddlePaddle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/paddle-debug of PaddlePaddle/Paddle.

  • SKILL.md
  • references/case-studies.md
  • references/cuda-debug.md

Open the folder on GitHubat commit 4793e33

Compare with similar skills

Paddle Debug next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paddle Debug compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paddle Debug this skillPaddlePaddle/Paddle24k—~1.4kAutomated safety check: PassApache-2.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Fix Envevo-design/proto-tools135—~2.5kAutomated safety check: NotesMIT
Migrate Workflow Ec2 To Osdcpytorch/test-infra113—~2kAutomated safety check: PassCustom licence
Hyperpod Version Checkerawslabs/agent-plugins9151 repos~910Automated safety check: PassApache-2.0
Quark Env Preflightamd/Quark181—~1.4kAutomated safety check: PassMIT

Similar skills

  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • Fix Env

    evo-design/proto-tools

    Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…

    135 GitHub stars~2.5k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed…

    113 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    915 GitHub starsUsed in 1 repo~910 tokens
    AI & LLM EngineeringAuto-check passed
  • Collect and normalize environment facts (OS, Python, GPU, CUDA/ROCm, container state) before Quark installation or PTQ planning.

    181 GitHub stars~1.4k tokensUpdated 10 days ago
    AI & LLM EngineeringAuto-check passed
  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    398 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from PaddlePaddle/Paddle

All 11 skills in this repo
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Design Distributed

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

    24k GitHub stars~660 tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Eager Graph

    PaddlePaddle/Paddle

    A skill your agent uses when navigating Paddle eager-mode (dynamic graph) source code, tracing forward/backward execution, debugging autograd issues, understanding PyLayer, or investigating…

    24k GitHub stars~562 tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Phi Kernel

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's PHI kernel system: registering new kernels, debugging kernel selection/dispatch, understanding code auto-generation from YAML, or implementing…

    24k GitHub stars~656 tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about Paddle Debug

What does Paddle Debug do?

在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。. Paddle Debug is an agent skill from PaddlePaddle/Paddle.

When should I use Paddle Debug?

Paddle Debug fits situations like: tasks that involve Deep learning.

How do I install Paddle Debug in Claude Code?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-debug -a claude-code`. Or copy the skill folder (.agents/skills/paddle-debug in PaddlePaddle/Paddle) into .claude/skills/paddle-debug in your project. Claude Code loads it when a task matches its description.

How do I install Paddle Debug in Codex?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-debug -a codex`. Or copy the skill folder (.agents/skills/paddle-debug in PaddlePaddle/Paddle) into .agents/skills/paddle-debug in your project. Codex loads it when a task matches its description.

Can I use Paddle Debug in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PaddlePaddle/Paddle --skill paddle-debug -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paddle-debug, .gemini/skills/paddle-debug, .github/skills/paddle-debug and .opencode/skills/paddle-debug in your project.

What does Paddle Debug need to run?

Going by SKILL.md and its folder, Paddle Debug needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Paddle Debug access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Paddle Debug safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paddle Debug use?

Paddle Debug is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paddle Debug use?

About 1.4k tokens (SKILL.md is roughly 5.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.2k tokens, read only when the agent opens those files.

What are the alternatives to Paddle Debug?

Skills that share tags, products or a category with Paddle Debug: Ako4all (TongmingLAIC/AKO4ALL, 369 stars), Fix Env (evo-design/proto-tools, 135 stars), Migrate Workflow Ec2 To Osdc (pytorch/test-infra, 113 stars) and Hyperpod Version Checker (awslabs/agent-plugins, 915 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paddle Debug?

PaddlePaddle (a GitHub organization) maintains it in PaddlePaddle/Paddle, which has 24,120 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on September 30, 2026.

Source: PaddlePaddle/Paddle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.