Agent skill

Paddle Op Dev

by PaddlePaddle in PaddlePaddle/Paddle

PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Paddle Op Dev

skills CLI
$ npx skills add PaddlePaddle/Paddle --skill paddle-op-dev -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PaddlePaddle/Paddle paddle-op-dev --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PaddlePaddle/Paddle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/paddle-op-dev .claude/skills/paddle-op-dev && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paddle-op-dev
GitHub stars
24k
Token cost
~1.3k tokens
SKILL.md length
267 words
Files
6 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

  • Tasks that involve Deep learning
  • SKILL.md covers 架构概览, 开发流程, 文件清单速查 and 显存优化, plus 1 more section
  • Calls python, pip and make

What it does

Paddle Op Dev is an agent skill from PaddlePaddle/Paddle. PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML 配置、InferMeta、Kernel、Python API 或单元测试 (4) 理解 Paddle 算子开发架构和流程 (5) 编译 Paddle 并验证算子正确性

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/infermeta.md`, `references/kernel-dev.md` and `references/python-api.md`).

It sits in AI & LLM Engineering, covering Deep learning. It works with Python, C++, CUDA and PyTorch. The repository describes itself as: PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署). The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Deep learning

Example prompts

  • “/paddle-op-dev”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4793e33. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paddle Op Dev loads about 1.3k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 267 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PaddlePaddle/Paddle at commit 4793e33, republished under its Apache-2.0 licence (© PaddlePaddle). 267 words, ~1,264 tokens.

Download SKILL.mdSave it as .claude/skills/paddle-op-dev/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
paddle-op-dev
description
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML 配置、InferMeta、Kernel、Python API 或单元测试 (4) 理解 Paddle 算子开发架构和流程 (5) 编译 Paddle 并验证算子正确性

Paddle 算子开发

架构概览

Python API (paddle.xxx)
    │
    ▼ (YAML 自动生成的调度代码)
算子 InferMeta ──→ 推导输出 shape/dtype
    │
    ▼
算子 Kernel ──→ 实际计算(CPU/GPU 分别实现)

YAML 配置是连接 Python API 与底层 Kernel 的桥梁,框架编译时自动生成调度代码。

开发流程

新增算子 xxx 需完成以下 6 步:

步骤 1:YAML 算子定义

在 paddle/phi/ops/yaml/ops.yaml 和 backward.yaml 中配置前向和反向算子。

关键配置项:op(名称)、args(输入)、output(输出)、infer_meta(推导函数)、kernel(计算函数)、backward(反向算子)。

快速示例:

yaml
# ops.yaml
- op: trace
  args: (Tensor x, int offset = 0, int axis1 = 0, int axis2 = 1)
  output: Tensor(out)
  infer_meta:
    func: TraceInferMeta
  kernel:
    func: trace
  backward: trace_grad

详细配置规则见 references/yaml-config.md。

步骤 2:InferMeta 函数

在 paddle/phi/infermeta/ 下实现,按输入 Tensor 个数分文件(unary/binary/multiary)。

职责:推导输出 shape 和 dtype,检查输入合法性。

函数签名模式:

cpp
void XxxInferMeta(const MetaTensor& x, int attr, MetaTensor* out) {
  // PADDLE_ENFORCE_XX 检查输入
  out->set_dims(...);
  out->set_dtype(x.dtype());
}

详细规则和示例见 references/infermeta.md。

步骤 3:Kernel 实现

在 paddle/phi/kernels/ 下实现。设备无关的放根目录,设备相关的分 cpu/ 和 gpu/ 子目录。

文件结构:

  • xxx_kernel.h — 前向声明
  • cpu/xxx_kernel.cc / gpu/xxx_kernel.cu — 设备实现
  • xxx_grad_kernel.h + 对应设备实现 — 反向

Kernel 函数模式:

cpp
template <typename T, typename Context>
void XxxKernel(const Context& dev_ctx, const DenseTensor& x,
               int attr, DenseTensor* out) {
  dev_ctx.template Alloc<T>(out);
  // 计算逻辑
}
// 注册
PD_REGISTER_KERNEL(xxx, CPU, ALL_LAYOUT, phi::XxxKernel, float, double) {}

参考 PyTorch 实现:Paddle 算子的计算逻辑与 PyTorch (ATen) 往往相似。实现 Kernel 时可参考 PyTorch 对应算子的实现(位于 aten/src/ATen/native/ 目录),理解核心算法后适配为 Paddle 的 API 风格(DenseTensor、dev_ctx 等)。注意不要直接复制代码,需根据 Paddle 框架规范进行适配。

详细规则和示例见 references/kernel-dev.md。

步骤 4:Python API 封装

在 python/paddle/ 对应子目录中实现 Python API。

关键要素:

  • 动态图:_C_ops.xxx(args...)
  • 静态图:LayerHelper + append_op
  • 完整的 docstring(含可运行的 Examples)

详细规则和示例见 references/python-api.md。

步骤 5:单元测试

在 test/legacy_test/test_xxx_op.py 中编写,继承 OpTest。

测试要点:

  • setUp() 设置 op_type、python_api、inputs、attrs、outputs
  • test_check_output(check_pir=True) 验证前向
  • test_check_grad(['X'], 'Out', check_pir=True) 验证反向
  • 多 Case 继承基类重写 init_config()

详细规则和示例见 references/unit-test.md。

步骤 6:编译与验证

完成代码开发后,需要编译 Paddle 并运行测试来验证算子的正确性。

6.1 增量编译

在 Paddle 的 build 目录下执行增量编译。新增算子通常只需重新编译 phi 相关目标,无需全量编译。

bash
cd build

# 增量编译(推荐使用 ninja,速度更快)
ninja -j$(nproc)

# 或使用 make
make -j$(nproc)

如果只修改了 Kernel/InferMeta 等 C++ 文件,可以缩小编译范围:

bash
# 只编译 phi 库(覆盖 Kernel + InferMeta)
ninja phi -j$(nproc)

# 编译完成后重新安装 Python 包
pip install -e . --no-build-isolation
# 或
cd python && pip install -e .
6.2 运行单元测试
bash
# 方式 1:直接运行测试文件
python test/legacy_test/test_xxx_op.py

# 方式 2:使用 pytest(支持更灵活的筛选)
python -m pytest test/legacy_test/test_xxx_op.py -v

# 方式 3:通过 ctest 运行(在 build 目录下)
cd build && ctest -R test_xxx_op -V
6.3 验证要点
  • 前向正确性:test_check_output 通过,算子输出与 NumPy 参考实现一致
  • 反向正确性:test_check_grad 通过,梯度通过数值微分法校验
  • PIR 模式:确认 check_pir=True 下测试通过
  • 多设备验证:有 GPU 环境时确认 CPU 和 GPU 结果一致
6.4 GPU 算子调试

GPU 算子出现 CUDA 错误时,使用以下环境变量辅助定位:

bash
# 启用 CUDA 同步错误检查 + 系统内存分配器
FLAGS_check_cuda_error=1 FLAGS_use_system_allocator=1 python test/legacy_test/test_xxx_op.py

# 强制 CUDA kernel 同步执行,定位出错的 kernel
CUDA_LAUNCH_BLOCKING=1 python test/legacy_test/test_xxx_op.py

常见问题:

  • CUDA error(9) — kernel 配置无效,检查 grid/block size 是否为 0(通常由空 Tensor 触发)
  • CUDA error(700) — 非法内存访问,检查数组越界或空指针
  • CUDA error(2) — 显存不足,减小测试数据规模或检查是否有显存泄漏
6.5 常见编译错误排查
现象排查方向
找不到 XxxInferMeta 符号检查 InferMeta 函数是否在 .h 中声明、YAML 中函数名是否拼写一致
找不到 xxx kernel检查 PD_REGISTER_KERNEL 注册名是否与 YAML kernel:func 一致
Python 端 _C_ops.xxx 不存在确认 YAML 配置正确且已重新编译,pip install -e . 已执行
参数数量不匹配对照 YAML args 与 InferMeta/Kernel 函数签名的参数列表

文件清单速查

内容文件位置
前向 YAMLpaddle/phi/ops/yaml/ops.yaml
反向 YAMLpaddle/phi/ops/yaml/backward.yaml
InferMetapaddle/phi/infermeta/{unary,binary,multiary}.{h,cc}
Kernel 头文件paddle/phi/kernels/xxx_kernel.h
CPU Kernelpaddle/phi/kernels/cpu/xxx_kernel.cc
GPU Kernelpaddle/phi/kernels/gpu/xxx_kernel.cu
Python APIpython/paddle/ 对应子目录
单元测试test/legacy_test/test_xxx_op.py

显存优化

  • inplace:输出复用输入显存,YAML 中配置 inplace: (x -> out)
  • no_need_buffer:反向不需要前向变量的内存数据时配置,提前释放内存
  • 减少反向无关变量:反向 args 中只包含实际需要的前向变量

参考文档

© PaddlePaddle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in .agents/skills/paddle-op-dev of PaddlePaddle/Paddle.

  • SKILL.md
  • references/infermeta.md
  • references/kernel-dev.md
  • references/python-api.md
  • references/unit-test.md
  • references/yaml-config.md

Open the folder on GitHubat commit 4793e33

Compare with similar skills

Paddle Op Dev next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paddle Op Dev compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paddle Op Dev this skillPaddlePaddle/Paddle24k—~1.3kAutomated safety check: PassApache-2.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
ExecuTorch Build Guidepytorch/executorch5.1k—~2.3kAutomated safety check: NotesCustom licence
Embedded AI Deploymentmatlab/agent-skills-playground1811 repos~3.4kAutomated safety check: PassCustom licence
Fix Envevo-design/proto-tools135—~2.5kAutomated safety check: NotesMIT
Migrate Workflow Ec2 To Osdcpytorch/test-infra113—~2kAutomated safety check: PassCustom licence

Similar skills

  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 23 days ago
    AI & LLM EngineeringAuto-check passed
  • ExecuTorch Build Guide

    pytorch/executorch

    Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.

    5.1k GitHub stars~2.3k tokensUpdated today
    DevelopmentAuto-check: notes
  • Embedded AI Deployment

    matlab/agent-skills-playground

    Deploy AI models to embedded hardware using MathWorks tools (MATLAB, Simulink, Embedded Coder).

    181 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Fix Env

    evo-design/proto-tools

    Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…

    135 GitHub stars~2.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check: notes
  • Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed…

    113 GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • At Dispatch V2

    intel/torch-xpu-ops

    Official

    Convert PyTorch ATDISPATCH macros to ATDISPATCHV2 format in ATen C++ code.

    115 GitHub starsUsed in 3 repos~2.2k tokens
    AI & LLM EngineeringAuto-check passed

More from PaddlePaddle/Paddle

All 11 skills in this repo
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Design Distributed

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

    24k GitHub stars~660 tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Eager Graph

    PaddlePaddle/Paddle

    A skill your agent uses when navigating Paddle eager-mode (dynamic graph) source code, tracing forward/backward execution, debugging autograd issues, understanding PyLayer, or investigating…

    24k GitHub stars~562 tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Phi Kernel

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's PHI kernel system: registering new kernels, debugging kernel selection/dispatch, understanding code auto-generation from YAML, or implementing…

    24k GitHub stars~656 tokensUpdated 8 days ago
    Auto-check passed
  • Paddle Debug

    PaddlePaddle/Paddle

    在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

    24k GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed

Questions about Paddle Op Dev

What does Paddle Op Dev do?

PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…. Paddle Op Dev is an agent skill from PaddlePaddle/Paddle.

When should I use Paddle Op Dev?

Paddle Op Dev fits situations like: tasks that involve Deep learning.

How do I install Paddle Op Dev in Claude Code?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-op-dev -a claude-code`. Or copy the skill folder (.agents/skills/paddle-op-dev in PaddlePaddle/Paddle) into .claude/skills/paddle-op-dev in your project. Claude Code loads it when a task matches its description.

How do I install Paddle Op Dev in Codex?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-op-dev -a codex`. Or copy the skill folder (.agents/skills/paddle-op-dev in PaddlePaddle/Paddle) into .agents/skills/paddle-op-dev in your project. Codex loads it when a task matches its description.

Can I use Paddle Op Dev in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PaddlePaddle/Paddle --skill paddle-op-dev -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paddle-op-dev, .gemini/skills/paddle-op-dev, .github/skills/paddle-op-dev and .opencode/skills/paddle-op-dev in your project.

What does Paddle Op Dev need to run?

Going by SKILL.md and its folder, Paddle Op Dev needs the command-line tools its instructions call (python, pip and make). Our summary lists: Python 3.

Does Paddle Op Dev access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Paddle Op Dev safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paddle Op Dev use?

Paddle Op Dev is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paddle Op Dev use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.

What are the alternatives to Paddle Op Dev?

Skills that share tags, products or a category with Paddle Op Dev: Ako4all (TongmingLAIC/AKO4ALL, 369 stars), ExecuTorch Build Guide (pytorch/executorch, 5.1k stars), Embedded AI Deployment (matlab/agent-skills-playground, 181 stars) and Fix Env (evo-design/proto-tools, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paddle Op Dev?

PaddlePaddle (a GitHub organization) maintains it in PaddlePaddle/Paddle, which has 24,120 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on September 30, 2026.

Source: PaddlePaddle/Paddle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.