Agent skill

Paddle Cross Ecosystem Custom Op

by PaddlePaddle in PaddlePaddle/Paddle

将原生 PyTorch 自定义算子库、Torch extension、生态库(TorchCodec/FlashInfer/DeepEP 等)以及 Kernel DSL 生态(Triton/TileLang/TVM FFI 等)以最小修改方式接入 PaddlePaddle。遇到以下场景务必使用:迁移外部算子库到 Paddle;分析 PFCCLab fork 与上游的兼容差异;处理…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Paddle Cross Ecosystem Custom Op

skills CLI
$ npx skills add PaddlePaddle/Paddle --skill paddle-cross-ecosystem-custom-op -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PaddlePaddle/Paddle paddle-cross-ecosystem-custom-op --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PaddlePaddle/Paddle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/paddle-cross-ecosystem-custom-op .claude/skills/paddle-cross-ecosystem-custom-op && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paddle-cross-ecosystem-custom-op
GitHub stars
24k
Token cost
~883 tokens
SKILL.md length
243 words
Files
21 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

将原生 PyTorch 自定义算子库、Torch extension、生态库(TorchCodec/FlashInfer/DeepEP 等)以及 Kernel DSL 生态(Triton/TileLang/TVM FFI 等)以最小修改方式接入 PaddlePaddle。遇到以下场景务必使用:迁移外部算子库到 Paddle;分析 PFCCLab fork 与上游的兼容差异;处理…

  • Works in 5 steps: 识别上游仓库、当前 fork、默认分支和实际迁移分支。PFCCLab… → 按控制面把仓库分成四层 → 如果任务是分析多个 PFCCLab… → …
  • Tasks that involve Deep learning
  • SKILL.md covers 任务定义, 核心约束, 工作顺序 and 默认改动边界, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Paddle Cross Ecosystem Custom Op is an agent skill from PaddlePaddle/Paddle. 将原生 PyTorch 自定义算子库、Torch extension、生态库(TorchCodec/FlashInfer/DeepEP 等)以及 Kernel DSL 生态(Triton/TileLang/TVM FFI 等)以最小修改方式接入 PaddlePaddle。遇到以下场景务必使用:迁移外部算子库到 Paddle;分析 PFCCLab fork 与上游的兼容差异;处理 paddle.enablecompat、paddle.utils.cppextension、TORCHLIBRARY、torch.ops、at::Tensor/c10 compat 问题;为 compat gap 设计最小 workaround 并准备 Paddle issue 最小复现;将已迁移的生态库集成进 PaddleFleet(paddlefleetops、eager import 约束)。

Its SKILL.md is about 880 tokens, which your agent loads only when the skill is triggered. The skill folder holds 23 other files, including reference files (for example `evals/evals.json`, `references/compat-gap-policy.md` and `references/ecosystem-cases/cudnn-frontend.md`).

It sits in AI & LLM Engineering, covering Deep learning. It works with C++, PyTorch and Python. The repository describes itself as: PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署). The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Deep learning

Example prompts

  • “/paddle-cross-ecosystem-custom-op”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. 识别上游仓库、当前 fork、默认分支和实际迁移分支。PFCCLab 适配仓库的默认分支通常是 paddle,比较前先确认 parent 和默认分支。
  2. 按控制面把仓库分成四层
  3. 如果任务是分析多个 PFCCLab fork,且用户明确要求并行,按仓库拆分并行分析;每个子任务都要输出 parent、比较分支、四层 diff 归类和可复用模式。
  4. 先确定第一轮改动的位置。首轮补丁通常集中在 build、runtime glue、device / stream / distributed 边界。
  5. 沿最小路径逐步推进验证:build → import → 最小功能测试 → 运行时对照。

What it can do on your machine

Read from SKILL.md and the folder at commit a48f1c3. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paddle Cross Ecosystem Custom Op loads about 883 tokens when it runs, and up to ~25k if it reads all its reference files. Until then it costs about 107 tokens; SKILL.md has 243 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~107
When it runs · the whole SKILL.md, loaded when a task matches
~883
With references · SKILL.md plus every file in references/, read only if the agent opens them
~25k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PaddlePaddle/Paddle at commit a48f1c3, republished under its Apache-2.0 licence (© PaddlePaddle). 243 words, ~883 tokens.

Download SKILL.mdSave it as .claude/skills/paddle-cross-ecosystem-custom-op/SKILL.md (or your agent's skills folder). This skill also uses 20 other files; get the full folder from GitHub.
name
paddle-cross-ecosystem-custom-op
description
将原生 PyTorch 自定义算子库、Torch extension、生态库(TorchCodec/FlashInfer/DeepEP 等)以及 Kernel DSL 生态(Triton/TileLang/TVM FFI 等)以最小修改方式接入 PaddlePaddle。遇到以下场景务必使用:迁移外部算子库到 Paddle;分析 PFCCLab fork 与上游的兼容差异;处理 paddle.enable_compat、paddle.utils.cpp_extension、TORCH_LIBRARY、torch.ops、at::Tensor/c10 compat 问题;为 compat gap 设计最小 workaround 并准备 Paddle issue 最小复现;将已迁移的生态库集成进 PaddleFleet(paddlefleet_ops、eager import 约束)。

Paddle 跨生态自定义算子迁移

任务定义

这个 skill 的主任务:让上游 PyTorch 自定义算子仓库在 Paddle 上按原来的调用路径跑起来,同时保持后续 rebase / sync upstream 的能力。迁移完成后,如果需要把库集成进 PaddleFleet,走 PaddleFleet 集成。

一次完整的输出应该覆盖四个方面:

  • 迁移方案
  • 最小修改边界
  • 验证路径
  • compat gap 处理策略

核心约束

  • 最小修改:不做额外格式化、优化、重构,不主动改公共 API。
  • 上游同步:所有改动都要考虑后续 rebase / sync upstream 的便利性。
  • compat 优先:优先使用 Paddle 现有的 compat 机制,让 compat 层承担兼容职责。
  • 缺口要明确:compat gap 要分类清楚、标明边界,并准备最小复现。
  • 验证要闭环:至少跑通一条最小 build/test 路径。

工作顺序

  1. 识别上游仓库、当前 fork、默认分支和实际迁移分支。PFCCLab 适配仓库的默认分支通常是 paddle,比较前先确认 parent 和默认分支。
  2. 按控制面把仓库分成四层:
    • 框架无关的内核 / 算法
    • 构建与打包
    • C++ compat API / 注册
    • Python 包装 / runtime glue / tests
  3. 如果任务是分析多个 PFCCLab fork,且用户明确要求并行,按仓库拆分并行分析;每个子任务都要输出 parent、比较分支、四层 diff 归类和可复用模式。
  4. 先确定第一轮改动的位置。首轮补丁通常集中在 build、runtime glue、device / stream / distributed 边界。
  5. 沿最小路径逐步推进验证:build → import → 最小功能测试 → 运行时对照。

默认改动边界

通常不需要动的部分:

  • CUDA/C++ 核心 kernel 与算法逻辑
  • 原有 schema 定义
  • 大部分 TORCH_LIBRARY / pybind11 注册代码
  • 上游目录结构与 Python package 形状

通常需要先检查的部分:

  • setup.py / pyproject.toml
  • 入口脚本、测试脚本、示例脚本
  • torch.ops / torch.library / torch._dynamo / torch.profiler 使用点
  • device / stream / distributed / DLPack / custom op registration glue

具体规则

  • setup.py / pyproject.toml:优先加 paddle.enable_compat(),保留原有 from torch.utils import cpp_extension 的写法;只有代理路径覆盖不到时,才最小化地切到 paddle.utils.cpp_extension 或局部调整 include / lib / flags。
  • TORCH_LIBRARY / TORCH_LIBRARY_IMPL / pybind11:默认先保持原样,等编译或运行时真正失败了再定位具体缺口。
  • at::Tensor / c10::TensorOptions / torch::empty 等 C++ API:优先依赖 compat headers;遇到缺口时只桥接单个 API 点。
  • Python 入口与测试:优先用 paddle.enable_compat(scope={...}) 限定代理范围;短生命周期的 build script 可以用全局 paddle.enable_compat()。PaddleFleet 集成场景按 PaddleFleet 集成 的既有模板写。
  • 分布式 / stream / device:先把运行时上下文边界接上,再看是否需要深入 phi::GPUContext、ProcessGroup、DLPack 或 stream wrapper。
  • 分析 PFCCLab fork:输出要提炼成可复用的模式,覆盖 build / C++ / Python / tests 四层。

按需读取参考材料

按当前任务选择参考材料:

当前任务读取文件
先理解跨生态机制和分层口径机制总览
实际迁移一个新仓库迁移手册
把错误定位到 Paddle 仓库内部Paddle 内部锚点
分析清单内的 PFCCLab fork先读 生态库案例索引,再只读索引指向的对应 case
为新仓库复用既有迁移经验先读 生态库案例索引,再按控制面最多选择一到两个相近 case
判断 compat gap、workaround、issue MREcompat 缺口处理
build/import 已通但运行时行为不一致运行时调试
把已迁移的生态库集成进 PaddleFleetPaddleFleet 集成

输出要求

  • 明确列出哪些文件不需要动、哪些文件需要改、每一处改动对应哪一层。
  • 如果需要 workaround,必须写清楚覆盖范围、删除条件,以及是否需要提 Paddle issue。
  • 如果问题进入运行时对照阶段,要指出第一次差异出现在哪一行、哪个调用点、属于哪一层。
  • 如果分析的是现有 fork,要总结出可复用的迁移顺序,并把 diff 提炼成稳定模式。

完成前检查

  • 没有无关的格式化、清理、重命名。
  • 保留了上游目录结构和主要 API 形状。
  • 运行时的 enable_compat 已尽量限定 scope;build script 的全局 compat 只用在构建入口。
  • build/test 至少跑通了一条最小路径。
  • compat gap 已经准备了 issue MRE,或在结果中明确写出了缺口与临时 workaround。

© PaddlePaddle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 20 other files (references) in .agents/skills/paddle-cross-ecosystem-custom-op of PaddlePaddle/Paddle.

  • SKILL.md
  • evals/evals.json
  • references/compat-gap-policy.md
  • references/ecosystem-cases/cudnn-frontend.md
  • references/ecosystem-cases/deep-ep.md
  • references/ecosystem-cases/deepgemm.md
  • references/ecosystem-cases/fast-hadamard-transform.md
  • references/ecosystem-cases/flash-linear-attention.md
  • references/ecosystem-cases/flash-mla.md
  • references/ecosystem-cases/flashinfer.md
  • references/ecosystem-cases/moonep.md
  • references/ecosystem-cases/paddlecodec.md
  • references/ecosystem-cases/quack.md
  • references/ecosystem-cases/sonic-moe.md
  • references/ecosystem-cases/tilelang.md
  • references/ecosystem-diff-patterns.md
  • references/mechanism-overview.md
  • references/migration-playbook.md
  • … and 3 more

Open the folder on GitHubat commit a48f1c3

Compare with similar skills

Paddle Cross Ecosystem Custom Op next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paddle Cross Ecosystem Custom Op compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paddle Cross Ecosystem Custom Op this skillPaddlePaddle/Paddle24k—~883Automated safety check: PassApache-2.0
Ako4allTongmingLAIC/AKO4ALL369—~4kAutomated safety check: PassMIT
Quark Installamd/Quark182—~1.8kAutomated safety check: NotesMIT
Software Harnessexeex/edge-cores110—~1kAutomated safety check: PassApache-2.0
ExecuTorch Build Guidepytorch/executorch5.1k—~2.3kAutomated safety check: NotesCustom licence
Onnxtxtonnx/onnx22k—~1.3kAutomated safety check: PassApache-2.0

Similar skills

  • Ako4all

    TongmingLAIC/AKO4ALL

    Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.

    369 GitHub stars~4k tokensUpdated 26 days ago
    AI & LLM EngineeringAuto-check passed
  • Quark Install

    amd/Quark

    Install or verify the AMD Quark package and its dependencies.

    182 GitHub stars~1.8k tokensUpdated 13 days ago
    AI & LLM EngineeringAuto-check: notes
  • Software Harness

    exeex/edge-cores

    Run, extend, debug, or review the public edge-e3 bare-metal software harness, including encrypted Verilator builds, hello and tensor examples, all example/llama/model smoke cases, PyTorch BF16…

    110 GitHub stars~1k tokensUpdated 15 days ago
    AI & LLM EngineeringAuto-check passed
  • ExecuTorch Build Guide

    pytorch/executorch

    Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.

    5.1k GitHub stars~2.3k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Onnxtxt

    onnx/onnx

    Read or write ONNX text format ("onnxtxt"). An agent skill from onnx/onnx.

    22k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Benchmark Pyrefly

    facebook/pyrefly

    Official

    Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.

    7.1k GitHub stars~1.8k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from PaddlePaddle/Paddle

All 11 skills in this repo
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated yesterday
    Auto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Paddle Design Distributed

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

    24k GitHub stars~660 tokensUpdated yesterday
    Auto-check passed
  • Paddle Eager Graph

    PaddlePaddle/Paddle

    A skill your agent uses when navigating Paddle eager-mode (dynamic graph) source code, tracing forward/backward execution, debugging autograd issues, understanding PyLayer, or investigating…

    24k GitHub stars~562 tokensUpdated yesterday
    Auto-check passed
  • Paddle Phi Kernel

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's PHI kernel system: registering new kernels, debugging kernel selection/dispatch, understanding code auto-generation from YAML, or implementing…

    24k GitHub stars~656 tokensUpdated yesterday
    Auto-check passed
  • Paddle Debug

    PaddlePaddle/Paddle

    在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

    24k GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed

Questions about Paddle Cross Ecosystem Custom Op

What does Paddle Cross Ecosystem Custom Op do?

将原生 PyTorch 自定义算子库、Torch extension、生态库(TorchCodec/FlashInfer/DeepEP 等)以及 Kernel DSL 生态(Triton/TileLang/TVM FFI 等)以最小修改方式接入 PaddlePaddle。遇到以下场景务必使用:迁移外部算子库到 Paddle;分析 PFCCLab fork 与上游的兼容差异;处理…. Paddle Cross Ecosystem Custom Op is an agent skill from PaddlePaddle/Paddle.

When should I use Paddle Cross Ecosystem Custom Op?

Paddle Cross Ecosystem Custom Op fits situations like: tasks that involve Deep learning.

How do I install Paddle Cross Ecosystem Custom Op in Claude Code?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-cross-ecosystem-custom-op -a claude-code`. Or copy the skill folder (.agents/skills/paddle-cross-ecosystem-custom-op in PaddlePaddle/Paddle) into .claude/skills/paddle-cross-ecosystem-custom-op in your project. Claude Code loads it when a task matches its description.

How do I install Paddle Cross Ecosystem Custom Op in Codex?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-cross-ecosystem-custom-op -a codex`. Or copy the skill folder (.agents/skills/paddle-cross-ecosystem-custom-op in PaddlePaddle/Paddle) into .agents/skills/paddle-cross-ecosystem-custom-op in your project. Codex loads it when a task matches its description.

Can I use Paddle Cross Ecosystem Custom Op in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PaddlePaddle/Paddle --skill paddle-cross-ecosystem-custom-op -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paddle-cross-ecosystem-custom-op, .gemini/skills/paddle-cross-ecosystem-custom-op, .github/skills/paddle-cross-ecosystem-custom-op and .opencode/skills/paddle-cross-ecosystem-custom-op in your project.

What does Paddle Cross Ecosystem Custom Op need to run?

SKILL.md names no scripts, command-line tools or credentials: Paddle Cross Ecosystem Custom Op is instructions for the agent only. Our summary lists: Python 3.

Does Paddle Cross Ecosystem Custom Op access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Paddle Cross Ecosystem Custom Op safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paddle Cross Ecosystem Custom Op use?

Paddle Cross Ecosystem Custom Op is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paddle Cross Ecosystem Custom Op use?

About 883 tokens (SKILL.md is roughly 3.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 24k tokens, read only when the agent opens those files.

What are the alternatives to Paddle Cross Ecosystem Custom Op?

Skills that share tags, products or a category with Paddle Cross Ecosystem Custom Op: Ako4all (TongmingLAIC/AKO4ALL, 369 stars), Quark Install (amd/Quark, 182 stars), Software Harness (exeex/edge-cores, 110 stars) and ExecuTorch Build Guide (pytorch/executorch, 5.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paddle Cross Ecosystem Custom Op?

PaddlePaddle (a GitHub organization) maintains it in PaddlePaddle/Paddle, which has 24,120 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 10, 2026.

Source: PaddlePaddle/Paddle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.