Agent skill

Paddle Design Compiler

by PaddlePaddle in PaddlePaddle/Paddle

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Paddle Design Compiler

skills CLI
$ npx skills add PaddlePaddle/Paddle --skill paddle-design-compiler -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install PaddlePaddle/Paddle paddle-design-compiler --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/PaddlePaddle/Paddle.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/paddle-design-compiler .claude/skills/paddle-design-compiler && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
paddle-design-compiler
GitHub stars
24k
Token cost
~3.6k tokens
SKILL.md length
666 words
Files
7 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

  • Working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture
  • SKILL.md covers 全链路概览, SOT(Symbolic Opcode Translator), PIR(Paddle Intermediate… and CINN 编译与执行, plus 3 more sections
  • Calls python
  • PIR (Paddle IR) for SSA-based intermediate representation

What it does

Paddle Design Compiler is an agent skill from PaddlePaddle/Paddle. Use when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate representation, CINN for fused CUDA kernel generation, operator decomposition (Prim), or the end-to-end flow from Python eager code to optimized GPU execution.

Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `references/cinn-pipeline.md`, `references/control-flow.md` and `references/executor.md`).

It sits in AI & LLM Engineering, covering Translation, GPU and accelerator computing and Deep learning. It works with Python and CUDA. The repository describes itself as: PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署). The licence is Apache-2.0.

When your agent uses it

  • Working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture
  • PIR (Paddle IR) for SSA-based intermediate representation
  • CINN for fused CUDA kernel generation
  • Operator decomposition (Prim)

Example prompts

  • “/paddle-design-compiler”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 5434c21. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Paddle Design Compiler loads about 3.6k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 666 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~89
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from PaddlePaddle/Paddle at commit 5434c21, republished under its Apache-2.0 licence (© PaddlePaddle). 666 words, ~3,619 tokens.

Download SKILL.mdSave it as .claude/skills/paddle-design-compiler/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
paddle-design-compiler
description
Use when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate representation, CINN for fused CUDA kernel generation, operator decomposition (Prim), or the end-to-end flow from Python eager code to optimized GPU execution.

Paddle 3.0 编译器全链路

Paddle 3.0 的编译器体系通过 SOT → PIR → CINN 三阶段流水线,将用户的动态图 Python 代码编译为高性能 GPU Kernel,实现「动态图编写、编译器加速」的开发体验。

全链路概览

用户 Python 代码(动态图 eager mode)
  │
  ▼  Stage 0: SOT 图捕获
  PEP 523 eval_frame 拦截 → OpcodeExecutor 字节码模拟
  → FunctionGraph / StatementIR
  → paddle.jit.to_static(full_graph=True) 编译子图
  → pir::Program
  │
  ▼  Stage 1: PIR Pass 优化
  pir::Program(SSA 形式的 pd_op.* 算子图)
  │  ├── ShapeOptimizationPass(InferSymbolicShape 动态 shape 符号推导)
  │  ├── 组合算子分解 (DecompInterface → primitive operators)
  │  └── 通用 Pass 优化(常量折叠、死代码消除等)
  │
  ▼  Stage 2: CINN 编译
  ├── PdOpToCinnOpPass / PdOpToDynamicShapeCinnOpPass(算子映射)
  ├── add_cinn_pass → cinn_op.group(算子融合)
  ├── OpLower(Compute + Schedule)→ LoweredFunc
  ├── CodeGenCUDA_Dev → CUDA source → NVRTC → CUfunction
  ├── CompilationCache(编译缓存,相同子图复用已编译 Kernel)
  │
  ▼  Stage 3: 执行
  PirInterpreter 调度 → CinnJitInstruction → cuLaunchKernel

SOT → PIR 的衔接:SOT 捕获的 StatementIR 被包装为 Python 函数后,通过 paddle.jit.to_static(full_graph=True) 再次走 AST Transformer 路径编译为 pir::Program(参见 python/paddle/jit/sot/symbolic/compile_cache.py)。这意味着 SOT 负责"图捕获",而 to_static 负责"图编译"。

SOT(Symbolic Opcode Translator)

SOT 是 Paddle 3.0 的动转静前端,在 Python VM 字节码层面拦截和模拟执行用户代码,精确捕获 Tensor 计算子图。相比旧的 AST Transformer 方案,SOT 能处理 numpy/Tensor 互操作、动态控制流、第三方库调用等复杂场景。

核心机制
Python Frame
  │
  ▼  PEP 523 eval_frame 拦截
PyInterpreterState.eval_frame
  │
  ▼
OpcodeExecutor(模拟 Python VM 字节码执行)
  │  ├─ Variable 体系 (TensorVariable, ConstantVariable, ContainerVariable, ...)
  │  ├─ Tracker 追踪来源 → 生成 Guard(缓存有效性校验)
  │  └─ SideEffect 记录副作用(全局变量修改、可变对象修改)
  ▼
FunctionGraph / StatementIR(记录算子调用)
  │
  ├─ 无 fallback: 完整子图 → to_static(full_graph=True) → pir::Program
  └─ fallback: 子图切分 → 可静态化部分编译 + 不可静态化部分 Python 执行
组件说明
OpcodeExecutor模拟 Python VM 执行字节码,不真正计算,而是追踪 Tensor 操作
Variable 体系将 Python 对象包装为 Variable(TensorVariable / ConstantVariable / ContainerVariable / CallableVariable)
Tracker记录 Variable 来源(provenance),形成 DAG,用于生成 Guard
GuardCallable[[FrameType], bool],判断当前帧输入是否满足编译假设,用于缓存命中判断
FunctionGraph收集 Tensor 相关操作,输出 StatementIR
StatementIR4 种语句类型(call_api / call_method / call_sir / call_layer),最终经 to_static(full_graph=True) 编译为 Program
SideEffect记录并回放模拟执行中对全局变量和可变对象的修改,保证语义等价
OpcodeInlineExecutor跨函数边界模拟执行,实现子图跨函数融合
Fallback 场景
缩写全称场景
DDCFData-Dependent Control Flow控制流条件依赖 Tensor 值(如 if x.sum() > 0)
UNSPSUnsupported Simulation无法模拟的 Python 操作(如某些 C 扩展、.numpy())
CDBLCustom Blacklist用户或框架标记的不转换函数(如产生 -1 shape 的算子)
UNIMPUnimplemented Opcode尚未实现模拟的字节码指令

Fallback 是安全兜底:任何无法处理的情况退化为部分子图编译 + 部分 Python 执行,不会导致报错。

使用方式
python
# full_graph=False(默认):启用 SOT 模式(字节码级别捕获 + 自动 fallback)
# full_graph=True:使用传统 AST Transformer(要求整图可转)
net = paddle.jit.to_static(net)  # 默认 full_graph=False,即 SOT 模式
output = net(x)

PIR(Paddle Intermediate Representation)

PIR 是 Paddle 3.0 的统一中间表示,采用 MLIR 风格的 SSA 设计,替代旧的 ProgramDesc/OpDesc 体系。

核心概念
概念关键类说明
TypeTypeID / AbstractType / TypeStorage / Type统一类型系统:TypeID 用 static 变量地址做唯一标识,Type 本质是指向 TypeStorage 的指针,相等性通过指针比较 O(1)
ValueValueImpl / OpResultImpl / OpOperandImplSSA 值系统:OpResult 是算子输出(inline 0-5 / out-of-line),OpOperand 通过侵入式双向链表管理 use-chain
OperationOperation(连续内存布局)核心执行单元:`[OutOfLineResults
Block/RegionBlock / RegionBlock 持有 Operation 列表 + BlockArgument + terminator;Region 是 Block 的容器,约束 Value 作用域
DialectBuiltinDialect / PaddleDialect / CinnDialect模块化容器:聚合一组 Type、Attribute、Op 定义,支持独立注册与扩展
Trait/InterfaceOpTraitBase / concept-model 多态Trait 是静态标记,Interface 通过 concept-model 实现多态分派,替代 C++ 虚函数
核心 Dialect
Dialect职责典型内容
BuiltinDialectPIR 内置基础类型Float32Type, Int64Type, VectorType, DenseTensorType
PaddleDialectPaddle 算子定义pd_op.matmul, pd_op.relu, pd_op.conv2d
CinnDialectCINN 编译器专用cinn_op.group, cinn_op.yield, cinn_op.generate_shape
ControlFlowDialect控制流辅助cf.yield, cf.stack_create, cf.tuple_push, cf.tuple_pop
PaddleDialect(控制流部分)控制流算子pd_op.if, pd_op.while
PIR Program 结构
Program
├── weights: unordered_map<string, shared_ptr<Parameter>>
└── ModuleOp (顶层 Operation)
    └── Region[0]
        └── Block[0]
            ├── builtin.parameter("w")      → %0  (从权重表读取参数)
            ├── pd_op.matmul(%input, %0)    → %1
            ├── pd_op.if(%cond)              → %2
            │   ├── Region[0] (then)
            │   │   └── Block[0]: pd_op.relu(%1) → cf.yield
            │   └── Region[1] (else)
            │       └── Block[0]: pd_op.tanh(%1) → cf.yield
            └── builtin.set_parameter(%2, "out")
组合算子分解(Prim)

将高层算子分解为基础算子(primitive operators),降低编译器 / 分布式 / 新硬件适配成本:

  • 前向分解:DecompInterface → call_decomp_rule() → composite.h
  • 反向分解(VJP):两条路径——VjpInterface 经 call_vjp() 处理前向 op 的反向;DecompVjpInterface 经 call_decomp_vjp() 分解反向 op。规则实现均在 details.h
  • CustomVJP:为 sigmoid、log_softmax 等数值敏感算子提供手写反向
PIR Pass 框架

PIR 提供 MLIR 风格的 Pass 基础设施,用于图优化:

  • Pass:单个优化 Pass 基类,通过 Run(Operation*) 执行
  • PassManager:管理 Pass 执行顺序,支持嵌套 Pipeline
  • PatternRewritePass:基于 Pattern Matching 的重写 Pass,通过 RewritePattern 定义匹配和替换规则

CINN 编译与执行

CINN(Compiler Infrastructure for Neural Networks)将 PIR Program 中的算子子图编译为高性能 CUDA Kernel,由 PirInterpreter 调度执行。当前默认走动态 shape主线。

编译流水线(含动态 shape)
PIR Program (pd_op.*)
  │
  ▼  Stage 1: Frontend(前端)
  ├── ShapeOptimizationPass(InferSymbolicShape 符号推导)
  ├── PdOpToCinnOpPass / PdOpToDynamicShapeCinnOpPass(算子映射)
  ├── add_broadcast_to_elementwise_pass(显式 broadcast 插入)
  └── add_cinn_pass → cinn_op.group(按 OpPatternKind 融合)
  │
  ▼  Stage 2: Lowering(后端下降)
  ├── PirCompiler → CompilationTask(per GroupOp)
  │   ├── CompilationCache 查询(命中则跳过编译)
  │   ├── OpLower:Compute → AST IR → Schedule
  │   ├── DynamicShapeGroupScheduler(动态 shape 调度)
  │   └── LowerToAstVec → LoweredFunc
  │
  ▼  Stage 3: CodeGen(代码生成)
  ├── ir::Module → CodeGenCUDA_Dev → CUDA __global__ source
  └── nvrtc::Compiler → PTX → cubin → CUfunction
  │
  ▼  Stage 4: Execution(执行)
  └── cinn_runtime.jit_kernel (CINNKernelInfo: fn_ptr + symbol_args_map)
      └── CinnJitInstruction → cuLaunchKernel
OpPatternKind 融合规则
Kind含义典型算子
kElementWise逐元素计算relu, add, multiply
kBroadcast含广播语义broadcast_to
kInjective单射映射reshape, transpose, slice
kReduction规约操作reduce_sum, reduce_max
kOutFusible规约但输出可继续融合softmax 中间步骤
kNonFusible不可融合custom_call, sort
Group-level Schedule(DynamicShapeGroupScheduler)
步骤说明
DoLoopAlignment对齐各算子的循环范围
DoComputeInline将简单计算内联到消费者
OptimizeReduction优化规约算子的并行策略
DoHorizontalLoopFusion水平融合:合并独立的并行循环
DoVerticalLoopFusion垂直融合:合并生产者-消费者循环
BindCudaAxis绑定循环到 CUDA threadIdx/blockIdx
AllocateStorage分配 shared memory 和 local buffer
编译缓存(CompilationCache)

CINN 对已编译的 GroupOp 结果进行缓存(基于 FusionInfo hash),相同结构的子图可直接复用已编译的 Kernel,避免重复编译开销。

执行(PirInterpreter)

编译完成的 Kernel 最终由 PirInterpreter 调度执行:

StandaloneExecutor
  └─ PirInterpreter (per Job)
       │
       ├─ Build(首次 Run,结果缓存)
       │   ├── 为每个 Op 构建 Instruction(Kernel 选择 + 数据传输插入)
       │   ├── 构建算子依赖 DAG → 传递性边消除
       │   ├── PirStreamAnalyzer 流调度分类(direct / event / sync)
       │   └── Variable 引用计数 → GC 生命周期管理
       │
       └─ Scheduling(每次 Run)
            ├── dep_count=0 的 Instruction 推入 work queue
            ├── 线程池并行派发 → kernel launch
            ├── 跨 stream 同步:cudaEventRecord + cudaEventWait
            └── ref_count=0 时回收 Variable 内存

CINN 编译产物通过 CinnJitInstruction 执行:从 CINNKernelInfo 获取 fn_ptr,收集输入输出 device pointer,调用 cuLaunchKernel。非 CINN 算子则通过 PHI Kernel 常规路径执行。

Show full SKILL.md (261 more words)Show less

调试速查

场景应关注的文件
SOT 捕获失败 / fallback 过多python/paddle/jit/sot/opcode_translator/executor/opcode_executor.py — 检查未支持的 opcode
SOT SIR 到 Program 编译失败python/paddle/jit/sot/symbolic/compile_cache.py — to_static(full_graph=True) 环节
PIR 动态 shape 推导错误paddle/pir/src/dialect/shape/transforms/shape_optimization_pass.cc
CINN 融合策略问题paddle/cinn/hlir/dialect/operator/transforms/add_cinn_pass.cc
CINN 动态 shape 算子映射paddle/cinn/hlir/dialect/operator/transforms/pd_to_cinn_pass.cc — PdOpToDynamicShapeCinnOpPass
CINN 编译缓存命中 / 未命中paddle/cinn/hlir/framework/pir/compilation_cache.cc
CINN Schedule 调试paddle/cinn/ir/group_schedule/dy_shape_group_scheduler.cc
CINN CodeGen CUDA 源码paddle/cinn/backends/codegen_cuda_dev.cc
执行器 Kernel 启动paddle/fluid/framework/new_executor/pir_interpreter.cc
执行器依赖分析 / 调度paddle/fluid/framework/new_executor/interpreter/stream_analyzer.cc
执行器 Variable 内存泄漏paddle/fluid/framework/new_executor/garbage_collector/

什么场景看什么文件

场景参考文档
SOT 架构设计(eval_frame / OpcodeExecutor / Guard / Fallback)references/sot-design.md
PIR 类型系统、Dialect、Trait/Interface 设计references/pir-basics.md
PIR Program/Value/Operation 内存结构、ProgramTranslatorreferences/pir-program.md
CINN 从 GroupOp 到 CUDA Kernel 的完整编译流程references/cinn-pipeline.md
PIR 控制流(IfOp/WhileOp)、反向 Stack 机制references/control-flow.md
PIR 执行器(PirInterpreter)、Instruction 调度、Stream 分析、GCreferences/executor.md

源码入口

SOT
模块路径
to_static 入口(full_graph 分发)python/paddle/jit/api.py
eval_frame 入口python/paddle/jit/sot/opcode_translator/eval_frame_callback.py
OpcodeExecutorpython/paddle/jit/sot/opcode_translator/executor/opcode_executor.py
OpcodeInlineExecutorpython/paddle/jit/sot/opcode_translator/executor/opcode_inline_executor.py
Variable 体系python/paddle/jit/sot/opcode_translator/executor/variables/
Trackerpython/paddle/jit/sot/opcode_translator/executor/tracker.py
Guardpython/paddle/jit/sot/opcode_translator/executor/guard.py
FunctionGraphpython/paddle/jit/sot/opcode_translator/executor/function_graph.py
StatementIRpython/paddle/jit/sot/symbolic/statement_ir.py
SIR 编译缓存python/paddle/jit/sot/symbolic/compile_cache.py
SideEffectpython/paddle/jit/sot/opcode_translator/executor/side_effects.py
符号 Shape 推导python/paddle/jit/sot/symbolic_shape/
PIR
模块路径
PIR 核心paddle/pir/include/core/ — type.h, value.h, operation.h, block.h, program.h
IRContext / StorageManagerpaddle/pir/src/core/ir_context.cc, storage_manager.cc
Dialect 基类paddle/pir/include/core/dialect.h
PaddleDialectpaddle/fluid/pir/dialect/operator/ir/op_dialect.h
控制流 Dialectpaddle/pir/include/dialect/control_flow/ir/cf_op.h, cf_type.h
控制流 Op 实现paddle/fluid/pir/dialect/operator/ir/control_flow_op.h
Shape Dialectpaddle/pir/include/dialect/shape/
ShapeOptimizationPasspaddle/pir/src/dialect/shape/transforms/shape_optimization_pass.cc
InferSymbolicShape 接口paddle/pir/include/dialect/shape/interface/infer_symbolic_shape/
Pass 框架paddle/pir/include/pass/pass.h, pass_manager.h
Pattern Rewritepaddle/pir/include/pattern_rewrite/pattern_match.h
DecompInterface(Prim 前向分解接口)paddle/fluid/pir/dialect/operator/interface/decomp.h
组合算子(Prim)
模块路径
前向分解规则paddle/fluid/primitive/decomp_rule/decomp_rule/composite.h
反向分解规则(VJP)paddle/fluid/primitive/decomp_rule/decomp_vjp/details.h
分解调度入口paddle/fluid/primitive/base/decomp_trans.cc
Primitive 基础算子paddle/fluid/primitive/primitive/primitive.h
VJP 接口paddle/fluid/primitive/vjp_interface/vjp.h
Backend 适配paddle/fluid/primitive/backend/backend.h
CINN
模块路径
CINN 总入口 Passpaddle/cinn/hlir/dialect/operator/transforms/add_cinn_pass.cc
算子映射(含动态 shape)paddle/cinn/hlir/dialect/operator/transforms/pd_to_cinn_pass.cc
算子融合paddle/cinn/hlir/dialect/operator/transforms/cinn_group_cluster_pass.cc
PirCompilerpaddle/cinn/hlir/framework/pir_compiler.cc
OpLower 实现paddle/cinn/hlir/framework/pir/op_lowering_impl.cc
编译任务paddle/cinn/hlir/framework/pir/compilation_task.cc
编译缓存paddle/cinn/hlir/framework/pir/compilation_cache.cc
DynamicShapeGroupSchedulerpaddle/cinn/ir/group_schedule/dy_shape_group_scheduler.cc
CodeGenpaddle/cinn/backends/codegen_cuda_dev.cc
NVRTC 编译paddle/cinn/backends/nvrtc/nvrtc_util.cc
CINNKernelInfo 定义paddle/cinn/hlir/framework/pir/utils.h
JitKernelOp 定义paddle/cinn/hlir/dialect/runtime/ir/jit_kernel_op.h
AST IR 节点paddle/cinn/ir/
Schedule 原语paddle/cinn/ir/schedule/
执行器(PIR-based)
模块路径
Python Executor 入口python/paddle/base/executor.py
StandaloneExecutorpaddle/fluid/framework/new_executor/standalone_executor.cc
InterpreterCore 统一入口paddle/fluid/framework/new_executor/interpretercore.cc
PirInterpreterpaddle/fluid/framework/new_executor/pir_interpreter.cc
ProgramInterpreter(旧 IR 兼容)paddle/fluid/framework/new_executor/program_interpreter.cc
PirStreamAnalyzerpaddle/fluid/framework/new_executor/interpreter/stream_analyzer.cc
Instruction 定义paddle/fluid/framework/new_executor/instruction/
CinnJitInstructionpaddle/fluid/framework/new_executor/instruction/
Scope(变量容器)paddle/fluid/framework/scope.cc
GC 实现paddle/fluid/framework/new_executor/garbage_collector/

© PaddlePaddle, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in .agents/skills/paddle-design-compiler of PaddlePaddle/Paddle.

  • SKILL.md
  • references/cinn-pipeline.md
  • references/control-flow.md
  • references/executor.md
  • references/pir-basics.md
  • references/pir-program.md
  • references/sot-design.md

Open the folder on GitHubat commit 5434c21

Compare with similar skills

Paddle Design Compiler next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Paddle Design Compiler compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Paddle Design Compiler this skillPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Magpie Kernel Evaluatoramd/skills406—~2.3kAutomated safety check: PassMIT
Hyperpod Version Checkerawslabs/agent-plugins915—~910Automated safety check: PassApache-2.0
Triton Langmohitmishra786/low-level-dev-skills253—~1.8kAutomated safety check: PassMIT
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Cuda Index Widthpytorch/pytorch104k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Benchmarks LLM inference and drives GPU kernel optimization with Magpie.

    406 GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Hyperpod Version Checker

    awslabs/agent-plugins

    Official

    Check and compare software component versions on SageMaker HyperPod cluster nodes - NVIDIA drivers, CUDA toolkit, cuDNN, NCCL, EFA, AWS OFI NCCL, GDRCopy, MPI, Neuron SDK (Trainium/Inferentia)…

    915 GitHub stars~910 tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Triton Lang

    mohitmishra786/low-level-dev-skills

    Triton language skill for Python GPU kernel authoring. An agent skill from mohitmishra786/low-level-dev-skills.

    253 GitHub stars~1.8k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Cuda Index Width

    pytorch/pytorch

    Choose 32-bit vs 64-bit index math in PyTorch CUDA kernels. An agent skill from pytorch/pytorch.

    104k GitHub stars~1.6k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Graphsignal

    graphsignal/graphsignal

    Profile AI inference workloads (vLLM, SGLang, TensorRT-LLM, PyTorch, any GPU application) with the Graphsignal profiler and read the results from its local /signals JSON endpoint.

    257 GitHub stars~6.2k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed

More from PaddlePaddle/Paddle

All 11 skills in this repo
  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Paddle Design Distributed

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

    24k GitHub stars~660 tokensUpdated today
    Auto-check passed
  • Paddle Eager Graph

    PaddlePaddle/Paddle

    A skill your agent uses when navigating Paddle eager-mode (dynamic graph) source code, tracing forward/backward execution, debugging autograd issues, understanding PyLayer, or investigating…

    24k GitHub stars~562 tokensUpdated today
    Auto-check passed
  • Paddle Phi Kernel

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's PHI kernel system: registering new kernels, debugging kernel selection/dispatch, understanding code auto-generation from YAML, or implementing…

    24k GitHub stars~656 tokensUpdated today
    Auto-check passed
  • Paddle Debug

    PaddlePaddle/Paddle

    在 Paddle 代码库中定位问题并输出高质量调试报告的专用技能。当遇到以下场景时优先使用:(1) Paddle 框架 bug 调试,(2) 算子实现问题排查,(3) 训练脚本异常诊断,(4) 分布式训练故障定位,(5) CUDA/GPU 相关错误处理,(6) 需要生成结构化调试报告。

    24k GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Paddle Op Dev

    PaddlePaddle/Paddle

    PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…

    24k GitHub stars~1.3k tokensUpdated today
    Auto-check passed

Works with

Questions about Paddle Design Compiler

What does Paddle Design Compiler do?

A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…. Paddle Design Compiler is an agent skill from PaddlePaddle/Paddle.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate representation, CINN for fused CUDA kernel generation, operator decomposition (Prim), or the end-to-end flow from Python eager code to optimized GPU execution.

When should I use Paddle Design Compiler?

Paddle Design Compiler fits situations like: working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture; PIR (Paddle IR) for SSA-based intermediate representation; CINN for fused CUDA kernel generation; operator decomposition (Prim).

How do I install Paddle Design Compiler in Claude Code?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-design-compiler -a claude-code`. Or copy the skill folder (.agents/skills/paddle-design-compiler in PaddlePaddle/Paddle) into .claude/skills/paddle-design-compiler in your project. Claude Code loads it when a task matches its description.

How do I install Paddle Design Compiler in Codex?

Run `npx skills add PaddlePaddle/Paddle --skill paddle-design-compiler -a codex`. Or copy the skill folder (.agents/skills/paddle-design-compiler in PaddlePaddle/Paddle) into .agents/skills/paddle-design-compiler in your project. Codex loads it when a task matches its description.

Can I use Paddle Design Compiler in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add PaddlePaddle/Paddle --skill paddle-design-compiler -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/paddle-design-compiler, .gemini/skills/paddle-design-compiler, .github/skills/paddle-design-compiler and .opencode/skills/paddle-design-compiler in your project.

What does Paddle Design Compiler need to run?

Going by SKILL.md and its folder, Paddle Design Compiler needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does Paddle Design Compiler access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Paddle Design Compiler safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Paddle Design Compiler use?

Paddle Design Compiler is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Paddle Design Compiler use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.7k tokens, read only when the agent opens those files.

What are the alternatives to Paddle Design Compiler?

Skills that share tags, products or a category with Paddle Design Compiler: Magpie Kernel Evaluator (amd/skills, 406 stars), Hyperpod Version Checker (awslabs/agent-plugins, 915 stars), Triton Lang (mohitmishra786/low-level-dev-skills, 253 stars) and Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Paddle Design Compiler?

PaddlePaddle (a GitHub organization) maintains it in PaddlePaddle/Paddle, which has 24,121 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 9, 2026.

Source: PaddlePaddle/Paddle on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.