Paddle Op Dev
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
$ npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install TongmingLAIC/AKO4ALL ako4all --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ako4all" agent skill from https://github.com/TongmingLAIC/AKO4ALL/tree/main into .claude/skills/ako4all/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ako4all", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install TongmingLAIC/AKO4ALL ako4all --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ako4all" agent skill from https://github.com/TongmingLAIC/AKO4ALL/tree/main into .agents/skills/ako4all/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ako4all", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install TongmingLAIC/AKO4ALL ako4all --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ako4all" agent skill from https://github.com/TongmingLAIC/AKO4ALL/tree/main into .cursor/skills/ako4all/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ako4all", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install TongmingLAIC/AKO4ALL ako4all --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ako4all" agent skill from https://github.com/TongmingLAIC/AKO4ALL/tree/main into .gemini/skills/ako4all/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ako4all", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install TongmingLAIC/AKO4ALL ako4allInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ako4all" agent skill from https://github.com/TongmingLAIC/AKO4ALL/tree/main into .github/skills/ako4all/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ako4all", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install TongmingLAIC/AKO4ALL ako4all --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ako4all" agent skill from https://github.com/TongmingLAIC/AKO4ALL/tree/main into .opencode/skills/ako4all/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ako4all", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ako4allDrive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
Ako4all is an agent skill from TongmingLAIC/AKO4ALL. Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. Use this skill whenever the user wants to optimize / speed up / benchmark a GPU kernel (CUDA, Triton, TileLang, C++, Python), mentions AKO / AKO4ALL / AKO4X / agentic kernel optimization, asks to "make this kernel faster", or has a kernel they want measured against a PyTorch reference. The skill handles setup, profiling (ncu), correctness checking, iteration logging, and git commits. Bootstraps a workspace in any directory the user…
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 22 other files, including assets (for example `HINTS.md`, `ITERATIONS.md` and `README.md`).
It sits in AI & LLM Engineering, covering Deep learning and Commit messages. It works with CUDA, PyTorch, C++ and Python. The repository describes itself as: Agentic Kernel Optimization for All — automated GPU kernel optimization for any kernel, any hardware, any language. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit bbd0e1c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Shell and Python, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
gitbashpythoncondaFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ako4all loads about 4k tokens when it runs. Until then it costs about 134 tokens; SKILL.md has 2,146 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from TongmingLAIC/AKO4ALL at commit bbd0e1c, republished under its MIT licence (© TongmingLAIC). 2,146 words, ~4,005 tokens.
.claude/skills/ako4all/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.Drive a profile → modify → benchmark → log → commit loop on a GPU kernel until it runs faster than the reference. The user provides at minimum a kernel; everything else (reference, inputs, bench script, hints) is optional.
ncu, kernel profiling, GPU speedup targetDoes NOT apply when:
codex:rescue instead.Before doing anything else, establish the workspace — the directory the loop runs in. It is typically the user's CWD, or a subdirectory / path they name in the prompt.
Browse the workspace (don't run a fixed checklist — look around) and read the user's prompt to identify what the loop needs:
.npz, .bin, shape lists, custom formats, etc.)knowledge/ but anywhere the user points at.bench/kernelbench/ evaluatorbench-wrapper.sh, HINTS.md, ITERATIONS.md, bench/kernelbench/ are already at workspace rootWhether the workspace follows AKO4ALL's source/ / knowledge/ / bench/ naming or some entirely different layout is not the signal. What matters is whether you can identify each item above with confidence.
If the user's prompt + filesystem give you confidence about every required item, don't ask — skip straight to presenting the plan. Ask only when a piece's role is genuinely ambiguous (a kernel-shaped file with no obvious reference, two files that could both be the kernel, an input data file in an unfamiliar format you need permission to wire up a custom way, etc.). When in doubt, asking is cheaper than guessing wrong.
Whether you asked the user or not, list back what you decided — so the user can correct you even when you didn't think you needed to ask.
Use the format below. Bold field labels + inline-code path values + the leading emoji marker make the plan visually scannable in any terminal theme (don't flatten to a wall of prose):
📋 Resolved Plan
<path><path><path> (or none — will use original kernel)<path> (or inline in ref, or none)<path> (or none)<path>)<list of missing files> (or none — already present)If anything still feels uncertain at this point, stop and ask. Otherwise proceed to Workflow.
When copying scaffold (bench-wrapper.sh, bench/kernelbench/, starter HINTS.md / ITERATIONS.md, workspace.gitignore → as .gitignore in the workspace) from this skill's own directory into the workspace, do not overwrite files that already exist — the user may have edited HINTS.md, or ITERATIONS.md may carry prior iteration history. Copy only what's missing.
The user may supply behavior directives in two ways:
In both cases, merge those directives into HINTS.md. It's the persistence
layer — directives that only live in the current session's plan are lost on resume.
Whenever you merge directives into HINTS.md, tell the user explicitly what
happened. Example phrasings:
"I added your 'avoid shared memory' directive from the prompt to HINTS.md." "I added the 3 rules from /tmp/user-hints.md to HINTS.md."
Without this acknowledgment the user can't tell from your reply whether you added, replaced, or silently dropped their directives. Always name the source ("from your prompt" / "from /tmp/x.md").
Analyze inputs. Building on the inventory above, confirm class Model and get_inputs() can be assembled for default bench mode; if not, stop and ask the user. See bench/kernelbench/GUIDE.md for the input assembly contract (KernelBench-format input / raw kernel / kernel + separate data file / external path patterns).
Create branch. git checkout -b opt/<kernel-name>. If the workspace isn't a git repo, init one first.
Initialize solution. Create solution/ and scripts/. Copy the kernel implementation files into solution/ (only the kernel itself — reference / inputs helper files stay at their resolved locations). Do not copy or mkdir canonical directories (source/, input/, etc.) when the user's files already exist elsewhere. Point bench.sh's --ref and --inputs flags at the resolved paths in place. solution/ is the only directory the loop owns.
Generate bench.sh. Build the bench command with adjusted paths, pipe through 2>&1 | tee _bench_output.txt. Replace {{BENCH_COMMAND}} in bench-wrapper.sh to produce scripts/bench.sh. For default bench mode the command is python bench/kernelbench/bench.py --ref <ref> --solution solution/<kernel> [--inputs <inputs-file>] --verbose — include --inputs only when inputs are defined outside the ref file. Do not hardcode --backend in the rendered command; bench.py auto-detects backend from solution source. Add --backend only to override the sniff (explicit HIP labelling or mixed-backend solutions). scripts/bench.sh is a starting template — when the bench env needs setup (conda activate, sub-env python paths, multi-CUDA toolkit selection), edit it freely; preserve only the trajectory section (LABEL/TIMESTAMP handling and cp -r solution/* "$TRAJ_DIR/").
Common env friction: base shell often has no python on PATH when python lives in a sub-env (e.g. ~/anaconda3/envs/py312/bin/python). Tools that internally subprocess python (sol-execbench CLI, torch cpp_extension.load_inline) will then fail with command not found. Workaround: put PATH=<env-bin>:$PATH at the top of scripts/bench.sh (or source <conda>/etc/profile.d/conda.sh && conda activate <env>). Discover sub-envs via ls /home/*/anaconda3/envs/*/bin /root/*/envs/*/bin /opt/conda/envs/*/bin 2>/dev/null.
Verify baseline. Run bash scripts/bench.sh. Expect CORRECT=True. If not, diagnose and fix before iterating. Commit: git add -A && git commit -m "[baseline] Initialize solution and benchmark". Then run ncu once on the baseline to inform iter-1 direction.
Every modification to solution/ followed by a bench run = one iteration. Number sequentially (1, 2, 3, …). Each iter is exactly three steps:
bash scripts/bench.sh iter-N — label is required, must match iter-N format.ITERATIONS.md (template inside that file).git commit -m "[iter N] <short description of optimization direction>".Steps 2 and 3 MUST be the next two tool calls after step 1 — no ncu, no probes, no reads, no planning the next iter between them. A failed or partial bench is still an iter; log + commit first, debug after. This is the most-missed step in practice: agents read the bench result and telescope into next-iter analysis (probes, ncu, hypothesis forming) without closing out the current one, leaving commit gaps with ITERATIONS.md entries written from memory later.
Backstop: if you catch yourself starting a new iter (Editing solution/, or running ncu/probes for the next direction) and git log -1 doesn't show [iter N] ..., stop and finish the prior iter's steps 2 and 3 first. Related experiments that belong together narratively get grouped in ITERATIONS.md analysis prose, not in batched git commits.
Profile to identify bottlenecks — see "ncu profiling" below for the ncu workflow and analytical fallback. Do not optimize blindly.
A bench run must be cheap enough to iterate against (seconds to low minutes). When it isn't, the cause is almost always an expensive reference — it's re-run for every correctness trial and re-timed for the speedup denominator, yet it's invariant across solution edits, so most of that cost is wasted. This is an eval-time problem, not a metric problem: never change what you compare against to make the bench cheaper.
Separate the per-iteration signal from the final verdict:
RUNTIME (lower is better) — the reference contributes nothing to comparing two solutions.final): a full run — full trial counts, reference measured, real SPEEDUP, reward-hack check.Whose eval is it determines what you may touch:
bench/kernelbench/bench.py — the skill owns it): pull levers freely, cheapest-and-safest first — --no-ref (skip reference timing; REF_RUNTIME/SPEEDUP → -1, COMPILED/CORRECT/RUNTIME unaffected) → trim --num-perf-trials (e.g. 50→20; latency noise only) → trim --num-correct-trials (higher risk: weakens the fresh-input anti-cheat and still runs the ref once per trial, so keep ≥1 in the loop, full at the gate). Keep the input regime fixed across the whole run: --fresh-inputs (the default) and --no-fresh-inputs measure different quantities, so switching mid-run makes iterations incomparable — and the final verdict must use the same regime the loop ranked under.{{BENCH_COMMAND}}): the trial counts, correctness rounds, and reference handling are the user's contract — do not inject --no-ref or cut counts on a script you didn't author (it may have no such flag, break the interface, or silently invalidate the measurement / a leaderboard's required N). Use only the fast-iteration switches the user exposed (flags / env vars documented in the prompt or HINTS.md). If iteration is too slow and none exist, raise it with the user — don't fabricate one.Caching the reference's runtime across iterations is sound only on a clock-locked GPU; on unlocked clocks (ref and solution timed in different clock states) prefer ranking by the solution's own latency.
When 3 consecutive iterations show no improvement (≥3% over current best), pause the loop and re-assess before iter N+1. Re-assessment combines:
ncu if available, or re-read runtime stats from ITERATIONS.md (median vs min/mean, distribution shape) if not.ITERATIONS.md for patterns (which axes have been tried, which haven't, where prior wins came from).Default outcome: pick a new direction and continue. Only escalate to stop (see next section) if re-assessment produces concrete evidence the current state is at a physical floor.
Legitimate triggers:
HINTS.md).ITERATIONS.md.ITERATIONS.md before invoking this trigger, to prevent premature stops.Do not stop silently because tooling is unavailable — that's a re-assessment input, not a stop reason.
After deciding to stop, leave HEAD at the best-performing iter — not necessarily the latest. Procedure:
Identify the best iter by reading ITERATIONS.md Summary, the bench output for each iter under trajectory/, and your own reasoning notes. Useful signals from KernelBench output: median speedup, runtime std (consistency), min runtime (tail), CORRECT flag. Other bench harnesses expose different shapes — use what's available. Justify your pick in the commit message (e.g., "iter 4: best mean AND lowest min, while iter 6 ties on mean but has higher std").
If best iter ≠ latest iter:
git checkout <best-iter-sha> -- solution/ — verbatim copy, do NOT hand-reconstruct from memory or earlier notes.bash scripts/bench.sh final to sanity-verify on a fresh run.git commit -m "[final] Restore iter-K (X.XXx) — <one-sentence why>".The git checkout step is mandatory. Manual reconstruction risks introducing silent drift from the actually-benched code.
Probe ncu once after baseline. If it fails (driver mismatch, missing toolkit, user opt-out via free-text HINTS.md directive), proceed analytically for the rest of the loop without re-probing within this session. Don't gate iteration progress on ncu availability; analytical reasoning + runtime stats from the bench harness are a valid substitute for the optimizer's direction picking.
get_inputs / get_init_inputs. The bench script strips the solution's module-level tail before eval as an anti-cheat boundary. Inputs come from the reference or --inputs file, never the solution.get_inputs() must produce fresh data each call. Bench calls it many times — every correctness trial, and (under the default fresh regime) before every timed trial. Module-level cached tensors make correctness checks trivially pass and let timing measure cache-warm performance. Use torch.randn or reload from disk on every call.bench/kernelbench/GUIDE.md — full input assembly patterns, CLI flags, timing methods, tolerances. Read this before writing the bench command if anything about input shape, precision, or backend is non-obvious.HINTS.md (workspace) — user-editable behavior directives. Read at session start; respect any constraint named there.ITERATIONS.md (workspace) — your own iteration log. Write to it every iteration.© TongmingLAIC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 18 other files (assets) in the repository root of TongmingLAIC/AKO4ALL.
Open the folder on GitHubat commit bbd0e1c
Ako4all next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ako4all this skillTongmingLAIC/AKO4ALL | 369 | — | ~4k | Automated safety check: Pass | MIT | |
| Paddle Op DevPaddlePaddle/Paddle | 24k | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| ExecuTorch Build Guidepytorch/executorch | 5.1k | — | ~2.3k | Automated safety check: Notes | Custom licence | |
| Paddle BuildPaddlePaddle/Paddle | 24k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Fix Envevo-design/proto-tools | 135 | — | ~2.5k | Automated safety check: Notes | MIT | |
| Migrate Workflow Ec2 To Osdcpytorch/test-infra | 113 | — | ~2k | Automated safety check: Pass | Custom licence |
PaddlePaddle/Paddle
PaddlePaddle (飞桨) C++ 算子开发指南。提供从 YAML 配置、InferMeta 函数、Kernel 实现、Python API 封装、单元测试到编译验证的完整算子开发流程指导。在以下场景使用此 skill:(1) 为 Paddle 框架新增 C++ 算子 (2) 修改或调试已有 Paddle 算子 (3) 编写算子的 YAML…
pytorch/executorch
Builds ExecuTorch from source: the Python package, C++ runtime, model runners, Android and iOS cross-compilation and backend-specific builds, with environment checks.
PaddlePaddle/Paddle
A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.
evo-design/proto-tools
Fixes tool environment setup failures in proto-tools, either just for the current machine (eject the tool's standalone dir, patch it, and point PROTO<TOOLKITSTANDALONEDIR at it; works for any…
pytorch/test-infra
Step-by-step playbook for migrating a pytorch/pytorch .github/workflows/.yml from EC2 to OSDC (ARC) runners — covers both dial-up and 100% opt-in patterns, with the inputs that must be plumbed…
PaddlePaddle/Paddle
将原生 PyTorch 自定义算子库、Torch extension、生态库(TorchCodec/FlashInfer/DeepEP 等)以及 Kernel DSL 生态(Triton/TileLang/TVM FFI 等)以最小修改方式接入 PaddlePaddle。遇到以下场景务必使用:迁移外部算子库到 Paddle;分析 PFCCLab fork 与上游的兼容差异;处理…
Categories
Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup. Ako4all is an agent skill from TongmingLAIC/AKO4ALL. Drive an agentic loop that iteratively optimizes a GPU kernel for maximum speedup.
Ako4all fits situations like: the user wants to optimize / speed up / benchmark a GPU kernel (CUDA; mentions AKO / AKO4ALL / AKO4X / agentic kernel optimization; asks to make this kernel faster; has a kernel they want measured against a PyTorch reference.
Run `npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a claude-code`. Or copy the skill folder (the TongmingLAIC/AKO4ALL repository) into .claude/skills/ako4all in your project. Claude Code loads it when a task matches its description.
Run `npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a codex`. Or copy the skill folder (the TongmingLAIC/AKO4ALL repository) into .agents/skills/ako4all in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add TongmingLAIC/AKO4ALL --skill ako4all -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ako4all, .gemini/skills/ako4all, .github/skills/ako4all and .opencode/skills/ako4all in your project.
Going by SKILL.md and its folder, Ako4all needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (git, bash, python and conda). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ako4all is published under the MIT licence (from the LICENSE file in the skill folder). It allows redistribution, so the full SKILL.md is shown on this page.
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ako4all: Paddle Op Dev (PaddlePaddle/Paddle, 24k stars), ExecuTorch Build Guide (pytorch/executorch, 5.1k stars), Paddle Build (PaddlePaddle/Paddle, 24k stars) and Fix Env (evo-design/proto-tools, 135 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
TongmingLAIC (a GitHub organization) maintains it in TongmingLAIC/AKO4ALL, which has 369 GitHub stars. The repository was last updated on September 15, 2026.
Source: TongmingLAIC/AKO4ALL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.