Agent Builder
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
Run the evidence loop of one long-horizon GPU kernel optimization episode.
$ npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install alibaba/atrex-kernel-agent gpu-kernel-episode-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu-kernel-episode-loop .claude/skills/gpu-kernel-episode-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "gpu-kernel-episode-loop" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/gpu-kernel-episode-loop into .claude/skills/gpu-kernel-episode-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kernel-episode-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/gpu-kernel-episode-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install alibaba/atrex-kernel-agent gpu-kernel-episode-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/gpu-kernel-episode-loop .agents/skills/gpu-kernel-episode-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "gpu-kernel-episode-loop" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/gpu-kernel-episode-loop into .agents/skills/gpu-kernel-episode-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kernel-episode-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install alibaba/atrex-kernel-agent gpu-kernel-episode-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/gpu-kernel-episode-loop .cursor/skills/gpu-kernel-episode-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "gpu-kernel-episode-loop" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/gpu-kernel-episode-loop into .cursor/skills/gpu-kernel-episode-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kernel-episode-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/alibaba/atrex-kernel-agent.git --path skills/gpu-kernel-episode-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install alibaba/atrex-kernel-agent gpu-kernel-episode-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/gpu-kernel-episode-loop .gemini/skills/gpu-kernel-episode-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "gpu-kernel-episode-loop" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/gpu-kernel-episode-loop into .gemini/skills/gpu-kernel-episode-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kernel-episode-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install alibaba/atrex-kernel-agent gpu-kernel-episode-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/gpu-kernel-episode-loop .github/skills/gpu-kernel-episode-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "gpu-kernel-episode-loop" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/gpu-kernel-episode-loop into .github/skills/gpu-kernel-episode-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kernel-episode-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install alibaba/atrex-kernel-agent gpu-kernel-episode-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/gpu-kernel-episode-loop .opencode/skills/gpu-kernel-episode-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "gpu-kernel-episode-loop" agent skill from https://github.com/alibaba/atrex-kernel-agent/tree/main/skills/gpu-kernel-episode-loop into .opencode/skills/gpu-kernel-episode-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "gpu-kernel-episode-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
gpu-kernel-episode-loopRun the evidence loop of one long-horizon GPU kernel optimization episode.
GPU Kernel Episode Loop is an agent skill from alibaba/atrex-kernel-agent. Run the evidence loop of one long-horizon GPU kernel optimization episode. Use this skill to reconstruct the incumbent, profile and localize a bottleneck, research progressively, plan one coherent direction, implement and repair, validate development correctness and performance, and record every decisive experiment in the episode journal.
Its SKILL.md is about 3.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering. The repository describes itself as: An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a… The licence is Apache-2.0.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3d27c1e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonpython3bashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
GPU Kernel Episode Loop loads about 3.6k tokens when it runs. Until then it costs about 91 tokens; SKILL.md has 1,670 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from alibaba/atrex-kernel-agent at commit 3d27c1e, republished under its Apache-2.0 licence (© alibaba). 1,670 words, ~3,602 tokens.
.claude/skills/gpu-kernel-episode-loop/SKILL.md (or your agent's skills folder).Use this skill when an orchestrator episode prompt hands you one optimization episode in an isolated
Git worktree and points here for the evidence loop. It does not apply to the V0 baseline session
(gpu-kernel-baseline) or to a workspace without an episode journal.
The episode prompt supplies a concrete value for every ALL-CAPS bracketed name below. Substitute them before running any command; never invent a path. Lowercase bracketed names in the command examples are values you fill in from the campaign, or a choice among the listed alternatives, not bindings.
| Placeholder | Meaning |
|---|---|
<PROFILE_DIR> | profile output directory for this episode |
<PLAN_DRAFT> | evidence draft for the canonical version |
<PLAN_FILE> | generated plan for the canonical version |
<PLAN_GENERATOR> | backend-native plan generator invocation |
<JOURNAL_CLI> | episode journal command prefix |
<JOURNAL_PATH> | episode journal path, already shell-quoted |
The episode prompt's ownership rules, execution boundary, mode policy, and framework-escalation directive outrank this skill. Where they conflict, follow the prompt.
Use the Python-supplied ATREX_EPISODE_MODE (default: full), as defined in
skills/gen-plan/SKILL.md. Fast/full episodes advance one coherent engineering direction.
Goal episodes can repeat this loop across a roadmap of directions, preserving the best
validated checkpoint until the roadmap is complete or exhausted.
Telemetry is best-effort and must not block engineering work. Mark phase boundaries with standalone commands and keep at most one phase active:
python3 tools/iteration_trace.py phase-start <profile|research|planning|implementation|correctness|benchmark|recording>
python3 tools/iteration_trace.py phase-end <profile|research|planning|implementation|correctness|benchmark|recording>
python3 tools/iteration_trace.py source-read <gpu_wiki|reference_projects|workspace|public_web> <safe-relative-reference>Never put credentials, private URL parameters, absolute user paths, raw tool output, or transcript text into telemetry.
Repeat this evidence loop until the direction (fast/full) or roadmap (goal) yields a mature candidate or is exhausted. The numbered
steps map onto the telemetry phases above: profile, research, planning, implementation,
correctness/benchmark, and recording.
Read the workspace goal, unmasked memory/v*.json, and prior plans/profiles. Prior-episode summaries
are carried only by canonical memory and are not injected into the episode prompt. Identify attempted
dead ends and open directions from those records, including each record's compact
experience.experiments. For PPU, also inspect
profile_evidence.accepted_ppu_diagnostics and reuse a conclusion only when its recorded identities
remain comparable and none of its invalidation_conditions holds. Start with one falsifiable
hypothesis tied to the current bottleneck.
For a PPU target, do not apply the NVIDIA/AMD default below. Read
skills/ppu-acu-joint-profile/SKILL.md first and let its PPU-specific per-iteration rule decide
whether new PPU profiler evidence is needed. That route may use source or compiler inspection, the
probe-free benchmark, or still-valid PPU evidence instead of collecting a new profile.
Reuse a profile only when it matches the current committed kernel. Otherwise profile through the
sandbox using the vendor-appropriate tooling. Both wrappers run python <file>, so the profiled file
is the immutable profile_driver.py seeded next to kernel.py — never kernel.py itself, which the
evaluator only ever imports:
# NVIDIA
python tools/sandbox.py --kind profile --sync <PROFILE_DIR> -- \
bash tools/profile_nvidia.sh profile_driver.py --output-dir <PROFILE_DIR> --source
# AMD
python tools/sandbox.py --kind profile --sync <PROFILE_DIR> -- \
bash tools/profile_kernel.sh profile_driver.py --output-dir <PROFILE_DIR>profile_driver.py imports the current kernel.py, builds real inputs from the campaign contract
(definition.json + workload.jsonl, a privately injected generalized Atrex-Bench real shape,
or legacy shapes.json + input.py), warms up, and then
invokes the candidate repeatedly. Select what it drives with environment variables rather than editing
it — it is a protected path and a candidate that modifies it is rejected:
PROFILE_ITERS=30 PROFILE_WORKLOAD_IDX=2 python tools/sandbox.py ... # SOL: one workload
PROFILE_ITERS=30 PROFILE_SHAPE_ID=3 python tools/sandbox.py ... # generalized or legacy Atrex-BenchWhen several shapes or workloads need profiling, run one sandbox command per id in waves of at most
four concurrent jobs. Give every job its own <PROFILE_DIR>/shape-<index> sync/output directory and
wait for the whole wave before starting the next one.
For generalized Atrex-Bench tasks, choose PROFILE_SHAPE_ID from the previous canonical memory's
complete opaque-id performance.latency_us_by_shape map. The sandbox privately resolves that id and
injects only its real input case into the ephemeral remote profile job; the driver deletes the case
JSON before importing candidate code. Profile the highest-cost ids and additional ids representing
distinct latency regimes, but do not infer or reconstruct the complete hidden input table.
Extract a concrete bottleneck and source-level target. Use PTX/SASS/TTGIR inspection when compiler lowering or instruction selection is part of the hypothesis. Do not make speculative optimization changes before obtaining usable evidence.
When ordinary profiling has isolated one kernel but cannot distinguish a specific in-kernel timing
hypothesis, read skills/autonomous-gpu-kernel-timeline/SKILL.md and run its autonomous loop. Use
standalone CUDA/inline PTX through its CUDA backend and CuTe DSL through IKeT. Keep every attempt
under <PROFILE_DIR>/timeline/attempt-N; when the remote command reads backend files, pass that
specific skill path with sandbox --input and sync only the attempt output directory.
For PPU, use the routing and capture contracts in the PPU skill linked above.
Timeline instrumentation is a temporary working snapshot on this episode's single HEAD line, not a
candidate. Preserve the clean source and each useful instrumented source or reversible patch before
replacing it. After the evidence answers the question, restore or rewrite a probe-free kernel.py
before correctness/performance validation, commit, journal finalization, and handoff. Never submit a
profiling snapshot as candidate_commit; its latency and failures do not count as promotion attempts
or framework-stall events.
Escalate through the typed profile funnel instead of collecting everything at once: --profile-level survey to enumerate kernels, sol (the default) for the bottleneck class, and deep --kernel-regex '^<exact_base_function_name>$' for one named kernel, especially a Triton @triton.jit entry. Take
that name verbatim from the survey/SOL result; never guess a substring. Raw .ncu-rep/ATT artifacts
stay remote unless --include-raw-profile is justified.
On NVIDIA, summary.txt carries a LOCALIZE line naming the analysis files that pin a symptom to
source lines. Those files exist only on a --source run: never pin a source-level claim to a profile
collected without it.
Build a local fallback driver at <PROFILE_DIR>/harness/profile_driver.py and profile that file
instead when the seeded driver cannot express the case — a multi-kernel sequence, a new synthetic
case inside the public domain, or a driver that needs sibling helper modules. It must import kernel.py plus the
immutable input module, select a representative workload, warm up, invoke the entry point repeatedly,
and never write memory files. Because it lives below <PROFILE_DIR>/harness/, Python does not put the
workspace root on its import path; add it before importing anything local:
import sys
from pathlib import Path
WORKSPACE_ROOT = Path(__file__).resolve().parents[3]
if str(WORKSPACE_ROOT) not in sys.path:
sys.path.insert(0, str(WORKSPACE_ROOT))The sandbox uploads that whole harness/ directory automatically, so sibling helpers need no extra
flags. Only a file opened dynamically by command code needs a repeatable --input <relative-path>
option before --, which routes the job through the dev interface.
Search in this order and stop when one actionable direction is supported:
.atrex_plugins/instructions.md and inspect
python3 tools/plugin.py list. Follow the session's phase-specific plugin instructions.
Profile first, then describe measured symptoms, exact product/runtime architecture, operator,
framework, shapes/dtypes, failed attempts and the fact needed for the next decision. On PPU,
use the decisive evidence selected by ppu-acu-joint-profile, even when no new profile was needed.
Preserve source IDs and attribution exactly. Respect architecture scope and context budgets.
If no suitable plugin is enabled, continue with available references; never invoke a disabled tool.reference-projects/ only when the enabled knowledge tools are absent or insufficient.After repeated rejected episodes, expand across DSLs targeting the same architecture instead of repeating local parameter tweaks. Record stable source ids and the evidence-to-action chain.
Write or update <PLAN_DRAFT> with profile evidence, research findings, concrete edits, risks,
rollback points, and measurable acceptance criteria. For a PPU iteration that did not need a new
profile, record the decisive PPU evidence selected by its routing skill instead. Then produce
<PLAN_FILE> with the backend-native plan generator <PLAN_GENERATOR>.
Fast/full episodes may contain multiple related experiments, but they must advance one coherent engineering direction. Goal episodes may plan and validate multiple directions. Checkpoint useful intermediate states so failed sub-steps can be reverted without losing the whole direction.
Modify only candidate source/metadata files allowed by policy. Compile and probe through the sandbox. On compile or correctness failure, diagnose and repair while the direction remains viable. Do not publish an intermediate checkpoint as a candidate.
Land one optimization category per edit — vectorized load, swizzle, double buffering, tiling change,
and so on — and attribute each edit as evidence -> inference -> action. Do not mix unrelated
refactors, formatting, or cleanup into the same change: a bundled edit makes a regression
unattributable. When the evidence localizes a symptom to specific lines, change those lines only.
Use the immutable evaluator for development measurements:
python tools/sandbox.py --kind run --no-sync -- \
python test_kernel.py --version vlong --no-memory
python tools/sandbox.py --kind run --no-sync -- \
python test_kernel.py --version vlong --multi-seed 5 --no-memoryAll workloads and all additional seeds must pass. Never depend on tensor values, pointer identity, cached outputs, evaluator ordering, or hidden workload IDs. Shape/dtype/layout dispatch is allowed, and pre-converting stable weights (transpose, contiguous) is allowed because weights do not change during evaluation.
Before trusting a large delta — especially a regression beyond roughly 30% — re-run the same command
on the same sandbox hardware and compare. GPU selection belongs to the gateway; never set a local
CUDA_VISIBLE_DEVICES to steer it. Repeated development measurements are not promotion authority;
the supervisor reruns incumbent and candidate in one ABBA allocation.
Immediately after each decisive experiment, append it to the single episode journal. Do not batch
these writes at the end of the episode: every append refreshes the non-canonical memory/live.json
progress view in the incumbent workspace.
<JOURNAL_CLI> append --path <JOURNAL_PATH> \
--experiment-json '{"name":"...","hypothesis":"...","change":"...","evidence":"...","result":"...","evaluation":{"correctness":"pass|fail|unknown","performance":"improved|not_improved|unknown","latency_us":null,"kernel_hash":""},"decision":"keep_as_best|promote|reject_and_continue|revert|pivot|blocked","wiki_usage_status":"not_queried"}'Use declared only with non-empty wiki_usage. Include wiki_query_ids for both declared and
no_material_use; omit both arrays for not_queried.
Do not invent ids and do not collapse repeated use across experiments. Invalid Wiki rows are omitted
with wiki_usage_errors; this diagnostic field never blocks the experiment or handoff.
In fast/full mode, leave the loop as soon as one coherent candidate passes the full development correctness check and
has credible performance evidence, or as soon as the direction is exhausted or blocked. In goal mode, finish or exhaust the roadmap
and restore the best validated checkpoint, or report a blocker. Then follow
the episode prompt's terminal contract for finalizing the journal and publishing the handoff. For a
PPU full episode, include outcome.accepted_ppu_diagnostics using the schema in
skills/ppu-acu-joint-profile/SKILL.md; retain only evidence that still applies to the terminal
probe-free kernel. Each retained row must bind an accepted decision-grade artifact by path, SHA-256,
schema, and evidence id. This is optional when no reusable PPU profiler evidence exists.
© alibaba, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/gpu-kernel-episode-loop of alibaba/atrex-kernel-agent.
Open the folder on GitHubat commit 3d27c1e
GPU Kernel Episode Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| GPU Kernel Episode Loop this skillalibaba/atrex-kernel-agent | 154 | — | ~3.6k | Automated safety check: Pass | Apache-2.0 | |
| Agent BuildershareAI-lab/learn-claude-code | 78k | 6 repos | ~1.2k | Automated safety check: Pass | MIT | |
| Add Uint Supportpytorch/pytorch | 104k | 2 repos | ~2.3k | Automated safety check: Pass | Custom licence | |
| Peft Fine TuningOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.1k | Automated safety check: Pass | MIT | |
| Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs | 13k | 9 repos | ~3.3k | Automated safety check: Pass | MIT | |
| 1passwordtrpc-group/trpc-agent-go | 1.8k | 13 repos | ~656 | Automated safety check: Pass | Apache-2.0 |
shareAI-lab/learn-claude-code
Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.
pytorch/pytorch
Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.
Orchestra-Research/AI-Research-SKILLs
Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.
Orchestra-Research/AI-Research-SKILLs
Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.
trpc-group/trpc-agent-go
Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.
Orchestra-Research/AI-Research-SKILLs
Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.
alibaba/atrex-kernel-agent
Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU…
alibaba/atrex-kernel-agent
Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki.
alibaba/atrex-kernel-agent
Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.
alibaba/atrex-kernel-agent
Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…
alibaba/atrex-kernel-agent
Generate a structured implementation plan from an evidence draft.
alibaba/atrex-kernel-agent
Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel.
Categories
Run the evidence loop of one long-horizon GPU kernel optimization episode. GPU Kernel Episode Loop is an agent skill from alibaba/atrex-kernel-agent. Run the evidence loop of one long-horizon GPU kernel optimization episode.
GPU Kernel Episode Loop fits situations like: reconstruct the incumbent; profile and localize a bottleneck; research progressively; plan one coherent direction.
Run `npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a claude-code`. Or copy the skill folder (skills/gpu-kernel-episode-loop in alibaba/atrex-kernel-agent) into .claude/skills/gpu-kernel-episode-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a codex`. Or copy the skill folder (skills/gpu-kernel-episode-loop in alibaba/atrex-kernel-agent) into .agents/skills/gpu-kernel-episode-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-episode-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-kernel-episode-loop, .gemini/skills/gpu-kernel-episode-loop, .github/skills/gpu-kernel-episode-loop and .opencode/skills/gpu-kernel-episode-loop in your project.
Going by SKILL.md and its folder, GPU Kernel Episode Loop needs the command-line tools its instructions call (python, python3 and bash). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
GPU Kernel Episode Loop is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with GPU Kernel Episode Loop: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
alibaba (a GitHub organization) maintains it in alibaba/atrex-kernel-agent, which has 154 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.
Source: alibaba/atrex-kernel-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.