Agent skill

GPU Kernel Baseline

by alibaba in alibaba/atrex-kernel-agent

Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel.

Apache-2.0Auto-check passedAI & LLM Engineering

Install GPU Kernel Baseline

skills CLI
$ npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-baseline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alibaba/atrex-kernel-agent gpu-kernel-baseline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/gpu-kernel-baseline .claude/skills/gpu-kernel-baseline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
gpu-kernel-baseline
GitHub stars
154
Token cost
~2.1k tokens
SKILL.md length
783 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel.

  • Works in 4 steps: Understand PyTorch Semantics → Learn Framework APIs from enabled… → Implement Baseline Kernel and… → …
  • Understand compute semantics
  • SKILL.md covers When to Use, Workflow, Phase 1: Understand PyTorch… and Phase 2: Learn Framework APIs…, plus 5 more sections
  • Calls python, git and python3

What it does

GPU Kernel Baseline is an agent skill from alibaba/atrex-kernel-agent. Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel. Use this skill to understand compute semantics, determine the target platform and framework, search reference implementations, and produce a correct V0 baseline with performance records for later profile-driven optimization.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. It works with PyTorch. The repository describes itself as: An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a… The licence is Apache-2.0.

When your agent uses it

  • Understand compute semantics
  • Determine the target platform and framework
  • Search reference implementations
  • Produce a correct V0 baseline with performance records for later profile-driven optimization

Example prompts

  • “/gpu-kernel-baseline”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Understand PyTorch Semantics
  2. Learn Framework APIs from enabled knowledge tools
  3. Implement Baseline Kernel and Correctness Tests
  4. Performance, Correctness, and Quality Gate

What it can do on your machine

Read from SKILL.md and the folder at commit 3d27c1e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

GPU Kernel Baseline loads about 2.1k tokens when it runs. Until then it costs about 84 tokens; SKILL.md has 783 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~84
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from alibaba/atrex-kernel-agent at commit 3d27c1e, republished under its Apache-2.0 licence (© alibaba). 783 words, ~2,125 tokens.

Download SKILL.mdSave it as .claude/skills/gpu-kernel-baseline/SKILL.md (or your agent's skills folder).
name
gpu-kernel-baseline
description
Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel. Use this skill to understand compute semantics, determine the target platform and framework, search reference implementations, and produce a correct V0 baseline with performance records for later profile-driven optimization.

GPU Kernel Baseline

When to Use

Use this skill when the user provides PyTorch logic or a kernel demo and asks to:

  • Write a GPU kernel for the target platform.
  • Build a baseline from scratch.
  • Prepare kernel.py, reference.py, test_kernel.py, and baseline_report.md for later profile-driven optimization.

Workflow

This stage first understands the PyTorch semantics, then learns the framework APIs (CuteDSL or FlyDSL) through enabled knowledge tools, implements kernel.py and test_kernel.py, validates correctness, records performance, writes baseline_report.md, and writes memory/v0.json.

The orchestrator installs discovered plugin instructions at .atrex_plugins/instructions.md.

Phase 1: Understand PyTorch Semantics

  1. Read the user-provided PyTorch logic and kernel_demo.
  2. Extract and record:
    • Compute pattern, such as GEMM, Decode Attention, Reduction, or Elementwise.
    • Input/output shape, stride, dtype, layout, and device.
    • Data dependencies, broadcasting, masks, boundary handling, and write-back semantics.
    • Accuracy requirements, tolerance, accumulation dtype, and special-value handling.
  3. Determine target platform and framework:
    • H100/H20/H200 -> Hopper -> CuteDSL
    • MI300X/MI308X -> CDNA3 -> FlyDSL
    • MI355X -> CDNA4 -> FlyDSL
  4. If the PyTorch logic is ambiguous, first create a minimal runnable reference, then continue.

Phase 2: Learn Framework APIs from enabled knowledge tools

  1. Read .atrex_plugins/instructions.md and inspect python3 tools/plugin.py list for enabled tool names and input schemas. Follow the session's phase-specific plugin instructions.
  2. Query available knowledge tools with the exact product, authoritative runtime architecture, framework, operator, shapes, dtypes and missing implementation facts. If no suitable plugin is enabled, use available reference sources and record missing facts explicitly.
  3. Preserve returned source and attribution identifiers exactly. Respect hardware scope and context budgets. Never substitute another hardware identity or treat a fallback sample as a match.
  4. Prefer sources with the same framework and compute pattern. Record references and the constraints they establish in plans/v0_plan.md.

Phase 3: Implement Baseline Kernel and Correctness Tests

  1. Implement a correct baseline kernel.py based on PyTorch semantics and the learned framework APIs.Not only must the functionality be correct, but the framework implementation must also be correct, using either CuteDSL or FlyDSL.
  2. Write test_kernel.py using PyTorch logic directly as the correctness reference.
  3. Cover representative inputs, including normal shapes, boundary shapes, and relevant dtype or stride cases.
  4. Example correctness check:
python
ref = pytorch_reference(inputs)
out = kernel_v1(inputs)
rel_err = (out.float() - ref).norm() / ref.norm()
assert rel_err < 0.01
  1. The default BF16 threshold is rel_err < 0.01; lower precision formats may use task-specific relaxed thresholds.
  2. Add per-case timeout guard in test_kernel.py to prevent hanging:
python
import signal

def timeout_handler(signum, frame):
    raise TimeoutError("Test case exceeded timeout limit")

signal.signal(signal.SIGALRM, timeout_handler)

TIMEOUT_SEC = int(os.environ.get("TEST_TIMEOUT_SEC", "30"))

for case in test_cases:
    signal.alarm(TIMEOUT_SEC)
    try:
        run_test(case)
    except TimeoutError:
        record_failure(case, "TIMEOUT_FAIL")
    finally:
        signal.alarm(0)
  1. If API, compilation, accuracy, performance, or hardware issues appear, query enabled knowledge tools again with the exact failure and measured evidence, and then fix the implementation.
  2. Record the baseline configuration, including tile size, thread organization, grid/block design, and major data-movement patterns.
Show full SKILL.md (408 more words)Show less

Phase 4: Performance, Correctness, and Quality Gate

  1. Run exactly one full-workload base-seed V0 measurement through the mandatory sandbox. Do not pass --multi-seed and do not launch a separate robustness run for V0:
bash
python tools/sandbox.py --kind run --no-sync -- \
  python test_kernel.py --version v0 --no-memory

Parse the emitted [test_kernel] RESULT_JSON=..., use its performance result and accompanying correctness status for memory/v0.json, and avoid repeating the expensive baseline workload.

  • Each individual test case must complete within 30 seconds (configurable via TEST_TIMEOUT_SEC env var).
  • If a case exceeds the timeout, mark it as TIMEOUT_FAIL, kill the process, and record the failure in baseline_report.md.
  • Common timeout causes: infinite loops in index calculation, deadlocks in synchronization, or excessive compilation time. Consult enabled knowledge tools with the failure mode to diagnose.
  1. Verify all correctness cases pass and record max rel_err plus PASS/FAIL.
  2. Measure baseline performance and record:
text
latency(us) | TFLOPS | bandwidth(GB/s) | TFLOPS peak utilization(%) | bandwidth peak utilization(%)
  1. Use compute_utilization.py to calculate TFLOPS and bandwidth utilization:
bash
python tools/compute_utilization.py   --gpu <gpu> --dtype <dtype>   --flops-expr '<expr>' --bytes-expr '<expr>'   --time-ms <ms> --grid-blocks <blocks>
  1. Every theoretical peak, bandwidth, and utilization calculation must cite the auditable spec sources registered in Step 0.

  2. Write baseline_report.md with:

    • Baseline kernel path
    • Correctness test path
    • PyTorch reference logic description
    • Stable source record ids consulted
    • Baseline configuration summary
    • Correctness results: case list, max rel_err, PASS/FAIL (include any TIMEOUT_FAIL cases)
    • Baseline performance: latency(us), TFLOPS, bandwidth(GB/s), and peak utilization percentages
  3. Write baseline iteration data to memory/v0.json using tools/memory_manager.py:

    bash
    # Create the iteration file
    python tools/memory_manager.py create --workspace kernel_opt_<name> --version v0
    
    # Fill in performance and metadata
    python tools/memory_manager.py update --workspace kernel_opt_<name> --version v0 \
        --set 'performance.latency_us=<value>' \
        --set 'performance.tflops=<value>' \
        --set 'performance.bandwidth_gbps=<value>' \
        --set 'performance.tflops_peak_utilization_pct=<value>' \
        --set 'performance.bandwidth_peak_utilization_pct=<value>' \
        --set 'optimization.action_category=baseline' \
        --set 'optimization.action_description=<summary>' \
        --set 'correctness.rel_err=<value>' \
        --set 'correctness.status=PASS' \
        --set 'quality_gate.result=PASS'

    For array fields (pitfalls_and_fixes, references), update the JSON file directly or use read + manual edit + write-back. Fill in:

    • pitfalls_and_fixes: any errors encountered during implementation
    • references: stable source record ids and other docs referenced during learning
  4. After the quality gate passes, commit:

bash
git add kernel.py test_kernel.py baseline_report.md memory/v0.json README.md
git commit -m "V0: baseline kernel"

memory/ Requirements

Each iteration produces a memory/v<N>.json file following the schema defined in reference/v_iteration.schema.json. The JSON structure captures performance data, optimization actions, profile evidence, correctness results, ISA metric progress, search logs, pitfalls and fixes, and references.

Key rules:

  • The masked field defaults to false. When set to true, the file is skipped during reads.
  • ISA optimization target thresholds are stored in README.md and must be derived from documented best practices, hardware specs, and Step 0 Roofline conclusions. Do not fabricate thresholds from experience.

Deliverables

  • Runnable and correct(using either CuteDSL or FlyDSL) kernel.py
  • PyTorch reference.py
  • test_kernel.py
  • baseline_report.md
  • Created memory/v0.json
  • Git commit

Appendix: Prohibited Actions

  • Do not use unspecified programming frameworks or import external projects.

© alibaba, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/gpu-kernel-baseline of alibaba/atrex-kernel-agent.

Open the folder on GitHubat commit 3d27c1e

Compare with similar skills

GPU Kernel Baseline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

GPU Kernel Baseline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
GPU Kernel Baseline this skillalibaba/atrex-kernel-agent154—~2.1kAutomated safety check: PassApache-2.0
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Add Torch Shapes Examplefacebook/pyrefly7.1k—~1.3kAutomated safety check: PassMIT
Interview Cheatsheetwanshuiyin/ARIS-in-AI-Offer5741 repos~3.4kAutomated safety check: NotesMIT
Ghstack CIpytorch/pytorch104k—~1.4kAutomated safety check: PassCustom licence

Similar skills

  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Torch Shapes Example

    facebook/pyrefly

    Official

    A skill your agent uses when adding a new PyTorch model to Pyrefly's shape-tracking example corpus under tensor-shapes/pyrefly-torch-stubs/examples — i.e.

    7.1k GitHub stars~1.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Interview Cheatsheet

    wanshuiyin/ARIS-in-AI-Offer

    Generate a long-form Chinese interview-prep cheat sheet on a specific ML/LLM topic — formulas with derivations, from-scratch PyTorch code, comparison tables, and 25 高频面试题 (L1 必会 / L2 进阶 / L3 顶级 lab).

    574 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check: notes
  • Ghstack CI

    pytorch/pytorch

    Manage CI for PyTorch ghstack stacks by running CI where its results are useful now and deferring other PRs with [no-ci].

    104k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • MUSA GPU Training Optimizer

    open-infra-skills/infra-skills

    Profiles, benchmarks and tunes AI training workloads on Moore Threads MUSA GPUs with a measurement-first process that keeps model behavior unchanged.

    141 GitHub stars~1.7k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed

More from alibaba/atrex-kernel-agent

  • Opt Trace Mining

    alibaba/atrex-kernel-agent

    Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU…

    154 GitHub stars~4.4k tokensUpdated 8 days ago
    Auto-check passed
  • Session Trace Mining

    alibaba/atrex-kernel-agent

    Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki.

    154 GitHub stars~3.1k tokensUpdated 8 days ago
    Auto-check passed
  • Ppu Acu Joint Profile

    alibaba/atrex-kernel-agent

    Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.

    154 GitHub stars~5.2k tokensUpdated 8 days ago
    Auto-check passed
  • Autonomous GPU Kernel Timeline

    alibaba/atrex-kernel-agent

    Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…

    154 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Gen Plan

    alibaba/atrex-kernel-agent

    Generate a structured implementation plan from an evidence draft.

    154 GitHub stars~3.4k tokensUpdated 8 days ago
    Auto-check passed
  • GPU Kernel Episode Loop

    alibaba/atrex-kernel-agent

    Run the evidence loop of one long-horizon GPU kernel optimization episode.

    154 GitHub stars~3.6k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about GPU Kernel Baseline

What does GPU Kernel Baseline do?

Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel. GPU Kernel Baseline is an agent skill from alibaba/atrex-kernel-agent. Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel.

When should I use GPU Kernel Baseline?

GPU Kernel Baseline fits situations like: understand compute semantics; determine the target platform and framework; search reference implementations; produce a correct V0 baseline with performance records for later profile-driven optimization.

How do I install GPU Kernel Baseline in Claude Code?

Run `npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-baseline -a claude-code`. Or copy the skill folder (skills/gpu-kernel-baseline in alibaba/atrex-kernel-agent) into .claude/skills/gpu-kernel-baseline in your project. Claude Code loads it when a task matches its description.

How do I install GPU Kernel Baseline in Codex?

Run `npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-baseline -a codex`. Or copy the skill folder (skills/gpu-kernel-baseline in alibaba/atrex-kernel-agent) into .agents/skills/gpu-kernel-baseline in your project. Codex loads it when a task matches its description.

Can I use GPU Kernel Baseline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alibaba/atrex-kernel-agent --skill gpu-kernel-baseline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/gpu-kernel-baseline, .gemini/skills/gpu-kernel-baseline, .github/skills/gpu-kernel-baseline and .opencode/skills/gpu-kernel-baseline in your project.

What does GPU Kernel Baseline need to run?

Going by SKILL.md and its folder, GPU Kernel Baseline needs the command-line tools its instructions call (python, git and python3). Our summary lists: Python 3.

Does GPU Kernel Baseline access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is GPU Kernel Baseline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does GPU Kernel Baseline use?

GPU Kernel Baseline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does GPU Kernel Baseline use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to GPU Kernel Baseline?

Skills that share tags, products or a category with GPU Kernel Baseline: Add Uint Support (pytorch/pytorch, 104k stars), CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Torch Shapes Example (facebook/pyrefly, 7.1k stars) and Interview Cheatsheet (wanshuiyin/ARIS-in-AI-Offer, 574 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains GPU Kernel Baseline?

alibaba (a GitHub organization) maintains it in alibaba/atrex-kernel-agent, which has 154 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: alibaba/atrex-kernel-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.