Agent skill

Autonomous GPU Kernel Timeline

by alibaba in alibaba/atrex-kernel-agent

Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Autonomous GPU Kernel Timeline

skills CLI
$ npx skills add alibaba/atrex-kernel-agent --skill autonomous-gpu-kernel-timeline -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alibaba/atrex-kernel-agent autonomous-gpu-kernel-timeline --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/autonomous-gpu-kernel-timeline .claude/skills/autonomous-gpu-kernel-timeline && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
autonomous-gpu-kernel-timeline
GitHub stars
154
Token cost
~1.3k tokens
SKILL.md length
599 words
Files
8 (incl. scripts, references)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…

  • Works in 6 steps: State one falsifiable timing question… → Save the current clean kernel.py… → Run representative correctness through… → …
  • AI & LLM Engineering work in your project
  • SKILL.md covers Ownership, Route, Autonomous loop and Hard boundaries
  • Runs Python scripts from its folder

What it does

Autonomous GPU Kernel Timeline is an agent skill from alibaba/atrex-kernel-agent. Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific kernel-internal timing question.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts and reference files (for example `backends/cuda_backend/adapter.py`, `backends/cutedsl_backend/adapter.py` and `references/cuda-backend.md`).

It sits in AI & LLM Engineering. It works with CUDA. The repository describes itself as: An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a… The licence is Apache-2.0.

When your agent uses it

  • AI & LLM Engineering work in your project

Example prompts

  • “/autonomous-gpu-kernel-timeline”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. State one falsifiable timing question and begin with the fewest useful sites.
  2. Save the current clean kernel.py content, then create an instrumented working snapshot in the
  3. Run representative correctness through the immutable evaluator and capture through the campaign
  4. Read the summary and Perfetto trace. Remove, move, or refine probes and repeat when the evidence
  5. Before replacing the snapshot, preserve its source or reversible patch and its content hash.
  6. Restore a probe-free kernel, implement the optimization, and use the normal evaluator for final

What it can do on your machine

Read from SKILL.md and the folder at commit 3d27c1e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Autonomous GPU Kernel Timeline loads about 1.3k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 65 tokens; SKILL.md has 599 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~65
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alibaba/atrex-kernel-agent at commit 3d27c1e, republished under its Apache-2.0 licence (© alibaba). 599 words, ~1,250 tokens.

Download SKILL.mdSave it as .claude/skills/autonomous-gpu-kernel-timeline/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
autonomous-gpu-kernel-timeline
description
Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific kernel-internal timing question.

Autonomous GPU Kernel Timeline

Use this inside an AKA optimization episode only after a correct runnable kernel and representative workload exist. Do not trigger it when aggregate kernel timing or ordinary profiler evidence already answers the question.

Ownership

  • AKA chooses the hypothesis, sites, writer roles, density, boundary semantics, and number of retries.
  • The backend records and decodes those choices without silent identity, capacity, or pairing errors.
  • The final validator rejects known factual failures; it does not decide whether AKA chose the best algorithm boundary or interpretation.

Route

Autonomous loop

  1. State one falsifiable timing question and begin with the fewest useful sites.
  2. Save the current clean kernel.py content, then create an instrumented working snapshot in the same episode worktree. This snapshot is not a handoff candidate and must not enter promotion or stall accounting.
  3. Run representative correctness through the immutable evaluator and capture through the campaign sandbox. For final evidence, pass the sandbox-owned .atrex_long_horizon/evaluations.jsonl as --correctness-evidence; a model-supplied --correctness passed is only an exploration note and cannot produce decision_grade. Use scripts/timeline.py to validate and export the returned evidence. If the remote command reads backend files, pass --input skills/autonomous-gpu-kernel-timeline/backends/<backend> to tools/sandbox.py; sync only the attempt-specific directory.
  4. Read the summary and Perfetto trace. Remove, move, or refine probes and repeat when the evidence is insufficient, semantically misplaced, or too intrusive. Name every reported interval as start_site -> end_site and recompute its numbers from canonical events rather than trusting a prose label. A local timeline delta is mechanism evidence, not a substitute for probe-free end-to-end ABBA. Reviewer feedback is optional.
  5. Before replacing the snapshot, preserve its source or reversible patch and its content hash.
  6. Restore a probe-free kernel, implement the optimization, and use the normal evaluator for final correctness and performance. Re-instrument the new clean state only when confirmation is useful.

For a final perturbation claim, use scripts/timeline.py measure inside one GPU allocation. Give it baseline/instrumented commands as JSON argv arrays, the exact sources, and materialized binaries when the compiler exposes them. JIT-only CuTe DSL measurements need not invent a binary artifact. Each command must perform the requested warmup and iterations, synchronize the device, check the full representative output, and emit exactly one line of this form:

Show full SKILL.md (206 more words)Show less
text
__ATREX_TIMELINE_SAMPLE__={"latency_ms":1.0,"correctness":"passed","synchronized":true,"workload_identity":"...","device_identity":{"uuid":"..."},"warmup":10,"iterations":100}

The helper runs each sample in a fresh process, defaults to ABBA followed by BAAB, rejects workload or device drift, and writes the raw schedule and samples. Pass that artifact to capture/export with --measurement; validate rechecks its hashes and recomputes the medians and overhead. The sample's correctness field rejects a bad timing run but does not replace the immutable evaluator record required for final evidence.

Hard boundaries

  • Never modify profile_driver.py, evaluators, ground truth, or other protected paths.
  • Never hand off or promote an instrumented snapshot. Only a probe-free kernel may become the episode candidate_commit == HEAD.
  • Do not combine events from different launches into one apparent execution.
  • Construct exactly one recorder per selected owner per launch and reuse it for that owner's events; duplicate ownership is a capture failure, not a sampling policy.
  • Reject overflow, truncation, invalid owner/site identity, deterministic range mismatch, stale source/binary/workload provenance, and unrecomputable numeric claims.
  • Never classify final evidence as decision_grade without a matching sandbox evaluator record.
  • The default low-perturbation target is 10%. A higher-overhead trace may remain diagnostic when it is labelled honestly; the threshold is hard only when the task explicitly requires it.
  • Do not add fixed site catalogs, fixed coverage quotas, mandatory reviewer approval, or a separate orchestration state machine.

© alibaba, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in skills/autonomous-gpu-kernel-timeline of alibaba/atrex-kernel-agent.

  • SKILL.md
  • backends/cuda_backend/adapter.py
  • backends/cuda_backend/atrex_timeline.cuh
  • backends/cuda_backend/test_backend.cu
  • backends/cutedsl_backend/adapter.py
  • references/cuda-backend.md
  • references/iket-quickstart.md
  • scripts/timeline.py

Open the folder on GitHubat commit 3d27c1e

Compare with similar skills

Autonomous GPU Kernel Timeline next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Autonomous GPU Kernel Timeline compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Autonomous GPU Kernel Timeline this skillalibaba/atrex-kernel-agent154—~1.3kAutomated safety check: PassApache-2.0
Kernel VerificationZJLi2013/awesome-kernel-skills102—~702Automated safety check: PassNone
Esmfold2JimLiu/science-skills2274 repos~2.5kAutomated safety check: PassApache-2.0
Hugging Face Local Modelshuggingface/skills11k3 repos~945Automated safety check: PassApache-2.0
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent143—~5.2kAutomated safety check: PassApache-2.0
Make Op VerifyCVCUDA/CV-CUDA2.7k—~433Automated safety check: PassCustom licence

Similar skills

  • Kernel Verification

    ZJLi2013/awesome-kernel-skills

    5-stage kernel correctness verification protocol for Triton and CUDA kernels.

    102 GitHub stars~702 tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Esmfold2

    JimLiu/science-skills

    Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al.

    227 GitHub starsUsed in 4 repos~2.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Hugging Face Local Models

    huggingface/skills

    Official

    Finds llama.cpp-compatible GGUF models on the Hugging Face Hub, picks a quantization for your hardware and launches them with llama-cli or llama-server.

    11k GitHub starsUsed in 3 repos~945 tokens
    AI & LLM EngineeringAuto-check passed
  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    143 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Make Op Verify

    CVCUDA/CV-CUDA

    Verify a new CV-CUDA operator against the deterministic final regression checklist (the /make-op done-gate).

    2.7k GitHub stars~433 tokensUpdated 21 days ago
    AI & LLM EngineeringAuto-check passed
  • GPU Optimizer

    Mathews-Tom/armory

    GPU optimization for consumer NVIDIA GPUs (8-24GB VRAM) covering mixed precision, gradient checkpointing, XGBoost GPU, CuPy/cuDF migration, and torch.compile.

    327 GitHub stars~3.5k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check: notes

More from alibaba/atrex-kernel-agent

  • Opt Trace Mining

    alibaba/atrex-kernel-agent

    Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU…

    154 GitHub stars~4.4k tokensUpdated 8 days ago
    Auto-check passed
  • Session Trace Mining

    alibaba/atrex-kernel-agent

    Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki.

    154 GitHub stars~3.1k tokensUpdated 8 days ago
    Auto-check passed
  • Ppu Acu Joint Profile

    alibaba/atrex-kernel-agent

    Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.

    154 GitHub stars~5.2k tokensUpdated 8 days ago
    Auto-check passed
  • Gen Plan

    alibaba/atrex-kernel-agent

    Generate a structured implementation plan from an evidence draft.

    154 GitHub stars~3.4k tokensUpdated 8 days ago
    Auto-check passed
  • GPU Kernel Baseline

    alibaba/atrex-kernel-agent

    Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel.

    154 GitHub stars~2.1k tokensUpdated 8 days ago
    Auto-check passed
  • GPU Kernel Episode Loop

    alibaba/atrex-kernel-agent

    Run the evidence loop of one long-horizon GPU kernel optimization episode.

    154 GitHub stars~3.6k tokensUpdated 8 days ago
    Auto-check passed

Works with

Questions about Autonomous GPU Kernel Timeline

What does Autonomous GPU Kernel Timeline do?

Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…. Autonomous GPU Kernel Timeline is an agent skill from alibaba/atrex-kernel-agent. Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific kernel-internal timing question.

When should I use Autonomous GPU Kernel Timeline?

Autonomous GPU Kernel Timeline fits situations like: AI & LLM Engineering work in your project.

How do I install Autonomous GPU Kernel Timeline in Claude Code?

Run `npx skills add alibaba/atrex-kernel-agent --skill autonomous-gpu-kernel-timeline -a claude-code`. Or copy the skill folder (skills/autonomous-gpu-kernel-timeline in alibaba/atrex-kernel-agent) into .claude/skills/autonomous-gpu-kernel-timeline in your project. Claude Code loads it when a task matches its description.

How do I install Autonomous GPU Kernel Timeline in Codex?

Run `npx skills add alibaba/atrex-kernel-agent --skill autonomous-gpu-kernel-timeline -a codex`. Or copy the skill folder (skills/autonomous-gpu-kernel-timeline in alibaba/atrex-kernel-agent) into .agents/skills/autonomous-gpu-kernel-timeline in your project. Codex loads it when a task matches its description.

Can I use Autonomous GPU Kernel Timeline in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alibaba/atrex-kernel-agent --skill autonomous-gpu-kernel-timeline -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/autonomous-gpu-kernel-timeline, .gemini/skills/autonomous-gpu-kernel-timeline, .github/skills/autonomous-gpu-kernel-timeline and .opencode/skills/autonomous-gpu-kernel-timeline in your project.

What does Autonomous GPU Kernel Timeline need to run?

Going by SKILL.md and its folder, Autonomous GPU Kernel Timeline needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Autonomous GPU Kernel Timeline access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Autonomous GPU Kernel Timeline safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Autonomous GPU Kernel Timeline use?

Autonomous GPU Kernel Timeline is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Autonomous GPU Kernel Timeline use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2.4k tokens, read only when the agent opens those files.

What are the alternatives to Autonomous GPU Kernel Timeline?

Skills that share tags, products or a category with Autonomous GPU Kernel Timeline: Kernel Verification (ZJLi2013/awesome-kernel-skills, 102 stars), Esmfold2 (JimLiu/science-skills, 227 stars), Hugging Face Local Models (huggingface/skills, 11k stars) and Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 143 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Autonomous GPU Kernel Timeline?

alibaba (a GitHub organization) maintains it in alibaba/atrex-kernel-agent, which has 154 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: alibaba/atrex-kernel-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.