Agent skill

Session Trace Mining

by alibaba in alibaba/atrex-kernel-agent

Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Session Trace Mining

skills CLI
$ npx skills add alibaba/atrex-kernel-agent --skill session-trace-mining -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alibaba/atrex-kernel-agent session-trace-mining --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alibaba/atrex-kernel-agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/gpu-wiki/skills/session-trace-mining .claude/skills/session-trace-mining && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
session-trace-mining
GitHub stars
161
Token cost
~3.1k tokens
SKILL.md length
1,556 words
Files
22 (incl. scripts, references, assets)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki.

  • Works in 2 steps: point STM_ROOT at the directory holding… → register your sets in scripts/config.py…
  • Asked to turn vibe-coding sessions
  • SKILL.md covers Overview, Why every part of it lives here, Setup and Pipeline, plus 6 more sections
  • Runs Python scripts from its folder; calls python3 and git

What it does

Session Trace Mining is an agent skill from alibaba/atrex-kernel-agent. Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki. Use when asked to turn vibe-coding sessions, Codex rollout logs, or Claude Code project transcripts into wiki records; to summarise what a kernel-optimization session achieved; to build or extend a session-trace store; or to re-run and validate one. Also use when asked how a session-derived record's number, snippet, or provenance was established.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 25 other files, including scripts, reference files and assets (for example `assets/schema/clean-1.3.frozen.json`, `assets/schema/session-trace-1.0.schema.json` and `references/distill-brief.md`).

It sits in AI & LLM Engineering. The repository describes itself as: An end-to-end agent project for GPU kernel implementation, analysis, profiling, and iterative optimization. It helps an agent turn PyTorch logic or an existing kernel into a… The licence is Apache-2.0.

When your agent uses it

  • Asked to turn vibe-coding sessions
  • Codex rollout logs
  • Claude Code project transcripts into wiki records
  • Summarise what a kernel-optimization session achieved

Example prompts

  • “/session-trace-mining”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. point STM_ROOT at the directory holding your own transcript archive;
  2. register your sets in scripts/config.py SETS, replacing the single

What it can do on your machine

Read from SKILL.md and the folder at commit 3d27c1e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 12 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Session Trace Mining loads about 3.1k tokens when it runs, and up to ~8.9k if it reads all its reference files. Until then it costs about 122 tokens; SKILL.md has 1,556 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alibaba/atrex-kernel-agent at commit 3d27c1e, republished under its Apache-2.0 licence (© alibaba). 1,556 words, ~3,140 tokens.

Download SKILL.mdSave it as .claude/skills/session-trace-mining/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.
name
session-trace-mining
description
Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki. Use when asked to turn vibe-coding sessions, Codex rollout logs, or Claude Code project transcripts into wiki records; to summarise what a kernel-optimization session achieved; to build or extend a session-trace store; or to re-run and validate one. Also use when asked how a session-derived record's number, snippet, or provenance was established.

Session Trace Mining

Overview

Turns session transcripts into session-trace-1.0 records: one record per change that measurably moved a metric, answering which operator, what was wrong, what changed, why it worked at the machine level, the verbatim code, and the gain.

The pipeline splits three ways: scripts do deterministic extraction, an agent does semantic distillation, and sixteen mechanical gates catch fabrication. The gates are why the output can be trusted. A model filling fields is the weak link, so nothing enters the store that cannot be checked against the transcript.

What makes a transcript corpus different from a repository of commits: there is no repository to check a citation against. The transcript is the corpus. So provenance is a set name plus a set-relative path plus a per-line digest, and every number must trace to a span the packet carried — with the agent's own prose deliberately excluded from that span set.

Why every part of it lives here

The same shape of problem — extract deterministically, distil semantically, gate mechanically — applies to any corpus of optimisation history. What travels between such pipelines is design, not code: the three-way split, the declarative make_schema.py patch list, the packet contract, and the discipline of writing down every pitfall with the number behind it.

Nothing is imported. This skill is self-contained, deliberately:

concernherewhy not shared
operator naming, workload familiesfamilies.pythe layout gate derives a record's directory from these functions, so a change in another tree would silently refile records
milestone selection (the ratchet)ladder.pyits thresholds decide this store's record set; and an A/B corpus with no ladder pushes back on rules that have no business knowing about it
ranking modelscore.pya re-ranked store with no diff to show for it is the worst kind of drift

Each of the three pins its own behaviour with a self_test() that runs standalone (python3 families.py, python3 score.py). score.py reproduces the curve published by the shared ranking model in tools/wiki_score.py for the inputs this corpus produces — 4% → 0.35, 20% → 0.66, 99% → 1.0, and the 0.35 feedback band — so scores stay comparable with the committed store by verification rather than by coupling. If the shared model changes, that self-test is where the divergence surfaces.

The only paths outside the skill are the two data locations declared in config.py: where the product goes (kernel_wiki/session_trace/<set>/) and the dedup / overlap scan over kernel_wiki/records. No logic crosses the boundary.

Setup

One thing has to be configured, because no transcripts ship with this repository:

  1. point STM_ROOT at the directory holding your own transcript archive;
  2. register your sets in scripts/config.py SETS, replacing the single example-set placeholder entry. A set is named, never passed as a path, so a record's provenance survives moving the archive.

Both output locations are derived and need no configuration:

  • scratch — /tmp/session-trace-mining/<set>/: parsed candidates, segments, packets. Reproducible, never committed. Override with STM_WORKSPACE.
  • product — <gpu-wiki>/kernel_wiki/session_trace/<set>/: records and the reports that justify them. Override with STM_STORE. It sits beside the committed store but not inside kernel_wiki/records/: these records use the derived session-trace-1.0 schema, so filing them into the committed store would make that store fail its own schema gate. Promoting one means rewriting it against schema/kernel/schema.json, deliberately and one at a time.

Plain python3 is enough. The schema gate additionally needs jsonschema; without it that one gate reports SKIP and the other fifteen still run.

Pipeline

bash
S=<skill-dir>/scripts
export STM_ROOT=/path/to/your/transcript-archive
export STM_SET=example-set             # a key of config.SETS -- register your own

python3 $S/make_schema.py --check      # schema matches its patch list
python3 $S/ingest.py                   # transcripts -> work/versions.jsonl
python3 $S/recon.py                    # -> reports/recon.md   READ THIS FIRST
python3 $S/partition.py                # -> work/segments.jsonl, reports/partition.md
python3 $S/build_packets.py            # -> packets/<seg>.{json,diff}
# distil (see below), then:
python3 $S/validate_store.py --verbose          # sixteen gates
python3 $S/validate_store.py --injection-tests
python3 $S/score_records.py            # worth.rank.score + records/index.json
python3 $S/make_readme.py              # -> kernel_wiki/session_trace/<set>/README.md

Re-run the last two after every distillation batch.

Read reports/recon.md before distilling. It decides what the product can honestly be: the citable-number share, the diff-coverage share, and how much the committed store already covers. If a set turns out to have almost no numbers, its value is mechanism and anti-patterns — do not force a gain claim onto every record.

Two candidate units

The unit is a property of the corpus, not a preference, and it is declared per set in config.SETS. Getting it wrong yields either nothing or nonsense, so it was settled by a probe before any schema existed (see references/lessons.md §1–2).

unitwhenwhat one record is
version-ladderthe run keeps a numbered ladder (memory/vN.json + vN: commit subjects)one version, assembled across the whole set — one session is one version, so the ladder does not exist inside a single file
ab-comparisonno ladderone measured A/B: a variant comparison printed complete in one output, or the same benchmark run either side of an edit

Evidence tiers

The single most important design decision. Every span carries a tier, and only three of five may be cited:

tiercontentcitablecaps gain.basis at
T1benchmark / profiler stdoutyesmeasured
T2the agent reading back its own notes (cat NOTES.md, git log, Read memory/vN.json)yesreported
T3an agent-authored structured fieldyesreported
T4agent prose and thinkingno—
T5the orchestrator promptno—

T4 is excluded because admitting it makes the fabrication gate vacuous: the agent's invented number becomes its own proof. T5 is excluded because those prompts state the target percentage, which would license any number near it.

Show full SKILL.md (730 more words)Show less

The sixteen gates

schema · ids · layout · provenance · verbatim · no-fabrication · direction · raw-isolation · relations · index · evidence-tier · diff-coverage · unit-normalization · wiki-overlap · pairing-integrity · anonymization

The six that carry the weight:

  • provenance — re-resolves the set by name, the file by set-relative path, and every cited line by sha256(raw line)[:12]. Retargeting a citation fails; moving the whole archive to another absolute path still passes. A cited line with no digest is a failure too — without that clause the check silently does nothing, which is what the line-shift injection caught.
  • verbatim — implementation.snippet must appear literally in packets/<seg>.diff. The gate and the distiller read the same file on purpose.
  • no-fabrication — every number in worth.gain must be in the packet's evidence_text, or derivable from it by one of exactly two closed-form rules (before/after, or speedup=Nx). Derivation from arbitrary pairs of pool numbers is deliberately not allowed.
  • unit-normalization — the delta must be reproducible from the measured levels, and no absolute time may appear inside worth.gain.
  • wiki-overlap — nothing the committed store already covers: a colliding (operator, version) dedup key, a record id, or an episode_key. It scans kernel_wiki/records, and an empty scan is a failure, not a pass — an overlap gate whose index resolved to nothing would print OK forever.
  • anonymization — the served layers must name no person, host or corpus path. Session transcripts are full of /home/<user>/..., and payload is what gets served. The record id is checked separately, because a home path flattened into a slug has no slash left for the path pattern to catch.

Never weaken a gate to make records pass, and never let a distilling agent edit scripts/. When a gate looks wrong, verify by injection: validate_store.py --injection-tests mutates a record (and, where the error lives there, its packet) and asserts that the named gate complains. Adding a gate without an injection test is how a store ends up falsely green — two of the eleven cases here were asleep on their first run.

Distillation

Spawn agents with references/distill-brief.md verbatim, substituting the placeholders. Batch by set and record type so a failure has a small blast radius, and point every agent at the one record that already passes as the worked example.

Require each agent to run validate_store.py itself and iterate to green, and to report which fields the packet was too thin to fill and which gate blocked it. That report is the main signal for improving the pipeline; treat a batch that reports no difficulties with suspicion.

When several agents write into one store concurrently, tell them explicitly to ignore gate failures naming records they do not own.

Porting to another corpus

Everything is corpus-agnostic except two places:

  • scripts/config.py SETS — register the set: its path under the archive root, its transcript format, its candidate unit, and default scope. Defaults are fallbacks only; ingest.py detects hardware and DSL and records which happened in arch_basis / dsl_basis.
  • scripts/transcripts.py — the only file that knows how a session log is shaped. A third agent product means one new parse_* function returning the same Event stream, plus a branch in detect_format. Everything downstream sees events and never a raw line.

Two corpus-specific details that will need attention on a new corpus: how a long-running benchmark's output is linked back to the command that launched it (in Codex logs it is a SESSION_ID=N handshake), and which label words name a whole implementation rather than a knob (partition.IMPL_LABEL_RE).

Decide the unit before writing any schema. It is what ids, pairing, dedup_key and the whole worth.gain ladder key on, so getting it wrong means re-doing the schema, the packets and three gates. Three probes exist for exactly that decision and should be re-run on a new corpus:

bash
python3 $S/probe_versions.py <transcript> [...]   # per file: versions, edits, metrics, pairing
python3 $S/probe_set.py <set-root>                # does the ladder exist across the set?
python3 $S/probe_ab.py <set-root>                 # if not, how many measured A/Bs are there?

probe_versions.py answers "can I see versions in one file"; probe_set.py answers the question that actually matters for a ladder, since one session is one version; probe_ab.py sizes the fallback. Write the pass bars down before running them, and if a corpus fails its bar, change the unit rather than the bar.

Resources

  • references/lessons.md — read before starting. Every pitfall found while building this, with the measured numbers behind each: why one session is one version, why the codex sets cannot be paired before/after, the s-for-seconds trap, the table column that inherited a unit it did not have, and the two gates that were asleep.
  • references/distill-brief.md — the agent brief template.
  • assets/schema/session-trace-1.0.schema.json — the record schema, generated by make_schema.py from clean-1.3.frozen.json, which is a pinned byte copy of this repository's schema/kernel/schema.json.

© alibaba, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 21 other files (scripts, references, assets) in gpu-wiki/skills/session-trace-mining of alibaba/atrex-kernel-agent.

  • SKILL.md
  • assets/schema/clean-1.3.frozen.json
  • assets/schema/session-trace-1.0.schema.json
  • references/distill-brief.md
  • references/lessons.md
  • scripts/build_packets.py
  • scripts/config.py
  • scripts/families.py
  • scripts/ingest.py
  • scripts/ladder.py
  • scripts/make_readme.py
  • scripts/make_schema.py
  • scripts/metrics.py
  • scripts/partition.py
  • scripts/probe_ab.py
  • scripts/probe_set.py
  • scripts/probe_versions.py
  • … and 5 more

Open the folder on GitHubat commit 3d27c1e

Compare with similar skills

Session Trace Mining next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Session Trace Mining compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Session Trace Mining this skillalibaba/atrex-kernel-agent161—~3.1kAutomated safety check: PassApache-2.0
Agent BuildershareAI-lab/learn-claude-code78k6 repos~1.2kAutomated safety check: PassMIT
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Peft Fine TuningOrchestra-Research/AI-Research-SKILLs13k9 repos~3.1kAutomated safety check: PassMIT
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
1passwordtrpc-group/trpc-agent-go1.8k15 repos~656Automated safety check: PassApache-2.0

Similar skills

  • Agent Builder

    shareAI-lab/learn-claude-code

    Design and build AI agents for any domain. An agent skill from shareAI-lab/learn-claude-code.

    78k GitHub starsUsed in 6 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Peft Fine Tuning

    Orchestra-Research/AI-Research-SKILLs

    Parameter-efficient fine-tuning for LLMs using LoRA, QLoRA, and 25+ methods.

    13k GitHub starsUsed in 9 repos~3.1k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • 1password

    trpc-group/trpc-agent-go

    Set up and use 1Password CLI (op). An agent skill from trpc-group/trpc-agent-go.

    1.8k GitHub starsUsed in 15 repos~656 tokens
    AI & LLM EngineeringAuto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 8 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from alibaba/atrex-kernel-agent

  • Opt Trace Mining

    alibaba/atrex-kernel-agent

    Mine a per-kernel optimization trace — a git repository capturing successive versions of one kernel being optimized — into structured, gate-validated optimization-experience records for the GPU…

    161 GitHub stars~4.4k tokensUpdated 8 days ago
    Auto-check passed
  • Ppu Acu Joint Profile

    alibaba/atrex-kernel-agent

    Choose and run ACU-only, adaptive PPU in-kernel timeline, or optional bounded joint analysis for a PPU kernel.

    161 GitHub stars~5.2k tokensUpdated 8 days ago
    Auto-check passed
  • Autonomous GPU Kernel Timeline

    alibaba/atrex-kernel-agent

    Let AKA autonomously add, run, inspect, and revise intra-kernel timeline probes for standalone CUDA/inline PTX or CuTe DSL when ordinary benchmark, NSYS, or NCU evidence cannot answer a specific…

    161 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Gen Plan

    alibaba/atrex-kernel-agent

    Generate a structured implementation plan from an evidence draft.

    161 GitHub stars~3.4k tokensUpdated 8 days ago
    Auto-check passed
  • GPU Kernel Baseline

    alibaba/atrex-kernel-agent

    Learn the target framework from enabled knowledge tools and implement a baseline GPU kernel.

    161 GitHub stars~2.1k tokensUpdated 8 days ago
    Auto-check passed
  • GPU Kernel Episode Loop

    alibaba/atrex-kernel-agent

    Run the evidence loop of one long-horizon GPU kernel optimization episode.

    161 GitHub stars~3.6k tokensUpdated 8 days ago
    Auto-check passed

Questions about Session Trace Mining

What does Session Trace Mining do?

Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki. Session Trace Mining is an agent skill from alibaba/atrex-kernel-agent. Mine AI coding-agent session transcripts into structured, gate-validated GPU-kernel optimization records for the wiki.

When should I use Session Trace Mining?

Session Trace Mining fits situations like: asked to turn vibe-coding sessions; Codex rollout logs; Claude Code project transcripts into wiki records; summarise what a kernel-optimization session achieved.

How do I install Session Trace Mining in Claude Code?

Run `npx skills add alibaba/atrex-kernel-agent --skill session-trace-mining -a claude-code`. Or copy the skill folder (gpu-wiki/skills/session-trace-mining in alibaba/atrex-kernel-agent) into .claude/skills/session-trace-mining in your project. Claude Code loads it when a task matches its description.

How do I install Session Trace Mining in Codex?

Run `npx skills add alibaba/atrex-kernel-agent --skill session-trace-mining -a codex`. Or copy the skill folder (gpu-wiki/skills/session-trace-mining in alibaba/atrex-kernel-agent) into .agents/skills/session-trace-mining in your project. Codex loads it when a task matches its description.

Can I use Session Trace Mining in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alibaba/atrex-kernel-agent --skill session-trace-mining -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/session-trace-mining, .gemini/skills/session-trace-mining, .github/skills/session-trace-mining and .opencode/skills/session-trace-mining in your project.

What does Session Trace Mining need to run?

Going by SKILL.md and its folder, Session Trace Mining needs Python for the scripts in its folder and the command-line tools its instructions call (python3 and git). Our summary lists: Python 3.

Does Session Trace Mining access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Session Trace Mining safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Session Trace Mining use?

Session Trace Mining is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Session Trace Mining use?

About 3.1k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.7k tokens, read only when the agent opens those files.

What are the alternatives to Session Trace Mining?

Skills that share tags, products or a category with Session Trace Mining: Agent Builder (shareAI-lab/learn-claude-code, 78k stars), Add Uint Support (pytorch/pytorch, 104k stars), Peft Fine Tuning (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Session Trace Mining?

alibaba (a GitHub organization) maintains it in alibaba/atrex-kernel-agent, which has 161 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on September 29, 2026.

Source: alibaba/atrex-kernel-agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.