A-Evolve Agent Evolution
Orchestra-Research/AI-Research-SKILLs
Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
$ npx skills add amd/gaia --skill analyzing-claude-sessions -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install amd/gaia analyzing-claude-sessions --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/analyzing-claude-sessions .claude/skills/analyzing-claude-sessions && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "analyzing-claude-sessions" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/analyzing-claude-sessions into .claude/skills/analyzing-claude-sessions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyzing-claude-sessions", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/amd/gaia/tree/main/.claude/skills/analyzing-claude-sessionsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add amd/gaia --skill analyzing-claude-sessions -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install amd/gaia analyzing-claude-sessions --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/analyzing-claude-sessions .agents/skills/analyzing-claude-sessions && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "analyzing-claude-sessions" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/analyzing-claude-sessions into .agents/skills/analyzing-claude-sessions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyzing-claude-sessions", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill analyzing-claude-sessions -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install amd/gaia analyzing-claude-sessions --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/analyzing-claude-sessions .cursor/skills/analyzing-claude-sessions && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "analyzing-claude-sessions" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/analyzing-claude-sessions into .cursor/skills/analyzing-claude-sessions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyzing-claude-sessions", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/amd/gaia.git --path .claude/skills/analyzing-claude-sessions--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add amd/gaia --skill analyzing-claude-sessions -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install amd/gaia analyzing-claude-sessions --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/analyzing-claude-sessions .gemini/skills/analyzing-claude-sessions && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "analyzing-claude-sessions" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/analyzing-claude-sessions into .gemini/skills/analyzing-claude-sessions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyzing-claude-sessions", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install amd/gaia analyzing-claude-sessionsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add amd/gaia --skill analyzing-claude-sessions -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/analyzing-claude-sessions .github/skills/analyzing-claude-sessions && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "analyzing-claude-sessions" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/analyzing-claude-sessions into .github/skills/analyzing-claude-sessions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyzing-claude-sessions", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add amd/gaia --skill analyzing-claude-sessions -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install amd/gaia analyzing-claude-sessions --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/analyzing-claude-sessions .opencode/skills/analyzing-claude-sessions && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "analyzing-claude-sessions" agent skill from https://github.com/amd/gaia/tree/main/.claude/skills/analyzing-claude-sessions into .opencode/skills/analyzing-claude-sessions/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "analyzing-claude-sessions", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
analyzing-claude-sessionsMines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.
Claude Code writes a full JSONL transcript of every session to a local projects folder, recording every tool call, failure and cost. The core pipeline, gaia.factory.harvest, is deterministic Python with no model or network calls; two optional steps, classify and synthesize, do call a model through the local Claude Code CLI. Everything the pipeline derives is written to a cache directory outside any repository, since transcripts can contain absolute paths, branch names and anything pasted into a prompt, and the skill warns against redirecting output into a tracked working tree where it could be committed by accident.
Running it produces traces.jsonl (one normalized trace per session with steps, outcomes, tokens and subagents), intents.jsonl (one line per session with goal, steps, turns, tokens, cost and duration) and stats.json (aggregates including effectiveness and error-profile blocks). The skill documents five specific ways a naive analysis misleads: a delegated subagent run gets its own transcript file that a simple glob misses, so a walk helper is needed to include it; and keying a tool call's identity on only one argument, such as a file path, can make different edits to the same file look identical and badly overstate a thrash metric.
Read from SKILL.md and the folder at commit 05fb50b. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythongitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Analyzing Claude Code Sessions loads about 2.3k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 1,153 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from amd/gaia at commit 05fb50b, republished under its MIT licence (© amd). 1,153 words, ~2,323 tokens.
.claude/skills/analyzing-claude-sessions/SKILL.md (or your agent's skills folder).Claude Code writes a full JSONL transcript of every session to ~/.claude/projects/.
That is a complete record of what an agent was asked to do, every tool call it made,
what failed, and what it cost. This skill turns that exhaust into an evidence-backed
report.
The core pipeline is deterministic Python (gaia.factory.harvest) — no LLM, no
network. Two optional steps do call a model through the local Claude Code CLI:
classify (use-case labels) and synthesize (the written analysis). Everything derived is written to ~/.gaia/cache/factory/ and
must never be committed — transcripts contain absolute paths, branch names, and whatever
the user pasted into a prompt.
Redirect into the cache directory, never into a repository working tree. These tables
are built from real transcripts, and an untracked tables.md sitting at a repo root is
one git add -A away from being published.
FACTORY=~/.gaia/cache/factory
# 1. Extract. Deterministic, no LLM, no network. Roughly linear in transcript count.
python -m gaia.factory.harvest.scan
# 2. Tables. Absolute counts and share-of-total for every figure.
python -m gaia.factory.harvest.report > "$FACTORY/tables.md"
# 3. With use-case labels (see "Classifying intent" below):
python -m gaia.factory.harvest.report --labels "$FACTORY/labels.txt" > "$FACTORY/tables.md"
# 4. Per-request prompt size and local KV-cache memory. Re-reads the raw
# transcripts, because steps 1-2 aggregate per session and that hides how
# large any single request got.
python -m gaia.factory.harvest.context --labels "$FACTORY/labels.txt" > "$FACTORY/context.md"
# 5. What each proposed fix would actually save, in tokens and dollars.
python -m gaia.factory.harvest.savings > "$FACTORY/savings.md"
# 6. Optional, needs the Claude Code CLI: label use cases, then write the analysis.
python -m gaia.factory.harvest.classify
python -m gaia.factory.harvest.synthesize > "$FACTORY/analysis.md"Only scan takes --root (transcripts elsewhere); scan and classify write with
--out. Every other step reads the cache with --cache, and context and savings
also take --projects if the raw transcripts are not under ~/.claude/projects.
Outputs in ~/.gaia/cache/factory/:
| File | Contents |
|---|---|
traces.jsonl | One normalized trace per session: ordered steps, outcomes, tokens, subagents |
intents.jsonl | One line per session: goal, title, steps, turns, tokens, cost, duration |
stats.json | Every aggregate, including the effectiveness and error_profile blocks |
Learned by getting each one wrong first.
1. Subagents are not sessions. A delegated Task/Agent run gets its own transcript
under <session-uuid>/subagents/. A plain */*.jsonl glob misses them entirely, and they
can carry a large share of all tool calls in a delegation-heavy corpus. iter_traces
attaches them to the parent; use Trace.walk() to include them and say explicitly which
scope a number covers.
2. Identity must hash the full arguments. Keying a tool call on one argument (the file
path, say) makes consecutive different edits to one file look identical. That single
mistake overstated a "thrash" metric ~40×. Step.arg_hash covers the whole argument
object; arg_digest is for display only. Never compute identity from the digest.
3. Order your error taxonomy specific-before-generic. Matching is first-wins, and a
timeout also carries a non-zero exit code. With command_failed above timeout, real
timeouts — often the largest failure class — get filed as generic command failures.
4. Rank savings by carried tokens, not by call count. A prompt is the whole
conversation so far, so a result emitted at step 5 of a 50-step run is re-sent 45 more
times. That carry multiplier decides what a fix is worth. Expect the two rankings to
disagree: re-reads can be a large share of reads by count yet a small share of tokens,
while budgeting oversized results is worth several times more. savings.py
computes both; report the token figure and say which one a claim rests on.
5. Harness turns are not human turns. Hook feedback and system reminders appear as
user records. They match correction patterns like "stop" and inflated a "user corrected
the model" metric 8×. Filter on the isMeta field; a startswith("<") heuristic does not
catch them.
scan produces intents.jsonl but assigns no use-case. Run
python -m gaia.factory.harvest.classify, which batches the sessions, labels each
from its opening instruction against a closed taxonomy, and writes
$FACTORY/labels.txt as <8-char-session-prefix> <use-case> — one tag per line, the
format report --labels and context --labels read. Then re-run both with --labels.
Three properties matter, and they are enforced rather than advised:
labels.txt appears only once every batch has validated.no_instruction is assigned locally. A session whose transcript carries no
opening instruction cannot be classified from one, and folding those into other
hides them among sessions that did state a goal.Classify from the first user message, never the auto-generated title — the title summarises what happened, which leaks the outcome into the label.
To label by hand, write the same two-column file yourself. It carries session-id
prefixes, so it belongs in the cache directory like everything else derived. report
checks those prefixes against the corpus either way: nothing matching is an error, and
partial coverage is stated in every use-case table so the labelled subset is never read
as the whole corpus.
Raw counts alone mislead. The findings that carried signal in the reference corpus:
These are not optional; the analysis is worthless without them.
scan covers every Claude Code project on the machine, not just the current repo:
~/.claude/projects/ holds one subdirectory per project, and there is no per-project
filter — --root relocates the scan, it cannot narrow it. tables.md never breaks the
count down by project, so check top_projects in stats.json to see the mix before
sharing anything.
Transcripts contain absolute paths, branch names, repository content, and any secret
pasted into a prompt. The pipeline writes only to ~/.gaia/cache/factory/. If a report is
shared, put it somewhere private and scrub paths first. Nothing derived from a corpus
belongs in a public repository.
Before publishing, have a subagent adversarially review the extraction code against the corpus — not just the prose. Four of the reference report's numbers were wrong on first pass, including its headline, and every one was a bug in the metric rather than a mistake in the writing. Ask specifically: does each metric measure what its name claims, and is the denominator the one the sentence implies?
© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/analyzing-claude-sessions of amd/gaia.
Open the folder on GitHubat commit 05fb50b
Analyzing Claude Code Sessions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Analyzing Claude Code Sessions this skillamd/gaia | 1.6k | — | ~2.3k | Automated safety check: Pass | MIT | |
| A-Evolve Agent EvolutionOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.6k | Automated safety check: Pass | MIT | |
| AI Observabilityomer-metin/skills-for-antigravity | 162 | — | ~578 | Automated safety check: Pass | Apache-2.0 | |
| Caveman Workflow LabelerJuliusBrussee/caveman | 111k | 1 repos | ~1.3k | Automated safety check: Pass | Apache-2.0 | |
| Caveman Evidence ReviewJuliusBrussee/caveman | 111k | 1 repos | ~927 | Automated safety check: Pass | Apache-2.0 | |
| Copilot Session Failure Analysisdotnet/maui | 23k | — | ~3.4k | Automated safety check: Pass | MIT |
Orchestra-Research/AI-Research-SKILLs
Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.
omer-metin/skills-for-antigravity
Implement comprehensive observability for LLM applications including tracing (Langfuse/Helicone), cost tracking, token optimization, RAG evaluation metrics (RAGAS), hallucination detection, and…
JuliusBrussee/caveman
Finds every LLM workflow in a repository, proposes a labeling table and, once you agree, wires labels so Caveman Cloud groups spend per workflow.
JuliusBrussee/caveman
Read-only review of Caveman Cloud data to explain where LLM spend goes: cost, score, workflows, traces, latency, errors, routing and verified savings.
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
steipete/CodexBar
CodexBar read. Provider usage, limits, credits, config health. JSON. No writes.
amd/gaia
Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.
amd/gaia
Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.
amd/gaia
Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.
amd/gaia
Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.
amd/gaia
Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.
amd/gaia
Turns a source document such as a README or spec into an executive slide deck as one self-contained HTML file that prints to PDF, one slide per page.
Works with
Categories
Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs. Claude Code writes a full JSONL transcript of every session to a local projects folder, recording every tool call, failure and cost.harvest, is deterministic Python with no model or network calls; two optional steps, classify and synthesize, do call a model through the local Claude Code CLI.
Analyzing Claude Code Sessions fits situations like: finding out what tasks an agent is actually being used for across many sessions; measuring how often Claude Code sessions fail and why; estimating token and dollar cost across a corpus of sessions; producing an evidence-backed report on real agent usage.
Run `npx skills add amd/gaia --skill analyzing-claude-sessions -a claude-code`. Or copy the skill folder (.claude/skills/analyzing-claude-sessions in amd/gaia) into .claude/skills/analyzing-claude-sessions in your project. Claude Code loads it when a task matches its description.
Run `npx skills add amd/gaia --skill analyzing-claude-sessions -a codex`. Or copy the skill folder (.claude/skills/analyzing-claude-sessions in amd/gaia) into .agents/skills/analyzing-claude-sessions in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill analyzing-claude-sessions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyzing-claude-sessions, .gemini/skills/analyzing-claude-sessions, .github/skills/analyzing-claude-sessions and .opencode/skills/analyzing-claude-sessions in your project.
Going by SKILL.md and its folder, Analyzing Claude Code Sessions needs the command-line tools its instructions call (python and git). Our summary lists: Python, for the deterministic harvesting pipeline; The Claude Code CLI, for the optional classify and synthesize steps; Local Claude Code session transcripts under ~/.claude/projects/.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Analyzing Claude Code Sessions is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Analyzing Claude Code Sessions: A-Evolve Agent Evolution (Orchestra-Research/AI-Research-SKILLs, 13k stars), AI Observability (omer-metin/skills-for-antigravity, 162 stars), Caveman Workflow Labeler (JuliusBrussee/caveman, 111k stars) and Caveman Evidence Review (JuliusBrussee/caveman, 111k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
amd (a GitHub organization) maintains it in amd/gaia, which has 1,579 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 8, 2026.
Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.