Agent skill

Analyzing Claude Code Sessions

by amd in amd/gaia

Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

MITAuto-check passedAI & LLM Engineering

Install Analyzing Claude Code Sessions

skills CLI
$ npx skills add amd/gaia --skill analyzing-claude-sessions -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install amd/gaia analyzing-claude-sessions --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/amd/gaia.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/analyzing-claude-sessions .claude/skills/analyzing-claude-sessions && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
analyzing-claude-sessions
GitHub stars
1.6k
Token cost
~2.3k tokens
SKILL.md length
1,153 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
MIT

At a glance

Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

  • Finding out what tasks an agent is actually being used for across many sessions
  • SKILL.md covers Run it, The five things that will…, Classifying intent and What to actually report, plus 3 more sections
  • Calls python and git
  • Measuring how often Claude Code sessions fail and why

What it does

Claude Code writes a full JSONL transcript of every session to a local projects folder, recording every tool call, failure and cost. The core pipeline, gaia.factory.harvest, is deterministic Python with no model or network calls; two optional steps, classify and synthesize, do call a model through the local Claude Code CLI. Everything the pipeline derives is written to a cache directory outside any repository, since transcripts can contain absolute paths, branch names and anything pasted into a prompt, and the skill warns against redirecting output into a tracked working tree where it could be committed by accident.

Running it produces traces.jsonl (one normalized trace per session with steps, outcomes, tokens and subagents), intents.jsonl (one line per session with goal, steps, turns, tokens, cost and duration) and stats.json (aggregates including effectiveness and error-profile blocks). The skill documents five specific ways a naive analysis misleads: a delegated subagent run gets its own transcript file that a simple glob misses, so a walk helper is needed to include it; and keying a tool call's identity on only one argument, such as a file path, can make different edits to the same file look identical and badly overstate a thrash metric.

When your agent uses it

  • Finding out what tasks an agent is actually being used for across many sessions
  • Measuring how often Claude Code sessions fail and why
  • Estimating token and dollar cost across a corpus of sessions
  • Producing an evidence-backed report on real agent usage

Example prompts

  • “Scan my Claude Code projects folder and build the usage statistics cache.”
  • “Classify the use cases across my last month of sessions and summarize the error profile.”
  • “How much did our Claude Code usage cost last week, including subagent runs?”

Requirements

  • Python, for the deterministic harvesting pipeline
  • The Claude Code CLI, for the optional classify and synthesize steps
  • Local Claude Code session transcripts under ~/.claude/projects/

What it can do on your machine

Read from SKILL.md and the folder at commit 05fb50b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Analyzing Claude Code Sessions loads about 2.3k tokens when it runs. Until then it costs about 102 tokens; SKILL.md has 1,153 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~102
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from amd/gaia at commit 05fb50b, republished under its MIT licence (© amd). 1,153 words, ~2,323 tokens.

Download SKILL.mdSave it as .claude/skills/analyzing-claude-sessions/SKILL.md (or your agent's skills folder).
name
analyzing-claude-sessions
description
Analyze your own local Claude Code session transcripts to find what you actually use the agent for, how effective it is, where it fails, and what it costs. Use when asked to mine Claude Code sessions, extract use-cases or workflows from session history, measure agent effectiveness or error rates, analyze token/cost usage, or produce a report on how an agent is really being used.

Analyzing your Claude Code sessions

Claude Code writes a full JSONL transcript of every session to ~/.claude/projects/. That is a complete record of what an agent was asked to do, every tool call it made, what failed, and what it cost. This skill turns that exhaust into an evidence-backed report.

The core pipeline is deterministic Python (gaia.factory.harvest) — no LLM, no network. Two optional steps do call a model through the local Claude Code CLI: classify (use-case labels) and synthesize (the written analysis). Everything derived is written to ~/.gaia/cache/factory/ and must never be committed — transcripts contain absolute paths, branch names, and whatever the user pasted into a prompt.

Run it

Redirect into the cache directory, never into a repository working tree. These tables are built from real transcripts, and an untracked tables.md sitting at a repo root is one git add -A away from being published.

bash
FACTORY=~/.gaia/cache/factory

# 1. Extract. Deterministic, no LLM, no network. Roughly linear in transcript count.
python -m gaia.factory.harvest.scan

# 2. Tables. Absolute counts and share-of-total for every figure.
python -m gaia.factory.harvest.report > "$FACTORY/tables.md"

# 3. With use-case labels (see "Classifying intent" below):
python -m gaia.factory.harvest.report --labels "$FACTORY/labels.txt" > "$FACTORY/tables.md"

# 4. Per-request prompt size and local KV-cache memory. Re-reads the raw
#    transcripts, because steps 1-2 aggregate per session and that hides how
#    large any single request got.
python -m gaia.factory.harvest.context --labels "$FACTORY/labels.txt" > "$FACTORY/context.md"

# 5. What each proposed fix would actually save, in tokens and dollars.
python -m gaia.factory.harvest.savings > "$FACTORY/savings.md"

# 6. Optional, needs the Claude Code CLI: label use cases, then write the analysis.
python -m gaia.factory.harvest.classify
python -m gaia.factory.harvest.synthesize > "$FACTORY/analysis.md"

Only scan takes --root (transcripts elsewhere); scan and classify write with --out. Every other step reads the cache with --cache, and context and savings also take --projects if the raw transcripts are not under ~/.claude/projects.

Outputs in ~/.gaia/cache/factory/:

FileContents
traces.jsonlOne normalized trace per session: ordered steps, outcomes, tokens, subagents
intents.jsonlOne line per session: goal, title, steps, turns, tokens, cost, duration
stats.jsonEvery aggregate, including the effectiveness and error_profile blocks

The five things that will mislead you

Learned by getting each one wrong first.

1. Subagents are not sessions. A delegated Task/Agent run gets its own transcript under <session-uuid>/subagents/. A plain */*.jsonl glob misses them entirely, and they can carry a large share of all tool calls in a delegation-heavy corpus. iter_traces attaches them to the parent; use Trace.walk() to include them and say explicitly which scope a number covers.

2. Identity must hash the full arguments. Keying a tool call on one argument (the file path, say) makes consecutive different edits to one file look identical. That single mistake overstated a "thrash" metric ~40×. Step.arg_hash covers the whole argument object; arg_digest is for display only. Never compute identity from the digest.

3. Order your error taxonomy specific-before-generic. Matching is first-wins, and a timeout also carries a non-zero exit code. With command_failed above timeout, real timeouts — often the largest failure class — get filed as generic command failures.

4. Rank savings by carried tokens, not by call count. A prompt is the whole conversation so far, so a result emitted at step 5 of a 50-step run is re-sent 45 more times. That carry multiplier decides what a fix is worth. Expect the two rankings to disagree: re-reads can be a large share of reads by count yet a small share of tokens, while budgeting oversized results is worth several times more. savings.py computes both; report the token figure and say which one a claim rests on.

5. Harness turns are not human turns. Hook feedback and system reminders appear as user records. They match correction patterns like "stop" and inflated a "user corrected the model" metric 8×. Filter on the isMeta field; a startswith("<") heuristic does not catch them.

Classifying intent

scan produces intents.jsonl but assigns no use-case. Run python -m gaia.factory.harvest.classify, which batches the sessions, labels each from its opening instruction against a closed taxonomy, and writes $FACTORY/labels.txt as <8-char-session-prefix> <use-case> — one tag per line, the format report --labels and context --labels read. Then re-run both with --labels.

Three properties matter, and they are enforced rather than advised:

  • The taxonomy is closed. A batch that invents tags makes its per-use-case tables incomparable with every other batch.
  • Nothing partial is written. A batch whose labels do not cover it exactly is an error; labels.txt appears only once every batch has validated.
  • no_instruction is assigned locally. A session whose transcript carries no opening instruction cannot be classified from one, and folding those into other hides them among sessions that did state a goal.

Classify from the first user message, never the auto-generated title — the title summarises what happened, which leaks the outcome into the label.

To label by hand, write the same two-column file yourself. It carries session-id prefixes, so it belongs in the cache directory like everything else derived. report checks those prefixes against the corpus either way: nothing matching is an error, and partial coverage is stated in every use-case table so the labelled subset is never read as the whole corpus.

Show full SKILL.md (433 more words)Show less

What to actually report

Raw counts alone mislead. The findings that carried signal in the reference corpus:

  • Token composition, not just totals. Cache-read vs cache-write vs output. Agentic coding is overwhelmingly context, with output a rounding error.
  • Binary frequency inside shell commands. The single richest signal — it shows which tools the model actually reaches for versus which ones it was given.
  • Failure rate per tool, not corpus-wide. A corpus-wide average hides wide per-tool variation — always break failure rate out per tool.
  • Failure rate by position in the session. Tests whether reliability decays as context fills. Measure it rather than assuming decay; it may well be flat.
  • What happens after a failure — recovery rate and streak length separate "handles errors well" from "gets stuck".
  • Main-session vs subagent rates, which isolates the cost of write capability.

Honesty requirements

These are not optional; the analysis is worthless without them.

  • There is no ground truth for task success. Nothing in a transcript says whether the goal was met. Never present a tool-failure rate as a task-failure rate.
  • Friction signals are regex proxies. Corrections and interrupts are evidence, not proof. Compare rates across use-cases; do not quote absolutes as fact.
  • Cost is API-equivalent, not money spent, if the sessions ran on a subscription.
  • Duration is unusable past the median — a session left open overnight reports the whole night.
  • The corpus grows while you analyse it. The session doing the analysis is itself being recorded, so counts drift between runs. Timestamp the snapshot.
  • Every derived table needs a column glossary. A column nobody can define is a column nobody should trust.

Privacy

scan covers every Claude Code project on the machine, not just the current repo: ~/.claude/projects/ holds one subdirectory per project, and there is no per-project filter — --root relocates the scan, it cannot narrow it. tables.md never breaks the count down by project, so check top_projects in stats.json to see the mix before sharing anything.

Transcripts contain absolute paths, branch names, repository content, and any secret pasted into a prompt. The pipeline writes only to ~/.gaia/cache/factory/. If a report is shared, put it somewhere private and scrub paths first. Nothing derived from a corpus belongs in a public repository.

Verifying the analysis

Before publishing, have a subagent adversarially review the extraction code against the corpus — not just the prose. Four of the reference report's numbers were wrong on first pass, including its headline, and every one was a bug in the metric rather than a mistake in the writing. Ask specifically: does each metric measure what its name claims, and is the denominator the one the sentence implies?

© amd, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/analyzing-claude-sessions of amd/gaia.

Open the folder on GitHubat commit 05fb50b

Compare with similar skills

Analyzing Claude Code Sessions next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Analyzing Claude Code Sessions compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Analyzing Claude Code Sessions this skillamd/gaia1.6k—~2.3kAutomated safety check: PassMIT
A-Evolve Agent EvolutionOrchestra-Research/AI-Research-SKILLs13k—~3.6kAutomated safety check: PassMIT
AI Observabilityomer-metin/skills-for-antigravity162—~578Automated safety check: PassApache-2.0
Caveman Workflow LabelerJuliusBrussee/caveman111k1 repos~1.3kAutomated safety check: PassApache-2.0
Caveman Evidence ReviewJuliusBrussee/caveman111k1 repos~927Automated safety check: PassApache-2.0
Copilot Session Failure Analysisdotnet/maui23k—~3.4kAutomated safety check: PassMIT

Similar skills

  • A-Evolve Agent Evolution

    Orchestra-Research/AI-Research-SKILLs

    Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.

    13k GitHub stars~3.6k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • AI Observability

    omer-metin/skills-for-antigravity

    Implement comprehensive observability for LLM applications including tracing (Langfuse/Helicone), cost tracking, token optimization, RAG evaluation metrics (RAGAS), hallucination detection, and…

    162 GitHub stars~578 tokensUpdated 8 mo ago
    AI & LLM EngineeringAuto-check passed
  • Caveman Workflow Labeler

    JuliusBrussee/caveman

    Finds every LLM workflow in a repository, proposes a labeling table and, once you agree, wires labels so Caveman Cloud groups spend per workflow.

    111k GitHub starsUsed in 1 repo~1.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Caveman Evidence Review

    JuliusBrussee/caveman

    Read-only review of Caveman Cloud data to explain where LLM spend goes: cost, score, workflows, traces, latency, errors, routing and verified savings.

    111k GitHub starsUsed in 1 repo~927 tokens
    AI & LLM EngineeringAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • CodexBar Usage Reader

    steipete/CodexBar

    CodexBar read. Provider usage, limits, credits, config health. JSON. No writes.

    22k GitHub stars~320 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from amd/gaia

All 44 skills in this repo
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated yesterday
    Auto-check passed
  • Walks through releasing a GAIA sidecar agent as a frozen binary plus npm client through the tag-triggered Agent Hub CI pipeline, with a human gate before publishing.

    1.6k GitHub stars~3.6k tokensUpdated yesterday
    Auto-check passed
  • Benchmarks AMD's GAIA agent against Claude Code and across models on quality, honesty, steps, tokens, time and real cost, using gaia eval tasks.

    1.6k GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Guides safe code changes by finding the right file with grep or semantic search, reading before editing, reproducing bugs first, and proving a fix with a real test run.

    1.6k GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed
  • Walks through scaffolding, writing and testing a new GAIA agent as a Python class with the SDK, from the base Agent subclass to registered tool methods.

    1.6k GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Turns a source document such as a README or spec into an executive slide deck as one self-contained HTML file that prints to PDF, one slide per page.

    1.6k GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about Analyzing Claude Code Sessions

What does Analyzing Claude Code Sessions do?

Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs. Claude Code writes a full JSONL transcript of every session to a local projects folder, recording every tool call, failure and cost.harvest, is deterministic Python with no model or network calls; two optional steps, classify and synthesize, do call a model through the local Claude Code CLI.

When should I use Analyzing Claude Code Sessions?

Analyzing Claude Code Sessions fits situations like: finding out what tasks an agent is actually being used for across many sessions; measuring how often Claude Code sessions fail and why; estimating token and dollar cost across a corpus of sessions; producing an evidence-backed report on real agent usage.

How do I install Analyzing Claude Code Sessions in Claude Code?

Run `npx skills add amd/gaia --skill analyzing-claude-sessions -a claude-code`. Or copy the skill folder (.claude/skills/analyzing-claude-sessions in amd/gaia) into .claude/skills/analyzing-claude-sessions in your project. Claude Code loads it when a task matches its description.

How do I install Analyzing Claude Code Sessions in Codex?

Run `npx skills add amd/gaia --skill analyzing-claude-sessions -a codex`. Or copy the skill folder (.claude/skills/analyzing-claude-sessions in amd/gaia) into .agents/skills/analyzing-claude-sessions in your project. Codex loads it when a task matches its description.

Can I use Analyzing Claude Code Sessions in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add amd/gaia --skill analyzing-claude-sessions -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/analyzing-claude-sessions, .gemini/skills/analyzing-claude-sessions, .github/skills/analyzing-claude-sessions and .opencode/skills/analyzing-claude-sessions in your project.

What does Analyzing Claude Code Sessions need to run?

Going by SKILL.md and its folder, Analyzing Claude Code Sessions needs the command-line tools its instructions call (python and git). Our summary lists: Python, for the deterministic harvesting pipeline; The Claude Code CLI, for the optional classify and synthesize steps; Local Claude Code session transcripts under ~/.claude/projects/.

Does Analyzing Claude Code Sessions access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Analyzing Claude Code Sessions safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Analyzing Claude Code Sessions use?

Analyzing Claude Code Sessions is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Analyzing Claude Code Sessions use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Analyzing Claude Code Sessions?

Skills that share tags, products or a category with Analyzing Claude Code Sessions: A-Evolve Agent Evolution (Orchestra-Research/AI-Research-SKILLs, 13k stars), AI Observability (omer-metin/skills-for-antigravity, 162 stars), Caveman Workflow Labeler (JuliusBrussee/caveman, 111k stars) and Caveman Evidence Review (JuliusBrussee/caveman, 111k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Analyzing Claude Code Sessions?

amd (a GitHub organization) maintains it in amd/gaia, which has 1,579 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on October 8, 2026.

Source: amd/gaia on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.