Agent skill

Coding Trace Collect

by uw-syfi in uw-syfi/TraceLab

Collect, count, and extract Claude Code and Codex CLI local histories into normalized coding-trace JSONL files.

Apache-2.0Auto-check: notes

Install Coding Trace Collect

skills CLI
$ npx skills add uw-syfi/TraceLab --skill coding-trace-collect -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install uw-syfi/TraceLab coding-trace-collect --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/uw-syfi/TraceLab.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/coding-trace-collect .claude/skills/coding-trace-collect && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coding-trace-collect
GitHub stars
138
Token cost
~1.9k tokens
SKILL.md length
627 words
Files
2
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Collect, count, and extract Claude Code and Codex CLI local histories into normalized coding-trace JSONL files.

  • Works in 6 steps: Work from the coding-trace repo root,… → Run every Python entry point through uv… → Decide scan scope before running commands → …
  • Running collectllmtraces.py
  • SKILL.md covers Overview, First Steps, Main Collection Commands and Timing Fit CSV, plus 4 more sections
  • Calls uv

What it does

Coding Trace Collect is an agent skill from uw-syfi/TraceLab. Collect, count, and extract Claude Code and Codex CLI local histories into normalized coding-trace JSONL files. Use when running collectllmtraces.py, extractclauderounds.py, extractcodexrounds.py, collectalluserssudo.sh, scanning current-user or all-user homes, choosing output paths, using --fresh-extract or --append-dedup, or troubleshooting extraction counts.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

The repository describes itself as: An open toolkit and public dataset hub for collecting, sanitizing, analyzing, and visualizing coding agent traces. The licence is Apache-2.0.

When your agent uses it

  • Running collectllmtraces.py
  • Extractclauderounds.py
  • Extractcodexrounds.py
  • Collectalluserssudo.sh

Example prompts

  • “/coding-trace-collect”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Work from the coding-trace repo root, identified by pyproject.toml, README.md, and scripts/collect_llm_traces.py.
  2. Run every Python entry point through uv run python ....
  3. Decide scan scope before running commands
  4. Decide write behavior
  5. Do not publish private outputs until $coding-trace-sanitize has been applied.
  6. After sanitization, use $coding-trace-analyze for artifact and validator dispatchers. Collection should not own plotting outputs or…

What it can do on your machine

Read from SKILL.md and the folder at commit 11b8b14. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Coding Trace Collect loads about 1.9k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 627 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:18
    under `/home`: use `--all-user`, or the sudo wrapper when unreadable homes are expected. Use `--home-root PATH` for non
  • NoteRuns commands with sudoSKILL.md:21
    - The sudo wrapper's default dated output is a fresh snapshot of currently readable sessions.
  • NoteRuns commands with sudoSKILL.md:57
    Use the sudo wrapper for a fresh, named all-user snapshot while preserving final file

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from uw-syfi/TraceLab at commit 11b8b14, republished under its Apache-2.0 licence (© uw-syfi). 627 words, ~1,911 tokens.

Download SKILL.mdSave it as .claude/skills/coding-trace-collect/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
coding-trace-collect
description
Collect, count, and extract Claude Code and Codex CLI local histories into normalized coding-trace JSONL files. Use when running collect_llm_traces.py, extract_claude_rounds.py, extract_codex_rounds.py, collect_all_users_sudo.sh, scanning current-user or all-user homes, choosing output paths, using --fresh-extract or --append-dedup, or troubleshooting extraction counts.

Coding Trace Collect

Overview

Use this skill to collect local Claude Code and Codex history stores and write normalized round traces. Collection finds source files and counts sessions; extraction converts provider-specific histories into JSONL rows where each line is one LLM invocation.

First Steps

  1. Work from the coding-trace repo root, identified by pyproject.toml, README.md, and scripts/collect_llm_traces.py.
  2. Run every Python entry point through uv run python ....
  3. Decide scan scope before running commands:
    • Current user: default behavior.
    • All users under /home: use --all-user, or the sudo wrapper when unreadable homes are expected. Use --home-root PATH for nonstandard home roots.
  4. Decide write behavior:
    • Default extraction appends deduped rows.
    • The sudo wrapper's default dated output is a fresh snapshot of currently readable sessions.
    • For a corpus refresh, merge fresh dated host snapshots after trace/collections/current/merged.private.jsonl; this both refreshes normalizer fields and retains history whose raw sessions have been cleaned up.
    • Use an explicit cumulative --output without --fresh-extract only when intentionally maintaining one append-only file.
  5. Do not publish private outputs until $coding-trace-sanitize has been applied.
  6. After sanitization, use $coding-trace-analyze for artifact and validator dispatchers. Collection should not own plotting outputs or timing-fit CSVs.

Main Collection Commands

Count current-user Claude and Codex history without extraction:

bash
uv run python scripts/collect_llm_traces.py

Extract a combined current-user trace:

bash
uv run python scripts/collect_llm_traces.py --extract-rounds

Write to a specific file and start fresh. This is the recommended current-user command only when intentionally creating a new, non-cumulative dataset:

bash
uv run python scripts/collect_llm_traces.py --extract-rounds trace/llm_round_trace.jsonl --fresh-extract

Scan all homes under /home without sudo:

bash
uv run python scripts/collect_llm_traces.py --all-user --extract-rounds

Use the sudo wrapper for a fresh, named all-user snapshot while preserving final file ownership. Reuse one collection_id on every host:

bash
collection_id="20260724-100b"
scripts/collect_all_users_sudo.sh \
  --collection-id "$collection_id" \
  --fresh-extract

The wrapper writes trace/collections/<collection_id>/hosts/llm_round_trace_v2.<collection_id>.<host>.all_users.jsonl and a sibling .collection_report.json. With --sanitize, it also writes a sibling .public.jsonl. It prints an overview summary by default; add --no-summary for quiet batch runs. Add --quiet-progress only to suppress collector progress messages.

To combine older and newer private normalized traces while letting newer normalizer fields win on overlap:

bash
uv run python scripts/merge_round_traces.py \
  trace/collections/current/merged.private.jsonl \
  "trace/collections/${collection_id}/hosts/"*.all_users.jsonl \
  -o "trace/collections/${collection_id}/merged.private.jsonl"

For a private-to-public current-user flow:

bash
collection_id="20260724-100b"

uv run python scripts/collect_llm_traces.py \
  --extract-rounds "trace/collections/${collection_id}/merged.private.jsonl" \
  --fresh-extract

uv run python scripts/sanitize_round_trace.py \
  "trace/collections/${collection_id}/merged.private.jsonl" \
  -o "trace/collections/${collection_id}/merged.public.jsonl"

Timing Fit CSV

The long-form timing-segment CSV is now owned by the timing-fit artifact, not the collection pipeline. artifacts/run_all.py --db <trace.duckdb> builds it automatically before timing analyses. Do not precompute it during normal collection. To build it directly for a timing-only manual run:

bash
uv run python artifacts/llm_generation/timing_fit/collect_timing_fit_trace.py \
  --db trace/collections/current/merged.public.duckdb \
  -o artifacts/llm_generation/timing_fit/timing_fit_trace.csv

The output has one timing segment per row. Claude segments use message-level output accounting; Codex can also emit reasoning-split segments when reasoning markers and exact reasoning-token counts are available.

Show full SKILL.md (238 more words)Show less

Output Paths

Important output names:

  • Current-user combined trace: trace/llm_round_trace.jsonl.
  • Current-user Claude-only trace: trace/claude_round_trace.jsonl.
  • Current-user Codex-only trace: trace/codex_round_trace.jsonl.
  • Sudo-wrapper all-user trace: trace/collections/<collection_id>/hosts/llm_round_trace_v2.<collection_id>.<host>.all_users.jsonl.
  • Sudo-wrapper report: the sibling .collection_report.json.
  • Sudo-wrapper sanitized output with --sanitize: the sibling .public.jsonl.
  • Selected historical collection: trace/collections/current.
  • Timing-fit artifact CSV: artifacts/llm_generation/timing_fit/timing_fit_trace.csv.

Use --trace-dir DIR to change the default extraction directory for omitted output paths. Passing an explicit path to --extract-rounds, --extract-claude-rounds, or --extract-codex-rounds overrides both the built-in default and --trace-dir.

Provider-Specific Extraction

Use direct extractors when the user already has a provider source directory:

bash
uv run python scripts/extract_claude_rounds.py ~/.claude/projects/PROJECT_DIR -o trace/claude_round_trace.jsonl --append-dedup
bash
uv run python scripts/extract_codex_rounds.py ~/.codex/sessions -o trace/codex_round_trace.jsonl --append-dedup

Prefer collect_llm_traces.py for normal use because it scans both providers, handles .claude.back, tracks skipped paths, and combines extraction stats.

Important Options

  • --json: emit a machine-readable collection report.
  • --no-claude-back: skip .claude.back/projects.
  • --quiet-host-progress: collector option that suppresses progress messages during host scanning and extraction.
  • --quiet-progress: sudo-wrapper option that passes the collector quiet mode.
  • --collection-id: safe collection slug used in both the collection directory and file names; prefer a UTC-date prefix such as 20260724-100b.
  • --no-summary: sudo-wrapper option that skips the post-collection overview summary.
  • --no-sudo: sudo-wrapper option that runs all-user collection without sudo.
  • --extract-project-filter TEXT: only extract Claude projects whose directory name contains TEXT; repeat the option for multiple filters.
  • --fresh-extract: remove selected extraction outputs before writing.
  • --append-dedup: direct extractor option that appends only unseen trace_key rows.

Validation After Collection

After extraction, check basic shape before analysis:

bash
uv run python artifacts/trace_facts/overview_summary/analyze.py -i trace/collections/current/merged.private.jsonl

For JSON output:

bash
uv run python artifacts/trace_facts/overview_summary/analyze.py -i trace/collections/current/merged.private.jsonl --json

If outputs are unexpectedly small, inspect the collection report for skipped_paths, source counts, candidate_rounds, written_rounds, and skipped_duplicate_rounds.

© uw-syfi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/coding-trace-collect of uw-syfi/TraceLab.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 11b8b14

Compare with similar skills

Coding Trace Collect next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coding Trace Collect compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coding Trace Collect this skilluw-syfi/TraceLab138—~1.9kAutomated safety check: NotesApache-2.0
Observe Traceruvnet/ruflo74k—~522Automated safety check: NotesMIT
Dotnet Trace Collectdotnet/skills5.6k2 repos~6.1kAutomated safety check: PassMIT
Extractalirezarezvani/claude-skills28k—~1.4kAutomated safety check: PassMIT
Brand Extractnexu-io/open-design100k—~3.1kAutomated safety check: PassApache-2.0
Design Extractnexu-io/open-design100k—~549Automated safety check: PassApache-2.0

Similar skills

  • Observe Trace

    ruvnet/ruflo

    Trace agent execution by collecting spans and building a trace tree for a task

    74k GitHub stars~522 tokensUpdated today
    Auto-check: notes
  • Official

    Guide developers through capturing diagnostic artifacts to diagnose production .NET performance issues.

    5.6k GitHub starsUsed in 2 repos~6.1k tokens
    Auto-check passed
  • Extract

    alirezarezvani/claude-skills

    Turn a proven pattern or debugging solution into a standalone reusable skill with SKILL.md, reference docs, and examples.

    28k GitHub stars~1.4k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Brand Extract

    nexu-io/open-design

    Extract a complete Brand Kit from a live website by driving the in-app browser.

    100k GitHub stars~3.1k tokensUpdated today
    Productivity & AutomationAuto-check passed
  • Design Extract

    nexu-io/open-design

    Extract design tokens (color / typography / spacing) from imported source code, screenshots, or Figma exports into the canonical token bag token-map consumes.

    100k GitHub stars~549 tokensUpdated today
    Frontend & DesignAuto-check passed
  • Kg Extract

    ruvnet/ruflo

    Extract entities and relations from source files to build a knowledge graph

    74k GitHub stars~751 tokensUpdated today
    Knowledge ManagementAuto-check: notes

More from uw-syfi/TraceLab

  • Prepare, publish, refresh, or validate the TraceLab public dataset on Hugging Face under UW-SyFI/TraceLab.

    138 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Coding Trace Analyze

    uw-syfi/TraceLab

    Run, summarize, plot, validate, and export normalized coding-trace JSONL files.

    138 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Coding Trace Normalize

    uw-syfi/TraceLab

    Explain and work with the normalized coding-trace JSONL row format produced by extractclauderounds.py, extractcodexrounds.py, and collectllmtraces.py.

    138 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Coding Trace Raw

    uw-syfi/TraceLab

    Read and explain raw Claude Code and Codex CLI session logs in this coding-trace repo.

    138 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Coding Trace Sanitize

    uw-syfi/TraceLab

    Sanitize normalized coding-trace JSONL rows for public sharing.

    138 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check: notes
  • Tracelab Public Release

    uw-syfi/TraceLab

    Publish TraceLab snapshots from the internal repo to the public uw-syfi/TraceLab GitHub repo as clean, mergeable, incremental pull requests from a persistent public mirror branch.

    138 GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Coding Trace Collect

What does Coding Trace Collect do?

Collect, count, and extract Claude Code and Codex CLI local histories into normalized coding-trace JSONL files. Coding Trace Collect is an agent skill from uw-syfi/TraceLab. Collect, count, and extract Claude Code and Codex CLI local histories into normalized coding-trace JSONL files.

When should I use Coding Trace Collect?

Coding Trace Collect fits situations like: running collectllmtraces.py; extractclauderounds.py; extractcodexrounds.py; collectalluserssudo.sh.

How do I install Coding Trace Collect in Claude Code?

Run `npx skills add uw-syfi/TraceLab --skill coding-trace-collect -a claude-code`. Or copy the skill folder (skills/coding-trace-collect in uw-syfi/TraceLab) into .claude/skills/coding-trace-collect in your project. Claude Code loads it when a task matches its description.

How do I install Coding Trace Collect in Codex?

Run `npx skills add uw-syfi/TraceLab --skill coding-trace-collect -a codex`. Or copy the skill folder (skills/coding-trace-collect in uw-syfi/TraceLab) into .agents/skills/coding-trace-collect in your project. Codex loads it when a task matches its description.

Can I use Coding Trace Collect in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add uw-syfi/TraceLab --skill coding-trace-collect -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coding-trace-collect, .gemini/skills/coding-trace-collect, .github/skills/coding-trace-collect and .opencode/skills/coding-trace-collect in your project.

What does Coding Trace Collect need to run?

Going by SKILL.md and its folder, Coding Trace Collect needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Coding Trace Collect access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Coding Trace Collect safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Coding Trace Collect use?

Coding Trace Collect is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coding Trace Collect use?

About 1.9k tokens (SKILL.md is roughly 7.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coding Trace Collect?

Skills that share tags, products or a category with Coding Trace Collect: Observe Trace (ruvnet/ruflo, 74k stars), Dotnet Trace Collect (dotnet/skills, 5.6k stars), Extract (alirezarezvani/claude-skills, 28k stars) and Brand Extract (nexu-io/open-design, 100k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coding Trace Collect?

uw-syfi (a GitHub organization) maintains it in uw-syfi/TraceLab, which has 138 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on August 22, 2026.

Source: uw-syfi/TraceLab on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.