Agent skill

Coding Trace Normalize

by uw-syfi in uw-syfi/TraceLab

Explain and work with the normalized coding-trace JSONL row format produced by extractclauderounds.py, extractcodexrounds.py, and collectllmtraces.py.

Apache-2.0Auto-check passedBusiness, Finance & HR

Install Coding Trace Normalize

skills CLI
$ npx skills add uw-syfi/TraceLab --skill coding-trace-normalize -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install uw-syfi/TraceLab coding-trace-normalize --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/uw-syfi/TraceLab.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/coding-trace-normalize .claude/skills/coding-trace-normalize && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coding-trace-normalize
GitHub stars
138
Token cost
~1.6k tokens
SKILL.md length
649 words
Files
2
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Explain and work with the normalized coding-trace JSONL row format produced by extractclauderounds.py, extractcodexrounds.py, and collectllmtraces.py.

  • Works in 4 steps: Work from the coding-trace repo root,… → Read README.md section "Normalized Rows"… → Read example_sessions/derived/README.md… → …
  • Interpreting normalized rows
  • SKILL.md covers Overview, First Steps, Row Identity and Token Semantics, plus 3 more sections
  • Calls uv

What it does

Coding Trace Normalize is an agent skill from uw-syfi/TraceLab. Explain and work with the normalized coding-trace JSONL row format produced by extractclauderounds.py, extractcodexrounds.py, and collectllmtraces.py. Use when interpreting normalized rows, provider differences, token semantics, Claude cache creation/read/uncached fields, prefix versus append accounting, input-event summaries, timingevents, tools metadata, tracekey identity, or converting raw Claude/Codex concepts into the common analysis schema.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Business, Finance & HR, covering Accounting and bookkeeping. The repository describes itself as: An open toolkit and public dataset hub for collecting, sanitizing, analyzing, and visualizing coding agent traces. The licence is Apache-2.0.

When your agent uses it

  • Interpreting normalized rows
  • Provider differences
  • Token semantics
  • Claude cache creation/read/uncached fields

Example prompts

  • “/coding-trace-normalize”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Work from the coding-trace repo root, identified by pyproject.toml, README.md, and scripts/collect_llm_traces.py.
  2. Read README.md section "Normalized Rows" for the public contract.
  3. Read example_sessions/derived/README.md and inspect example_sessions/derived/round_trace.expanded.json for a small worked example.
  4. If a field is ambiguous, inspect the extractor that writes it: scripts/extract_claude_rounds.py, scripts/extract_codex_rounds.py, or…

What it can do on your machine

Read from SKILL.md and the folder at commit 11b8b14. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Coding Trace Normalize loads about 1.6k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 649 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from uw-syfi/TraceLab at commit 11b8b14, republished under its Apache-2.0 licence (© uw-syfi). 649 words, ~1,623 tokens.

Download SKILL.mdSave it as .claude/skills/coding-trace-normalize/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
coding-trace-normalize
description
Explain and work with the normalized coding-trace JSONL row format produced by extract_claude_rounds.py, extract_codex_rounds.py, and collect_llm_traces.py. Use when interpreting normalized rows, provider differences, token semantics, Claude cache creation/read/uncached fields, prefix versus append accounting, input-event summaries, timing_events, tools metadata, trace_key identity, or converting raw Claude/Codex concepts into the common analysis schema.

Coding Trace Normalize

Overview

Use this skill to interpret the common JSONL format produced by this repo. Each normalized row represents one LLM invocation, even though the raw Claude and Codex logs use different event schemas.

First Steps

  1. Work from the coding-trace repo root, identified by pyproject.toml, README.md, and scripts/collect_llm_traces.py.
  2. Read README.md section "Normalized Rows" for the public contract.
  3. Read example_sessions/derived/README.md and inspect example_sessions/derived/round_trace.expanded.json for a small worked example.
  4. If a field is ambiguous, inspect the extractor that writes it: scripts/extract_claude_rounds.py, scripts/extract_codex_rounds.py, or scripts/collect_llm_traces.py.

Row Identity

Common top-level fields:

  • provider: claude or codex.
  • session_id: provider session identifier.
  • round_index: zero-based order within the extracted session when available.
  • round_id: provider-specific round identifier.
  • trace_key: stable dedupe key, built as {provider}:{session_id}:{round_id}. Some session_id values already include a provider prefix, so keys such as claude:claude:...:msg_... are expected.
  • model: model reported by the provider trace.
  • project, store, user, home, turn_id, cwd, or session_file: private extracted traces may include local context fields. Sanitized traces pseudonymize project, user, and ids, and remove local path/source fields.
  • current_input_event_count, current_user_message_count, current_tool_result_count, current_user_message_chars, current_tool_result_chars, current_input_chars, and first_input_event_type: summaries derived from the row's input-side timing_events.

Token Semantics

Normalized input fields follow this invariant:

text
input_tokens_total = prefix_tokens + newly_append_tokens

Provider mappings:

  • Claude prefix_tokens comes from cache_read_input_tokens.
  • Claude newly_append_tokens comes from input_tokens + cache_creation_input_tokens.
  • Claude input_tokens_total is cache_read_input_tokens + cache_creation_input_tokens + input_tokens.
  • Claude rows also keep the provider-specific split as claude_uncached_input_tokens, claude_cache_creation_input_tokens, and claude_cache_read_input_tokens.
  • Claude reasoning_output_tokens is normally null because Claude examples do not expose a separate reasoning-token field in this schema.
  • Codex prefix_tokens comes from cached_input_tokens.
  • Codex newly_append_tokens is input_tokens - cached_input_tokens, clamped at zero by the extractor.
  • Codex input_tokens_total comes from input_tokens.
  • Codex rows set the Claude-specific cache fields to null.
  • Codex reasoning_output_tokens is a subset of output_tokens, not extra output.

Interpretation guidance:

  • prefix_tokens approximates prompt-cache hit tokens.
  • newly_append_tokens approximates prompt-side work not served from prompt cache.
  • These are provider prompt-cache semantics, not a direct dump of local serving-engine KV-cache objects.
  • output_tokens is inclusive generated output for the round. Do not add reasoning_output_tokens to it.
Show full SKILL.md (322 more words)Show less

Timing Events

timing_events[] preserves the ordered trace-observed timeline for the normalized round. Common event types:

  • user_message: visible user input that triggered or preceded a model round.
  • tool_result: tool output returned to the model before a later model round.
  • reasoning: existence of a reasoning or thinking output item, usually without readable private content.
  • text: visible assistant text.
  • tool_call: model-emitted tool request.
  • usage_report: Codex token accounting event.

For an "input ready to next tool input" latency proxy, use the latest user_message or tool_result before the first following tool_call, then subtract timestamps. For reasoning timing, the summarizer distinguishes latest input to reasoning marker/end from the exact-reasoning TPOT subset. TPOT and TTFT residual estimates require exact numeric reasoning_output_tokens; otherwise they should remain null.

The current-input summary fields are derived only from user_message and tool_result timing events in the row. They are convenience counters for analysis, not replacements for timing_events[] when exact ordering matters.

Tool Metadata

Each row can contain tools[], with one object per model-emitted tool call:

  • tool_index: order within the row.
  • tool_name: provider/tool runner name such as Bash, exec_command, Read, or apply_patch.
  • tool_call_id: call id, paired with timing events when available.
  • emitted_at: timestamp when the model emitted the tool call.
  • result_at: timestamp when a tool result was observed, or null.
  • input: raw tool input in private extracted traces. Sanitized traces remove it.
  • input_chars: serialized input size retained even when input is removed.
  • result_chars: result size; full outputs are not kept in normalized rows.
  • tool_wall_latency_ms: trace-observed result_at - emitted_at.
  • tool_internal_latency_ms: tool/runner-reported duration when available.
  • is_error: error status if inferable.

Analysis experiments use tool_internal_latency_ms when present, then fall back to tool_wall_latency_ms. The multi-round CSV exporter uses tool_wall_latency_ms by default unless --tool-latency-source internal is passed.

Quick Inspection

Inspect a small public normalized example:

bash
sed -n '1,120p' example_sessions/derived/round_trace.expanded.json

Summarize a normalized file:

bash
uv run python artifacts/trace_facts/overview_summary/analyze.py \
  --db trace/syfi_coding_trace.duckdb --json

When answering field questions, cite the relevant extractor or docs and avoid exposing private tools[].input or local context from unsanitized rows unless the user explicitly asks to inspect their own private data.

© uw-syfi, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/coding-trace-normalize of uw-syfi/TraceLab.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 11b8b14

Compare with similar skills

Coding Trace Normalize next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coding Trace Normalize compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coding Trace Normalize this skilluw-syfi/TraceLab138—~1.6kAutomated safety check: PassApache-2.0
Sync Upstreamnyaruka/phonenumbers1.6k—~2.8kAutomated safety check: PassMIT
Longbridge Value Investinghelsome/folio2692 repos~1.2kAutomated safety check: PassMIT
Radiology Tablehuang-sir1/radiology-skills1.9k—~1.3kAutomated safety check: PassCustom licence
Odoo Agency Fleet Reviewerpipe-org/mcp-odoo420—~699Automated safety check: PassMIT
Beancount Closebex-co/beancount-io295—~1.4kAutomated safety check: PassMIT

Similar skills

  • Sync Upstream

    nyaruka/phonenumbers

    Sync this Go port with a new upstream google/libphonenumber release — regenerate the embedded metadata and reconcile the ported Java logic.

    1.6k GitHub stars~2.8k tokensUpdated 6 days ago
    Business, Finance & HRAuto-check passed
  • Value investing analysis using Graham (NCAV/net-net/defensive-investor) and Buffett (economic moat/ROE/FCF) methodologies.

    269 GitHub starsUsed in 2 repos~1.2k tokens
    Business, Finance & HRAuto-check passed
  • Radiology Table

    huang-sir1/radiology-skills

    Create/audit editable publication tables with source reconciliation; not figures or statistical inference.

    1.9k GitHub stars~1.3k tokensUpdated 17 days ago
    Business, Finance & HRAuto-check passed
  • Odoo Agency Fleet Review

    erpipe-org/mcp-odoo

    Review many client Odoo databases at once through odoo-mcp's cross-instance tools — fleet-wide accounting health, per-client aging, partial-failure triage — for agencies and partners managing 5–50…

    420 GitHub stars~699 tokensUpdated 1 mo ago
    Business, Finance & HRAuto-check passed
  • Beancount Close

    bex-co/beancount-io

    Close an accounting period in a Beancount ledger by reconciling each active account through beancount-reconcile, checking assertions and recurring gaps, reviewing flags, then proposing a commit with…

    295 GitHub stars~1.4k tokensUpdated yesterday
    Business, Finance & HRAuto-check passed
  • ERPClaw ERP Controller

    avansaber/erpclaw

    Operates the ERPClaw self-hosted ERP in plain language: accounting, invoicing, inventory, purchasing, tax, HR, payroll and reports, treating the ERP as the single source of truth.

    114 GitHub stars~15k tokensUpdated 2 days ago
    Business, Finance & HRAuto-check passed

More from uw-syfi/TraceLab

  • Prepare, publish, refresh, or validate the TraceLab public dataset on Hugging Face under UW-SyFI/TraceLab.

    138 GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Coding Trace Analyze

    uw-syfi/TraceLab

    Run, summarize, plot, validate, and export normalized coding-trace JSONL files.

    138 GitHub stars~2.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Coding Trace Collect

    uw-syfi/TraceLab

    Collect, count, and extract Claude Code and Codex CLI local histories into normalized coding-trace JSONL files.

    138 GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check: notes
  • Coding Trace Raw

    uw-syfi/TraceLab

    Read and explain raw Claude Code and Codex CLI session logs in this coding-trace repo.

    138 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Coding Trace Sanitize

    uw-syfi/TraceLab

    Sanitize normalized coding-trace JSONL rows for public sharing.

    138 GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check: notes
  • Tracelab Public Release

    uw-syfi/TraceLab

    Publish TraceLab snapshots from the internal repo to the public uw-syfi/TraceLab GitHub repo as clean, mergeable, incremental pull requests from a persistent public mirror branch.

    138 GitHub stars~2.7k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Coding Trace Normalize

What does Coding Trace Normalize do?

Explain and work with the normalized coding-trace JSONL row format produced by extractclauderounds.py, extractcodexrounds.py, and collectllmtraces.py. Coding Trace Normalize is an agent skill from uw-syfi/TraceLab.py.

When should I use Coding Trace Normalize?

Coding Trace Normalize fits situations like: interpreting normalized rows; provider differences; token semantics; Claude cache creation/read/uncached fields.

How do I install Coding Trace Normalize in Claude Code?

Run `npx skills add uw-syfi/TraceLab --skill coding-trace-normalize -a claude-code`. Or copy the skill folder (skills/coding-trace-normalize in uw-syfi/TraceLab) into .claude/skills/coding-trace-normalize in your project. Claude Code loads it when a task matches its description.

How do I install Coding Trace Normalize in Codex?

Run `npx skills add uw-syfi/TraceLab --skill coding-trace-normalize -a codex`. Or copy the skill folder (skills/coding-trace-normalize in uw-syfi/TraceLab) into .agents/skills/coding-trace-normalize in your project. Codex loads it when a task matches its description.

Can I use Coding Trace Normalize in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add uw-syfi/TraceLab --skill coding-trace-normalize -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coding-trace-normalize, .gemini/skills/coding-trace-normalize, .github/skills/coding-trace-normalize and .opencode/skills/coding-trace-normalize in your project.

What does Coding Trace Normalize need to run?

Going by SKILL.md and its folder, Coding Trace Normalize needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Coding Trace Normalize access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Coding Trace Normalize safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Coding Trace Normalize use?

Coding Trace Normalize is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coding Trace Normalize use?

About 1.6k tokens (SKILL.md is roughly 6.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coding Trace Normalize?

Skills that share tags, products or a category with Coding Trace Normalize: Sync Upstream (nyaruka/phonenumbers, 1.6k stars), Longbridge Value Investing (helsome/folio, 269 stars), Radiology Table (huang-sir1/radiology-skills, 1.9k stars) and Odoo Agency Fleet Review (erpipe-org/mcp-odoo, 420 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coding Trace Normalize?

uw-syfi (a GitHub organization) maintains it in uw-syfi/TraceLab, which has 138 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on August 22, 2026.

Source: uw-syfi/TraceLab on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.