Agent skill

Braintrust Tracing

by parcadei in parcadei/Continuous-Claude-v3

Braintrust tracing for Claude Code - hook architecture, sub-agent correlation, debugging

MITAuto-check passedAgent Workflows

Install Braintrust Tracing

skills CLI
$ npx skills add parcadei/Continuous-Claude-v3 --skill braintrust-tracing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install parcadei/Continuous-Claude-v3 braintrust-tracing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/parcadei/Continuous-Claude-v3.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/braintrust-tracing .claude/skills/braintrust-tracing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
braintrust-tracing
GitHub stars
3.9k
Used in
1 other repo
Token cost
~3.2k tokens
SKILL.md length
873 words
Files
1
Skills in repo
141
Repo updated
First seen
Licence
MIT

At a glance

Braintrust tracing for Claude Code - hook architecture, sub-agent correlation, debugging

  • Works in 3 steps: Query parent session's Task spans for… → Match agentId or timing with orphaned… → Sub-agent's session_id = its trace's…
  • Tasks that involve Subagents
  • SKILL.md covers Architecture Overview, Hook Event Flow, Trace Hierarchy and Sub-Agent Tracing: What Works…, plus 6 more sections
  • Calls uv, bash and jq; reaches api.braintrust.dev; needs BRAINTRUST_API_KEY

What it does

Braintrust Tracing is an agent skill from parcadei/Continuous-Claude-v3. Braintrust tracing for Claude Code - hook architecture, sub-agent correlation, debugging

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Subagents, Hooks and plugins and Debugging. The repository describes itself as: Context management for Claude Code. Hooks maintain state via ledgers and handoffs. MCP execution without context pollution. Agent orchestration with isolated context windows. The licence is MIT.

When your agent uses it

  • Tasks that involve Subagents
  • Tasks that involve Hooks and plugins
  • Tasks that involve Debugging

Example prompts

  • “/braintrust-tracing”

Requirements

  • Python 3
  • A credential in BRAINTRUST_API_KEY

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Query parent session's Task spans for agent metadata
  2. Match agentId or timing with orphaned traces
  3. Sub-agent's session_id = its trace's root_span_id

What it can do on your machine

Read from SKILL.md and the folder at commit d07ff4b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • bash
    • jq

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.braintrust.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • BRAINTRUST_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Braintrust Tracing loads about 3.2k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 873 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from parcadei/Continuous-Claude-v3 at commit d07ff4b, republished under its MIT licence (© parcadei). 873 words, ~3,219 tokens.

Download SKILL.mdSave it as .claude/skills/braintrust-tracing/SKILL.md (or your agent's skills folder).
name
braintrust-tracing
description
Braintrust tracing for Claude Code - hook architecture, sub-agent correlation, debugging
user-invocable
false

Braintrust Tracing for Claude Code

Comprehensive guide to tracing Claude Code sessions in Braintrust, including sub-agent correlation.

Architecture Overview

                         PARENT SESSION
                    +---------------------+
                    |  SessionStart       |
                    |  (creates root)     |
                    +----------+----------+
                               |
                    +----------v----------+
                    |  UserPromptSubmit   |
                    |  (creates Turn)     |
                    +----------+----------+
                               |
          +--------------------+--------------------+
          |                    |                    |
+---------v--------+  +--------v--------+  +--------v--------+
| PostToolUse      |  | PostToolUse     |  | PreToolUse      |
| (Read span)      |  | (Edit span)     |  | (Task - inject) |
+------------------+  +-----------------+  +--------+--------+
                                                    |
                                         +----------v----------+
                                         |   SUB-AGENT         |
                                         |   SessionStart      |
                                         |   (NEW root_span_id)|
                                         +----------+----------+
                                                    |
                                         +----------v----------+
                                         |   SubagentStop      |
                                         |   (has session_id)  |
                                         +---------------------+

Hook Event Flow

HookTriggerCreatesKey Fields
SessionStartSession beginsRoot spansession_id, root_span_id
UserPromptSubmitUser sends promptTurn spanprompt, turn_number
PreToolUseBefore tool runs(modifies Task prompts)tool_input.prompt
PostToolUseAfter tool runsTool spantool_name, input, output
StopTurn completesLLM spansmodel, tokens, tool_calls
SubagentStopSub-agent finishes(no span)session_id of sub-agent
SessionEndSession ends(finalizes root)turn_count, tool_count

Trace Hierarchy

Session (task span) - root_span_id = session_id
|
+-- Turn 1 (task span)
|   |
|   +-- claude-sonnet (llm span) - model call with tool_use
|   +-- Read (tool span)
|   +-- Edit (tool span)
|   +-- claude-sonnet (llm span) - response after tools
|
+-- Turn 2 (task span)
|   |
|   +-- claude-sonnet (llm span)
|   +-- Task (tool span) -----> [Sub-agent session - SEPARATE trace]
|   +-- claude-sonnet (llm span)
|
+-- Turn 3 ...

Sub-Agent Tracing: What Works and What Doesn't

What Doesn't Work

SessionStart doesn't receive the Task prompt.

We tried injecting trace context into Task prompts via PreToolUse:

bash
# PreToolUse hook injects:
[BRAINTRUST_TRACE_CONTEXT]
{"root_span_id": "abc", "parent_span_id": "xyz", "project_id": "123"}
[/BRAINTRUST_TRACE_CONTEXT]

But SessionStart only receives session metadata, not the modified prompt. The injected context is lost.

What DOES Work

Task spans in parent session contain everything:

  • agentId - identifier for the sub-agent run
  • totalTokens, totalToolUseCount - metrics
  • content - full agent response/summary
  • tool_input.prompt - original task prompt
  • tool_input.subagent_type - agent type (e.g., "oracle")

SubagentStop hook receives the sub-agent's session_id:

  • This equals the sub-agent's orphaned trace root_span_id
  • Allows correlation between parent Task span and child trace
The Correlation Pattern

Current state: Sub-agents create orphaned traces (new root_span_id).

Correlation method:

  1. Query parent session's Task spans for agent metadata
  2. Match agentId or timing with orphaned traces
  3. Sub-agent's session_id = its trace's root_span_id

Future solution (not yet implemented):

SubagentStop fires -> writes session_id to temp file
PostToolUse (Task) -> reads temp file -> adds child_session_id to Task span metadata

This would link: Task.agentId + Task.child_session_id -> orphaned trace root_span_id

State Management

Per-Session State Files
~/.claude/state/braintrust_sessions/
  {session_id}.json       # Per-session state

Each session file contains:

json
{
  "root_span_id": "abc-123",
  "project_id": "proj-456",
  "turn_count": 5,
  "tool_count": 23,
  "current_turn_span_id": "turn-789",
  "current_turn_start": 1703456789,
  "started": "2025-12-24T10:00:00.000Z",
  "is_subagent": false
}
Global State
~/.claude/state/braintrust_global.json   # Cached project_id
~/.claude/state/braintrust_hook.log      # Debug log

Debugging Commands

Check if Tracing is Active
bash
# View hook logs in real-time
tail -f ~/.claude/state/braintrust_hook.log

# Check if session has state
cat ~/.claude/state/braintrust_sessions/*.json | jq -s '.'

# Verify environment
echo "TRACE_TO_BRAINTRUST=$TRACE_TO_BRAINTRUST"
echo "BRAINTRUST_API_KEY=${BRAINTRUST_API_KEY:+set}"
Query Braintrust Directly
bash
# List recent sessions
uv run python -m runtime.harness scripts/braintrust_analyze.py --sessions 5

# Analyze last session
uv run python -m runtime.harness scripts/braintrust_analyze.py --last-session

# Replay specific session
uv run python -m runtime.harness scripts/braintrust_analyze.py --replay <session-id>

# Find sub-agent traces (orphaned roots)
uv run python -m runtime.harness scripts/braintrust_analyze.py --agent-stats
Debug Hook Execution
bash
# Enable verbose logging
export BRAINTRUST_CC_DEBUG=true

# Test hooks manually
echo '{"session_id":"test-123","type":"resume"}' | \
  bash "$CLAUDE_PROJECT_DIR/.claude/plugins/braintrust-tracing/hooks/session_start.sh"

# Test PreToolUse (Task injection)
echo '{"session_id":"test-123","tool_name":"Task","tool_input":{"prompt":"test"}}' | \
  bash "$CLAUDE_PROJECT_DIR/.claude/plugins/braintrust-tracing/hooks/pre_tool_use.sh"
Troubleshooting Checklist
  1. No traces appearing:

    • Check TRACE_TO_BRAINTRUST=true in .claude/settings.local.json
    • Verify API key: echo $BRAINTRUST_API_KEY
    • Check logs: tail -20 ~/.claude/state/braintrust_hook.log
  2. Sub-agents not linking:

    • This is expected - sub-agents create orphaned traces
    • Use --agent-stats to find agent activity
    • Correlate via timing or agentId in parent Task span
  3. Missing spans:

    • Check current_turn_span_id in session state
    • Ensure Stop hook runs (turn finalization)
    • Look for "Failed to create" errors in log
  4. State corruption:

    • Remove session state: rm ~/.claude/state/braintrust_sessions/*.json
    • Clear global cache: rm ~/.claude/state/braintrust_global.json

Key Files

FilePurpose
.claude/plugins/braintrust-tracing/hooks/common.shShared utilities, API, state management
.claude/plugins/braintrust-tracing/hooks/session_start.shCreates root span, handles sub-agent context
.claude/plugins/braintrust-tracing/hooks/user_prompt_submit.shCreates Turn spans per user message
.claude/plugins/braintrust-tracing/hooks/pre_tool_use.shInjects trace context into Task prompts
.claude/plugins/braintrust-tracing/hooks/post_tool_use.shCreates tool spans, captures agent/skill metadata
.claude/plugins/braintrust-tracing/hooks/stop_hook.shCreates LLM spans, finalizes Turns
.claude/plugins/braintrust-tracing/hooks/session_end.shFinalizes session, triggers learning extraction
scripts/braintrust_analyze.pyQuery and analyze traced sessions
~/.claude/state/braintrust_sessions/Per-session state files
~/.claude/state/braintrust_hook.logDebug log

Environment Variables

VariableRequiredDefaultDescription
TRACE_TO_BRAINTRUSTYes-Set to "true" to enable
BRAINTRUST_API_KEYYes-API key for Braintrust
BRAINTRUST_CC_PROJECTNoclaude-codeProject name
BRAINTRUST_CC_DEBUGNofalseVerbose logging
BRAINTRUST_API_URLNohttps://api.braintrust.devAPI endpoint

Session Learnings

What We Learned About Sub-Agent Tracing (Dec 2025)

Attempted: Inject trace context via PreToolUse into Task prompts.

Result: Failed - SessionStart only receives session metadata, not the prompt.

Discovery: Task spans already contain rich sub-agent data:

  • metadata.agent_type - agent type from subagent_type
  • metadata.skill_name - skill from Skill tool
  • tool_input - full prompt sent to agent
  • tool_output - agent response

Current correlation path:

  1. Parent session Task span has agentId and timing
  2. Sub-agent creates orphaned trace with root_span_id = session_id
  3. SubagentStop provides the sub-agent's session_id
  4. Manual correlation: match timing or use session_id link

Future work: Write child_session_id to Task span metadata from PostToolUse after SubagentStop.

Show full SKILL.md (340 more words)Show less

What We Learned About Sub-Agent Correlation

The Problem
  • Sub-agents spawned via Task tool create orphaned Braintrust traces
  • Parent session has Task spans with agentId, sub-agent has separate session_id
  • No built-in link between them
What DOESN'T Work

1. Prompt injection via PreToolUse

SessionStart hook only receives session metadata (session_id, type, cwd), NOT the prompt. Injected trace context is never seen.

The hook receives:

json
{
  "session_id": "...",
  "type": "start|resume|compact|clear",
  "cwd": "...",
  "env": {...}
}

No prompt field exists - context injection is impossible at SessionStart.

2. SubagentStop → PostToolUse file handoff

Race condition. These are independent async hooks with no timing guarantees:

  • SubagentStop fires when sub-agent session ends
  • PostToolUse (Task) fires when Task tool completes
  • No ordering guarantee between them
  • Writing to a correlation file creates a race

3. PreToolUse correlation files

SessionStart can't access the task_span_id because it has no context about which Task spawned it. PreToolUse modifies prompts but doesn't create a reliably accessible state file that SessionStart can find.

What DOES Work

Post-hoc matching for dataset building:

Parent session Task spans contain:

  • agentId - identifier for the sub-agent run
  • totalTokens, totalToolUseCount - aggregated metrics
  • content - full agent response/summary
  • tool_input.prompt - original task prompt
  • tool_input.subagent_type - agent type (e.g., "oracle")
  • Start/end timestamps

Sub-agent sessions contain:

  • session_id (equals orphaned trace root_span_id)
  • Start/end timestamps
  • All internal spans and tool calls

Correlation strategy:

  1. Export parent session traces (query parent root_span_id)
  2. Export sub-agent traces (query all sessions created within parent's time window)
  3. Match by:
    • Timing: Task span end ≈ sub-agent session end
    • Metadata: subagent_type from Task prompt
    • IDs: SubagentStop hook provides session_id (can be captured and logged)
Architecture Insight

SessionStart input is intentionally minimal - it contains no prompt or tool context:

typescript
interface SessionStartInput {
  session_id: string;
  type: "start" | "resume" | "compact" | "clear";
  cwd: string;
  env: { [key: string]: string };
  // NO: prompt, tool_context, task_span_id, parent_span_id
}

This design boundary prevents real-time correlation at hook time.

Recommendation

For building agent run datasets with sub-agent correlation:

  1. In-session logging: Capture SubagentStop session_id in logs or state
  2. Post-session export: Query Braintrust API for parent and sub-agent traces
  3. Offline correlation: Match traces by timing and metadata in a script
  4. Don't try real-time linking: Hooks don't have necessary context

Example script pattern:

bash
# 1. Export parent session
braintrust_analyze.py --replay <parent-session-id> > parent_traces.json

# 2. Query for orphaned sub-agent traces (those created during parent's time window)
braintrust_analyze.py --agent-stats > all_agent_traces.json

# 3. Correlate in Python:
#    - Parent Task spans -> agentId, timestamps, subagent_type
#    - Orphaned traces -> root_span_id, timestamps
#    - Match by timing and type

This approach is reliable, testable, and doesn't require hooks to maintain implicit state.

© parcadei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/braintrust-tracing of parcadei/Continuous-Claude-v3.

Open the folder on GitHubat commit d07ff4b

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in parcadei/Continuous-Claude-v3, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Braintrust Tracing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Braintrust Tracing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Braintrust Tracing this skillparcadei/Continuous-Claude-v33.9k1 repos~3.2kAutomated safety check: PassMIT
Claude Code Agent Developmentanthropics/claude-plugins-official38k7 repos~2.8kAutomated safety check: PassApache-2.0
Claude Automation Recommenderanthropics/claude-plugins-official38k3 repos~2.7kAutomated safety check: NotesApache-2.0
Analyze Trajectoryyologdev/yoyo-evolve1.9k—~3.6kAutomated safety check: PassMIT
Token-Saver Configurationppgranger/token-saver153—~1kAutomated safety check: PassApache-2.0
Weaveentropyvortex/meta-llm-charter249—~2.9kAutomated safety check: PassMIT

Similar skills

  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Claude Automation Recommender

    anthropics/claude-plugins-official

    Official

    Scans a codebase and suggests which Claude Code hooks, subagents, skills, plugins and MCP servers fit its stack, without changing any files.

    38k GitHub starsUsed in 3 repos~2.7k tokens
    Agent WorkflowsAuto-check: notes
  • Analyze Trajectory

    yologdev/yoyo-evolve

    Diagnoses a recurring failure such as a stuck task, repeated CI error or frequent reverts by sending sub-agents through the logs and returning one root-cause diagnosis.

    1.9k GitHub stars~3.6k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Token-Saver Configuration

    ppgranger/token-saver

    Checks and tunes token-saver output compression: reads stats, explains why a command was or was not compressed, and edits config files or environment variables.

    153 GitHub stars~1k tokensUpdated 19 days ago
    Agent WorkflowsAuto-check passed
  • Weave

    entropyvortex/meta-llm-charter

    Parallel strand orchestration — decompose a task into 3+ independently scoped strands, fan out real subagents (worktree-isolated when they write files), and coordinate through a session file and…

    249 GitHub stars~2.9k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Verify Setup Health Check

    diet103/claude-code-infrastructure-showcase

    Run the infrastructure health check and fix anything that fails

    10k GitHub stars~363 tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed

More from parcadei/Continuous-Claude-v3

All 141 skills in this repo
  • Compound Learnings

    parcadei/Continuous-Claude-v3

    Transform session learnings into permanent capabilities (skills, rules, agents).

    3.9k GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check: notes
  • Debug Hooks

    parcadei/Continuous-Claude-v3

    Systematic hook debugging workflow. An agent skill from parcadei/Continuous-Claude-v3.

    3.9k GitHub starsUsed in 1 repo~863 tokens
    Auto-check: notes
  • Tldr Deep

    parcadei/Continuous-Claude-v3

    Full 5-layer analysis of a specific function. An agent skill from parcadei/Continuous-Claude-v3.

    3.9k GitHub starsUsed in 1 repo~677 tokens
    Auto-check passed
  • Gradient Methods

    parcadei/Continuous-Claude-v3

    Problem-solving strategies for gradient methods in optimization

    3.9k GitHub starsUsed in 2 repos~1k tokens
    Auto-check: notes
  • Math

    parcadei/Continuous-Claude-v3

    Unified math capabilities - computation, solving, and explanation.

    3.9k GitHub starsUsed in 2 repos~1.6k tokens
    Auto-check: notes
  • Math Model Selector

    parcadei/Continuous-Claude-v3

    Routes problems to appropriate mathematical frameworks using expert heuristics

    3.9k GitHub starsUsed in 2 repos~841 tokens
    Auto-check passed

Categories

Questions about Braintrust Tracing

What does Braintrust Tracing do?

Braintrust tracing for Claude Code - hook architecture, sub-agent correlation, debugging. Braintrust Tracing is an agent skill from parcadei/Continuous-Claude-v3.

When should I use Braintrust Tracing?

Braintrust Tracing fits situations like: tasks that involve Subagents; tasks that involve Hooks and plugins; tasks that involve Debugging.

How do I install Braintrust Tracing in Claude Code?

Run `npx skills add parcadei/Continuous-Claude-v3 --skill braintrust-tracing -a claude-code`. Or copy the skill folder (.claude/skills/braintrust-tracing in parcadei/Continuous-Claude-v3) into .claude/skills/braintrust-tracing in your project. Claude Code loads it when a task matches its description.

How do I install Braintrust Tracing in Codex?

Run `npx skills add parcadei/Continuous-Claude-v3 --skill braintrust-tracing -a codex`. Or copy the skill folder (.claude/skills/braintrust-tracing in parcadei/Continuous-Claude-v3) into .agents/skills/braintrust-tracing in your project. Codex loads it when a task matches its description.

Can I use Braintrust Tracing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add parcadei/Continuous-Claude-v3 --skill braintrust-tracing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/braintrust-tracing, .gemini/skills/braintrust-tracing, .github/skills/braintrust-tracing and .opencode/skills/braintrust-tracing in your project.

What does Braintrust Tracing need to run?

Going by SKILL.md and its folder, Braintrust Tracing needs the command-line tools its instructions call (uv, bash and jq) and credentials named BRAINTRUST_API_KEY. Our summary lists: Python 3; A credential in BRAINTRUST_API_KEY.

Does Braintrust Tracing access the network?

SKILL.md names 1 domain. In commands or code: api.braintrust.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Braintrust Tracing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Braintrust Tracing use?

Braintrust Tracing is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Braintrust Tracing use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Braintrust Tracing?

Skills that share tags, products or a category with Braintrust Tracing: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Claude Automation Recommender (anthropics/claude-plugins-official, 38k stars), Analyze Trajectory (yologdev/yoyo-evolve, 1.9k stars) and Token-Saver Configuration (ppgranger/token-saver, 153 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Braintrust Tracing?

parcadei (a GitHub user) maintains it in parcadei/Continuous-Claude-v3, which has 3,943 GitHub stars. The repository holds 141 skills in this directory. The repository was last updated on January 26, 2026.

Source: parcadei/Continuous-Claude-v3 on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.