Agent skill

Evidence-Driven Trace

by Yeachan-Heo in Yeachan-Heo/oh-my-claudecode

Explains why something happened by generating competing hypotheses, gathering evidence in parallel, ranking explanations and proposing the next discriminating probe.

MITAuto-check passedDevelopment

Install Evidence-Driven Trace

skills CLI
$ npx skills add Yeachan-Heo/oh-my-claudecode --skill trace -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Yeachan-Heo/oh-my-claudecode trace --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Yeachan-Heo/oh-my-claudecode.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/trace .claude/skills/trace && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
trace
GitHub stars
40k
Token cost
~2.6k tokens
SKILL.md length
1,322 words
Files
1
Skills in repo
47
Repo updated
First seen
Licence
MIT

At a glance

Explains why something happened by generating competing hypotheses, gathering evidence in parallel, ranking explanations and proposing the next discriminating probe.

  • Works in 7 steps: Observation -- what was actually observed → Hypotheses -- competing explanations → Evidence For -- what supports each… → …
  • Working out why a regression or intermittent bug appeared before touching code
  • SKILL.md covers Good entry cases, Core tracing contract, Evidence strength hierarchy and Strong falsification /…, plus 10 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

This is an orchestration layer over a built-in tracer agent, meant for ambiguous, causal, evidence-heavy questions such as runtime regressions, latency, post-mortems or why a configuration behaved a certain way. It does not jump into fixing code. It restates the observation, builds competing explanations, collects evidence in parallel in Claude's built-in team mode, ranks the explanations and names the next probe that would settle the question fastest.

Every report keeps seven parts apart: the observation, hypotheses, evidence for, evidence against and gaps, the current best explanation, the critical unknown and a discriminating probe. Evidence is ranked from controlled reproductions down to intuition, weaker tiers are down-ranked when stronger evidence contradicts them, and each run must try to falsify its own favorite explanation. The skill warns against turning into a fix-it loop, a debugger summary or false certainty.

When your agent uses it

  • Working out why a regression or intermittent bug appeared before touching code
  • Investigating a latency or resource anomaly with several plausible causes
  • Writing a postmortem that weighs competing explanations

Example prompts

  • “Trace why the nightly export got three times slower after Tuesday's deploy.”
  • “Use trace on this stack output and rank the likely causes with evidence for and against.”
  • “Why does the router send these requests to the fallback model? Give hypotheses and the next probe.”

Requirements

  • Claude's built-in team mode for parallel tracer agents

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Observation -- what was actually observed
  2. Hypotheses -- competing explanations
  3. Evidence For -- what supports each explanation
  4. Evidence Against / Gaps -- what contradicts it or is still missing
  5. Current Best Explanation -- the leading explanation right now
  6. Critical Unknown -- the missing fact keeping the top explanations apart
  7. Discriminating Probe -- the highest-value next step to collapse uncertainty

What it can do on your machine

Read from SKILL.md and the folder at commit 454bae0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evidence-Driven Trace loads about 2.6k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 1,322 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~27
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Yeachan-Heo/oh-my-claudecode at commit 454bae0, republished under its MIT licence (© Yeachan-Heo). 1,322 words, ~2,585 tokens.

Download SKILL.mdSave it as .claude/skills/trace/SKILL.md (or your agent's skills folder).
name
trace
description
Evidence-driven tracing lane that orchestrates competing tracer hypotheses in Claude built-in team mode
argument-hint
<observation to trace>
agent
tracer
level
2

Trace Skill

Use this skill for ambiguous, causal, evidence-heavy questions where the goal is to explain why an observed result happened, not to jump directly into fixing or rewriting code.

This is the orchestration layer on top of the built-in tracer agent. The goal is to make tracing feel like a reusable OMC operating lane: restate the observation, generate competing explanations, gather evidence in parallel, rank the explanations, and propose the next probe that would collapse uncertainty fastest.

Good entry cases

Use /oh-my-claudecode:trace when the problem is:

  • ambiguous
  • causal
  • evidence-heavy
  • best answered by exploring competing explanations in parallel

Examples:

  • runtime bugs and regressions
  • performance / latency / resource behavior
  • architecture / premortem / postmortem analysis
  • scientific or experimental result tracing
  • config / routing / orchestration behavior explanation
  • “given this output, trace back the likely causes”

Core tracing contract

Always preserve these distinctions:

  1. Observation -- what was actually observed
  2. Hypotheses -- competing explanations
  3. Evidence For -- what supports each explanation
  4. Evidence Against / Gaps -- what contradicts it or is still missing
  5. Current Best Explanation -- the leading explanation right now
  6. Critical Unknown -- the missing fact keeping the top explanations apart
  7. Discriminating Probe -- the highest-value next step to collapse uncertainty

Do not collapse into:

  • a generic fix-it coding loop
  • a generic debugger summary
  • a raw dump of worker output
  • fake certainty when evidence is incomplete

Evidence strength hierarchy

Treat evidence as ranked, not flat.

From strongest to weakest:

  1. Controlled reproductions / direct experiments / uniquely discriminating artifacts
  2. Primary source artifacts with tight provenance (trace events, logs, metrics, benchmark outputs, configs, git history, file:line behavior)
  3. Multiple independent sources converging on the same explanation
  4. Single-source code-path or behavioral inference
  5. Weak circumstantial clues (timing, naming, stack order, resemblance to prior bugs)
  6. Intuition / analogy / speculation

Explicitly down-rank hypotheses that depend mostly on lower tiers when stronger contradictory evidence exists.

Strong falsification / disconfirmation rules

Every serious /trace run must try to falsify its own favorite explanation.

For each top hypothesis:

  • collect evidence for it
  • collect evidence against it
  • state what distinctive prediction it makes
  • state what observation would be hard to reconcile with it
  • identify the cheapest probe that would discriminate it from the next-best alternative

Down-rank a hypothesis when:

  • direct evidence contradicts it
  • it survives only by adding new unverified assumptions
  • it makes no distinctive prediction compared with rivals
  • a stronger alternative explains the same facts with fewer assumptions
  • its support is mostly circumstantial while the rival has stronger evidence tiers

Team-mode orchestration shape

Use Claude built-in team mode for /trace.

The lead should:

  1. Restate the observed result or “why” question precisely
  2. Extract the tracing target
  3. Generate multiple deliberately different candidate hypotheses
  4. Spawn 3 tracer lanes by default in team mode
  5. Assign one tracer worker per lane
  6. Instruct each tracer worker to gather evidence for and against its lane
  7. Run a rebuttal round between the leading hypothesis and the strongest remaining alternative
  8. Detect whether the top lanes genuinely differ or actually converge on the same root cause
  9. Merge findings into a ranked synthesis with an explicit critical unknown and discriminating probe

Important: workers should pursue deliberately different explanations, not the same explanation in parallel.

Default hypothesis lanes for v1

Unless the prompt strongly suggests a better partition, use these 3 default lanes:

  1. Code-path / implementation cause
  2. Config / environment / orchestration cause
  3. Measurement / artifact / assumption mismatch cause — covers verification-method defects, not just system defects. Examples: the verification query reuses a single dimensional key across distinct entities, tenants, streams, or groups; the comparison filter shape does not match the schema grain; or the catalog or column name was assumed portable across runtimes without enumeration. This includes multi-entity premise/key-assumption mismatches.

For lane 3, cross-entity discrepancies need a premise audit before escalation: enumerate entity dimensions and check whether a zero-row or mismatch result came from applying one key across multiple entities rather than from a system defect; the result may be a verification-methodology defect.

These defaults are intentionally broad so the first slice works across bug, performance, architecture, and experiment tracing.

Mandatory cross-check lenses

After the initial evidence pass, pressure-test the leaders with these lenses when relevant:

  • Systems lens -- queues, retries, backpressure, feedback loops, upstream/downstream dependencies, boundary failures, coordination effects
  • Premortem lens -- assume the current best explanation is incomplete or wrong; what failure mode would embarrass the trace later?
  • Science lens -- controls, confounders, measurement bias, alternative variables, falsifiable predictions

These lenses are not filler. Use them when they can surface a missed explanation, hidden dependency, or weak inference.

Show full SKILL.md (574 more words)Show less

Worker contract

Each worker should be a tracer lane owner, not a generic executor.

Each worker must:

  • own exactly one hypothesis lane
  • restate its lane hypothesis explicitly
  • gather evidence for the lane
  • gather evidence against the lane
  • rank the evidence strength behind its case
  • call out missing evidence, failed predictions, and remaining uncertainty
  • name the critical unknown for the lane
  • recommend the best lane-specific discriminating probe
  • avoid collapsing into implementation unless explicitly told to do so

Useful evidence sources include:

  • relevant code, tests, configs, docs, logs, outputs, and benchmark artifacts
  • existing trace artifacts via trace_timeline
  • existing aggregate trace evidence via trace_summary

Recommended worker return structure:

  1. Lane
  2. Hypothesis
  3. Evidence For
  4. Evidence Against / Gaps
  5. Evidence Strength
  6. Critical Unknown
  7. Best Discriminating Probe
  8. Confidence

Leader synthesis contract

The final /trace answer should synthesize, not just concatenate.

Return:

  1. Observed Result
  2. Ranked Hypotheses
  3. Evidence Summary by Hypothesis
  4. Evidence Against / Missing Evidence
  5. Rebuttal Round
  6. Convergence / Separation Notes
  7. Most Likely Explanation
  8. Critical Unknown
  9. Recommended Discriminating Probe
  10. Additional Trace Lanes (optional, only if uncertainty remains high)

Preserve a ranked shortlist even if one explanation is currently dominant.

Rebuttal round and convergence detection

Before closing the trace:

  • let the strongest non-leading lane present its best rebuttal to the current leader
  • force the leader to answer the rebuttal with evidence, not assertion
  • if the rebuttal materially weakens the leader, re-rank the table
  • if two “different” hypotheses reduce to the same underlying mechanism, merge them and say so explicitly
  • if two hypotheses still imply different next probes, keep them separate even if they sound similar

Do not claim convergence just because multiple workers use similar language. Convergence requires either:

  • the same root causal mechanism, or
  • independent evidence streams pointing to the same explanation

Explicit down-ranking guidance

The lead should explicitly say why a hypothesis moved down:

  • contradicted by stronger evidence
  • lacks the observation it predicted
  • requires extra ad hoc assumptions
  • explains fewer facts than the leader
  • lost the rebuttal round
  • converged into a stronger parent explanation

This is important because /trace should teach the reader why one explanation outranks another, not just present a final table.

Suggested lead prompt skeleton

Use a team-oriented orchestration prompt along these lines:

  1. “Restate the observation exactly.”
  2. “Generate 3 deliberately different hypotheses.”
  3. “Create one tracer lane per hypothesis using Claude built-in team mode.”
  4. “For each lane, gather evidence for and against, rank evidence strength, and name the critical unknown plus best discriminating probe.”
  5. “Apply systems, premortem, and science lenses to the leaders if useful.”
  6. “Run a rebuttal round between the top two explanations.”
  7. “Return a ranked explanation table, convergence notes, the critical unknown, and the single best discriminating probe.”

Output quality bar

Good /trace output is:

  • evidence-backed
  • concise but rigorous
  • skeptical of premature certainty
  • explicit about missing evidence
  • practical about the next action
  • explicit about why weaker explanations were down-ranked

Example final synthesis shape

Observed Result

[What happened]

Ranked Hypotheses
RankHypothesisConfidenceEvidence StrengthWhy it leads
1...High / Medium / LowStrong / Moderate / Weak...
Evidence Summary by Hypothesis
  • Hypothesis 1: ...
  • Hypothesis 2: ...
  • Hypothesis 3: ...
Evidence Against / Missing Evidence
  • Hypothesis 1: ...
  • Hypothesis 2: ...
  • Hypothesis 3: ...
Rebuttal Round
  • Best rebuttal to leader: ...
  • Why leader held / failed: ...
Convergence / Separation Notes
  • ...
Most Likely Explanation

[Current best explanation]

Critical Unknown

[Single missing fact keeping uncertainty open]

[Single next probe]

Additional Trace Lanes

[Only if uncertainty remains high]

© Yeachan-Heo, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/trace of Yeachan-Heo/oh-my-claudecode.

Open the folder on GitHubat commit 454bae0

Compare with similar skills

Evidence-Driven Trace next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evidence-Driven Trace compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evidence-Driven Trace this skillYeachan-Heo/oh-my-claudecode40k—~2.6kAutomated safety check: PassMIT
Octocode Code Researchbgauryy/octocode949—~1.5kAutomated safety check: PassMIT
Muse Code Product Doctorasgeirtj/system_prompts_leaks69k—~3.5kAutomated safety check: PassCC0-1.0
Root Cause Investigationgarrytan/gstack136k—~12kAutomated safety check: NotesMIT
Targeted Emergency Bug FixVeryGoodOpenSource/vgv-wingspan109—~1.9kAutomated safety check: PassMIT
Flowstudio Power Automate Debuggithub/awesome-copilot40k2 repos~5kAutomated safety check: PassMIT

Similar skills

  • Octocode Code Research

    bgauryy/octocode

    Researches code with evidence: traces callers, imports and cross-repo links, diagnoses failures and reports findings with exact file and line references and a confidence label.

    949 GitHub stars~1.5k tokensUpdated today
    DevelopmentAuto-check passed
  • Muse Code Product Doctor

    asgeirtj/system_prompts_leaks

    Diagnoses a Muse Code installation's own failures from binary and session evidence, instead of treating the report as an ordinary repository bug.

    69k GitHub stars~3.5k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Debugs in four phases (investigate, analyze, hypothesize, implement) under one rule: no fix is made until the root cause is found.

    136k GitHub stars~12k tokensUpdated today
    DevelopmentAuto-check: notes
  • Targeted Emergency Bug Fix

    VeryGoodOpenSource/vgv-wingspan

    Applies a minimal fix to an emergency bug through triage, root-cause location, a hotfix branch and a blast-radius check, with tests and review still required.

    109 GitHub stars~1.9k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Flowstudio Power Automate Debug

    github/awesome-copilot

    Official

    Debug failing Power Automate cloud flows using the FlowStudio MCP server.

    40k GitHub starsUsed in 2 repos~5k tokens
    DevelopmentAuto-check passed
  • Vibe Reflect And Compound

    ash1794/vibe-engineering

    Captures reusable knowledge after work is done — learnings from feedback or failures, resolved bugs (symptom, root cause, prevention), and recurring code patterns — into the project's persistent…

    163 GitHub stars~864 tokensUpdated 2 days ago
    DevelopmentAuto-check passed

More from Yeachan-Heo/oh-my-claudecode

All 47 skills in this repo
  • Ask Advisor Routing

    Yeachan-Heo/oh-my-claudecode

    Sends a question or task to another locally installed agent CLI, such as Codex or Gemini, through omc ask and saves the answer as a file.

    40k GitHub stars~572 tokensUpdated yesterday
    Auto-check passed
  • Ask Navigator

    Yeachan-Heo/oh-my-claudecode

    Charts a foggy effort into a map of decision tickets on the repo's issue tracker and works through them one per session, producing decisions rather than deliverables.

    40k GitHub stars~4.1k tokensUpdated yesterday
    Auto-check passed
  • Self-Improve Evolutionary Loop

    Yeachan-Heo/oh-my-claudecode

    Runs an autonomous improvement loop on a repository: agents propose and execute plans, a tournament picks the winner by benchmark, and each round is recorded and plotted.

    40k GitHub stars~5.3k tokensUpdated yesterday
    Auto-check: warnings
  • Autopilot

    Yeachan-Heo/oh-my-claudecode

    Takes a short product idea through requirements, design, planning, parallel implementation, QA cycles and multi-reviewer validation to produce working code.

    40k GitHub stars~4.4k tokensUpdated yesterday
    Auto-check passed
  • OMC Mode Cancellation

    Yeachan-Heo/oh-my-claudecode

    Detects and gracefully cancels whichever OMC mode, autopilot, ralph, swarm, pipeline, or team, is currently active, then clears its state.

    40k GitHub stars~4.6k tokensUpdated yesterday
    Auto-check passed
  • Hierarchical AGENTS.md Generator

    Yeachan-Heo/oh-my-claudecode

    Maps a codebase directory by directory and writes linked AGENTS.md files, each pointing to its parent, to document what each area contains.

    40k GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed

Questions about Evidence-Driven Trace

What does Evidence-Driven Trace do?

Explains why something happened by generating competing hypotheses, gathering evidence in parallel, ranking explanations and proposing the next discriminating probe. This is an orchestration layer over a built-in tracer agent, meant for ambiguous, causal, evidence-heavy questions such as runtime regressions, latency, post-mortems or why a configuration behaved a certain way. It does not jump into fixing code.

When should I use Evidence-Driven Trace?

Evidence-Driven Trace fits situations like: working out why a regression or intermittent bug appeared before touching code; investigating a latency or resource anomaly with several plausible causes; writing a postmortem that weighs competing explanations.

How do I install Evidence-Driven Trace in Claude Code?

Run `npx skills add Yeachan-Heo/oh-my-claudecode --skill trace -a claude-code`. Or copy the skill folder (skills/trace in Yeachan-Heo/oh-my-claudecode) into .claude/skills/trace in your project. Claude Code loads it when a task matches its description.

How do I install Evidence-Driven Trace in Codex?

Run `npx skills add Yeachan-Heo/oh-my-claudecode --skill trace -a codex`. Or copy the skill folder (skills/trace in Yeachan-Heo/oh-my-claudecode) into .agents/skills/trace in your project. Codex loads it when a task matches its description.

Can I use Evidence-Driven Trace in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Yeachan-Heo/oh-my-claudecode --skill trace -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/trace, .gemini/skills/trace, .github/skills/trace and .opencode/skills/trace in your project.

What does Evidence-Driven Trace need to run?

SKILL.md names no scripts, command-line tools or credentials: Evidence-Driven Trace is instructions for the agent only. Our summary lists: Claude's built-in team mode for parallel tracer agents.

Does Evidence-Driven Trace access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evidence-Driven Trace safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evidence-Driven Trace use?

Evidence-Driven Trace is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evidence-Driven Trace use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evidence-Driven Trace?

Skills that share tags, products or a category with Evidence-Driven Trace: Octocode Code Research (bgauryy/octocode, 949 stars), Muse Code Product Doctor (asgeirtj/system_prompts_leaks, 69k stars), Root Cause Investigation (garrytan/gstack, 136k stars) and Targeted Emergency Bug Fix (VeryGoodOpenSource/vgv-wingspan, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evidence-Driven Trace?

Yeachan-Heo (a GitHub user) maintains it in Yeachan-Heo/oh-my-claudecode, which has 39,720 GitHub stars. The repository holds 47 skills in this directory. The repository was last updated on October 8, 2026.

Source: Yeachan-Heo/oh-my-claudecode on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.