Agent skill

Investigate

by tobihagemann in tobihagemann/turbo

Systematically investigate bugs, test failures, build errors, performance issues, or unexpected behavior by cycling through characterize-isolate-hypothesize-test steps.

MITAuto-check passedDevelopment

Install Investigate

skills CLI
$ npx skills add tobihagemann/turbo --skill investigate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tobihagemann/turbo investigate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/codex/skills/investigate .claude/skills/investigate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
investigate
GitHub stars
408
Token cost
~3.2k tokens
SKILL.md length
1,613 words
Files
2 (incl. references)
Skills in repo
81
Repo updated
First seen
Licence
MIT

At a glance

Systematically investigate bugs, test failures, build errors, performance issues, or unexpected behavior by cycling through characterize-isolate-hypothesize-test steps.

  • Works in 4 steps: Characterize → Isolate → Hypothesize → …
  • The user asks to investigate this bug
  • SKILL.md covers Step 1: Characterize, Step 2: Isolate, Step 3: Hypothesize and Step 4: Test, plus 3 more sections
  • Calls git, npm and pip3

What it does

Investigate is an agent skill from tobihagemann/turbo. Systematically investigate bugs, test failures, build errors, performance issues, or unexpected behavior by cycling through characterize-isolate-hypothesize-test steps. Use when the user asks to "investigate this bug", "debug this", "figure out why this fails", "find the root cause", "why is this broken", "troubleshoot this", "diagnose the issue", "what's causing this error", "look into this failure", "why is this test failing", or "track down this bug".

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/problem-type-playbooks.md`).

It sits in Development, covering Failing and flaky tests and Root cause analysis. The repository describes itself as: Reusable workflows for planning, building, reviewing, and shipping with Claude Code and Codex. The licence is MIT.

When your agent uses it

  • The user asks to investigate this bug
  • Figure out why this fails
  • Find the root cause
  • Why is this broken

Example prompts

  • “investigate this bug”
  • “debug this”
  • “figure out why this fails”
  • “/investigate”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Characterize
  2. Isolate
  3. Hypothesize
  4. Test

What it can do on your machine

Read from SKILL.md and the folder at commit 931eda5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • npm
    • pip3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, npm and pip3, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Investigate loads about 3.2k tokens when it runs, and up to ~4.1k if it reads all its reference files. Until then it costs about 118 tokens; SKILL.md has 1,613 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~4.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tobihagemann/turbo at commit 931eda5, republished under its MIT licence (© tobihagemann). 1,613 words, ~3,166 tokens.

Download SKILL.mdSave it as .claude/skills/investigate/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
investigate
description
Systematically investigate bugs, test failures, build errors, performance issues, or unexpected behavior by cycling through characterize-isolate-hypothesize-test steps. Use when the user asks to "investigate this bug", "debug this", "figure out why this fails", "find the root cause", "why is this broken", "troubleshoot this", "diagnose the issue", "what's causing this error", "look into this failure", "why is this test failing", or "track down this bug".

Investigate

Systematic methodology for finding the root cause of bugs, failures, and unexpected behavior. Cycle through characterize-isolate-hypothesize-test steps, with oracle escalation for hard problems. Diagnose the root cause — do not apply fixes.

Optional: $ARGUMENTS contains the problem description or error message.

Step 1: Characterize

Gather the symptom and establish what is actually happening:

  1. Collect evidence — error message, stack trace, test output, log entries, or user description of unexpected behavior
  2. Classify the problem type:
SignalType
Stack trace / exceptionRuntime error
Test assertion failureTest failure
Compilation / bundler / build errorBuild failure
Type checker error (tsc, mypy, pyright)Type error
Slow response / high CPU / memory growthPerformance
"It does X instead of Y" / no errorUnexpected behavior
  1. Establish reproduction — run the failing command, test, or operation. If the problem cannot be reproduced (intermittent, environment-specific), document the constraints and proceed with historical evidence.

Record the exact reproduction command and its output for verification. For intermittent or long-running reproductions, tail logs in a background shell, filtered for relevant signals (errors, stack traces, specific identifiers) so failures surface live while you work.

Step 2: Isolate

Narrow from "something is wrong" to "the problem is in this area." Read references/problem-type-playbooks.md for type-specific first moves and tool sequences.

Git Archeology

For all problem types, check what changed recently near the failure point:

bash
git log --oneline -20 -- <file>
git blame -L <start>,<end> <file>

If a known-good state exists (e.g., "this worked yesterday"), consider git bisect to pinpoint the breaking commit.

When the problem description names a version, tag, or build that differs from the ref under investigation, resolve it to a ref and read what changed on the failing path between the two before generating hypotheses:

bash
git diff <reported-ref>..HEAD -- <failing-path-files>

Read the full diff rather than its --stat summary. Carry each difference that could produce the symptom forward as a ranked hypothesis.

When the failure surfaces inside a third-party dependency, search its issue tracker for a distinctive string from the error before reading deeper into the dependency's code. An issue whose symptom matches often names the cause and the fix outright. Carry a match forward as a ranked hypothesis and test it.

Scope Narrowing
  • Stack traces: Read the throwing function and its callers — full functions, not just the flagged line
  • Test failures: Read both the test and the system under test
  • Build errors: Read the config file and the referenced source
  • Unexpected behavior: Trace the data flow from input to the unexpected output

Before treating a record, file, or build artifact as evidence of the system's behavior, confirm the system under test produced it: check creator, source metadata, or generation time. Suspect imported, seeded, hand-edited, and leftover data from an earlier run, which reads identically to generated output. A checkout of another repository is the same trap: confirm it is current before reading it as evidence, since a stale one reads identically to the authoritative source.

Step 3: Hypothesize

Before forming a hypothesis about the machinery around a failure, such as a toolchain version, a configuration policy, or an environment difference, read the failing line, identify every path, package, symbol, or resource it names, and confirm each one resolves. Error text often names the site that consumed a missing input rather than the input itself, so the surrounding machinery looks responsible when it is not. Rank a machinery hypothesis only after every named reference checks out.

Once every named reference resolves and the operation has never once succeeded, rank a refusal ahead of any race or resource-exhaustion hypothesis: a denied permission, a firewall rule, an allowlist, an expired or missing credential. Intermittent failure is what a race or a contended resource usually looks like, so a run of attempts with zero successes ranks both below a refusal. Carry the hard blocks on the failing path into Step 4 as the first hypotheses to test, ahead of any measurement or instrumentation.

Generate 2-4 hypotheses ranked by likelihood. Each hypothesis must be falsifiable — specify what evidence would confirm or refute it.

Format:

H1 (most likely): [description] — confirmed if [X], refuted if [Y]
H2: [description] — confirmed if [X], refuted if [Y]
H3: [description] — confirmed if [X], refuted if [Y]

Check that the observed case can discriminate: when confirming and refuting evidence would look identical in it, the case is degenerate and any verdict drawn from it is inconclusive. Degenerate cases hide the difference they are supposed to reveal, such as a scaling factor of 1, a single-element collection, or an identity transform. Find a non-degenerate case, or construct one as a Step 4 experiment.

Parallel Investigation

For complex problems with 3+ hypotheses and a non-obvious root cause, spawn parallel investigators simultaneously.

Spawn condition: 3+ hypotheses AND the problem is not a simple typo, missing import, or syntax error.

Skip when 1-2 hypotheses are obvious (e.g., stack trace points directly to the bug).

Before dispatching, read the project's test configuration and CI workflow to identify any test tier that resets a shared external resource between tests, such as a database, a fixed port, or a cache. Such tiers have no cross-process interlock, so branches running them concurrently wipe each other's state and return failures that look like real defects. Name any such tier to every branch as off-limits.

When the evidence lives in a repository other than this one, including a submodule or a vendored clone with its own remote, establish the authoritative ref before dispatching and bring it up to date, fetching it or reading it through the forge API. A local checkout may be behind its remote, and an investigation branch reading a stale one returns findings the current code has already resolved. Name that ref and how to read it in every branch prompt, including the text the Claude consultation branch forwards.

Launch all investigation branches with spawn_agent / wait_agent using inherited model defaults, issuing every call in one batch. Do not issue one and await its result before issuing the rest. Expect one branch per hypothesis plus one Claude consultation branch. Every branch prompt must direct it to treat the shared working tree and its git index as read-only and to gather evidence by reading and reasoning; experiments that mutate code wait for Step 4, where they run one at a time. HEAD stays where it is: read other refs with git show <ref>:<path> rather than git checkout or git switch.

  • Hypothesis branch (one per hypothesis): Each receives the hypothesis, relevant file paths, what evidence to look for, and instructions to report confirmed / refuted / inconclusive with evidence. Budget: max 5 tool calls per branch.
  • Claude consultation branch: Run the $consult-claude skill with a focused prompt describing the problem, reproduction, and files examined. The external perspective can dig into patterns the hypothesis-driven branches miss. Run the $evaluate-findings skill on its output after the consultation returns.

Once wait_agent has returned every hypothesis branch's report and only the Claude consultation branch has not reported, run the Step 4 actions that change nothing in the working tree, its git index, or anything the consultation's prompt points it at, and state each result as it lands. Then call wait_agent until the consultation's report arrives. The merge, every other Step 4 action, the Iteration decision, and the Investigation Report wait for that report.

After all investigators complete, merge results. Claude findings that overlap with a confirmed hypothesis reinforce confidence. Novel Claude findings become additional hypotheses to test in Step 4.

Show full SKILL.md (421 more words)Show less

Step 4: Test

Verify each hypothesis with minimal, targeted actions:

Action TypeTool
Find usage or patternGrep
Read surrounding codeRead
Check recent changesBash (git log, git blame, git diff)
Run isolated testBash (specific test command)
Check dependency versionBash (npm ls, pip3 show, etc.)
Inspect runtime stateBash (add temporary logging, run, check output)
Vary one suspected variableBash (construct a throwaway fixture, run, compare)

When read-only evidence cannot discriminate, construct minimal throwaway fixtures that vary one suspected variable at a time. Exercise the system's inputs, and leave the working tree and its git index unchanged. Label each fixture clearly, delete them once the experiment concludes, and report anything that could not be deleted. When a check edits a tracked file instead, such as adding temporary logging, remove the edit once the check concludes and confirm with git diff -- <file> that the file is back to its pre-check state before recording the result. Write to an external or live system only after explicit user approval via request_user_input, stating the target system, every record the write will touch including those reached through triggers, cascades, and hooks, and the cleanup plan. When request_user_input does not reach the user, write nothing and report the approval as unresolved.

Record each result:

HypothesisVerdictEvidence
H1confirmed / refuted / inconclusive[what was found]
H2confirmed / refuted / inconclusive[what was found]
Iteration

If all hypotheses are refuted or inconclusive:

  1. Document what was learned — each refuted hypothesis eliminates a possibility and narrows the search
  2. Return to Step 2 with the new information to re-isolate
  3. Generate new hypotheses in Step 3 based on updated understanding

Cycle budget: maximum 2 full cycles (hypothesize → test → learn → repeat) before escalating.

Escalation

After 2 failed hypothesis cycles, offer escalation to $consult-oracle via request_user_input:

Investigation stalled after [N] hypothesis cycles.

Tested: [summary of hypotheses and evidence]
Remaining unknowns: [what is still unclear]

Escalate to Oracle? (consults external model with full context)

Proceed only if the user approves.

Investigation Report

Output results as text:

Investigation Report:

Problem: [one-line description]
Type: [runtime error | test failure | build failure | type error | performance | unexpected behavior]
Root cause: [confirmed cause, or "unresolved" with best hypothesis]

Evidence:
- [what confirmed the root cause]

Suggested fix: [description of what to change, or "needs further investigation"]
Reproduction command: [command to verify the fix once applied]

Hypotheses tested:
1. [hypothesis] — [confirmed/refuted/inconclusive] — [evidence]
2. [hypothesis] — [confirmed/refuted/inconclusive] — [evidence]

Escalation: [none | oracle]

Then call update_plan to mark this step completed and continue with the next step of the active workflow.

Rules

  • If the problem turns out to be environmental (wrong language runtime version, a declared dependency not installed locally, OS-specific), report that clearly — it may not require a code fix. A dependency the project never declared is a manifest defect, so report that as a code fix instead.
  • If the problem is in a dependency (not the project's code), document the dependency issue and suggest the options that leave the dependency's code unmodified: workarounds, and a report upstream unless an existing issue covers the defect or the dependency's latest release no longer has it.

© tobihagemann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in codex/skills/investigate of tobihagemann/turbo.

  • SKILL.md
  • references/problem-type-playbooks.md

Open the folder on GitHubat commit 931eda5

Compare with similar skills

Investigate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Investigate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Investigate this skilltobihagemann/turbo408—~3.2kAutomated safety check: PassMIT
Minimal Code Fixcobusgreyling/loop-engineering11k1 repos~671Automated safety check: NotesMIT
Backprop: Bug-to-Spec ProtocolJuliusBrussee/cavekit1.2k—~653Automated safety check: PassMIT
CI TriageMentra-Community/MentraOS2.4k—~582Automated safety check: PassApache-2.0
Root Cause Debuggingjsmastery-pro/skills1.5k—~1.8kAutomated safety check: NotesMIT
Superpowers Systematic Debuggingchristopherarter/superpowers-reasonix102—~2kAutomated safety check: PassMIT

Similar skills

  • Minimal Code Fix

    cobusgreyling/loop-engineering

    Makes the smallest code change that fixes one well-scoped problem, such as a CI failure, review comment or typo, without refactoring anything unrelated.

    11k GitHub starsUsed in 1 repo~671 tokens
    DevelopmentAuto-check: notes
  • Backprop: Bug-to-Spec Protocol

    JuliusBrussee/cavekit

    After a bug is found, traces its root cause and feeds a new testable invariant back into the project spec so the bug class can't recur.

    1.2k GitHub stars~653 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • CI Triage

    Mentra-Community/MentraOS

    Triage failing GitHub PR checks: list failures with gh, fetch capped Actions logs, skip non-Actions checks, and summarize root cause.

    2.4k GitHub stars~582 tokensUpdated today
    DevelopmentAuto-check passed
  • Root Cause Debugging

    jsmastery-pro/skills

    Runs a reproduce, localize, hypothesize, test, fix and verify loop to find a bug's root cause, applies the minimal fix and hands off a regression test.

    1.5k GitHub stars~1.8k tokensUpdated 2 mo ago
    DevelopmentAuto-check: notes
  • Superpowers Systematic Debugging

    christopherarter/superpowers-reasonix

    Any bug, failing or flaky test, or surprise behavior?. An agent skill from christopherarter/superpowers-reasonix.

    102 GitHub stars~2k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • CI Fix

    warpdotdev/oz-skills

    Diagnose and fix GitHub Actions CI failures. An agent skill from warpdotdev/oz-skills.

    824 GitHub stars~790 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from tobihagemann/turbo

All 81 skills in this repo
  • Consult Oracle

    tobihagemann/turbo

    Consult ChatGPT Pro via ChatGPT browser automation for problems that resist standard approaches.

    408 GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Fetch PR Comments

    tobihagemann/turbo

    Fetch and summarize review feedback and conversation from a GitHub PR (unresolved review threads, review bodies, and PR conversation comments) without making changes.

    408 GitHub stars~967 tokensUpdated 2 days ago
    Auto-check passed
  • Recall Rationale

    tobihagemann/turbo

    Recall why a past change was made by locating the Claude Code transcript that produced it.

    408 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Resolve PR Comments

    tobihagemann/turbo

    Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.

    408 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Resolve PR Comments

    tobihagemann/turbo

    Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.

    408 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Assess Technical Debt

    tobihagemann/turbo

    Assess project-wide structural technical debt: complexity hotspots, deprecated API usage, duplication clusters, architecture rot, and low-value tests.

    408 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed

Questions about Investigate

What does Investigate do?

Systematically investigate bugs, test failures, build errors, performance issues, or unexpected behavior by cycling through characterize-isolate-hypothesize-test steps. Investigate is an agent skill from tobihagemann/turbo. Systematically investigate bugs, test failures, build errors, performance issues, or unexpected behavior by cycling through characterize-isolate-hypothesize-test steps.

When should I use Investigate?

Investigate fits situations like: the user asks to investigate this bug; figure out why this fails; find the root cause; why is this broken.

How do I install Investigate in Claude Code?

Run `npx skills add tobihagemann/turbo --skill investigate -a claude-code`. Or copy the skill folder (codex/skills/investigate in tobihagemann/turbo) into .claude/skills/investigate in your project. Claude Code loads it when a task matches its description.

How do I install Investigate in Codex?

Run `npx skills add tobihagemann/turbo --skill investigate -a codex`. Or copy the skill folder (codex/skills/investigate in tobihagemann/turbo) into .agents/skills/investigate in your project. Codex loads it when a task matches its description.

Can I use Investigate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tobihagemann/turbo --skill investigate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/investigate, .gemini/skills/investigate, .github/skills/investigate and .opencode/skills/investigate in your project.

What does Investigate need to run?

Going by SKILL.md and its folder, Investigate needs the command-line tools its instructions call (git, npm and pip3).

Does Investigate access the network?

SKILL.md contains no URLs. Its commands use git and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Investigate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Investigate use?

Investigate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Investigate use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 906 tokens, read only when the agent opens those files.

What are the alternatives to Investigate?

Skills that share tags, products or a category with Investigate: Minimal Code Fix (cobusgreyling/loop-engineering, 11k stars), Backprop: Bug-to-Spec Protocol (JuliusBrussee/cavekit, 1.2k stars), CI Triage (Mentra-Community/MentraOS, 2.4k stars) and Root Cause Debugging (jsmastery-pro/skills, 1.5k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Investigate?

tobihagemann (a GitHub user) maintains it in tobihagemann/turbo, which has 408 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on October 9, 2026.

Source: tobihagemann/turbo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.