Agent skill

Hypothesis Testing

by rohitg00 in rohitg00/skillkit

Applies the scientific method to debugging by helping users form specific, testable hypotheses, design targeted experiments, and systematically confirm or reject theories to find root causes.

Apache-2.0Auto-check passedDevelopment

Install Hypothesis Testing

skills CLI
$ npx skills add rohitg00/skillkit --skill hypothesis-testing -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rohitg00/skillkit hypothesis-testing --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rohitg00/skillkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/packages/core/src/methodology/packs/debugging/hypothesis-testing .claude/skills/hypothesis-testing && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hypothesis-testing
GitHub stars
1.5k
Token cost
~1.5k tokens
SKILL.md length
232 words
Files
1
Skills in repo
14
Repo updated
First seen
Licence
Apache-2.0

At a glance

Applies the scientific method to debugging by helping users form specific, testable hypotheses, design targeted experiments, and systematically confirm or reject theories to find root causes.

  • Works in 5 steps: Observe - Gather Facts → Hypothesize - Form Testable Theories → Predict - Define Expected Results → …
  • A user says their code isnt working
  • SKILL.md covers Core Principle, The Scientific Debugging Method, Hypothesis Tracking Template and Testing Techniques by…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Hypothesis Testing is an agent skill from rohitg00/skillkit. Applies the scientific method to debugging by helping users form specific, testable hypotheses, design targeted experiments, and systematically confirm or reject theories to find root causes. Use when a user says their code isn't working, they're getting an error, something broke, they want to troubleshoot a bug, or they're trying to figure out what's causing an issue. Concrete actions include isolating failing components, forming and testing hypotheses, analyzing error messages, tracing execution paths, and…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Root cause analysis and Debugging. The repository describes itself as: Supercharge AI coding agents with portable skills. Install, translate & share skills across Claude Code, Cursor, Codex, Copilot & 40 more. The licence is Apache-2.0.

When your agent uses it

  • A user says their code isnt working
  • Theyre getting an error
  • Something broke
  • They want to troubleshoot a bug

Example prompts

  • “t working, they”
  • “re trying to figure out what”
  • “Use the hypothesis-testing skill to apply the scientific method to debugging by helping users form specific, testable hypotheses, design targeted…”
  • “/hypothesis-testing”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Observe - Gather Facts
  2. Hypothesize - Form Testable Theories
  3. Predict - Define Expected Results
  4. Test - Experiment Systematically
  5. Analyze - Interpret Results

What it can do on your machine

Read from SKILL.md and the folder at commit d2e5c34. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are typescript and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hypothesis Testing loads about 1.5k tokens when it runs. Until then it costs about 147 tokens; SKILL.md has 232 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~147
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rohitg00/skillkit at commit d2e5c34, republished under its Apache-2.0 licence (© rohitg00). 232 words, ~1,549 tokens.

Download SKILL.mdSave it as .claude/skills/hypothesis-testing/SKILL.md (or your agent's skills folder).
name
hypothesis-testing
description
Applies the scientific method to debugging by helping users form specific, testable hypotheses, design targeted experiments, and systematically confirm or reject theories to find root causes. Use when a user says their code isn't working, they're getting an error, something broke, they want to troubleshoot a bug, or they're trying to figure out what's causing an issue. Concrete actions include isolating failing components, forming and testing hypotheses, analyzing error messages, tracing execution paths, and interpreting test results to narrow down root causes.
version
1.0.0
triggers
hypothesis, theory about bug, might be caused by, test theory, prove theory
tags
debugging, scientific-method, investigation, validation
difficulty
intermediate
estimatedTime
15
relatedSkills
debugging/root-cause-analysis, debugging/trace-and-isolate

Hypothesis-Driven Debugging

You are applying the scientific method to debugging. Form clear hypotheses, design tests that can definitively confirm or reject them, and systematically narrow down to the truth.

Core Principle

Every debugging action should test a specific hypothesis. Random changes are not debugging.

The Scientific Debugging Method

1. Observe - Gather Facts

Before forming hypotheses, collect observations:

  • What exactly happens? (specific symptoms)
  • When does it happen? (timing, frequency)
  • Where does it happen? (environment, component)
  • What changed recently? (code, config, data)

Write down observations objectively:

Observations:
- API returns 500 error on POST /orders
- Happens only when cart has > 10 items
- Started after deployment on 2024-01-15
- Works fine in staging environment
- Error logs show "connection refused" to inventory service
2. Hypothesize - Form Testable Theories

Examples (bad → good):

  • "Something is wrong with the network" → "The inventory service connection pool is exhausted when processing orders with >10 items"
  • "There might be a race condition" → "The order processing timeout (5s) is insufficient for large orders"
3. Predict - Define Expected Results

For each hypothesis, define what you expect to observe if it is true versus false:

Hypothesis: Connection pool exhausted for large orders

If TRUE:
- Active connections should hit max (20) during large orders
- Small orders should still work during this time
- Increasing pool size should fix the issue

If FALSE:
- Connection count stays well below max
- Small orders also fail during the issue
- Pool size change has no effect
4. Test - Experiment Systematically

Design tests that definitively confirm or reject:

Test Plan for Connection Pool Hypothesis:

1. Add connection pool monitoring
   - Log active connections before/after each request
   - Expected if true: Count reaches 20 during failures

2. Artificial stress test
   - Send 5 large orders simultaneously
   - Expected if true: Failures start when pool exhausted

3. Increase pool size to 50
   - Repeat stress test
   - Expected if true: Failures stop or threshold moves

4. Control test with small orders
   - Send 20 small orders simultaneously
   - Expected if true: No failures (faster processing)
5. Analyze - Interpret Results

After testing:

  • Did results match predictions for TRUE or FALSE?
  • Are results conclusive or ambiguous?
  • Do results suggest a different hypothesis?
Results:
- Connection count reached 20/20 during failures ✓
- Small orders succeeded during same period ✓
- Pool size increase to 50 → failures stopped ✓

Conclusion: Hypothesis CONFIRMED
Connection pool exhaustion is the proximate cause.

New question: Why do large orders exhaust the pool?
New hypothesis: Large orders make multiple inventory calls per item

Hypothesis Tracking Template

markdown
## Bug: [Description]

### Hypothesis 1: [Theory]
**Status:** Testing | Confirmed | Rejected
**Probability:** High | Medium | Low

**Evidence For:**
- [Evidence 1]
- [Evidence 2]

**Evidence Against:**
- [Evidence 1]

**Test Plan:**
1. [Test 1] - Expected result if true
2. [Test 2] - Expected result if false

**Test Results:**
- [Result 1]: [Supports/Contradicts]
- [Result 2]: [Supports/Contradicts]

**Conclusion:** [Confirmed/Rejected] because [reasoning]

---

### Hypothesis 2: [Next Theory]
...

Testing Techniques by Hypothesis Type

Testing Timing Hypotheses
typescript
// Add timing instrumentation
const start = performance.now();
await suspectedSlowOperation();
const duration = performance.now() - start;
console.log(`Operation took ${duration}ms`);
// Hypothesis confirmed if duration > expected
Testing Data Hypotheses
typescript
// Validate data at key points
function processWithValidation(data) {
  console.assert(data.id != null, 'Missing id');
  console.assert(data.items?.length > 0, 'Empty items');
  console.assert(typeof data.total === 'number', 'Invalid total');
  // If assertions fail, data hypothesis likely true
}
Testing State Hypotheses
typescript
// Snapshot state before and after
const stateBefore = JSON.stringify(currentState);
suspectedStateMutation();
const stateAfter = JSON.stringify(currentState);
if (stateBefore !== stateAfter) {
  console.log('State changed:', diff(stateBefore, stateAfter));
}

Decision Tree

Is the hypothesis testable?
├── NO → Refine it to be more specific
└── YES → Can I test it without side effects?
    ├── NO → Design a safe test (staging, logs-only)
    └── YES → Run the test
        └── Results conclusive?
            ├── NO → Design a better test
            └── YES → Hypothesis confirmed or rejected?
                ├── CONFIRMED → Root cause found?
                │   ├── YES → Fix and verify
                │   └── NO → Form next hypothesis (why?)
                └── REJECTED → Form next hypothesis

Integration with Other Skills

  • root-cause-analysis: Hypothesis testing is a key technique within RCA
  • trace-and-isolate: Use tracing to gather evidence for hypotheses
  • testing/red-green-refactor: Write test that confirms the bug before fixing

© rohitg00, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in packages/core/src/methodology/packs/debugging/hypothesis-testing of rohitg00/skillkit.

Open the folder on GitHubat commit d2e5c34

Compare with similar skills

Hypothesis Testing next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hypothesis Testing compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hypothesis Testing this skillrohitg00/skillkit1.5k—~1.5kAutomated safety check: PassApache-2.0
OpenLogi macOS Permissions TriageAprilNEA/OpenLogi23k—~2.5kAutomated safety check: NotesApache-2.0
Bug Finder for daisyUIsaadeghi/daisyui43k—~2.3kAutomated safety check: PassMIT
Root Cause Debugginggarrytan/gstack136k—~1.4kAutomated safety check: PassMIT
Graph-Based Bug Tracingtirth8205/code-review-graph32k1 repos~287Automated safety check: PassMIT
Systematic DebuggingChrisWiles/claude-code-showcase6.1k3 repos~1.2kAutomated safety check: PassNone

Similar skills

  • Decides whether an OpenLogi device problem on macOS is a privacy-permission (TCC) problem, using agent log lines, and says which identity needs which grant.

    23k GitHub stars~2.5k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes
  • Bug Finder for daisyUI

    saadeghi/daisyui

    Investigates suspected bugs in the daisyUI monorepo through read-only analysis, then writes a decision-ready fix plan in tmp/bugs without changing any product code.

    43k GitHub stars~2.3k tokensUpdated 7 days ago
    DevelopmentAuto-check passed
  • Root Cause Debugging

    garrytan/gstack

    Investigates bugs, errors and stack traces in phases and requires a root-cause hypothesis to be confirmed before any fix is written.

    136k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Graph-Based Bug Tracing

    tirth8205/code-review-graph

    Traces a bug through a code knowledge graph, following callers, callees and execution flow before opening source files, within a small token budget.

    32k GitHub starsUsed in 1 repo~287 tokens
    DevelopmentAuto-check passed
  • Systematic Debugging

    ChrisWiles/claude-code-showcase

    Applies a four-phase debugging routine that finds the root cause of a bug or failing test before any fix is written.

    6.1k GitHub starsUsed in 3 repos~1.2k tokens
    DevelopmentAuto-check passed
  • Debugging and Error Recovery

    addyosmani/agent-skills

    Applies a stop-the-line rule and a step-by-step triage when tests fail, builds break or something stops working, aiming at the root cause instead of guesses.

    102k GitHub starsUsed in 1 repo~2.6k tokens
    DevelopmentAuto-check passed

More from rohitg00/skillkit

All 14 skills in this repo
  • Design First

    rohitg00/skillkit

    Guides the creation of technical design documents before writing code, producing architecture diagrams, data models, API interface definitions, implementation plans, and multi-option trade-off…

    1.5k GitHub stars~1.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Find Skills

    rohitg00/skillkit

    Discovers, searches, and installs skills from multiple AI agent skill marketplaces (400K+ skills) using the SkillKit CLI.

    1.5k GitHub stars~1.2k tokensUpdated 4 mo ago
    Auto-check passed
  • Handoff Protocols

    rohitg00/skillkit

    Manages work transitions between team members or agents by creating structured handoff documents, summarizing project status, documenting key decisions, blockers, and open questions, and generating…

    1.5k GitHub stars~1.5k tokensUpdated 4 mo ago
    Auto-check passed
  • Parallel Investigation

    rohitg00/skillkit

    Coordinates parallel investigation threads to simultaneously explore multiple hypotheses or root causes across different system areas.

    1.5k GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed
  • Red Green Refactor

    rohitg00/skillkit

    Guides the red-green-refactor TDD workflow: write a failing test first, implement the minimum code to make it pass, then refactor while keeping tests green.

    1.5k GitHub stars~1.3k tokensUpdated 4 mo ago
    Auto-check passed
  • Root Cause Analysis

    rohitg00/skillkit

    Performs systematic root cause analysis to identify the true source of bugs, errors, and unexpected behavior through structured investigation phases — not just treating symptoms.

    1.5k GitHub stars~1.4k tokensUpdated 4 mo ago
    Auto-check passed

Categories

Questions about Hypothesis Testing

What does Hypothesis Testing do?

Applies the scientific method to debugging by helping users form specific, testable hypotheses, design targeted experiments, and systematically confirm or reject theories to find root causes. Hypothesis Testing is an agent skill from rohitg00/skillkit. Applies the scientific method to debugging by helping users form specific, testable hypotheses, design targeted experiments, and systematically confirm or reject theories to find root causes.

When should I use Hypothesis Testing?

Hypothesis Testing fits situations like: A user says their code isnt working; theyre getting an error; something broke; they want to troubleshoot a bug.

How do I install Hypothesis Testing in Claude Code?

Run `npx skills add rohitg00/skillkit --skill hypothesis-testing -a claude-code`. Or copy the skill folder (packages/core/src/methodology/packs/debugging/hypothesis-testing in rohitg00/skillkit) into .claude/skills/hypothesis-testing in your project. Claude Code loads it when a task matches its description.

How do I install Hypothesis Testing in Codex?

Run `npx skills add rohitg00/skillkit --skill hypothesis-testing -a codex`. Or copy the skill folder (packages/core/src/methodology/packs/debugging/hypothesis-testing in rohitg00/skillkit) into .agents/skills/hypothesis-testing in your project. Codex loads it when a task matches its description.

Can I use Hypothesis Testing in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rohitg00/skillkit --skill hypothesis-testing -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hypothesis-testing, .gemini/skills/hypothesis-testing, .github/skills/hypothesis-testing and .opencode/skills/hypothesis-testing in your project.

What does Hypothesis Testing need to run?

SKILL.md names no scripts, command-line tools or credentials: Hypothesis Testing is instructions for the agent only.

Does Hypothesis Testing access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Hypothesis Testing safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hypothesis Testing use?

Hypothesis Testing is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hypothesis Testing use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hypothesis Testing?

Skills that share tags, products or a category with Hypothesis Testing: OpenLogi macOS Permissions Triage (AprilNEA/OpenLogi, 23k stars), Bug Finder for daisyUI (saadeghi/daisyui, 43k stars), Root Cause Debugging (garrytan/gstack, 136k stars) and Graph-Based Bug Tracing (tirth8205/code-review-graph, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hypothesis Testing?

rohitg00 (a GitHub user) maintains it in rohitg00/skillkit, which has 1,546 GitHub stars. The repository holds 14 skills in this directory. The repository was last updated on June 2, 2026.

Source: rohitg00/skillkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.