Agent skill

Hypothesis Testing Guide

by wentorai in wentorai/research-plugins

Statistical hypothesis testing, power analysis, and significance reporting

MITAuto-check passedResearch & Science

Install Hypothesis Testing Guide

skills CLI
$ npx skills add wentorai/research-plugins --skill hypothesis-testing-guide -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wentorai/research-plugins hypothesis-testing-guide --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wentorai/research-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/analysis/statistics/hypothesis-testing-guide .claude/skills/hypothesis-testing-guide && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
hypothesis-testing-guide
GitHub stars
298
Used in
1 other repo
Token cost
~2k tokens
SKILL.md length
526 words
Files
1
Skills in repo
428
Repo updated
First seen
Licence
MIT

At a glance

Statistical hypothesis testing, power analysis, and significance reporting

  • Works in 7 steps: State hypotheses. Define H0 (null: no… → Choose significance level. Typically… → Select the appropriate test. Based on… → …
  • Tasks that involve Experimental design
  • SKILL.md covers Overview, The Hypothesis Testing Framework, Test Selection Guide and Running Tests in Python, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Hypothesis Testing Guide is an agent skill from wentorai/research-plugins. Statistical hypothesis testing, power analysis, and significance reporting

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Experimental design and Statistics. The repository describes itself as: 350+ academic research skills, MCP configs, and plugins for Research-Claw and AI agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Experimental design
  • Tasks that involve Statistics

Example prompts

  • “/hypothesis-testing-guide”

Requirements

  • Python 3

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. State hypotheses. Define H0 (null: no effect) and H1 (alternative: effect exists).
  2. Choose significance level. Typically alpha = 0.05, but justify your choice.
  3. Select the appropriate test. Based on data type, distribution, and design.
  4. Check assumptions. Normality, homogeneity of variance, independence.
  5. Compute test statistic and p-value.
  6. Report effect size and confidence interval. p-values alone are insufficient.
  7. Make a decision. Reject or fail to reject H0, with practical interpretation.

What it can do on your machine

Read from SKILL.md and the folder at commit bf44b3c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.scipy.org
    • pingouin-stats.org
    • statsmodels.org
    • apastyle.apa.org
    • machinelearningmastery.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Hypothesis Testing Guide loads about 2k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 526 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wentorai/research-plugins at commit bf44b3c, republished under its MIT licence (© wentorai). 526 words, ~1,995 tokens.

Download SKILL.mdSave it as .claude/skills/hypothesis-testing-guide/SKILL.md (or your agent's skills folder).
name
hypothesis-testing-guide
description
Statistical hypothesis testing, power analysis, and significance reporting

Hypothesis Testing Guide

Overview

Hypothesis testing is the backbone of empirical research. It provides a principled framework for deciding whether observed differences in data reflect genuine effects or merely random variation. Misuse of hypothesis tests -- p-hacking, ignoring assumptions, confusing statistical and practical significance -- is a leading cause of irreproducible findings.

This guide covers the core hypothesis testing framework, the most commonly used tests across disciplines, assumption checking, effect size reporting, power analysis for sample size planning, and multiple comparison corrections. Each test is accompanied by Python code using scipy, statsmodels, and pingouin, ready to integrate into research workflows.

The goal is not just to help you run tests, but to help you run the right test correctly and report results following modern standards (APA 7th edition, journal best practices).

The Hypothesis Testing Framework

Step-by-Step Procedure
  1. State hypotheses. Define H0 (null: no effect) and H1 (alternative: effect exists).
  2. Choose significance level. Typically alpha = 0.05, but justify your choice.
  3. Select the appropriate test. Based on data type, distribution, and design.
  4. Check assumptions. Normality, homogeneity of variance, independence.
  5. Compute test statistic and p-value.
  6. Report effect size and confidence interval. p-values alone are insufficient.
  7. Make a decision. Reject or fail to reject H0, with practical interpretation.
Common Errors
Error TypeDefinitionProbability
Type I (False Positive)Reject H0 when it is truealpha (usually 0.05)
Type II (False Negative)Fail to reject H0 when it is falsebeta (usually 0.20)
PowerProbability of correctly detecting an effect1 - beta (target: 0.80)

Test Selection Guide

Research QuestionData TypeGroupsTest
Two group means differ?Continuous, normal2 independentIndependent t-test
Before/after difference?Continuous, normal2 pairedPaired t-test
Multiple group means differ?Continuous, normal3+ independentOne-way ANOVA
Two group medians differ?Ordinal / non-normal2 independentMann-Whitney U
Before/after (non-normal)?Ordinal / non-normal2 pairedWilcoxon signed-rank
Multiple groups (non-normal)?Ordinal / non-normal3+ independentKruskal-Wallis
Association between categories?Categorical2 variablesChi-square test
Correlation?Continuous2 variablesPearson or Spearman
Show full SKILL.md (194 more words)Show less

Running Tests in Python

Independent Samples t-Test
python
from scipy import stats
import numpy as np
import pingouin as pg

# Generate example data
control = np.random.normal(50, 10, n=30)
treatment = np.random.normal(55, 10, n=30)

# Check normality assumption
stat_c, p_c = stats.shapiro(control)
stat_t, p_t = stats.shapiro(treatment)
print(f"Normality p-values: control={p_c:.3f}, treatment={p_t:.3f}")

# Check homogeneity of variance
stat_l, p_l = stats.levene(control, treatment)
print(f"Levene's test p={p_l:.3f}")

# Run t-test
t_stat, p_val = stats.ttest_ind(control, treatment, equal_var=(p_l > 0.05))

# Effect size (Cohen's d)
cohens_d = (treatment.mean() - control.mean()) / np.sqrt(
    ((len(control)-1)*control.var() + (len(treatment)-1)*treatment.var())
    / (len(control) + len(treatment) - 2)
)

print(f"t={t_stat:.3f}, p={p_val:.4f}, Cohen's d={cohens_d:.3f}")
One-Way ANOVA with Post-Hoc Tests
python
import pandas as pd

df = pd.DataFrame({
    'score': np.concatenate([
        np.random.normal(50, 10, 30),
        np.random.normal(55, 10, 30),
        np.random.normal(60, 10, 30)
    ]),
    'group': np.repeat(['A', 'B', 'C'], 30)
})

# ANOVA
aov = pg.anova(data=df, dv='score', between='group', detailed=True)
print(aov)

# Post-hoc pairwise comparisons (Tukey HSD)
posthoc = pg.pairwise_tukey(data=df, dv='score', between='group')
print(posthoc[['A', 'B', 'diff', 'p-tukey', 'hedges']])
Chi-Square Test of Independence
python
# Contingency table
observed = pd.DataFrame(
    [[45, 30], [25, 50]],
    index=['Method A', 'Method B'],
    columns=['Success', 'Failure']
)

chi2, p, dof, expected = stats.chi2_contingency(observed)
cramers_v = np.sqrt(chi2 / (observed.values.sum() * (min(observed.shape) - 1)))

print(f"chi2={chi2:.3f}, p={p:.4f}, Cramer's V={cramers_v:.3f}")

Power Analysis and Sample Size

Power analysis answers: "How many participants do I need?"

python
from statsmodels.stats.power import TTestIndPower, FTestAnovaPower

# For a two-sample t-test
analysis = TTestIndPower()

# Calculate required sample size
n = analysis.solve_power(
    effect_size=0.5,   # Cohen's d (medium effect)
    alpha=0.05,
    power=0.80,
    ratio=1.0,         # Equal group sizes
    alternative='two-sided'
)
print(f"Required n per group: {int(np.ceil(n))}")

# Power curve
import matplotlib.pyplot as plt

sample_sizes = np.arange(10, 200, 5)
powers = [analysis.power(effect_size=0.5, nobs1=n, ratio=1.0, alpha=0.05)
          for n in sample_sizes]

fig, ax = plt.subplots()
ax.plot(sample_sizes, powers)
ax.axhline(0.8, color='red', linestyle='--', label='Power = 0.80')
ax.set_xlabel('Sample Size per Group')
ax.set_ylabel('Statistical Power')
ax.legend()
fig.savefig('power_curve.pdf')
Effect Size Reference Table
Effect SizeSmallMediumLarge
Cohen's d (t-test)0.20.50.8
eta-squared (ANOVA)0.010.060.14
Cramer's V (chi-square)0.10.30.5
Pearson r (correlation)0.10.30.5

Multiple Comparison Corrections

When running multiple tests, the family-wise error rate inflates. Use corrections:

python
from statsmodels.stats.multitest import multipletests

p_values = [0.01, 0.04, 0.03, 0.08, 0.002]

# Bonferroni (conservative)
reject_bonf, pvals_bonf, _, _ = multipletests(p_values, method='bonferroni')

# Benjamini-Hochberg FDR (less conservative)
reject_bh, pvals_bh, _, _ = multipletests(p_values, method='fdr_bh')

for i, p in enumerate(p_values):
    print(f"p={p:.3f} | Bonferroni: {pvals_bonf[i]:.3f} ({reject_bonf[i]}) "
          f"| BH-FDR: {pvals_bh[i]:.3f} ({reject_bh[i]})")

Best Practices

  • Always report effect sizes alongside p-values. A significant p-value with a tiny effect size is rarely meaningful.
  • Pre-register your analysis plan. This prevents p-hacking and HARKing (Hypothesizing After Results are Known).
  • Check assumptions before running parametric tests. Use non-parametric alternatives when assumptions are violated.
  • Use confidence intervals. They convey both effect magnitude and precision.
  • Report exact p-values (p = 0.032), not thresholds (p < 0.05). Except when p < 0.001.
  • Consider Bayesian alternatives. Bayes factors provide evidence for H0, not just against it.
  • Plan sample sizes a priori. Power analysis should be done before data collection, not after.

References

© wentorai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/analysis/statistics/hypothesis-testing-guide of wentorai/research-plugins.

Open the folder on GitHubat commit bf44b3c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in wentorai/research-plugins, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Hypothesis Testing Guide next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Hypothesis Testing Guide compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Hypothesis Testing Guide this skillwentorai/research-plugins2981 repos~2kAutomated safety check: PassMIT
Data Scientistmagnus919/hermes-profiles278—~3.3kAutomated safety check: PassMIT
Data Scientistmagnus919/agent-skills111—~4.1kAutomated safety check: PassMIT
Statistical PowerK-Dense-AI/scientific-agent-skills48k1 repos~4.4kAutomated safety check: NotesMIT
Algo Rank Wilsonasgard-ai-platform/skills241—~1.1kAutomated safety check: PassMIT
Analyze StatsAperivue/medsci-skills329—~7kAutomated safety check: PassMIT

Similar skills

  • Data Scientist

    magnus919/hermes-profiles

    PhD-level expertise in data science, statistics, and machine learning.

    278 GitHub stars~3.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    111 GitHub stars~4.1k tokensUpdated today
    Research & ScienceAuto-check passed
  • Statistical Power

    K-Dense-AI/scientific-agent-skills

    Calculates sample sizes and statistical power for study planning.

    48k GitHub starsUsed in 1 repo~4.4k tokens
    Research & ScienceAuto-check: notes
  • Algo Rank Wilson

    asgard-ai-platform/skills

    Calculate Wilson Score confidence intervals for ranking items by positive proportion with sample size correction.

    241 GitHub stars~1.1k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Analyze Stats

    Aperivue/medsci-skills

    A skill your agent uses when data needs statistical analysis.

    329 GitHub stars~7k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Hri Author Response

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when writing the single ACM/IEEE HRI full-paper rebuttal that responds to three external reviewers and two area chairs — triaging study-design and statistics critiques…

    1.2k GitHub stars~1.4k tokensUpdated 10 days ago
    Research & ScienceAuto-check passed

More from wentorai/research-plugins

All 428 skills in this repo
  • Abstract Writing Guide

    wentorai/research-plugins

    Craft structured research abstracts that maximize clarity and journal acceptance

    298 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Academic Citation Manager

    wentorai/research-plugins

    Manage academic citations across BibTeX, APA, MLA, and Chicago formats

    298 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Academic Paper Summarizer

    wentorai/research-plugins

    Summarize academic papers with structured extraction of key elements

    298 GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Academic Study Methods

    wentorai/research-plugins

    Evidence-based study techniques for academic learning and retention

    298 GitHub starsUsed in 1 repo~1.8k tokens
    Auto-check passed
  • Academic Tone Guide

    wentorai/research-plugins

    Adjust writing tone and register for academic audiences and venues

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Academic Translation Guide

    wentorai/research-plugins

    Academic translation, post-editing, and Chinglish correction guide

    298 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Hypothesis Testing Guide

What does Hypothesis Testing Guide do?

Statistical hypothesis testing, power analysis, and significance reporting. Hypothesis Testing Guide is an agent skill from wentorai/research-plugins.

When should I use Hypothesis Testing Guide?

Hypothesis Testing Guide fits situations like: tasks that involve Experimental design; tasks that involve Statistics.

How do I install Hypothesis Testing Guide in Claude Code?

Run `npx skills add wentorai/research-plugins --skill hypothesis-testing-guide -a claude-code`. Or copy the skill folder (skills/analysis/statistics/hypothesis-testing-guide in wentorai/research-plugins) into .claude/skills/hypothesis-testing-guide in your project. Claude Code loads it when a task matches its description.

How do I install Hypothesis Testing Guide in Codex?

Run `npx skills add wentorai/research-plugins --skill hypothesis-testing-guide -a codex`. Or copy the skill folder (skills/analysis/statistics/hypothesis-testing-guide in wentorai/research-plugins) into .agents/skills/hypothesis-testing-guide in your project. Codex loads it when a task matches its description.

Can I use Hypothesis Testing Guide in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wentorai/research-plugins --skill hypothesis-testing-guide -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/hypothesis-testing-guide, .gemini/skills/hypothesis-testing-guide, .github/skills/hypothesis-testing-guide and .opencode/skills/hypothesis-testing-guide in your project.

What does Hypothesis Testing Guide need to run?

SKILL.md names no scripts, command-line tools or credentials: Hypothesis Testing Guide is instructions for the agent only. Our summary lists: Python 3.

Does Hypothesis Testing Guide access the network?

SKILL.md names 5 domains. As links in the text: docs.scipy.org, pingouin-stats.org, statsmodels.org, apastyle.apa.org and machinelearningmastery.com. This is read from the text; nothing was executed.

Is Hypothesis Testing Guide safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Hypothesis Testing Guide use?

Hypothesis Testing Guide is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Hypothesis Testing Guide use?

About 2k tokens (SKILL.md is roughly 8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Hypothesis Testing Guide?

Skills that share tags, products or a category with Hypothesis Testing Guide: Data Scientist (magnus919/hermes-profiles, 278 stars), Data Scientist (magnus919/agent-skills, 111 stars), Statistical Power (K-Dense-AI/scientific-agent-skills, 48k stars) and Algo Rank Wilson (asgard-ai-platform/skills, 241 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Hypothesis Testing Guide?

wentorai (a GitHub user) maintains it in wentorai/research-plugins, which has 298 GitHub stars. The repository holds 428 skills in this directory. The repository was last updated on June 19, 2026.

Source: wentorai/research-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.