Agent skill

Rigorous Experiments

by glebis in glebis/claude-skills

This skill should be used when designing, running, validating, or auditing statistical experiments on personal or observational time-series data (health metrics, speech/text corpora, behavioral…

MITAuto-check passedData & Analytics

Install Rigorous Experiments

skills CLI
$ npx skills add glebis/claude-skills --skill rigorous-experiments -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install glebis/claude-skills rigorous-experiments --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/glebis/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/rigorous-experiments .claude/skills/rigorous-experiments && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
rigorous-experiments
GitHub stars
388
Token cost
~1.7k tokens
SKILL.md length
780 words
Files
16 (incl. scripts, references)
Skills in repo
91
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when designing, running, validating, or auditing statistical experiments on personal or observational time-series data (health metrics, speech/text corpora, behavioral…

  • Works in 8 steps: Pre-register before computing.… → Exact permutation, never sampled, on… → Permute over the full calendar, not the… → …
  • Design an experiment
  • SKILL.md covers Modes, Non-negotiable core (all modes), Workflow (full study) and Viewing results, plus 1 more section
  • Runs Python scripts from its folder; calls python3

What it does

Rigorous Experiments is an agent skill from glebis/claude-skills. This skill should be used when designing, running, validating, or auditing statistical experiments on personal or observational time-series data (health metrics, speech/text corpora, behavioral logs, diaries, n-of-1 self-tracking). It enforces pre-registration, exact permutation tests, FDR discipline, data-validation gates, adversarial code review, and cross-validation with external models. Triggers on "design an experiment", "test this hypothesis on my data", "is this correlation real", "audit these findings"…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 19 other files, including scripts and reference files (for example `evals/cases/bad_exp.py`, `evals/cases/bad_results.json` and `evals/cases/cases.json`).

It sits in Data & Analytics, covering Machine learning, Forecasting and time series and Journaling and reflection. The repository describes itself as: Collection of Claude Code skills for enhanced AI workflows. The licence is MIT.

When your agent uses it

  • Design an experiment
  • Test this hypothesis on my data
  • Is this correlation real
  • Audit these findings

Example prompts

  • “design an experiment”
  • “test this hypothesis on my data”
  • “is this correlation real”
  • “/rigorous-experiments”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Pre-register before computing. Hypotheses, exact tests, family size
  2. Exact permutation, never sampled, on small n. A session sequence of
  3. Permute over the full calendar, not the compressed series. Shifting
  4. BH with FIXED family size m, a LITERAL CONSTANT declared at design
  5. Stationarity check before correlating trending series. Exact
  6. Stratify before pooling (Simpson check): within group (e.g.
  7. Controls can re-describe a finding, not just kill it. When a control
  8. Honest statuses: confirmed (q<0.10 exact) ≠ lead (p<0.06) ≠ null ≠

What it can do on your machine

Read from SKILL.md and the folder at commit 7524dff. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Rigorous Experiments loads about 1.7k tokens when it runs, and up to ~6.2k if it reads all its reference files. Until then it costs about 157 tokens; SKILL.md has 780 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~157
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from glebis/claude-skills at commit 7524dff, republished under its MIT licence (© glebis). 780 words, ~1,731 tokens.

Download SKILL.mdSave it as .claude/skills/rigorous-experiments/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
rigorous-experiments
description
This skill should be used when designing, running, validating, or auditing statistical experiments on personal or observational time-series data (health metrics, speech/text corpora, behavioral logs, diaries, n-of-1 self-tracking). It enforces pre-registration, exact permutation tests, FDR discipline, data-validation gates, adversarial code review, and cross-validation with external models. Triggers on "design an experiment", "test this hypothesis on my data", "is this correlation real", "audit these findings", "pre-register", "validate this dataset", or any n-of-1 / quantified-self analysis request.

Rigorous Experiments

Run statistical experiments on observational/personal time-series data that survive scrutiny. Distilled from a 54-experiment n-of-1 program in which sampled permutation tests, missing-data artifacts, app-categorization bugs and collinear mechanisms repeatedly manufactured — and then destroyed — "findings". Every rule here exists because its absence once produced a wrong conclusion.

Modes

Pick the mode matching the request; chain them for a full study.

ModeWhenReference
designNew hypothesis or studyreferences/design.md
conductImplementing + running the experimentreferences/statistics.md
validate-dataBefore trusting ANY new data sourcereferences/data-validation.md
cross-validateFindings worth defending; code review; external model review (e.g. GPT Pro)references/cross-validation.md
investigate-leadsA sweep/run produced leads (p<0.06, not FDR-confirmed)references/lead-investigation.md
auditRe-examining past claims, registries of findingsreferences/statistics.md §Audit

Non-negotiable core (all modes)

  1. Pre-register before computing. Hypotheses, exact tests, family size m, and the acceptance threshold go in the script docstring BEFORE the first run. Post-hoc tests are reported as descriptive, never promoted.
  2. Exact permutation, never sampled, on small n. A session sequence of n=19 has 18 circular shifts: the minimum honest p is ~1/19≈0.05. Sampling 2000 shifts with replacement fabricates precision (this killed a flagship "q=0.028" finding). Use scripts/perm_stats.py.
  3. Permute over the full calendar, not the compressed series. Shifting a gap-compressed series breaks the timeline; keep missingness as NaN masks re-applied per shift. Event indicators must be pure 0/1 with no gaps — missingness lives only in the outcome series.
  4. BH with FIXED family size m, a LITERAL CONSTANT declared at design time — never len(tests) (that defeats pre-registration; the linter rejects it). Assert the run matches the declared m. Confirmatory families small and separate from exploratory sweeps; pooling everything into one BH buries true effects, cherry-picking families manufactures them. Plain BH assumes independent/positively-dependent tests; for strongly dependent lag families use BH-Yekutieli or maxT resampling.
  5. Stationarity check before correlating trending series. Exact circular shift on a trending series is "exactly, reproducibly wrong": report prewhitened-r (AR1 residuals) and stationary bootstrap alongside.
  6. Stratify before pooling (Simpson check): within group (e.g. therapy/coaching) and within regime (pre/post known breaks). A pooled r=−0.25 once hid therapy −0.64 vs coaching +0.53.
  7. Controls can re-describe a finding, not just kill it. When a control collapses an effect, check collinearity of control and predictor — r(self-focus, session-length)=0.79 meant "mechanism ambiguous", not "effect fake". Report the decomposition.
  8. Honest statuses: confirmed (q<0.10 exact) ≠ lead (p<0.06) ≠ null ≠ descriptive. Status flips are recorded, never silently edited. Nulls with adequate power are findings. Robust ≠ significant: a lead surviving leave-one-out at small n is still underpowered — a candidate for prospective test, not a finding. 8b. Series scope is part of the test. A lagged "[t+1]" means the next unit in the series the hypothesis is about, not the next pooled row; define scope before lagging (it once flipped a sign). When recomputing a prior result, reproduce a stored artifact on that scope first.
  9. Privacy: raw text/audio never enters output files or external uploads — statistics, rates and embedding-derived scores only.
  10. Plain-language reporting: every statistic carries its practical meaning inline; define r/p/q/n once per report; no untranslated jargon calques. Narrative first, numbers as support.
Show full SKILL.md (268 more words)Show less

Workflow (full study)

  1. validate-data gate on any new source (see reference — the checklist has caught: zero-vs-missing conflation, dedup semantics, substring category bugs, rolling purge windows, timezone conventions).
  2. design: pre-registered hypotheses + family + power sanity.
  3. conduct: implement with scripts/perm_stats.py; run; write results JSON with tests, statuses, and caveats including known limitations.
  4. cross-validate: adversarial code review (e.g. Codex read-only) BEFORE trusting results; fix findings; re-run. For major claims, external model review with a privacy-screened archive.
  5. investigate-leads on anything that surfaced as a lead (not at the same scale — the triage battery: LOO, directionality, detrend-vs-step, within-cycle, prewhiten+bootstrap; consolidate same-direction leads into one composite). Mark diagnostic runs descriptive_only: true.
  6. Verdicts in honest prose (mixed/rejected allowed); report; registry update with status provenance.

Viewing results

Launch the bundled explorer over any directory of results JSONs:

bash
python3 scripts/explorer.py <results_dir> [--port 8799] [--pattern "exp*.json"] [--sort newest|oldest]

Generates explorer.html in the directory, starts (or reuses) a loopback http server on the port, and opens the browser: experiment list with confirmed/lead badges, filter, sortable test tables color-coded by status, verdicts, caveats, raw JSON. The page fetches result files live — re-running experiments updates the view; re-run the script only when new result files appear. Serve over localhost, never file:// (CDN fonts) and never on a non-loopback interface (results may contain personal statistics).

Evals

Run python3 evals/run_evals.py (from the skill directory) to lint an experiment script/results pair against the standards (pre-registration present, fixed literal m, exact perm usage, caveats, no raw text in outputs). A diagnostic/triage run that intentionally mints no new tests sets descriptive_only: true in its results JSON to satisfy the "has tests" check. Eval cases in evals/cases/ document expected pass/fail examples.

© glebis, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (scripts, references) in rigorous-experiments of glebis/claude-skills.

  • SKILL.md
  • dist.zip
  • evals/cases/bad_exp.py
  • evals/cases/bad_results.json
  • evals/cases/cases.json
  • evals/cases/good_exp.py
  • evals/cases/good_results.json
  • evals/run_evals.py
  • references/cross-validation.md
  • references/data-validation.md
  • references/design.md
  • references/lead-investigation.md
  • references/statistics.md
  • scripts/explorer.py
  • scripts/perm_stats.py
  • scripts/triage.py

Open the folder on GitHubat commit 7524dff

Compare with similar skills

Rigorous Experiments next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Rigorous Experiments compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Rigorous Experiments this skillglebis/claude-skills388—~1.7kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries168—~3.1kAutomated safety check: PassApache-2.0
Aeon Time Series Machine Learningdavila7/claude-code-templates32k14 repos~2.6kAutomated safety check: PassMIT
Data Sciencetravisjneuman/.claude1011 repos~2.3kAutomated safety check: PassMIT
Data Scientistdavila7/claude-code-templates32k8 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    168 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Aeon Time Series Machine Learning

    davila7/claude-code-templates

    Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.

    32k GitHub starsUsed in 14 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Data Science

    travisjneuman/.claude

    Data science and analytics expertise for statistical analysis, machine learning pipelines, data governance, business intelligence, predictive modeling, and analytics strategy.

    101 GitHub starsUsed in 1 repo~2.3k tokens
    Data & AnalyticsAuto-check passed
  • Data Scientist

    davila7/claude-code-templates

    Expert data scientist for advanced analytics, machine learning, and statistical modeling.

    32k GitHub starsUsed in 8 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Longbridge Quant

    helsome/folio

    Quantitative strategy frameworks: pairs trading/cointegration, volatility regime strategies, seasonality/calendar effects, multi-factor models (IC/IR), factor research and screening, correlation…

    269 GitHub starsUsed in 1 repo~1.6k tokens
    Data & AnalyticsAuto-check passed

More from glebis/claude-skills

All 91 skills in this repo
  • Runs a human-first workflow for labeling PII spans in a transcript, then scores inter-annotator agreement and drafts an adjudicated gold set.

    388 GitHub stars~1.3k tokensUpdated 11 days ago
    Auto-check passed
  • Automates a dedicated, logged-in Chrome instance per profile without ever closing the user's own open tabs or browser windows.

    388 GitHub stars~973 tokensUpdated 11 days ago
    Auto-check passed
  • Deep Research

    glebis/claude-skills

    This skill should be used when conducting comprehensive research on any topic using the OpenAI Deep Research API.

    388 GitHub stars~2.6k tokensUpdated 11 days ago
    Auto-check: notes
  • Elimination Research

    glebis/claude-skills

    This skill should be used for elimination-style research where the user wants to choose from a shortlist of products, tools, services, vendors, or other options using explicit criteria, numeric…

    388 GitHub stars~1.6k tokensUpdated 11 days ago
    Auto-check passed
  • Narrated HTML Presentations

    glebis/claude-skills

    Generates a self-contained HTML presentation with article and slides modes, ElevenLabs voiceover narration and optional GPT Image 2 illustrations.

    388 GitHub stars~2.3k tokensUpdated 11 days ago
    Auto-check: notes
  • Writes fictional but realistic coaching or therapy session transcripts for evals, demos and few-shot examples, in several modalities and export formats.

    388 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed

Questions about Rigorous Experiments

What does Rigorous Experiments do?

This skill should be used when designing, running, validating, or auditing statistical experiments on personal or observational time-series data (health metrics, speech/text corpora, behavioral…. Rigorous Experiments is an agent skill from glebis/claude-skills. This skill should be used when designing, running, validating, or auditing statistical experiments on personal or observational time-series data (health metrics, speech/text corpora, behavioral logs, diaries, n-of-1 self-tracking).

When should I use Rigorous Experiments?

Rigorous Experiments fits situations like: design an experiment; test this hypothesis on my data; is this correlation real; audit these findings.

How do I install Rigorous Experiments in Claude Code?

Run `npx skills add glebis/claude-skills --skill rigorous-experiments -a claude-code`. Or copy the skill folder (rigorous-experiments in glebis/claude-skills) into .claude/skills/rigorous-experiments in your project. Claude Code loads it when a task matches its description.

How do I install Rigorous Experiments in Codex?

Run `npx skills add glebis/claude-skills --skill rigorous-experiments -a codex`. Or copy the skill folder (rigorous-experiments in glebis/claude-skills) into .agents/skills/rigorous-experiments in your project. Codex loads it when a task matches its description.

Can I use Rigorous Experiments in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add glebis/claude-skills --skill rigorous-experiments -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/rigorous-experiments, .gemini/skills/rigorous-experiments, .github/skills/rigorous-experiments and .opencode/skills/rigorous-experiments in your project.

What does Rigorous Experiments need to run?

Going by SKILL.md and its folder, Rigorous Experiments needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Rigorous Experiments access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Rigorous Experiments safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Rigorous Experiments use?

Rigorous Experiments is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Rigorous Experiments use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Rigorous Experiments?

Skills that share tags, products or a category with Rigorous Experiments: Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars), Aeon Time Series Machine Learning (davila7/claude-code-templates, 32k stars) and Data Science (travisjneuman/.claude, 101 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Rigorous Experiments?

glebis (a GitHub user) maintains it in glebis/claude-skills, which has 388 GitHub stars. The repository holds 91 skills in this directory. The repository was last updated on September 26, 2026.

Source: glebis/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.