Agent skill

Calc Sample Size

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when planning how many patients or cases a study needs before data collection (power analysis, IRB justification).

MITAuto-check passedResearch & Science

Install Calc Sample Size

skills CLI
$ npx skills add Aperivue/medsci-skills --skill calc-sample-size -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills calc-sample-size --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/calc-sample-size .claude/skills/calc-sample-size && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
calc-sample-size
GitHub stars
329
Token cost
~2.9k tokens
SKILL.md length
1,159 words
Files
11 (incl. references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when planning how many patients or cases a study needs before data collection (power analysis, IRB justification).

  • Works in 4 steps: Understand the Study → Collect Parameters → Calculate and Report → …
  • Planning how many patients
  • SKILL.md covers Decision Tree, Tests 1–11, Tests 12–17 (specialised… and Out of Scope, plus 2 more sections
  • Runs Python scripts from its folder

What it does

Calc Sample Size is an agent skill from Aperivue/medsci-skills. Use when planning how many patients or cases a study needs before data collection (power analysis, IRB justification). Walks a decision tree to the right test and returns reproducible R/Python code and IRB-ready justification text. Analyzing collected data is /analyze-stats.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including reference files (for example `references/formulas.md`, `references/justification_examples.md` and `references/mrmc_reader_study_sample_size.md`).

It sits in Research & Science, covering Experimental design. It works with Python. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • Planning how many patients
  • Cases a study needs before data collection (power analysis
  • IRB justification)

Example prompts

  • “/calc-sample-size”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Understand the Study
  2. Collect Parameters
  3. Calculate and Report
  4. Sensitivity Analysis (Optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • psychologie.hhu.de

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Calc Sample Size loads about 2.9k tokens when it runs, and up to ~27k if it reads all its reference files. Until then it costs about 73 tokens; SKILL.md has 1,159 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~27k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 1,159 words, ~2,908 tokens.

Download SKILL.mdSave it as .claude/skills/calc-sample-size/SKILL.md (or your agent's skills folder). This skill also uses 10 other files; get the full folder from GitHub.
name
calc-sample-size
description
Use when planning how many patients or cases a study needs before data collection (power analysis, IRB justification). Walks a decision tree to the right test and returns reproducible R/Python code and IRB-ready justification text. Analyzing collected data is /analyze-stats.
metadata.triggers
sample size, power analysis, power calculation, how many patients, how many subjects, IRB sample size

Calc-Sample-Size Skill

Decision Tree

Walk the user through this tree one question at a time; do not assume answers. Reader studies, segmentation and model-comparison designs sit outside the tree: see Tests 14–17.

What is your primary outcome?
|
+-- Binary (yes/no, positive/negative)
|   |
|   +-- Paired data (same subjects, two methods)?
|   |   +-- YES --> [5] McNemar test
|   |   +-- NO  --> How many groups?
|   |       +-- 2 groups, superiority     --> [4] Two-proportion comparison (chi-square)
|   |       +-- 2 groups, non-inferiority --> [10] Non-inferiority / equivalence
|   |       +-- Multivariable model       --> single-predictor hypothesis test? --> [9] Logistic regression
|   |                                     --> clinical prediction / AI model for use?
|   |                                         +-- developing the model  --> [12] Prediction-model development (Riley)
|   |                                         +-- externally validating  --> [13] External-validation (Riley)
|   |
+-- Continuous (measurement, score)
|   |
|   +-- How many groups?
|       +-- 2 groups  --> [6] Independent t-test
|       +-- 3+ groups --> [8] One-way ANOVA
|
+-- Time-to-event (survival, recurrence)
|   |
|   +-- Two groups, unadjusted      --> [7] Log-rank test
|   +-- Multivariable / adjusted HR  --> [7] Log-rank (Schoenfeld) + [11] Cox EPV
|
+-- Agreement (inter-rater, reproducibility)
|   |
|   +-- Continuous measurements --> [2] ICC
|   +-- Categorical ratings     --> [3] Kappa
|
+-- Diagnostic accuracy (Se, Sp, AUC precision)
    |
    +--> [1] Diagnostic accuracy (precision-based)

Tests 1–11

Once the test is chosen, read ${CLAUDE_SKILL_DIR}/references/formulas.md § Test N — the parameter table with defaults, the effect-size interpretation, the formula, R/Python code and the methodological reference.

#TestUse when
1Diagnostic accuracy — Se/Sp precisiondesired 95% CI half-width for sensitivity or specificity
2ICC agreement (Walter 1998 test; Bonett 2002 CI width)inter-/intra-rater agreement on continuous measurements (tumor size, angle)
3Kappa agreement (Donner & Eliasziw 1992; needs the trait prevalence)agreement on categorical ratings (BI-RADS category, lesion present/absent)
4Two-proportion comparison (chi-square)two independent groups (AI vs conventional detection rate)
5McNemar (paired proportions)paired binary outcomes (two readers on the same cases, before/after)
6Independent t-testmeans in two independent groups (lesion size, malignant vs benign)
7Survival / log-rank (Schoenfeld events, then patients)time-to-event between two groups
8One-way ANOVAmeans across 3+ independent groups
9Logistic regression (Peduzzi EPV + Hsieh 1998, continuous or binary predictor)multivariable binary outcome, single-predictor hypothesis test
10Non-inferiority / equivalencenew method not worse than standard by more than a pre-specified margin, or equivalent within it
11Cox regression EPVmultivariable Cox model — enough events for stable estimates
  • Test 9: Peduzzi EPV ≥ 10 is a minimum baseline for a single-predictor hypothesis test only. For a clinical prediction / medical-AI model intended for use, EPV-10 is outdated and reviewer-vulnerable — use Test 12 (development) / Test 13 (validation). Always report both Peduzzi and Hsieh and recommend the larger N.
  • Test 10: NI alpha is one-sided (typically 0.025); orient the difference so > 0 favours the new method. Equivalence is TOST, powered jointly. The margin must be clinically justified (see formulas.md § Margin selection).
  • Test 11: EPV ensures model stability, not power for a specific HR. If an HR is available, also run Test 7 (Schoenfeld) and recommend the larger N.

Tests 12–17 (specialised designs)

Each has its own reference file (parameters, method, reporting); read it once the test is chosen.

  • Test 12 — Prediction-model development (Riley). Developing a clinical prediction / classification model (including a medical-AI model evaluated as one) for use. EPV-10 does not apply: N is the largest satisfying all Riley criteria — three for a binary or time-to-event outcome, four for a continuous one (R pmsampsize).
  • Test 13 — External validation (Riley). Validating an existing prediction/AI model: size for the CI width of the C-statistic, calibration slope, O:E and (if claimed) net benefit (R pmvalsampsize); ≥ 100 events and ≥ 100 non-events is only a floor. For Tests 12–13 read ${CLAUDE_SKILL_DIR}/references/prediction_model_sample_size.md.
  • Test 14 — MRMC reader study (Obuchowski–Rockette). Readers with vs without the AI, or AI non-inferior to readers. Test 1 under-sizes it because readers are a random effect: size readers J × cases from pilot/literature variance components with RJafroc / MRMCaov / iMRMC (do not hand-roll the OR algebra) and report the J × N power grid. Read ${CLAUDE_SKILL_DIR}/references/mrmc_reader_study_sample_size.md.
  • Test 15 — Segmentation-metric precision (Dice / HD95 / NSD). The outcome is a per-case score, not a proportion: n ≈ (1.96·SD/δ)² from the pilot SD of per-case Dice, sized on the worst structure; CI by patient-level bootstrap (BCa). This is precision; a comparison is Test 16. Read ${CLAUDE_SKILL_DIR}/references/segmentation_metric_sample_size.md.
  • Test 16 — Between-model comparison. Model A beats B, C, …: power the paired per-case difference, n = ((z₁₋α/₂ + z₁₋β)·SD_Δ/Δ)² (a CI sized to just exclude zero has ~50% power), not each model's precision; for > 2 models pre-specify one primary contrast or pay the family-wise correction; a ranking claim needs multiple seeds. Read ${CLAUDE_SKILL_DIR}/references/multi_model_comparison_sample_size.md.
  • Test 17 — Segmentation usability. Clinicians can use it (acceptability rate, catastrophic-failure bound, edit time): acceptability is a proportion sized per structure class; DE ≈ 1 + (m−1)ρ holds only when each case has its own readers — the same readers on every case (crossed) add a reader term more cases cannot shrink; bounding failures at ≤ 1% needs ~300 clean cases (rule of three). Read ${CLAUDE_SKILL_DIR}/references/segmentation_acceptability_sample_size.md.

Out of Scope

Do not compute adaptive trials (group-sequential, sample size re-estimation), cluster-randomized trials (design effect, ICC-based inflation), Bayesian sample size determination, crossover designs, or multi-endpoint correction (mention Bonferroni if asked, but do not compute corrected sample sizes). Say the design is beyond this skill and point to G*Power (free, https://www.psychologie.hhu.de/gpower), PASS, or a biostatistician.

Show full SKILL.md (461 more words)Show less

Workflow

Phase 1: Understand the Study
  1. Ask the user to describe the study briefly (design, primary outcome, groups).
  2. Walk the decision tree to a test, and confirm it with the user before proceeding.
Phase 2: Collect Parameters
  1. Present the selected test's parameter table.
  2. For each parameter without a user-provided value, explain it and offer the default.
  3. Estimate effect sizes from prior literature (ask for the references) or pilot data; use Cohen's conventions only as a last resort, noting that convention-based estimates are less precise.
Phase 2b: Retrospective Studies

When the dataset already exists, formal power analysis is often impractical. Offer:

  • Fixed extract: read ${CLAUDE_SKILL_DIR}/references/observational_cohort.md and report event budget / confidence-interval precision instead of forcing a prospective recruitment-style power calculation.
  • Experience-based justification (acceptable for IRB and many journals):
    • Institution volume: total exams in period × prevalence × (1 − exclusion rate) = expected N. Ask for the annual exam volume for the modality, study period, prevalence and exclusion rate. This gives a realistic upper bound for N.
    • Prior studies: report the N of 3–5 comparable published studies and cite them; the user's N should be in the same range or larger.
    • IRB templates for both: ${CLAUDE_SKILL_DIR}/references/justification_examples.md § Retrospective.

Use a formal calculation (Phase 3) even for a retrospective study when a subset is enrolled prospectively, the primary analysis tests a hypothesis (not just estimation), the journal's Instructions for Authors require a power analysis, or the IRB requires it.

Phase 3: Calculate and Report
  1. Generate R code (primary) and Python code (alternative) from the reference formula.
  2. Run the R code via Bash; the reported N is the number it prints, not a hand calculation.
  3. Present the result in the Output Format below. Cite methodological sources only from formulas.md or the test's reference file; any other reference needs a DOI/PMID confirmed via /search-lit, otherwise mark it [UNVERIFIED - NEEDS MANUAL CHECK]. Mark an effect size, clinical definition or threshold you could not confirm [VERIFY].
  4. In a project, save the IRB text as protocol/sample_size_justification.md and the scripts as protocol/sample_size_calc.R / .py: /write-protocol and /write-paper embed that text verbatim, so the numbers are never retyped.
Phase 4: Sensitivity Analysis (Optional)

If a parameter is uncertain or the effect-size estimate is vague, flag it and offer a table of N across plausible values (e.g., varying effect size, or power from 0.80 to 0.90).

Output Format

Always structure the final output as follows:

markdown
## Sample Size Calculation Report

### Study Design
[1-2 sentence summary of the design and test selected]

### Parameters
| Parameter | Value | Source |
|-----------|-------|--------|
| ... | ... | user / literature / convention |

### Result
- **Required sample size**: N = [value]
- **With [X]% attrition adjustment**: N_adj = [value]

### R Code (Reproducible)
```r
# [complete, self-contained R script]
# Dependencies: [list packages]
# Run: Rscript sample_size_calc.R
Python Code (Alternative)
python
# [complete, self-contained Python script]
# Dependencies: [list packages]
# Run: python sample_size_calc.py
IRB Justification Text

A sample of [N] participants is required to detect [effect description] with [power]% power at a [one/two]-sided significance level of [alpha], assuming [key assumptions]. Accounting for an estimated [X]% attrition rate, we plan to enroll [N_adj] participants. This calculation is based on [formula/method reference].

Effect Size Interpretation

[Cohen's benchmark classification + clinical meaning in the context of this study]


The IRB text must state N; name the test and its formula source; give every assumed parameter
(effect size, alpha, power); state the attrition adjustment and final enrollment target; cite the
methodological reference (e.g., "Schoenfeld, 1981"); and use formal, third-person language. Read
`${CLAUDE_SKILL_DIR}/references/justification_examples.md` for per-design exemplars when writing it.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 10 other files (references) in skills/calc-sample-size of Aperivue/medsci-skills.

  • SKILL.md
  • references/formulas.md
  • references/justification_examples.md
  • references/mrmc_reader_study_sample_size.md
  • references/multi_model_comparison_sample_size.md
  • references/observational_cohort.md
  • references/prediction_model_sample_size.md
  • references/segmentation_acceptability_sample_size.md
  • references/segmentation_metric_sample_size.md
  • skill.yml
  • tests/test_worked_examples.py

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Calc Sample Size next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Calc Sample Size compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Calc Sample Size this skillAperivue/medsci-skills329—~2.9kAutomated safety check: PassMIT
Analytical Method Validation PlannerK-Dense-AI/scientific-agent-skills48k1 repos~4.9kAutomated safety check: NotesMIT
Light Experiment CodingLight0305/Light-skills641—~2.3kAutomated safety check: PassMIT
Bio Experimental Design Multiple TestingGPTomics/bioSkills1.2k1 repos~3.5kAutomated safety check: PassMIT
Adaptyvmajiayu000/claude-skill-registry6661 repos~1.9kAutomated safety check: NotesMIT
Meta Forest Binary Plotaipoch/medical-research-skills2k—~2.1kAutomated safety check: PassMIT

Similar skills

  • Analytical Method Validation Planner

    K-Dense-AI/scientific-agent-skills

    Plans, runs, and documents analytical method validation, verification, or transfer studies under ICH Q2(R2)/Q14, USP, ICH M10, CLSI EP, or ISO/IEC 17025.

    48k GitHub starsUsed in 1 repo~4.9k tokens
    Research & ScienceAuto-check: notes
  • Light Experiment Coding

    Light0305/Light-skills

    Builds the code for a frozen research experiment test-first, with leakage controls, seed handling and saved evidence so results can be rerun and audited.

    641 GitHub stars~2.3k tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Controls error rates across thousands of simultaneous tests in genomics discovery using false-discovery-rate methods (Benjamini-Hochberg 1995; Benjamini-Yekutieli 2001 for arbitrary dependence…

    1.2k GitHub starsUsed in 1 repo~3.5k tokens
    Research & ScienceAuto-check passed
  • Adaptyv

    majiayu000/claude-skill-registry

    How to use the Adaptyv Bio Foundry API and Python SDK for protein experiment design, submission, and results retrieval.

    666 GitHub starsUsed in 1 repo~1.9k tokens
    Research & ScienceAuto-check: notes
  • Meta Forest Binary Plot

    aipoch/medical-research-skills

    Generate meta-analysis forest plots for binary classification data.

    2k GitHub stars~2.1k tokensUpdated 21 days ago
    Documents & OfficeAuto-check passed
  • Builds a scoring-aligned outline for a mathematical modeling paper and a model selection plan with baseline, improvement and validation experiments.

    452 GitHub starsUsed in 1 repo~1.8k tokens
    EducationAuto-check passed

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Model Assessment

    Aperivue/medsci-skills

    A skill your agent uses when validating or evaluating a trained medical-imaging model.

    329 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 3 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Verify Refs

    Aperivue/medsci-skills

    A skill your agent uses when checking whether a manuscript's references are real.

    329 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    329 GitHub stars~3.9k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Calc Sample Size

What does Calc Sample Size do?

A skill your agent uses when planning how many patients or cases a study needs before data collection (power analysis, IRB justification). Calc Sample Size is an agent skill from Aperivue/medsci-skills. Use when planning how many patients or cases a study needs before data collection (power analysis, IRB justification).

When should I use Calc Sample Size?

Calc Sample Size fits situations like: planning how many patients; cases a study needs before data collection (power analysis; IRB justification).

How do I install Calc Sample Size in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill calc-sample-size -a claude-code`. Or copy the skill folder (skills/calc-sample-size in Aperivue/medsci-skills) into .claude/skills/calc-sample-size in your project. Claude Code loads it when a task matches its description.

How do I install Calc Sample Size in Codex?

Run `npx skills add Aperivue/medsci-skills --skill calc-sample-size -a codex`. Or copy the skill folder (skills/calc-sample-size in Aperivue/medsci-skills) into .agents/skills/calc-sample-size in your project. Codex loads it when a task matches its description.

Can I use Calc Sample Size in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill calc-sample-size -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/calc-sample-size, .gemini/skills/calc-sample-size, .github/skills/calc-sample-size and .opencode/skills/calc-sample-size in your project.

What does Calc Sample Size need to run?

Going by SKILL.md and its folder, Calc Sample Size needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Calc Sample Size access the network?

SKILL.md names 1 domain. As links in the text: psychologie.hhu.de. This is read from the text; nothing was executed.

Is Calc Sample Size safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Calc Sample Size use?

Calc Sample Size is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Calc Sample Size use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 24k tokens, read only when the agent opens those files.

What are the alternatives to Calc Sample Size?

Skills that share tags, products or a category with Calc Sample Size: Analytical Method Validation Planner (K-Dense-AI/scientific-agent-skills, 48k stars), Light Experiment Coding (Light0305/Light-skills, 641 stars), Bio Experimental Design Multiple Testing (GPTomics/bioSkills, 1.2k stars) and Adaptyv (majiayu000/claude-skill-registry, 666 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Calc Sample Size?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.