Agent skill

Batch Cohort

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when one validated cohort analysis must be repeated across many exposure/outcome pairs.

MITAuto-check passedData & Analytics

Install Batch Cohort

skills CLI
$ npx skills add Aperivue/medsci-skills --skill batch-cohort -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills batch-cohort --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/batch-cohort .claude/skills/batch-cohort && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
batch-cohort
GitHub stars
329
Token cost
~2.4k tokens
SKILL.md length
916 words
Files
2
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when one validated cohort analysis must be repeated across many exposure/outcome pairs.

  • Works in 5 steps: Template Validation → Variable Specification → Batch Code Generation → …
  • One validated cohort analysis must be repeated across many exposure/outcome pairs
  • SKILL.md covers Inputs, Workflow, Output Files and Cross-National Batch Mode
  • Calls python3

What it does

Batch Cohort is an agent skill from Aperivue/medsci-skills. Use when one validated cohort analysis must be repeated across many exposure/outcome pairs. Generates one R/Python script per combination from a single methodology template, changing only the variables, and aggregates the results into a summary matrix.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `skill.yml`).

It sits in Data & Analytics, covering Product analytics. It works with Python. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • One validated cohort analysis must be repeated across many exposure/outcome pairs
  • Tasks that involve Product analytics

Example prompts

  • “/batch-cohort”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Template Validation
  2. Variable Specification
  3. Batch Code Generation
  4. Batch Execution (if execute or full mode)
  5. Summary Matrix

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Batch Cohort loads about 2.4k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 916 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 916 words, ~2,410 tokens.

Download SKILL.mdSave it as .claude/skills/batch-cohort/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
batch-cohort
description
Use when one validated cohort analysis must be repeated across many exposure/outcome pairs. Generates one R/Python script per combination from a single methodology template, changing only the variables, and aggregates the results into a summary matrix.
model
opus
metadata.triggers
batch cohort, batch analysis, 대량 분석, 변수 교체, variable swap, mass production, 80명 팀, batch generate, 일괄 코드 생성, exposure outcome matrix, combinatorial analysis

Batch Cohort Analysis Skill

Generate one analysis script per exposure/outcome combination from a single validated methodology template; only the variables change between scripts.

Inputs

  1. Database path(s): CSV/SAS data files (KNHANES, NHANES, NHIS, or any cleaned cohort)
  2. Methodology template: One of:
    • Path to a validated R/Python analysis script (from /replicate-study or /cross-national)
    • A paper type template name: nhis_cohort, cross_national, survey_weighted
    • A source paper to extract methodology from (falls back to /replicate-study Phase 1)
  3. Combination spec: A list of exposure/outcome pairs, provided as:
    • Inline list: exposures: [depression, obesity, smoking]; outcomes: [diabetes, hypertension, CVD]
    • CSV file with columns: exposure, outcome, (optional) subgroup_vars
    • "all" keyword: generates all pairwise combinations from the lists
Optional Inputs
  • Covariate set: Fixed covariate list for all analyses (default: use template's set)
  • Subgroup variables: Variables to stratify by (default: sex, age group)
  • Output format: code_only (just scripts) | execute (run + collect results) | full (code + results + summary)
  • Cross-national mode: cross_national: true generates paired scripts for both countries per combination (see Cross-National Batch Mode)

Example:

/batch-cohort

DB Korea: /path/to/knhanes/HN18.csv
DB US: /path/to/nhanes/
Template: cross_national
Exposures: [depression, obesity, smoking]
Outcomes: [diabetes, hypertension, metabolic_syndrome]
cross_national: true
Mode: execute

Workflow

Phase 1: Template Validation
  1. Read the methodology template (R script or paper type reference).
  2. Identify the slot variables — parts that change per combination: EXPOSURE_VAR (raw variable name in the database), EXPOSURE_LABEL (label for tables/figures), EXPOSURE_CODING (derivation of the binary/categorical exposure), and OUTCOME_VAR, OUTCOME_LABEL, OUTCOME_CODING likewise.
  3. Confirm the survey design: weighted analysis is mandatory for KNHANES/NHANES and is inherited from the template; NHIS and other claims databases have no sampling weights (standard regression).
  4. Verify the template runs successfully on at least one combination before batch generation.
  5. Output: template summary with identified slots → user approval.
Phase 2: Variable Specification

For each exposure and outcome in the combination spec:

  1. Look up the variable in the database: KNHANES — the name exists in the CSV header; NHANES — which table contains it (codebook.csv if available); NHIS — claims code or variable name. For ICD-10 claims algorithms read ${CLAUDE_SKILL_DIR}/../analyze-stats/references/analysis_guides/nhis_icd10_mapping.md; for survey variable coding, survey_weighted.md in the same folder. If a mapping is uncertain, write [VERIFY: variable_name] and ask the user to confirm it against the data dictionary; never guess a variable name, column name or coding.
  2. Define coding: binary as a threshold or category mapping (e.g., HE_glu >= 126 → diabetes = 1); categorical as level definitions (e.g., smoking: current/former/never). Outcome definitions MUST include physician diagnosis, because lab-only definitions systematically overestimate exposure→outcome associations: Diabetes = FPG≥126 OR HbA1c≥6.5 OR physician-diagnosed (KNHANES: DE1_dg=1, NHANES: DIQ010="Yes"); Hypertension = SBP≥140 OR DBP≥90 OR physician-diagnosed (KNHANES: DI1_dg=1, NHANES: BPQ020="Yes").
  3. Set covariates: the full 8-covariate set (age, sex, education, income, smoking, alcohol, obesity, CVD) is the default unless explicitly justified — minimal models (age+sex+BMI only) leave residual confounding, which can bias effects in either direction. Remove self-adjustment: if the exposure is (or derives from) a covariate, drop that covariate (exposure = BMI → drop obesity; exposure = education/income → drop the same variable); if outcome = MetS, consider dropping obesity. Document every removal in the matrix Notes.
  4. Output: combination matrix (combination_matrix.csv) with all variable specifications.
| # | Exposure | Exposure Coding | Outcome | Outcome Coding | Covariates (adjusted) | Notes |
|---|----------|-----------------|---------|----------------|----------------------|-------|
| 1 | Depression (PHQ≥10) | BP_PHQ sum ≥10 | Diabetes | HE_glu≥126|HbA1c≥6.5|DE1_dg=1 | age,sex,edu,income,smoking,alcohol,obesity,CVD | — |
| 2 | Obesity (BMI≥25) | HE_obe ≥4 | Diabetes | same | age,sex,edu,income,smoking,alcohol,depression,CVD | obesity removed from covariates |
Show full SKILL.md (430 more words)Show less
Phase 3: Batch Code Generation

Never modify the core methodology across combinations — only swap exposure/outcome/covariates. For each combination in the matrix:

  1. Clone the template script.
  2. Replace slot variables with the combination-specific values.
  3. Apply the combination's adjusted covariate set from the matrix.
  4. Set output paths: each combination gets its own results subdirectory.
  5. Generate a master runner script (run_all.R or run_all.sh) that executes all N scripts sequentially (or in parallel via future/parallel), captures errors per script without stopping the batch, and logs execution time per analysis.
  6. Generated-code quality gate: one reproducibility slip (a missing seed, an absolute path, a hand-typed data literal) replicates across the whole batch, so lint the generated scripts with the /analyze-stats code-quality gate (python3 ${CLAUDE_SKILL_DIR}/../analyze-stats/scripts/check_generated_code.py --code-dir {batch_dir} --strict) and clear every Major (MISSING_SEED, HARDCODED_DATA_LITERAL, HARDCODED_ABS_PATH, INPLACE_SOURCE_OVERWRITE) before batch execution.
Phase 4: Batch Execution (if execute or full mode)
  1. Event count check: before running, verify ≥10 outcome events per model parameter (count the exposure and every dummy/spline term; use min(events, non-events)), in the full sample and in each subgroup that is modelled. Flag underpowered combinations.
  2. Run the master script.
  3. Collect results from each combination's output directory.
  4. Log which combinations failed and why (common: convergence issues, too few events, empty subgroups) in failed_runs.csv, and suggest fixes.
Phase 5: Summary Matrix

Aggregate all results into a single summary. Every number comes from a combination's executed results files; never fill a cell by hand.

Main Results Matrix (summary_matrix.csv):

ExposureOutcomeNEventsModel 1 OR (95% CI)Model 2 OR (95% CI)Model 3 OR (95% CI)p-valueSignificant

Subgroup Summary (subgroup_matrix.csv): Same format, stratified by subgroup variables.

Heatmap (optional): Visual matrix of effect sizes × significance, exposure on Y-axis, outcome on X-axis.

  • Multiple comparisons: whenever more than one combination is tested, include a Bonferroni- (or Holm-) corrected significance column in the summary matrix, with m = the number of tests in the reported family (combinations × models × subgroups), state m, and add a note on exploratory vs confirmatory framing.
  • No p-hacking framing: the summary matrix is for hypothesis generation, not confirmation. State this explicitly in README and any manuscript output. Any citation there needs a /search-lit-confirmed DOI/PMID; otherwise mark it [UNVERIFIED - NEEDS MANUAL CHECK].

Output Files

{working_dir}/batch_{timestamp}/
├── README.md                    — Batch run summary (N combinations, template used, date)
├── combination_matrix.csv       — All exposure/outcome specs with coding
├── template/
│   └── base_template.R          — The validated template (frozen copy)
├── scripts/
│   ├── 01_depression_diabetes.R
│   ├── 02_obesity_diabetes.R
│   ├── ...
│   └── run_all.R                — Master execution script
├── results/
│   ├── 01_depression_diabetes/
│   │   ├── table1.csv
│   │   ├── main_results.csv
│   │   └── subgroup_results.csv
│   └── ...
├── summary/
│   ├── summary_matrix.csv       — Main results across all combinations
│   ├── subgroup_matrix.csv      — Subgroup results across all combinations
│   ├── failed_runs.csv          — Combinations that failed + error messages
│   └── heatmap.png              — Optional effect size × significance visual
└── logs/
    └── batch_execution.log      — Timing + error log

Reproducibility: freeze the template version (template/) and include a SHA256 hash of the data file in README.

Cross-National Batch Mode

When cross_national: true:

  • Generate paired scripts for each combination (Korea + US), using /cross-national's dual-survey-design approach
  • Summary matrix includes both countries side-by-side, with a ratio-of-ORs column (95% CI; P) computed as in /cross-national Phase 4, not a direction-agreement tick

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/batch-cohort of Aperivue/medsci-skills.

  • SKILL.md
  • skill.yml

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Batch Cohort next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Batch Cohort compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Batch Cohort this skillAperivue/medsci-skills329—~2.4kAutomated safety check: PassMIT
Retentioneering Contributingretentioneering/retentioneering-tools920—~1.8kAutomated safety check: PassApache-2.0
Retentioneering Product Analyticsretentioneering/retentioneering-tools920—~1.6kAutomated safety check: PassApache-2.0
A/B Test Analysisphuryn/pm-skills27k—~893Automated safety check: PassMIT
Cohort Analysiskillvxk/pm-skills-zh168—~562Automated safety check: PassMIT
Opik Analytics Instrumentationcomet-ml/opik22k—~4.4kAutomated safety check: PassApache-2.0

Similar skills

  • Retentioneering Contributing

    retentioneering/retentioneering-tools

    Help the user turn their Retentioneering ideas, friction reports, bug findings, or feature needs into high-quality upstream contributions: from capturing and validating the idea, through minimal…

    920 GitHub stars~1.8k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Retentioneering Product Analytics

    retentioneering/retentioneering-tools

    Analyze event logs, clickstreams, user paths, product funnels, retention, behavioral segments, transition graphs, step matrices, sequence patterns, and customer journeys using Retentioneering.

    920 GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • A/B Test Analysis

    phuryn/pm-skills

    Validates an experiment's setup, works out lift, p-value and confidence interval from A/B test data, and recommends whether to ship, extend or stop.

    27k GitHub stars~893 tokensUpdated 24 days ago
    Data & AnalyticsAuto-check passed
  • Cohort Analysis

    killvxk/pm-skills-zh

    对用户参与度数据执行同期群分析——留存曲线、功能采用趋势及分层洞察。适用于按同期群分析用户留存、研究功能随时间的采用情况、调查流失规律,或识别参与度趋势。

    168 GitHub stars~562 tokensUpdated 6 mo ago
    Data & AnalyticsAuto-check passed
  • Shows how to add product analytics events to Opik's frontend, Java backend and Python SDK, all reporting through Segment to PostHog with an opik_ name prefix.

    22k GitHub stars~4.4k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Model Assessment

    Aperivue/medsci-skills

    A skill your agent uses when validating or evaluating a trained medical-imaging model.

    329 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 3 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Verify Refs

    Aperivue/medsci-skills

    A skill your agent uses when checking whether a manuscript's references are real.

    329 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    329 GitHub stars~3.9k tokensUpdated 3 days ago
    Auto-check passed

Works with

Questions about Batch Cohort

What does Batch Cohort do?

A skill your agent uses when one validated cohort analysis must be repeated across many exposure/outcome pairs. Batch Cohort is an agent skill from Aperivue/medsci-skills. Use when one validated cohort analysis must be repeated across many exposure/outcome pairs.

When should I use Batch Cohort?

Batch Cohort fits situations like: one validated cohort analysis must be repeated across many exposure/outcome pairs; tasks that involve Product analytics.

How do I install Batch Cohort in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill batch-cohort -a claude-code`. Or copy the skill folder (skills/batch-cohort in Aperivue/medsci-skills) into .claude/skills/batch-cohort in your project. Claude Code loads it when a task matches its description.

How do I install Batch Cohort in Codex?

Run `npx skills add Aperivue/medsci-skills --skill batch-cohort -a codex`. Or copy the skill folder (skills/batch-cohort in Aperivue/medsci-skills) into .agents/skills/batch-cohort in your project. Codex loads it when a task matches its description.

Can I use Batch Cohort in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill batch-cohort -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/batch-cohort, .gemini/skills/batch-cohort, .github/skills/batch-cohort and .opencode/skills/batch-cohort in your project.

What does Batch Cohort need to run?

Going by SKILL.md and its folder, Batch Cohort needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Batch Cohort access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Batch Cohort safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Batch Cohort use?

Batch Cohort is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Batch Cohort use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Batch Cohort?

Skills that share tags, products or a category with Batch Cohort: Retentioneering Contributing (retentioneering/retentioneering-tools, 920 stars), Retentioneering Product Analytics (retentioneering/retentioneering-tools, 920 stars), A/B Test Analysis (phuryn/pm-skills, 27k stars) and Cohort Analysis (killvxk/pm-skills-zh, 168 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Batch Cohort?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.