Agent skill

Find Cohort Gap

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry).

MITAuto-check passedResearch & Science

Install Find Cohort Gap

skills CLI
$ npx skills add Aperivue/medsci-skills --skill find-cohort-gap -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills find-cohort-gap --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/find-cohort-gap .claude/skills/find-cohort-gap && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
find-cohort-gap
GitHub stars
333
Token cost
~2.9k tokens
SKILL.md length
1,387 words
Files
8 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry).

  • Works in 7 steps: Cohort Intake → PI/CA Profiling → Intersection Matrix → …
  • Looking for research topics a longitudinal cohort database can answer (NHIS
  • SKILL.md covers Phase 0: Cohort Intake, Phase 1: PI/CA Profiling, Phase 2: Intersection Matrix and Phase 3: Literature Saturation…, plus 3 more sections
  • Runs Python and Shell scripts from its folder; calls python3, go and bash

What it does

Find Cohort Gap is an agent skill from Aperivue/medsci-skills. Use when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry). Profiles the cohort, matches PI expertise, scans literature saturation and returns ranked topic proposals with gap evidence.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts and reference files (for example `references/cohort_profile_template.md`, `references/onepager_template.md` and `references/pattern_scoring_rubric.md`).

It sits in Research & Science, covering Proposals and quotes. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • Looking for research topics a longitudinal cohort database can answer (NHIS
  • An institutional EMR

Example prompts

  • “/find-cohort-gap”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Cohort Intake
  2. PI/CA Profiling
  3. Intersection Matrix
  4. Literature Saturation Scan
  5. 6-Pattern Scoring + Comparison Table
  6. Feasibility Gate
  7. Output — Ranked Proposals + One-Pagers

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • go
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Find Cohort Gap loads about 2.9k tokens when it runs, and up to ~7.4k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 1,387 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 1,387 words, ~2,930 tokens.

Download SKILL.mdSave it as .claude/skills/find-cohort-gap/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.
name
find-cohort-gap
description
Use when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry). Profiles the cohort, matches PI expertise, scans literature saturation and returns ranked topic proposals with gap evidence.
model
opus
metadata.triggers
cohort gap, research topic, DB 주제, 코호트 갭, gap analysis, 연구주제 찾기, find research gap, 주제 발굴

Find-Cohort-Gap Skill

Output directory: user-specified (default: current working directory).

Phase 0: Cohort Intake

The cohort does not have to be one this skill has heard of. Route on what the user actually has.

The user has…Do this
A named public cohort (NHIS, UK Biobank, KNHANES, …)Fill the profile from published documentation. Cite the source for every field.
A codebook / data dictionary / CSV export of their own registry or EMR extractRun the input adapter below. This is the common case — an institutional registry or single-centre export that no public documentation describes.
A review, guideline, or preprint defining the clinical domainAttach it as domain context (--context), as a file or a URL.
Input adapter (local codebook / documents)
bash
python3 "${CLAUDE_SKILL_DIR}/scripts/build_cohort_profile.py" \
  --codebook data_dictionary.csv \
  --context narrative_review.pdf --context https://example.org/guideline \
  --cohort-name "Institutional CT registry" --out-dir .

Formats: .csv / .tsv / .json / .md / .txt (stdlib), .xlsx (needs openpyxl), .pdf (needs pdftotext). A .csv is auto-detected as a codebook (rows are variables) or a data export (the header row is the variable list). Writes cohort_profile.md + cohort_profile.json (+ context_extract.md).

Do not read the codebook yourself and summarise it. A paraphrased, merged, or invented variable poisons every downstream claim — the intersection matrix, the feasibility gate, and eventually the manuscript's Methods. The adapter enumerates variables verbatim with provenance (file:row). Read cohort_profile.md; do not re-derive it.

The adapter infers, and shows its work for, the variable cluster map, serial / repeated-measure groups (evidence for P1 Longitudinal Advantage), and endpoint candidates (evidence for P2 Endpoint Upgrade). A variable matching no cluster keyword is left unclassified — review those, since the lexicon is not exhaustive.

What the adapter cannot know — ASK, never guess

A codebook does not state any of the following; each is emitted as [UNKNOWN - ask the user]:

  1. Sample size (N at baseline, N with follow-up)
  2. Time span (enrollment period, follow-up duration, measurement intervals)
  3. Known limitations (healthy volunteer bias, attrition, missing-data patterns)
  4. Existing publications from this cohort (to avoid duplicating them)
  5. IRB status and data-access route

Collect these from the user before Phase 2, because a guessed N flows into the Phase 5 feasibility gate and makes it pass or fail for a reason unrelated to the cohort. Also confirm the setting (institution type, country, population type) and any special strengths the variable names cannot reveal — registry linkage, biobank availability, a distinctive population.

Gate: Present the cohort profile summary, including the [UNKNOWN] list and the unclassified variables. Confirm before proceeding.

Phase 1: PI/CA Profiling

Profile the intended PI or corresponding author to find topic-expertise alignment. If no PI is specified, skip this phase and use variable clusters alone in Phase 2.

  1. Search PubMed for the PI's recent publications (last 5 years) with /search-lit's E-utilities script: bash "${CLAUDE_SKILL_DIR}/../search-lit/references/pubmed_eutils.sh" search "AuthorLastName AuthorFirstInitial[Author]" 30 Extract top keyword clusters from titles/abstracts.
  2. Identify specialty signals: academic society positions (president, board member, editor), subspecialty focus areas, preferred journal tiers.
  3. Build a PI keyword map: 5-10 keyword clusters ranked by publication frequency.

Output: PI profile card (name, affiliation, top keywords, society roles, preferred journals).

Phase 2: Intersection Matrix

Cross cohort variable clusters with PI expertise to generate candidate topics.

Method

Create a matrix: rows = DB variable clusters, columns = PI keyword clusters. Score each cell 0-3:

  • 3: PI has published in this exact intersection (direct match)
  • 2: PI's subspecialty covers this area (strong relevance)
  • 1: Tangential connection (possible but needs framing)
  • 0: No connection
Candidate Generation
  1. Extract all cells scoring 2-3 as primary candidates.
  2. For cells scoring 1, apply the A-B substitution test: "Has someone published [this analysis] with [a different exposure/outcome] in a similar cohort?" If yes, substituting the PI's specialty variable creates a viable candidate.
  3. Generate 20-40 candidate topic statements in PICO format — P population from the cohort, E exposure/predictor variable(s), C comparison group, O outcome (preferably hard endpoint).
Discipline Alignment Filter

Before saturation scanning, identify the intended first author's department/specialty. The primary exposure variable must belong to that discipline (e.g., radiology first author → imaging variable as the primary exposure). Kill candidates where the primary exposure is outside the first author's discipline — a strong PI match alone is insufficient if the first author cannot claim ownership of the core variable.

Gate: Present the intersection matrix and top 20 candidates (post-discipline filter). User selects 8-12 for saturation scanning.

Phase 3: Literature Saturation Scan

For each selected candidate, determine how saturated the literature is.

Search Strategy

For each candidate:

  1. Build a PubMed query: (exposure terms) AND (outcome terms) AND (cohort OR longitudinal OR prospective)
  2. Execute search via /search-lit E-utilities.
  3. Count total results, then ask the Critical Filter question — "Has anyone published this with serial/repeated measurements?" — and check whether a meta-analysis exists. Grade:
PapersLongitudinal papersMA exists?GradeInterpretation
0-20NoBlue OceanFirst report possible. Verify the topic has audience interest.
3-100NoGreen FieldOptimal zone — established interest, longitudinal gap wide open.
11-300NoGreen Field (upgraded from Yellow)As above.
1-301+NoYellowViable only with very specific angle (unique population, novel endpoint).
>30AnyNoYellow (borderline Red)As above.
AnyAnyYes, outdated (>5 yr) or limited scopeYellowAs above.
AnyAnyYes, recentRedAvoid unless doing NMA or using truly unique data.
Show full SKILL.md (535 more words)Show less
"So What" Test

For each candidate, articulate 2-3 potential clinical implications of the findings. If you cannot state why a clinician or policymaker would care about the result, the topic fails regardless of gap score.

Output: Saturation table with grade, paper count, longitudinal gap status, and "So What" statement for each candidate.

Gate: Present saturation results. User selects 3-5 finalists for deep scoring.

Phase 4: 6-Pattern Scoring + Comparison Table

Apply the 6-Pattern framework to each finalist. Score each pattern 0 or 1.

6 Patterns (Universal)

Read the detailed rubric at ${CLAUDE_SKILL_DIR}/references/pattern_scoring_rubric.md and score each finalist on P1 Longitudinal Advantage, P2 Endpoint Upgrade, P3 Cohort Uniqueness, P4 PI-Topic Alignment (skip if no PI specified), P5 Comparison Table Gaps (3+ features unique to THIS STUDY in the table below), and P6 Complementary Design.

Comparison Table Construction

For each finalist, build a table comparing the top 3-5 existing papers against THIS STUDY, one row per feature (design, N, serial data, hard endpoint, population, ethnicity, subgroup analysis, …) — the rubric's P5 lists the construction steps and differentiator categories. Cite existing papers only with a /search-lit-confirmed DOI or PMID; otherwise mark the reference [UNVERIFIED - NEEDS MANUAL CHECK].

Score Interpretation
Total ScoreRecommendation
5-6Top-tier journal target (Lancet sub, JACC, J Hepatol level)
3-4Specialty journal target (solid publication)
0-2Restructure or kill — find a stronger angle before proceeding

Gate: Present scoring results and comparison tables. User approves final ranking.

Phase 5: Feasibility Gate

For each scored finalist, verify practical feasibility.

Checks
  1. Sample size adequacy:
    • Cox and logistic regression: minimum 10 events per predictor variable (EPV rule)
    • For large cohorts (N>100K) with many outcome events: warn that small effects become statistically significant (power depends on the number of events, not N — check the event count first), so focus on effect size thresholds (e.g., HR >1.2 or <0.8 for clinical relevance)
    • Consider negative control strategy (EPCV) for very large samples
  2. Missing data: key exposure variable <20% missing acceptable; key outcome <5% missing; if serial data, assess the attrition pattern (MCAR/MAR/MNAR)
  3. Follow-up adequacy: the outcome must have plausible latency within available follow-up — cancer outcomes minimum 5 years, CVD events minimum 3 years, mortality minimum 5 years
  4. Operational definition:
    • Can the exposure be defined from available variables?
    • For claims data: ICD codes alone = 40-60% accuracy. Require combination strategy (diagnosis + prescription + visit frequency + special codes)
    • Cross-check expected prevalence against known epidemiological data
    • Flag any clinical definition, diagnostic criterion, or guideline claim you cannot source with [VERIFY] and ask the user.
  5. IRB/ethics: is the data already IRB-approved for this type of analysis? Any additional approvals needed for data linkage?
  6. Disease Novelty Bonus (informational, not Go/No-Go): idiopathic etiology or debated mechanism → higher journal interest; established mechanism → needs stronger methodological novelty
Decision
  • Go: All checks pass.
  • Conditional Go: Minor issues solvable (e.g., missing data manageable with imputation).
  • No-Go: Fatal flaw (insufficient events, no valid endpoint, key variable unavailable).

Output: Feasibility report for each finalist with Go/Conditional/No-Go status.

Phase 6: Output — Ranked Proposals + One-Pagers

Ranked Summary Table
| Rank | Topic (PICO) | Saturation | 6-Pattern Score | Feasibility | Target Journal | Timeline |
|------|--------------|------------|-----------------|-------------|----------------|----------|
| 1 | ... | Green (0 longitudinal) | 5/6 | Go | JACC | 6 months |
| 2 | ... | Blue (0 papers) | 3/6 | Conditional | Radiology | 8 months |
One-Pager for Each Finalist

Fill every section of the template at ${CLAUDE_SKILL_DIR}/references/onepager_template.md; give the target journal with its rationale (PI alignment, scope match, gap fit). Save one-pagers as markdown files: {output_dir}/gap_proposal_{rank}_{short_topic}.md

Downstream: output feeds into /design-study → /write-paper.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 7 other files (scripts, references) in skills/find-cohort-gap of Aperivue/medsci-skills.

  • SKILL.md
  • references/cohort_profile_template.md
  • references/onepager_template.md
  • references/pattern_scoring_rubric.md
  • references/saturation_query_templates.md
  • scripts/build_cohort_profile.py
  • skill.yml
  • tests/test_cohort_profile.sh

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Find Cohort Gap next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Find Cohort Gap compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Find Cohort Gap this skillAperivue/medsci-skills333—~2.9kAutomated safety check: PassMIT
Research GrantsK-Dense-AI/claude-scientific-writer2.4k16 repos~3.6kAutomated safety check: NotesMIT
Verify Citationssickn33/agentic-awesome-skills47k1 repos~1kAutomated safety check: PassApache-2.0
Grantsborghei/Claude-Skills891—~1.8kAutomated safety check: PassMIT
Research Grantsaipoch/medical-research-skills1.9k—~2.2kAutomated safety check: PassMIT
Grant Writing Guidewentorai/research-plugins2981 repos~1.9kAutomated safety check: PassMIT

Similar skills

  • Research Grants

    K-Dense-AI/claude-scientific-writer

    Write competitive research proposals for NSF, NIH, DOE, DARPA, and Taiwan NSTC.

    2.4k GitHub starsUsed in 16 repos~3.6k tokens
    Research & ScienceAuto-check: notes
  • Verify Citations

    sickn33/agentic-awesome-skills

    Verify citations and references in a document, report, or article against real sources.

    47k GitHub starsUsed in 1 repo~1k tokens
    Research & ScienceAuto-check passed
  • Grants

    borghei/Claude-Skills

    Grant writing and proposal architecture: funder fit, proposal structure, budget design, and success-factor scoring.

    891 GitHub stars~1.8k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Research Grants

    aipoch/medical-research-skills

    Write competitive research proposals for NSF, NIH, DOE, DARPA, and Taiwan's NSTC when you need agency-compliant narratives, budgets, and review-criteria alignment for a specific solicitation/FOA/BAA.

    1.9k GitHub stars~2.2k tokensUpdated 24 days ago
    Research & ScienceAuto-check passed
  • Grant Writing Guide

    wentorai/research-plugins

    Write competitive research proposals with clear objectives and budgets

    298 GitHub starsUsed in 1 repo~1.9k tokens
    Research & ScienceAuto-check passed
  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    333 GitHub stars~1.6k tokensUpdated 5 days ago
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    333 GitHub stars~2k tokensUpdated 5 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    333 GitHub stars~3.9k tokensUpdated 5 days ago
    Auto-check passed
  • Fill Icmje Coi

    Aperivue/medsci-skills

    A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.

    333 GitHub stars~1.5k tokensUpdated 5 days ago
    Auto-check passed
  • Fill Protocol

    Aperivue/medsci-skills

    A skill your agent uses when an institutional Word form (.doc/.docx IRB protocol, ethics application, grant template) must be filled without breaking its styles, tables, fonts or page layout.

    333 GitHub stars~1.7k tokensUpdated 5 days ago
    Auto-check passed
  • Find Journal

    Aperivue/medsci-skills

    A skill your agent uses when choosing where to submit a manuscript.

    333 GitHub stars~3.4k tokensUpdated 5 days ago
    Auto-check passed

Questions about Find Cohort Gap

What does Find Cohort Gap do?

A skill your agent uses when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry). Find Cohort Gap is an agent skill from Aperivue/medsci-skills. Use when looking for research topics a longitudinal cohort database can answer (NHIS, UK Biobank, an institutional EMR or registry).

When should I use Find Cohort Gap?

Find Cohort Gap fits situations like: looking for research topics a longitudinal cohort database can answer (NHIS; an institutional EMR.

How do I install Find Cohort Gap in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill find-cohort-gap -a claude-code`. Or copy the skill folder (skills/find-cohort-gap in Aperivue/medsci-skills) into .claude/skills/find-cohort-gap in your project. Claude Code loads it when a task matches its description.

How do I install Find Cohort Gap in Codex?

Run `npx skills add Aperivue/medsci-skills --skill find-cohort-gap -a codex`. Or copy the skill folder (skills/find-cohort-gap in Aperivue/medsci-skills) into .agents/skills/find-cohort-gap in your project. Codex loads it when a task matches its description.

Can I use Find Cohort Gap in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill find-cohort-gap -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/find-cohort-gap, .gemini/skills/find-cohort-gap, .github/skills/find-cohort-gap and .opencode/skills/find-cohort-gap in your project.

What does Find Cohort Gap need to run?

Going by SKILL.md and its folder, Find Cohort Gap needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3, go and bash). Our summary lists: Python 3; A Bash shell.

Does Find Cohort Gap access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Find Cohort Gap safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Find Cohort Gap use?

Find Cohort Gap is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Find Cohort Gap use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.5k tokens, read only when the agent opens those files.

What are the alternatives to Find Cohort Gap?

Skills that share tags, products or a category with Find Cohort Gap: Research Grants (K-Dense-AI/claude-scientific-writer, 2.4k stars), Verify Citations (sickn33/agentic-awesome-skills, 47k stars), Grants (borghei/Claude-Skills, 891 stars) and Research Grants (aipoch/medical-research-skills, 1.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Find Cohort Gap?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 333 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.