Agent skill

Design Study

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

MITAuto-check passedResearch & Science

Install Design Study

skills CLI
$ npx skills add Aperivue/medsci-skills --skill design-study -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills design-study --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/design-study .claude/skills/design-study && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
design-study
GitHub stars
329
Token cost
~3.9k tokens
SKILL.md length
1,847 words
Files
18 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

  • Works in 4 steps: Reconstruct the study → Check structural validity → Clinical framing → …
  • Checking a radiology
  • SKILL.md covers Core Review Questions, Standard Output, Workflow and Frequent Failure Modes, plus 2 more sections
  • Runs Shell and Python scripts from its folder

What it does

Design Study is an agent skill from Aperivue/medsci-skills. Use when checking a radiology or medical AI study design before drafting or submission. Reviews the analysis unit, cohort logic, leakage risks, comparator, validation strategy and reporting-guideline fit. AI-vs-expert benchmarks are /design-ai-benchmarking.

Its SKILL.md is about 3.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 21 other files, including scripts and reference files (for example `references/combine_models_ablation_design.md`, `references/dag_adjustment.md` and `references/multi_model_comparison_design.md`).

It sits in Research & Science, covering Experimental design. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • Checking a radiology
  • Medical AI study design before drafting

Example prompts

  • “/design-study”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Reconstruct the study
  2. Check structural validity
  3. Clinical framing
  4. Reporting fit

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (Shell and Python, from the files we listed), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Design Study loads about 3.9k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 1,847 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~3.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 1,847 words, ~3,880 tokens.

Download SKILL.mdSave it as .claude/skills/design-study/SKILL.md (or your agent's skills folder). This skill also uses 17 other files; get the full folder from GitHub.
name
design-study
description
Use when checking a radiology or medical AI study design before drafting or submission. Reviews the analysis unit, cohort logic, leakage risks, comparator, validation strategy and reporting-guideline fit. AI-vs-expert benchmarks are /design-ai-benchmarking.
metadata.triggers
study design, leakage check, cohort design, analysis plan, validation strategy, comparator design, bias check

Design-Study Skill

Core Review Questions

Always inspect:

  1. The exact research question.
  2. The analysis unit: patient, lesion, exam, study, phase, report.
  3. The index date or decision point.
  4. How inclusion and exclusion criteria are applied.
  5. Any information leakage.
  6. The reference standard or endpoint definition.
  7. A clinically meaningful comparator.
  8. The validation strategy.
  9. The uncertainty reporting required.
  10. The best-fitting reporting guideline.
  11. Whether exposure/outcome/covariate definitions are literature-grounded or invented ad hoc from the data dictionary. If ad hoc, defer to /define-variables before drafting Methods.

Standard Output

text
## Study Design Review
Question: ...
Study type: ...
Analysis unit: ...
Index date / prediction timepoint: ...

### Strengths
- ...

### Major validity risks
1. ...
2. ...

### Minimal fixes
- ...

### Reporting fit
- Recommended guideline: ...

### Decision
- Ready for analysis / Needs redesign / Drafting can proceed with limitations

Cite a reference only with a /search-lit-confirmed DOI or PMID; mark any other [UNVERIFIED - NEEDS MANUAL CHECK]. Never invent a clinical definition, diagnostic criterion, or guideline recommendation: flag an unconfirmed one [VERIFY] and ask the user.


Workflow

Phase 1: Reconstruct the study

Extract from the protocol, draft, slides, tables, or notes: clinical problem, intended use case, population, inputs, outputs, outcome definition, and timing of variable availability.

Gate: Present the reconstructed study summary (question, analysis unit, intended use) to the user and confirm it before proceeding — a wrong reconstruction misdirects the entire validity review.

Phase 2: Check structural validity

Read a reference only when its condition holds:

FileRead it when
references/dag_adjustment.mdconfounding control needs an explicit adjustment set
references/target_trial_emulation.mdthe design emulates a target trial
references/venue_accept_recipe.mda clinical DL / AI-validation study must decide which venue tier the achievable design can be accepted at, and the one design move that reaches the tier above (the bridge into /find-journal); skip when there is no publication-tier decision
references/combine_models_ablation_design.mdthe model is built by combining / adapting / fine-tuning existing models (nnU-Net, TotalSegmentator, SAM/MedSAM, a pretrained backbone) and the comparator must be designed as an ablation; skip for a model trained de novo
references/multi_model_comparison_design.mdthe contribution is comparing several models / architectures head-to-head and the comparison must be fair; skip for a single-model study (one model's ablation → the row above; AI-vs-human → /design-ai-benchmarking)
references/segmentation_failure_characterization_design.mdthe claim is that a segmentation model is clinically usable, not that it scores well; skip when the endpoint is benchmark accuracy (metric choice → /model-assessment; abstention / risk–coverage → /model-assessment)
A. Analysis unit

Look for mismatches: a patient-level claim from lesion-level analysis; an exam-level split with patient overlap; phase-level samples treated as independent.

B. Leakage

Look for:

  • postoperative features used for preoperative prediction
  • normalization or thresholding performed before the data split
  • repeated exams across train/test
  • reader annotations derived from outcome information
  • input-text contamination for NLP/LLM extraction tasks: if the input includes clinical history, indication, impression, prior diagnosis, or referral text, confirm it does not name or strongly imply the target label. If it does, the task is information retrieval under label leakage, not phenotype inference: redesign the input mask, report a sensitivity analysis excluding those fields, or reframe.
  • construct dependence (a predictor that is a definitional component of the outcome): (i) an input that computes the outcome (fasting insulin and glucose when the outcome is HOMA-IR), or (ii) a near-tautological composite of the outcome's defining components (inflated, near-circular association). Test: "could this predictor be derived, in whole or part, from the outcome's definition or the same measurement?" If yes, exclude it, or keep it only as a labeled calibration probe.
F. Time origin & survivorship (incident / transition models)

For any time-to-event or incident/transition design, check before drafting:

  • Time origin per model. Each incident model starts its at-risk clock at the correct origin. Watch for immortal-time bias (a span in which the event cannot occur, misattributed to one group) and left-truncation / delayed entry.
  • Mediator-ascertainment-window survivorship. A "progressor" / transition label conditional on surviving to a later ascertainment (a second scan, a follow-up visit) is survivorship-biased; plan a landmark time or an explicit intermediate-state (multistate / illness-death) model.
  • Primary-analysis-set selection. If the primary is not the full cohort, pre-specify why, from the missingness mechanism relative to the outcome: complete-case regression is unbiased when missingness is independent of the outcome given the covariates (even with a covariate MNAR) and generally biased when it depends on the outcome, including under MAR, where multiple imputation is the usual primary (Hughes et al., Int J Epidemiol 2019; logistic-regression odds ratios are the exception when missingness depends on the outcome alone). Never pick it for significance.
  • A design that cannot yet answer these should say so — but at review time a Methods/Limitations admission that it was "not formally assessed" is escalated to MAJOR by the survival probe (S1).
C. Reference standard

Check who established ground truth, when, whether blinding was possible, and whether only a subset had gold-standard verification. Then:

  • Construct ↔ nominal-definition match. Restate each construct's nominal definition and confirm every included case satisfies it. An "incidentaloma" defined as an indeterminate finding must not include frank malignancy reads; a label that overshoots its definition inflates the cohort and breaks the κ.
  • Per-flag reference-standard concordance. Report concordance per flag category, not only overall; a construct where most flags (e.g., ~86%) miss the reference standard measures something else.
  • Manuscript definition ↔ variable_operationalization.md. Methods definitions must match the /define-variables operationalization table verbatim, and a blinded re-classification form must quote the analytic protocol's definition verbatim. A paraphrase or "common-sense extension" in the form is the documented cause of a low κ that is a definition mismatch, not real disagreement.
D. Validation

Classify: apparent only, internal split, cross-validation, temporal, external, or multi-center external.

E. Reader / expert-elicitation studies

When the study elicits expert ratings — a reader study, an annotation panel, an AI-output evaluation — read references/reader_elicitation_design.md before data collection. The acceptance ceiling of a perceptual / reader AI study is fixed at design time: no quality of execution lifts a ceiling baked into the comparator, the estimand, or the reader cohort. For an AI-system-versus-human-expert benchmark, route to /design-ai-benchmarking (arm definition, LLM-as-judge versus human-as-judge adjudication, structured export schema).

Show full SKILL.md (903 more words)Show less
Phase 3: Clinical framing

Ask whether the comparator and endpoint support the stated claim:

  • Is the model better than current practice, or just another model?
  • Is the endpoint clinically meaningful?
  • Does performance translate to action?
  • incremental value: if the model/marker is framed as adding value beyond / on top of / incremental to an existing tool (a clinical score, a routine test, a baseline model), pre-specify the baseline comparator built from the in-routine-use predictors and how added value will be shown: test it with the likelihood-ratio (or Wald) test of the new term in the development data; report ΔC-index / ΔAUC with a paired CI as the size of the gain, not as a second test (a DeLong test of nested models fitted and evaluated on the same data is invalid; DeLong is for two models frozen before a held-out evaluation); prefer the change in decision-curve net benefit at a pre-specified threshold; NRI only as a pre-specified categorical NRI reported separately for events and non-events, never the continuous NRI (analyze-stats incremental_value.md; Kerr et al., Epidemiology 2014). A standalone discrimination number does not support a "beyond X" claim, and the nested-model comparison cannot be added post hoc without the baseline model.
  • fine-tuning contribution baseline: if an NLP/LLM study claims that fine-tuning, LoRA, prompt engineering, or a multi-agent wrapper improves extraction/classification, pre-specify a same-backbone zero-shot or few-shot comparator on the identical input, output schema, and test split; a weaker or unrelated baseline cannot show the adaptation adds value. For imaging:
    • a model built by combining / adapting / fine-tuning existing models → the full ablation ladder (un-adapted base, best single component, direct-train vs transfer), per references/combine_models_ablation_design.md;
    • a head-to-head comparison of several models → comparison fairness: one frozen split/preprocessing through every model, a strong fairly-tuned baseline, a matched (or disclosed) compute budget, and a paired delta test, per references/multi_model_comparison_design.md;
    • a claim that a segmentation model is clinically usable → a pre-specified failure taxonomy, an acceptability endpoint with a named judge and adjudication rule, the tail beside the mean, and edit effort paired against manual-from-scratch, per references/segmentation_failure_characterization_design.md. A mean DSC cannot be converted into a usability claim after the fact.
  • endpoint↔conclusion scope: decide up front what kind of conclusion the design can support. A cross-sectional / single-visit / prevalence design cannot support a prognostic or surveillance claim (rescreen interval, disease progression) — that needs longitudinal follow-up. A binary surrogate endpoint (present/absent, >0, dichotomized) is risk stratification, not a patient-care directive (defer/withhold/initiate therapy). At review time /self-review §D + check_scope_coherence.py flag CROSS_SECTIONAL_PROGNOSTIC / SURROGATE_CARE_DIRECTIVE against the conclusion.
Phase 4: Reporting fit

Recommend one primary guideline — TRIPOD+AI, CLAIM, STARD, STROBE (TARGET for a target-trial emulation), PRISMA, CARE, or ARRIVE — plus journal-specific additions if needed.


Frequent Failure Modes

Diagnostic AI

No clinically relevant comparator; exam-level instead of patient-level split; unclear reference standard; AUROC-only reporting without threshold metrics.

Prognostic modeling

Unclear time zero; immortal time bias; feature timing mismatch; no calibration.

Retrospective cohort / screening database
  • Time zero misalignment: cohort entry ≠ follow-up start → immortal time bias.
  • Interval-censored outcomes treated as exact → underestimated event times.
  • Unacknowledged healthy-volunteer bias → inflated external-validity claims.
  • Surveillance bias from unequal follow-up frequency between groups.
  • Map each threat explicitly to one of the three bias classes (Hernán/Robins): selection bias (who enters), information bias (how measured), confounding (what else differs).
  • Comparative / causal question → emulate a target trial. For a treatment-vs-treatment, screening-vs-no-screening, or drug-A-vs-drug-B question on routinely-collected data, specify the seven target-trial components (eligibility, strategies, assignment, time zero, outcome, causal contrast, analysis plan) before extraction. This prevents immortal-time, prevalent-user, and confounding-by-indication bias and turns an association into a defensible causal contrast. New-user + active-comparator design, grace-period clone-censor-weight, and negative controls are in references/target_trial_emulation.md.
  • Confounding completeness: pre-specify the adjustment set from a DAG (not a Table-1 p < 0.05 rule), and plan an extended-adjustment sensitivity model to show whether any measured covariate that turns out imbalanced by exposure but lies outside the adjustment set leaves the primary estimate robust. Pre-screen the proposed covariates with scripts/adjustment_set_helper.py (flags mediator / collider / descendant adjustment and omitted confounders, and proposes a candidate backdoor set), then derive the minimal sufficient set with dagitty — see references/dag_adjustment.md. At review time /self-review Phase 2.5e and the O-probes in observational_confounding.md (O1–O18) check this against Table 1 — including O7 over-adjustment, O10 overlapping-subset-gradient discipline, O11 design-based weighting for complex-survey data, O12 data-driven-threshold mining, O13 (a cross-sectional mediation claim cannot order X→M→Y), and O14 (a synergy/joint-effect claim needs the additive interaction scale — RERI/AP/S — not a multiplicative-only test).
Multimodal LLM / report generation

No clear rubric for clinical correctness; benchmark labels derived from noisy reports without adjudication; unsupported claims about safety or workflow benefit; input text containing the target label or diagnosis being predicted; no same-backbone zero-shot/few-shot baseline for a fine-tuning or prompt-engineering claim.

Imaging meta-analysis

Overlapping cohorts; paired modalities analyzed as independent; heterogeneity metrics missing; zero-cell handling unspecified.


Minimal-Fix Principle

Recommend the smallest feasible repair first: clarify the claim, narrow the target population, add a limitation statement, add a clinically relevant baseline, re-run one key sensitivity analysis, or redefine the endpoint more explicitly. Escalate to redesign only when the central claim is not defensible otherwise.


Handoff Rules

  • route to analyze-stats when the design is basically sound but analysis details need refinement
  • route to check-reporting after the design is locked
  • route to self-review when the user wants a pre-submission quality check on their own manuscript
  • route back to write-paper only after the main validity risks are documented

This skill does not compute statistics, draft full manuscript prose, resolve raw data-engineering issues, or replace a full peer review when journal-facing tone is required.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 17 other files (scripts, references) in skills/design-study of Aperivue/medsci-skills.

  • SKILL.md
  • references/combine_models_ablation_design.md
  • references/dag_adjustment.md
  • references/multi_model_comparison_design.md
  • references/reader_elicitation_design.md
  • references/segmentation_failure_characterization_design.md
  • references/target_trial_emulation.md
  • references/venue_accept_recipe.md
  • scripts/adjustment_set_challenge/fixture/butterfly.json
  • scripts/adjustment_set_challenge/fixture/confounder.json
  • scripts/adjustment_set_challenge/fixture/instrument.json
  • scripts/adjustment_set_challenge/fixture/mbias.json
  • scripts/adjustment_set_challenge/fixture/mediator.json
  • scripts/adjustment_set_challenge/fixture/proxy.json
  • scripts/adjustment_set_challenge/problem.md
  • scripts/adjustment_set_challenge/verify.sh
  • scripts/adjustment_set_helper.py
  • … and 1 more

Open the folder on GitHubat commit 3b14ae2

Compare with similar skills

Design Study next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Design Study compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Design Study this skillAperivue/medsci-skills329—~3.9kAutomated safety check: PassMIT
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k23 repos~5.9kAutomated safety check: NotesMIT
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1287 repos~2.3kAutomated safety check: NotesNone
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.5k—~2.8kAutomated safety check: PassCC-BY-4.0
Research Refine PipelinezjYao36/Auto-Research-Refine1286 repos~1.4kAutomated safety check: NotesNone
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 23 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 7 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.5k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 6 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Experimental Design

    Oleafly/Oleafly

    Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so results are interpretable.

    206 GitHub starsUsed in 4 repos~3.5k tokens
    Research & ScienceAuto-check: notes

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Model Assessment

    Aperivue/medsci-skills

    A skill your agent uses when validating or evaluating a trained medical-imaging model.

    329 GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Verify Refs

    Aperivue/medsci-skills

    A skill your agent uses when checking whether a manuscript's references are real.

    329 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check passed
  • Fill Icmje Coi

    Aperivue/medsci-skills

    A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.

    329 GitHub stars~1.5k tokensUpdated 2 days ago
    Auto-check passed

Questions about Design Study

What does Design Study do?

A skill your agent uses when checking a radiology or medical AI study design before drafting or submission. Design Study is an agent skill from Aperivue/medsci-skills. Use when checking a radiology or medical AI study design before drafting or submission.

When should I use Design Study?

Design Study fits situations like: checking a radiology; medical AI study design before drafting.

How do I install Design Study in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill design-study -a claude-code`. Or copy the skill folder (skills/design-study in Aperivue/medsci-skills) into .claude/skills/design-study in your project. Claude Code loads it when a task matches its description.

How do I install Design Study in Codex?

Run `npx skills add Aperivue/medsci-skills --skill design-study -a codex`. Or copy the skill folder (skills/design-study in Aperivue/medsci-skills) into .agents/skills/design-study in your project. Codex loads it when a task matches its description.

Can I use Design Study in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill design-study -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/design-study, .gemini/skills/design-study, .github/skills/design-study and .opencode/skills/design-study in your project.

What does Design Study need to run?

Going by SKILL.md and its folder, Design Study needs a shell and Python for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does Design Study access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Design Study safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Design Study use?

Design Study is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Design Study use?

About 3.9k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 12k tokens, read only when the agent opens those files.

What are the alternatives to Design Study?

Skills that share tags, products or a category with Design Study: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.5k stars) and Research Refine Pipeline (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Design Study?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.