Agent skill

Model Assessment

by Aperivue in Aperivue/medsci-skills

A skill your agent uses when validating or evaluating a trained medical-imaging model.

MITAuto-check passedResearch & Science

Install Model Assessment

skills CLI
$ npx skills add Aperivue/medsci-skills --skill model-assessment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Aperivue/medsci-skills model-assessment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Aperivue/medsci-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/model-assessment .claude/skills/model-assessment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-assessment
GitHub stars
329
Used in
1 other repo
Token cost
~4.5k tokens
SKILL.md length
1,754 words
Files
77 (incl. scripts, references)
Skills in repo
54
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when validating or evaluating a trained medical-imaging model.

  • Works in 12 steps: Task, intended-use horizon, analysis unit → Leakage audit (run first) → Validation tier → …
  • Evaluating a trained medical-imaging model
  • SKILL.md covers Part A — Validation design, Part B — Held-out metrics, Part C — Uncertainty, OOD and… and Part D — Explainability, plus 1 more section
  • Runs Python and Shell scripts from its folder; calls python3

What it does

Model Assessment is an agent skill from Aperivue/medsci-skills. Use when validating or evaluating a trained medical-imaging model. Audits split leakage and validation design, computes task-correct held-out metrics (Dice + HD95, AUROC + AUPRC, FROC, calibration), and covers uncertainty/OOD and Grad-CAM explainability, each with a gate.

Its SKILL.md is about 4.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 81 other files, including scripts and reference files (for example `references/explainability_guide.md`, `references/metric_guide.md` and `references/metric_selection_grounding.md`).

It sits in Research & Science, covering Clinical and healthcare research and Performance reviews. The repository describes itself as: Agent Skills for medical research — literature search, reporting-guideline & citation checks, statistics, publication figures, submission. Works with Claude Code, Codex, Cursor &… The licence is MIT.

When your agent uses it

  • Evaluating a trained medical-imaging model
  • Tasks that involve Clinical and healthcare research
  • Tasks that involve Performance reviews

Example prompts

  • “/model-assessment”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Task, intended-use horizon, analysis unit
  2. Leakage audit (run first)
  3. Validation tier
  4. Comparator
  5. Test-set sizing
  6. Prospective evaluation and deployment-monitoring horizon
  7. Reporting-guideline fit
  8. Compute task-correct metrics
  9. Gate the metric choice
  10. Choose the uncertainty method, OOD guard and abstention rule
  11. Emit and gate the uncertainty manifest
  12. Produce, sanity-check and quantify the maps

What it can do on your machine

Read from SKILL.md and the folder at commit 3b14ae2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 9 files in scripts/ (Python and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Assessment loads about 4.5k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 1,754 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~4.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from Aperivue/medsci-skills at commit 3b14ae2, republished under its MIT licence (© Aperivue). 1,754 words, ~4,475 tokens.

Download SKILL.mdSave it as .claude/skills/model-assessment/SKILL.md (or your agent's skills folder). This skill also uses 76 other files; get the full folder from GitHub.
name
model-assessment
description
Use when validating or evaluating a trained medical-imaging model. Audits split leakage and validation design, computes task-correct held-out metrics (Dice + HD95, AUROC + AUPRC, FROC, calibration), and covers uncertainty/OOD and Grad-CAM explainability, each with a gate.
metadata.triggers
model validation, validate AI model, imaging model validation, data leakage, split leakage, train test split, patient-level split, internal validation…

Model-Assessment Skill

Assess a trained medical-imaging model (in-house, vendor or open-weights; segmentation, classification or detection) in the order the evidence is built: Part A designs and audits the validation study, Part B computes task-correct held-out metrics, Part C adds the uncertainty / OOD / abstention layer a deployment claim needs, and Part D makes an explainability analysis survive review. Run the parts the request needs (a Grad-CAM question starts at Part D), but no metric headline is reported before the Part A split gate is green. Each part ends in a stdlib gate whose verdict is reproduced from a file — never report a pass without running it.

Numbers come only from code executed on the supplied predictions, split table, or the researcher's executed UQ/XAI code; if predictions or ground truth are missing, say so and stop. Integrate MONAI / nnU-Net, MAPIE, captum, pytorch-grad-cam and pretrained OOD scorers by reference — never reimplement them, never build or train the model, never run a model on real patient data.

Elsewhere: building/training → /model-scaffold; choosing or vetting the model → /model-selection; data-stage preprocessing leakage → /imaging-data; paired model comparison / added value over a baseline / decision curves / MRMC / ICC / calibration tables → /analyze-stats (added value: incremental_value.md); AI-vs-expert reader rubric and IRR → /design-ai-benchmarking; LLM/MLLM → /mllm-eval; general validity → /design-study; classical-ML tabular calibration → /radiomics-ml; item-by-item reporting audit → /check-reporting; a finished manuscript → /self-review or /peer-review (the MD0–MD11 model_development.md probe).

Part A — Validation design

The rationale behind Phases 2–7 — the full leakage taxonomy, the internal-vs-external ladder, comparator design, run variance, test-set sizing, the reporting map — is in ${CLAUDE_SKILL_DIR}/references/validation_design.md (load on demand). Verify citations via /search-lit (confirmed DOI/PMID), else mark [UNVERIFIED - NEEDS MANUAL CHECK]; flag an uncertain CLAIM 2024 / TRIPOD+AI / Metrics Reloaded item [VERIFY] and ask.

Phase 1 — Task, intended-use horizon, analysis unit

State the task, the intended-use horizon (screening, triage, pre-procedure, post-hoc), the single headline metric the conclusion leans on, and the analysis unit (per-patient / per-lesion / per-image). Everything downstream is read against this; a per-lesion metric is never reported as per-patient.

Phase 2 — Leakage audit (run first)

Produce the emitted split-assignment table (patient_id,split) and run:

bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_split_leakage.py \
  --splits <split_assignment.csv> --out qc/split_leakage.json --strict

PATIENT_OVERLAP (a patient in ≥ 2 partitions) and MISSING_SEED (an unreproducible split) are Major, SINGLE_PARTITION Minor — proven by set arithmetic on the ID column the gate prints as id_col (pass --id-col with the patient identifier when it is not the one; an auto-picked column whose name is not patient-level, such as image_id, gets a Minor ID_COL_NOT_PATIENT_LEVEL). A design with patient overlap is never approved. Then walk the leakage the table cannot show (Kapoor & Narayanan, Patterns 2023): preprocessing fit before the split (normalisation, resampling, foundation-model embeddings, ComBat harmonisation over the whole cohort — /imaging-data gates the declared pipeline), site / scanner / burned-in-label shortcuts, and temporal leakage (a random split where future and past coexist). The decisive question: could any value used in training have been computed only with knowledge of a test case?

Phase 3 — Validation tier

Classify honestly: apparent → internal random split → cross-validation → temporal → geographic / external (different site, scanner, vendor) → multi-site external. Cross-validation and bootstrap are development-time optimism corrections, not external validation. Flag a generalisability or deployment claim that outruns the design, and "external validation" where the single external set was used for tuning. Confirm the test set was touched once — no architecture search, hyperparameter sweep, early stopping, or threshold choice read it.

Phase 4 — Comparator

Clinical-only baseline, incremental value over an existing score, or reader comparison (hand the rubric / inter-rater design to /design-ai-benchmarking).

Phase 5 — Test-set sizing

Count events per class in the test set, not the cohort total: a sparse positive set gives a CI spanning much of the usable range, and calibration needs roughly ≥ 100 events. Hand formal sizing to /calc-sample-size.

Phase 6 — Prospective evaluation and deployment-monitoring horizon

Retrospective external validation shows accuracy transfers, not that the model is safe and useful in the workflow. For a clinical-use claim design the higher tier explicitly — silent / shadow deployment → prospective comparative or impact study / RCT on a clinical endpoint → post-deployment monitoring with recalibration-or-withdrawal triggers and subgroup audit (references/validation_design.md §2b). Scope the claim to the tier reached: a retrospective external study never claims deployment readiness or outcome benefit.

Phase 7 — Reporting-guideline fit

Map via /check-reporting: CLAIM 2024 (diagnostic imaging AI), TRIPOD+AI (prediction model), STARD-AI (diagnostic accuracy), PROBAST+AI (risk of bias), and for a prospective/live evaluation DECIDE-AI or CONSORT-AI / SPIRIT-AI.

Part B — Held-out metrics

Phase 8 — Compute task-correct metrics

Generate and execute evaluation code on the held-out predictions (Metrics Reloaded — Maier-Hein & Reinke et al., Nat Methods 2024):

  • segmentation — Dice/IoU and a boundary metric (HD95 / NSD), per structure, 95% CIs by patient-level bootstrap (resample patients, not pixels or slices);
  • classification — AUROC and AUPRC with patient-level bootstrap CIs, sensitivity/specificity, and PPV/NPV at the deployment prevalence, never bare accuracy on a balanced set. Report AUPRC with the test-set prevalence, which is its no-skill value: AUPRC moves with prevalence, so a value from an enriched or case-control test set does not carry over to deployment or across datasets. For multiclass, state the aggregation (one-vs-rest / macro / micro / pairwise / Obuchowski);
  • detection — FROC / mAP with the IoU match criterion stated. Lesions and false positives cluster within patients, so CIs come from a patient-level bootstrap (resample patients, carrying all their lesions and false positives), not a Wilson/binomial interval over lesions, which is too narrow; compare two detectors' FROC curves with JAFROC (RJafroc), not per-lesion tests;
  • interactive / promptable segmentation (SAM2 / MedSAM2 / nnInteractive) — the segmentation metrics plus Dice-vs-interactions / number-of-clicks (NoC) to a target threshold, initial-vs-converged (or peak) Dice, and per-case interaction/inference time. With two arms (simulated prompting + human operator), record protocol fidelity — identical prompt types, stopping rule, target threshold, seeds — because arm-to-arm comparability is what lets the human arm validate the simulated one (human-arm design: /design-study);
  • generative / synthesis — full-reference (MSE/RMSE/PSNR/SSIM) or no-reference (SNR/CNR, visual scores) quality plus a downstream-task evaluation: image quality is not clinical utility (Park et al., Radiol Med 2024);
  • time-to-event discrimination (Harrell's C, time-dependent ROC) → /analyze-stats.

Report the headline as the point estimate with a patient-level bootstrap 95% CI over the test cases: that is the uncertainty of the test-set estimate. Seed-to-seed SD across training runs is a different quantity (training-run variability, usually smaller) — report it separately, over ≥ 5 runs, for a training-recipe or model-comparison claim, and never present it as the CI. A frozen vendor or open-weights model has no training runs to vary; its uncertainty is the test-set CI. Add calibration — for a binary risk or diagnostic output, calibration-in-the-large (intercept), the calibration slope and a flexible (loess) calibration curve, plus the Brier score; ECE only as a supplementary top-label summary for multi-class confidence, with its binning stated — and subgroup slices (the Model Card Factors). Emit results.md (metrics report) and a per-case CSV for /analyze-stats. Load ${CLAUDE_SKILL_DIR}/references/metric_guide.md for the per-task checklist and ${CLAUDE_SKILL_DIR}/references/metric_selection_grounding.md for why each pairing is required and the CLAIM 2024 fit map.

Show full SKILL.md (641 more words)Show less
Phase 9 — Gate the metric choice

Declare the reported metrics in metrics_manifest.json (copy ${CLAUDE_SKILL_DIR}/templates/metrics_manifest.json; fields and allowed values in references/metrics_manifest_schema.md), then:

bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_metric_reporting.py \
  --manifest metrics_manifest.json --out qc/metric_reporting.json --strict

PIXEL_ACCURACY_SEG / NO_BOUNDARY_METRIC / ACCURACY_ONLY / DETECTION_METRIC_MISSING / INTERACTIVE_NO_INTERACTION_COUNT / GENERATIVE_NO_DOWNSTREAM / CLASSIFICATION_METRIC_MISSING / SEGMENTATION_METRIC_MISSING (every Major) must be zero; the last two (no headline metric declared) are manifest-only. An off-list value exits 2; use "none" or "other:<description>". --report results.md --task <task> still runs the older keyword check on prose.

Known limits: manifest mode checks what is declared, not the reported numbers. Prose mode (--report) tests keyword presence with a short negation window: "MSD" counts as mean surface distance even when it names the Medical Segmentation Decathlon, "we did not compute the Hausdorff distance or HD95" still counts HD95, "sensitivity and specificity were not reported" or "FROC was not performed" still count as reported, a bare "map" ("saliency map") counts as mAP, and a wrapped "mean average\nprecision" is not seen.

Part C — Uncertainty, OOD and selective prediction (deployment claims)

A deployment-framed model must say what it does when unsure or off-distribution. Read ${CLAUDE_SKILL_DIR}/references/uncertainty_guide.md for method choice and the manifest schema.

Phase 10 — Choose the uncertainty method, OOD guard and abstention rule
  • Conformal (MAPIE) — prediction sets/intervals at nominal coverage; the strongest default with a calibration set. Its coverage guarantee is finite-sample but needs exchangeability, which can fail on clinical data, so measure achieved coverage on a test split disjoint from the calibration split and report it with its binomial CI — never report it as guaranteed.
  • Deep ensemble — K ≥ 2 independent members (distinct seeds/inits); shared seeds underestimate epistemic uncertainty.
  • MC-dropout — dropout active at inference, T passes; off, every pass is identical and the estimate collapses to a point prediction.
  • Bayesian / last-layer Laplace — a light option.
  • OOD guard — energy score, feature Mahalanobis, ODIN or max-softmax, evaluated on a held-out OOD set (different scanner/site/pathology) with detection AUROC and the operating point.
  • Selective prediction — abstain at a pre-specified coverage/risk target; report the risk–coverage curve. A post-hoc threshold inflates accuracy-at-coverage.
  • Under shift — report calibration/coverage on shifted or external data, not in-distribution only (Ovadia 2019).
Phase 11 — Emit and gate the uncertainty manifest

Write uncertainty_manifest.json:

json
{
  "task": "classification",
  "deployment_claim": true,
  "uncertainty_method": "conformal",
  "coverage_target": 0.90,
  "coverage_validated": true,
  "ood_method": "mahalanobis",
  "ood_heldout_set": "external-ood-cohort",
  "selective_prediction": true,
  "selective_target": 0.95,
  "calibration_under_shift": true
}
bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_uncertainty_reporting.py --manifest uncertainty_manifest.json \
  --out qc/uncertainty_reporting.json --strict

Verdicts: POINT_PREDICTION_NO_UNCERTAINTY, CONFORMAL_NO_COVERAGE_VALIDATION, OOD_NO_HELDOUT_SET (Major); ENSEMBLE_NOT_INDEPENDENT, MCDROPOUT_DISABLED_AT_INFERENCE, SELECTIVE_NO_TARGET, NO_CALIBRATION_UNDER_SHIFT (Minor). It audits the declared spec; it complements, not replaces, Phase 8's executed calibration. Report TRIPOD+AI / DECIDE-AI deployment-monitoring fit via /check-reporting.

Part D — Explainability

A saliency / Grad-CAM map is the most over-interpreted artifact in imaging AI: Adebayo et al. (NeurIPS 2018) showed many methods produce convincing maps independent of the model's weights and labels. Read ${CLAUDE_SKILL_DIR}/references/explainability_guide.md for method by architecture, sanity checks, localisation metrics and framing.

Phase 12 — Produce, sanity-check and quantify the maps

Choose the method for the architecture — Grad-CAM / Grad-CAM++ for CNNs, attention rollout for ViTs, integrated gradients / SHAP for attribution — wired through captum or pytorch-grad-cam. Run the Adebayo model-parameter and data (label) randomisation tests; a faithful map degrades when they are randomised, and both axes are the minimum bar. If the map is claimed to localise the finding, compute IoU / pointing game / Dice against ground-truth masks over the cohort — not eyeballed, cherry-picked cases. Frame a map as attribution ("where signal is attributed"), never as proof the model is correct or of causation.

Phase 13 — Emit and gate the explainability report

Write explainability_report.json:

json
{
  "method": "grad-cam++",
  "n_examples": 200,
  "cohort_level": true,
  "localization_metric": "iou",
  "localization_value": 0.63,
  "sanity_checks": ["model_randomization", "data_randomization"],
  "interpretation": "localization"
}

interpretation: attribution / localization / faithfulness — never validation / causal.

bash
python3 ${CLAUDE_SKILL_DIR}/scripts/check_explainability_report.py --manifest explainability_report.json \
  --out qc/explainability_report.json --strict

Verdicts: SALIENCY_AS_VALIDATION, NO_SANITY_CHECK, NO_LOCALIZATION_METRIC (Major); INSUFFICIENT_SANITY, CHERRY_PICKED_EXAMPLES, MISSING_METHOD (Minor).

Outputs and hand-off

  • Part A: validation-design decision notes (leakage, tier, comparator, metric, sizing handoff, reporting fit) and qc/split_leakage.json.
  • Part B: results.md, the per-case CSV, and qc/metric_reporting.json.
  • Part C: uncertainty_manifest.json + qc/uncertainty_reporting.json.
  • Part D: explainability_report.json + qc/explainability_report.json.

The per-case table → /analyze-stats (paired ΔAUC of frozen models on the same test patients — DeLong or bootstrap; added value over a baseline per incremental_value.md; decision curves; publication tables); figures → /make-figures; numbers and subgroup performance → /model-card; Methods/Results → /write-paper; compliance → /check-reporting; sizing → /calc-sample-size; the reviewer-side audit of the draft → /self-review, whose ai_overclaiming / image_synthesis probes also check saliency claims. Gate regression (${CLAUDE_SKILL_DIR}/): scripts/check_split_leakage_challenge/verify.sh, scripts/metric_reporting_challenge/verify.sh, scripts/check_uncertainty_reporting_challenge/verify.sh, scripts/check_explainability_report_challenge/verify.sh, and tests/test_*.sh.

© Aperivue, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 76 other files (scripts, references) in skills/model-assessment of Aperivue/medsci-skills.

  • SKILL.md
  • references/explainability_guide.md
  • references/metric_guide.md
  • references/metric_selection_grounding.md
  • references/metrics_manifest_schema.md
  • references/uncertainty_guide.md
  • references/validation_design.md
  • scripts/check_explainability_report.py
  • scripts/check_explainability_report_challenge/expected/strong.txt
  • scripts/check_explainability_report_challenge/expected/weak.txt
  • scripts/check_explainability_report_challenge/fixture/report_strong.json
  • scripts/check_explainability_report_challenge/fixture/report_weak.json
  • scripts/check_explainability_report_challenge/problem.md
  • scripts/check_explainability_report_challenge/verify.sh
  • scripts/check_metric_reporting.py
  • scripts/check_split_leakage.py
  • … and 61 more

Open the folder on GitHubat commit 3b14ae2

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in Aperivue/medsci-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Model Assessment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Assessment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Assessment this skillAperivue/medsci-skills3291 repos~4.5kAutomated safety check: PassMIT
Clinical Trials Databasegoogle-deepmind/science-skills3.2k3 repos~3.2kAutomated safety check: PassApache-2.0
CHARLS Paper Reproduction Guidexjtulyc/MedgeClaw6171 repos~1.8kAutomated safety check: PassNone
Biomedical Analysis Dispatchxjtulyc/MedgeClaw6171 repos~2kAutomated safety check: PassNone
Research Proposalluwill/research-skills858—~4.4kAutomated safety check: NotesNone
Medical Imaging ReviewLeonChaoX/qinyan-academic-skills9383 repos~1.1kAutomated safety check: NotesMIT

Similar skills

  • Clinical Trials Database

    google-deepmind/science-skills

    Query ClinicalTrials.gov via APIv2. An agent skill from google-deepmind/science-skills.

    3.2k GitHub starsUsed in 3 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed
  • Routes bioinformatics, drug discovery, clinical and multi-omics tasks from a chat interface to Claude Code sessions running K-Dense scientific skills, with a live dashboard per task.

    617 GitHub starsUsed in 1 repo~2k tokens
    Research & ScienceAuto-check passed
  • Research Proposal

    luwill/research-skills

    A skill your agent uses when the user asks to write or draft a PhD / doctoral research proposal, research plan, 研究计划书, or 开题报告 — a forward-looking plan of background, gap, research questions…

    858 GitHub stars~4.4k tokensUpdated 8 days ago
    Research & ScienceAuto-check: notes
  • Medical Imaging Review

    LeonChaoX/qinyan-academic-skills

    Write comprehensive literature reviews for medical imaging AI research.

    938 GitHub starsUsed in 3 repos~1.1k tokens
    Research & ScienceAuto-check: notes
  • Fhir API

    aehrc/pathling

    Expert guidance for implementing FHIR RESTful API servers and clients following the HL7 FHIR specification.

    137 GitHub starsUsed in 1 repo~1.5k tokens
    Research & ScienceAuto-check passed

More from Aperivue/medsci-skills

All 54 skills in this repo
  • Obsidian Paper Vault

    Aperivue/medsci-skills

    A skill your agent uses when turning a folder of research PDFs into Obsidian notes, even if Obsidian is not named.

    329 GitHub stars~1.6k tokensUpdated 3 days ago
    Auto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed
  • Verify Refs

    Aperivue/medsci-skills

    A skill your agent uses when checking whether a manuscript's references are real.

    329 GitHub starsUsed in 1 repo~3.1k tokens
    Auto-check passed
  • Clean Data

    Aperivue/medsci-skills

    A skill your agent uses when a clinical CSV/Excel dataset needs profiling and cleaning before analysis (missing values, outliers, duplicates, type mismatches).

    329 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Design Study

    Aperivue/medsci-skills

    A skill your agent uses when checking a radiology or medical AI study design before drafting or submission.

    329 GitHub stars~3.9k tokensUpdated 3 days ago
    Auto-check passed
  • Fill Icmje Coi

    Aperivue/medsci-skills

    A skill your agent uses when each author needs an ICMJE Conflict of Interest disclosure form (coidisclosure.docx) for submission.

    329 GitHub stars~1.5k tokensUpdated 3 days ago
    Auto-check passed

Questions about Model Assessment

What does Model Assessment do?

A skill your agent uses when validating or evaluating a trained medical-imaging model. Model Assessment is an agent skill from Aperivue/medsci-skills. Use when validating or evaluating a trained medical-imaging model.

When should I use Model Assessment?

Model Assessment fits situations like: evaluating a trained medical-imaging model; tasks that involve Clinical and healthcare research; tasks that involve Performance reviews.

How do I install Model Assessment in Claude Code?

Run `npx skills add Aperivue/medsci-skills --skill model-assessment -a claude-code`. Or copy the skill folder (skills/model-assessment in Aperivue/medsci-skills) into .claude/skills/model-assessment in your project. Claude Code loads it when a task matches its description.

How do I install Model Assessment in Codex?

Run `npx skills add Aperivue/medsci-skills --skill model-assessment -a codex`. Or copy the skill folder (skills/model-assessment in Aperivue/medsci-skills) into .agents/skills/model-assessment in your project. Codex loads it when a task matches its description.

Can I use Model Assessment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Aperivue/medsci-skills --skill model-assessment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-assessment, .gemini/skills/model-assessment, .github/skills/model-assessment and .opencode/skills/model-assessment in your project.

What does Model Assessment need to run?

Going by SKILL.md and its folder, Model Assessment needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3; A Bash shell.

Does Model Assessment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Assessment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Model Assessment use?

Model Assessment is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Assessment use?

About 4.5k tokens (SKILL.md is roughly 18k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to Model Assessment?

Skills that share tags, products or a category with Model Assessment: Clinical Trials Database (google-deepmind/science-skills, 3.2k stars), CHARLS Paper Reproduction Guide (xjtulyc/MedgeClaw, 617 stars), Biomedical Analysis Dispatch (xjtulyc/MedgeClaw, 617 stars) and Research Proposal (luwill/research-skills, 858 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Assessment?

Aperivue (a GitHub organization) maintains it in Aperivue/medsci-skills, which has 329 GitHub stars. The repository holds 54 skills in this directory. The repository was last updated on October 5, 2026.

Source: Aperivue/medsci-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.