Agent skill

Bio Machine Learning Model Validation

by GPTomics in GPTomics/bioSkills

Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net…

MITAuto-check passedData & Analytics

Install Bio Machine Learning Model Validation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-machine-learning-model-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/machine-learning/model-validation .claude/skills/bio-machine-learning-model-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-machine-learning-model-validation
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5k tokens
SKILL.md length
2,210 words
Files
3
Skills in repo
553
Repo updated
First seen
Licence
MIT

At a glance

Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net…

  • Estimating model performance honestly
  • SKILL.md covers Version Compatibility, The Single Most Important…, Leakage Taxonomy and Decision Tree by Scenario, plus 10 more sections
  • Runs Python scripts from its folder; calls pip
  • Choosing a CV scheme

What it does

Bio Machine Learning Model Validation is an agent skill from GPTomics/bioSkills. Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net benefit, optimism correction, sample-size planning, and TRIPOD+AI reporting. Use when estimating model performance honestly, choosing a CV scheme, detecting leakage, or judging whether reported discrimination means the model is actually useful. For feature selection itself see machine-learning/biomarker-discovery; for…

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/nested_cv_biomarker.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Machine learning. It works with scikit-learn. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Estimating model performance honestly
  • Choosing a CV scheme
  • Detecting leakage
  • Judging whether reported discrimination means the model is actually useful

Example prompts

  • “Use the bio-machine-learning-model-validation skill to validate predictive models on omics and biomedical data with nested cross-validation…”
  • “/bio-machine-learning-model-validation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Machine Learning Model Validation loads about 5k tokens when it runs. Until then it costs about 157 tokens; SKILL.md has 2,210 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~157
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,210 words, ~5,000 tokens.

Download SKILL.mdSave it as .claude/skills/bio-machine-learning-model-validation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-machine-learning-model-validation
description
Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net benefit, optimism correction, sample-size planning, and TRIPOD+AI reporting. Use when estimating model performance honestly, choosing a CV scheme, detecting leakage, or judging whether reported discrimination means the model is actually useful. For feature selection itself see machine-learning/biomarker-discovery; for confirmatory-trial inference see clinical-biostatistics/trial-reporting.
tool_type
python
primary_tool
sklearn

Version Compatibility

Reference examples tested with: numpy 1.26+, scikit-learn 1.4+ (note 1.6/1.8 API changes below).

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

scikit-learn drift to watch: CalibratedClassifierCV(cv='prefit') was deprecated in 1.6 and removed in 1.8 (it now raises; wrap a fitted model in sklearn.frozen.FrozenEstimator instead); ensemble default became 'auto' in 1.6; method='temperature' was added in 1.8. If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Model Validation for Biomedical and Omics Data

"Validate my omics classifier honestly" -> Keep every data-dependent step inside the resampling loop, never use the same data to both choose and grade, and report calibration and net benefit, not just AUC.

  • Nested CV: GridSearchCV (inner) wrapped by cross_val_score (outer)
  • Group/structured: StratifiedGroupKFold, TimeSeriesSplit
  • Calibration: calibration_curve, brier_score_loss, CalibratedClassifierCV

The Single Most Important Modern Insight -- A Reported Number Is an Honest Estimate Only If Nothing Leaked and Nothing Was Graded on What Was Chosen

A reported performance number is a claim about a data-generating process that will never recur. Almost every inflated result in ML-for-biology traces to one of two root causes: information from the test distribution leaked into model construction, or the same data was used to both choose and grade a decision. A clean train/test split is necessary but nowhere near sufficient -- the leakage has usually already contaminated the test set (a scaler fit on all data, a duplicate patient, ComBat run across the split). Leakage causes a reproducibility crisis across ML-based science (Kapoor 2023), and the bias is largest exactly when the true signal is weakest -- the omics regime.

A second, equally load-bearing insight: discrimination (AUC/C) and calibration (do predicted probabilities match observed frequencies) are orthogonal. AUC is invariant to any monotone transform of the score, so it is blind to calibration. For any decision that uses the probability itself, calibration -- not AUC -- is the property that matters, and it is the one routinely ignored (Van Calster 2019, "the Achilles heel").

Leakage Taxonomy

Leakage typeHow it happens in omicsSymptomPrevention
Preprocessing (most common, most missed)z-scoring, quantile/library normalization, ComBat/SVA, PCA, kNN/MICE imputation, VST fit on the full dataset before splittingTest performance suspiciously close to train; collapses on external dataFit every transform inside the CV fold via a Pipeline
Feature selection (severe special case)top-k DE genes / highest-variance / univariate filter chosen on all samples, then CV only the classifierNear-perfect CV from pure noise; unstable selected setSelection lives in the CV fold (Ambroise 2002)
Target / labela feature is a proxy for or downstream of the outcome (post-diagnosis labs, treatment-derived fields, a collection-site that tracks case/control)One feature dominates implausibly; fails when removedAudit temporal/causal admissibility; exclude post-outcome variables
Group / patient / replicatesame patient, tumor, organoid, or technical replicate in train and test; KFold scatters themInflated metrics that vanish under leave-one-group-outSplit by the highest independent unit (GroupKFold/StratifiedGroupKFold)
Batchbatch correlated with outcome and not respected in the split, or ComBat across the train/test boundaryModel discriminates batches not biology; external batch destroys itBlock the split by batch; never run unsupervised correction across the split
Temporalrandom-splitting time-ordered data; future-period statistics standardize the pastBacktest beats prospective deploymentTime-based split (TimeSeriesSplit); never shuffle first
Duplicate / homolognear-identical samples, augmented copies, public-dataset overlap, homologous sequences across the splitMemorization passes as generalizationDeduplicate / cluster-then-split before CV
Test-reuse / thresholdrepeatedly peeking to pick features, thresholds, "best epoch"; choosing the classification threshold on the test setIrreproducible SOTA; fragile configOne locked test set; all tuning + thresholds inside nested CV

Decision Tree by Scenario

Scenario / generalization questionRecommended schemeWhy
"A new sample like training" (and any tuning occurs)Nested CV: inner GridSearchCV, outer cross_val_score, Pipeline insideTuning and grading on the same CV is optimistic (Cawley-Talbot 2010)
"A new patient" (repeated measures)GroupKFold/StratifiedGroupKFold by patient/donorThe unit of independence is not the row
"A new hospital/site" (transportability)Leave-one-site-out (internal-external CV)Approximates external validation
"Next year" (time-ordered)TimeSeriesSplit forward-chainingRandom folds leak the future
Small n (dozens), need a stable estimateRepeatedStratifiedKFold (5x10) with an intervalA single CV is one high-variance draw
Probabilities will drive a decisionAdd calibration + decision-curve net benefitAUC is blind to calibration and utility
Final evidence for a clinical modelExternal/temporal validation + TRIPOD+AI reportInternal CV cannot detect a whole-dataset confound
Choosing the features themselves-> machine-learning/biomarker-discoverySelection is its own discipline (run it inside the fold)
Confirmatory trial inference (HR, p-value)-> clinical-biostatistics/trial-reportingEstimand is a treatment effect, not a prediction

Nested Cross-Validation

Goal: Estimate the performance of the whole procedure (tuning + fit) without optimistic bias.

Approach: The inner loop does all tuning, feature selection, and threshold choice; the outer loop grades the winning configuration once on a fold it never touched. The reported number is the aggregate over outer folds; it answers "if I run this pipeline on new data, what do I get?"

python
from sklearn.model_selection import cross_val_score, StratifiedKFold, GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.linear_model import LogisticRegression

pipe = Pipeline([('scaler', StandardScaler()),
                 ('select', SelectKBest(f_classif)),         # re-fit per inner fold -> no leakage
                 ('clf', LogisticRegression(max_iter=5000))])
grid = {'select__k': [10, 50, 200], 'clf__C': [0.01, 0.1, 1]}

inner = StratifiedKFold(5, shuffle=True, random_state=0)
outer = StratifiedKFold(5, shuffle=True, random_state=1)
search = GridSearchCV(pipe, grid, cv=inner, scoring='roc_auc')
scores = cross_val_score(search, X, y, cv=outer, scoring='roc_auc')   # unbiased estimate
print(f'Nested AUC: {scores.mean():.3f} +/- {scores.std():.3f}')

Nested CV is needed whenever model selection happens -- even informal "tried three options, kept the best." Flat CV with tuning is a known reviewer red flag (Varma-Simon 2006).

Group-Aware, Structured, and Small-Sample CV

Goal: Match the CV scheme to the real unit of independence and get a variance-aware estimate.

Approach: Pass a grouping vector so no group spans folds; for tiny n, repeat stratified k-fold and report the spread, not a bare number. Standard KFold assumes i.i.d. rows, which biomedical data almost never satisfy.

python
from sklearn.model_selection import StratifiedGroupKFold, RepeatedStratifiedKFold, cross_val_score

groups = meta['patient_id'].values                          # multiple samples per patient
gcv = StratifiedGroupKFold(n_splits=5)                      # group-disjoint AND class-balanced
g_auc = cross_val_score(pipe, X, y, cv=gcv, groups=groups, scoring='roc_auc')

rcv = RepeatedStratifiedKFold(n_splits=5, n_repeats=10, random_state=0)
r_auc = cross_val_score(pipe, X, y, cv=rcv, scoring='roc_auc')   # report an interval

Leave-one-out is high-variance and degenerate for ranking metrics (AUC is undefined on a size-1 test fold) -- prefer repeated stratified k-fold. The .632+ bootstrap (Efron-Tibshirani 1997) is a defensible alternative but is optimistic for zero-apparent-error learners; for internal validation of a single fixed model, bootstrap optimism correction is the cleaner choice.

Calibration vs Discrimination

Goal: Verify that predicted probabilities mean what they say, not just that they rank correctly.

Approach: Plot a reliability curve, score it with the proper Brier score, and recalibrate on a held-out fold if needed. AUC measures only ranking; the calibration slope (<1 signals overfitting) and the reliability curve localize the failure.

python
from sklearn.calibration import calibration_curve, CalibratedClassifierCV
from sklearn.metrics import brier_score_loss
from sklearn.frozen import FrozenEstimator                  # sklearn >=1.6

prob_true, prob_pred = calibration_curve(y_test, p_test, n_bins=10, strategy='quantile')
brier = brier_score_loss(y_test, p_test)                    # proper score: calibration + refinement

# Recalibrate a fitted model on a disjoint calibration fold (cv='prefit' deprecated in 1.6, removed in 1.8):
calibrated = CalibratedClassifierCV(FrozenEstimator(fitted_model), method='isotonic')
calibrated.fit(X_cal, y_cal)                                # X_cal disjoint from train and test

Calibration cautions: use strategy='quantile' (equal-mass bins) under imbalance; do not report a single Expected Calibration Error as ground truth -- equal-width ECE is biased and reports error even for perfectly calibrated models (Roelofs 2022). Use Platt (method='sigmoid') for small calibration sets, isotonic for hundreds-plus points. Recalibrating on the test set is leakage.

Net benefit / Decision Curve Analysis (Vickers-Elkin 2006): net_benefit = TP/n - (FP/n)*(pt/(1-pt)), where the threshold probability pt encodes the relative harm of a false positive. Plot it against treat-all and treat-none references; a model is clinically useful only where it sits above both. DCA requires good calibration to be valid and is the bridge from statistical performance to clinical usefulness -- a model can have high AUC yet zero net benefit at every plausible threshold.

Metric Selection in Imbalanced Data

MetricUseTrap
AccuracyAlmost never headline it under imbalanceAt 5% prevalence, "always negative" scores 95%
AUC / CDiscrimination, prevalence-independentBlind to calibration; not a usefulness measure
AUPRC (average precision)Rare-positive problemsBaseline is the prevalence, not 0.5 -- state it (Saito 2015)
Brier / log-lossWhen probabilities are usedProper; not comparable across prevalences without scaling
MCCBalanced single-threshold summaryStill threshold-dependent (Chicco 2020)
F1Retrieval-style problemsIgnores true negatives; assumes a cost ratio

The multiple-threshold problem: reporting the best F1/accuracy over thresholds is optimistic, and choosing that threshold on the test set is leakage. Pick the operating point on a separate fold (or by net benefit), then report the locked-threshold metric once; prefer threshold-free curves (ROC, PR, calibration) plus one pre-specified operating point.

Show full SKILL.md (946 more words)Show less

External Validation, Optimism, Sample Size, and TRIPOD+AI

  • Internal vs external. Internal validation (bootstrap optimism correction, repeated/nested CV) estimates reproducibility on new patients from the same source; external validation (different time, place, setting) estimates transportability and is the usual point of failure -- calibration degrades first (slope <1, intercept shift). Internal-external CV (leave-one-cluster-out) is the recommendation when multiple cohorts exist (Steyerberg 2001).
  • Optimism and shrinkage. Apparent performance overstates the future; the gap (optimism) grows with more predictors, more flexibility, smaller n. Remedy with a uniform shrinkage factor (the bootstrap calibration slope) or penalized estimation. A development-data calibration slope <1 is the optimism signal.
  • Sample size. The "10 events per variable" heuristic (Peduzzi 1996) is obsolete; the standard is Riley et al.'s minimum-sample-size framework (2019, Stat Med Parts I-II), which sizes for shrinkage >=0.9 and precise risk estimation (pmsampsize). For p>>n omics these formulas are out of regime, which is precisely why heavy penalization + nested validation, not unpenalized multivariable fits, are mandatory.
  • Reporting. TRIPOD+AI (Collins 2024, BMJ 385:e078378) supersedes TRIPOD 2015 and is the 2024+ target for any biomedical predictive-model claim -- it demands data-splitting and leakage controls, calibration (not just discrimination), fairness/subgroup performance, and uncertainty. PROBAST+AI is the companion risk-of-bias appraisal.

Per-Method Failure Modes

Preprocessing fit before the split
  • Trigger: StandardScaler().fit_transform(X) (or ComBat, PCA, imputation) on all data, then CV.
  • Mechanism: The fitted parameters encode the test rows.
  • Symptom: Test variance tiny; drop on external data.
  • Fix: Put every transform in the Pipeline so fit only sees training folds.
Threshold or best-of-many chosen on the test set
  • Trigger: Reporting the best F1 over thresholds, or the best of several CV runs.
  • Mechanism: Each peek leaks; over many tries the test set becomes a training set.
  • Symptom: Irreproducible "SOTA"; a fresh test set disappoints.
  • Fix: Lock one test set, pre-specify metric and threshold rule, choose thresholds on a separate fold.
SMOTE/resampling to fix imbalance breaks calibration
  • Trigger: Oversampling/SMOTE for a risk model.
  • Mechanism: Changing training prevalence inflates minority-class probabilities; no AUC gain (van den Goorbergh 2022).
  • Symptom: Good AUC, badly miscalibrated risks.
  • Fix: Do not resample for probability models; move the threshold on a calibrated model. If resampled, use imblearn.pipeline.Pipeline (train-fold only).
LOO for a ranking metric
  • Trigger: Leave-one-out with AUC.
  • Mechanism: AUC is undefined within a size-1 fold; pooling OOF predictions then scoring once is not equivalent to averaging fold scores for non-decomposable metrics.
  • Symptom: Unstable or misleading AUC.
  • Fix: Use repeated stratified k-fold; reserve cross_val_predict for visuals, not the headline metric.

Quantitative Thresholds

ThresholdSourceRationale
Selection/preprocessing inside every fold; nested CV for tuningAmbroise 2002; Varma-Simon 2006Same-data tune-and-grade is optimistic
Repeated 5-fold x ~10, report the spreadfield standardA single CV is one high-variance draw at small n
Calibration slope ~1; <1 means overfittingVan Calster 2019Basis of shrinkage
Sample size from Riley framework (shrinkage >=0.9)Riley 2019"10 EPV" is obsolete
Report per TRIPOD+AICollins 20242024+ standard: discrimination + calibration + fairness

Common Errors

Error / symptomCauseSolution
cv='prefit' warns or errorsDeprecated in 1.6, removed in 1.8 (now raises)Wrap the fitted model in FrozenEstimator
Calibration looks perfect on the test setCalibrated on the evaluation dataCalibrate on a disjoint fold
Scaler/selector fit outside the PipelinePreprocessing leakageMove into the Pipeline passed to cross_val_score/GridSearchCV
groups= ignoredNot threaded to the splitterPass groups= to cross_validate/GridSearchCV.fit
cross_val_predict used as the headline AUCNon-decomposable metric over pooled OOFAverage per-fold scores instead

References

  • Peduzzi P, Concato J, Kemper E, Holford TR, Feinstein AR. 1996. A simulation study of the number of events per variable in logistic regression analysis. J Clin Epidemiol 49:1373-1379.
  • Efron B, Tibshirani R. 1997. Improvements on cross-validation: the .632+ bootstrap method. J Am Stat Assoc 92:548-560.
  • Steyerberg EW, Harrell FE, Borsboom GJ, et al. 2001. Internal validation of predictive models. J Clin Epidemiol 54:774-781.
  • Ambroise C, McLachlan GJ. 2002. Selection bias in gene extraction on the basis of microarray gene-expression data. PNAS 99:6562-6566.
  • Vickers AJ, Elkin EB. 2006. Decision curve analysis: a novel method for evaluating prediction models. Med Decis Making 26:565-574.
  • Varma S, Simon R. 2006. Bias in error estimation when using cross-validation for model selection. BMC Bioinformatics 7:91.
  • Cawley GC, Talbot NLC. 2010. On over-fitting in model selection and subsequent selection bias in performance evaluation. J Mach Learn Res 11:2079-2107.
  • Saito T, Rehmsmeier M. 2015. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS ONE 10:e0118432.
  • Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. 2019. Calibration: the Achilles heel of predictive analytics. BMC Med 17:230.
  • Riley RD, Snell KIE, Ensor J, et al. 2019. Minimum sample size for developing a multivariable prediction model: Parts I-II. Stat Med 38:1262-1296.
  • Chicco D, Jurman G. 2020. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics 21:6.
  • Roelofs R, Cain N, Shlens J, Mozer MC. 2022. Mitigating bias in calibration error estimation. Proc AISTATS PMLR 151:4036-4054.
  • van den Goorbergh R, van Smeden M, Timmerman D, Van Calster B. 2022. The harm of class imbalance corrections for risk prediction models. J Am Med Inform Assoc 29:1525-1534.
  • Whalen S, Schreiber J, Noble WS, Pollard KS. 2022. Navigating the pitfalls of applying machine learning in genomics. Nat Rev Genet 23:169-181.
  • Kapoor S, Narayanan A. 2023. Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4:100804.
  • Collins GS, Moons KGM, Dhiman P, et al. 2024. TRIPOD+AI statement. BMJ 385:e078378.
  • machine-learning/biomarker-discovery - Feature selection run inside the CV fold
  • machine-learning/omics-classifiers - Model training, calibration directions, and imbalance handling
  • machine-learning/survival-analysis - Validation metrics for time-to-event models
  • experimental-design/batch-design - Designing out batch-outcome confounding before analysis
  • experimental-design/multiple-testing - FDR control for high-dimensional testing
  • clinical-biostatistics/trial-reporting - Confirmatory-trial reporting and the prediction-vs-inference boundary

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in machine-learning/model-validation of GPTomics/bioSkills.

  • SKILL.md
  • examples/nested_cv_biomarker.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Machine Learning Model Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Machine Learning Model Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Machine Learning Model Validation this skillGPTomics/bioSkills1.2k1 repos~5kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Time Series Analytics Useropen-edge-platform/edge-ai-libraries168—~3.1kAutomated safety check: PassApache-2.0
Estimate Online Covariancemicroprediction/precise336—~535Automated safety check: PassMIT
Aeon Time Series Machine Learningdavila7/claude-code-templates32k14 repos~2.6kAutomated safety check: PassMIT

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Time Series Analytics User

    open-edge-platform/edge-ai-libraries

    Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…

    168 GitHub stars~3.1k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Estimate Online Covariance

    microprediction/precise

    Estimate a covariance / correlation / precision matrix incrementally with precise.

    336 GitHub stars~535 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Aeon Time Series Machine Learning

    davila7/claude-code-templates

    Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.

    32k GitHub starsUsed in 14 repos~2.6k tokens
    Data & AnalyticsAuto-check passed
  • Precise

    microprediction/precise

    Online (incremental) covariance, correlation, and precision estimation in Python — the streaming complement to sklearn.covariance.

    336 GitHub stars~782 tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 553 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Works with

Questions about Bio Machine Learning Model Validation

What does Bio Machine Learning Model Validation do?

Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net…. Bio Machine Learning Model Validation is an agent skill from GPTomics/bioSkills. Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net benefit, optimism correction, sample-size planning, and TRIPOD+AI reporting.

When should I use Bio Machine Learning Model Validation?

Bio Machine Learning Model Validation fits situations like: estimating model performance honestly; choosing a CV scheme; detecting leakage; judging whether reported discrimination means the model is actually useful.

How do I install Bio Machine Learning Model Validation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a claude-code`. Or copy the skill folder (machine-learning/model-validation in GPTomics/bioSkills) into .claude/skills/bio-machine-learning-model-validation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Machine Learning Model Validation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a codex`. Or copy the skill folder (machine-learning/model-validation in GPTomics/bioSkills) into .agents/skills/bio-machine-learning-model-validation in your project. Codex loads it when a task matches its description.

Can I use Bio Machine Learning Model Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-machine-learning-model-validation, .gemini/skills/bio-machine-learning-model-validation, .github/skills/bio-machine-learning-model-validation and .opencode/skills/bio-machine-learning-model-validation in your project.

What does Bio Machine Learning Model Validation need to run?

Going by SKILL.md and its folder, Bio Machine Learning Model Validation needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Machine Learning Model Validation access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Machine Learning Model Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Machine Learning Model Validation use?

Bio Machine Learning Model Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Machine Learning Model Validation use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Machine Learning Model Validation?

Skills that share tags, products or a category with Bio Machine Learning Model Validation: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars) and Estimate Online Covariance (microprediction/precise, 336 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Machine Learning Model Validation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.