Scikit Learn
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net…
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-model-validation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/machine-learning/model-validation .claude/skills/bio-machine-learning-model-validation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-machine-learning-model-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/model-validation into .claude/skills/bio-machine-learning-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-model-validation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/machine-learning/model-validationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-model-validation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/machine-learning/model-validation .agents/skills/bio-machine-learning-model-validation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-machine-learning-model-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/model-validation into .agents/skills/bio-machine-learning-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-model-validation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-model-validation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/machine-learning/model-validation .cursor/skills/bio-machine-learning-model-validation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-machine-learning-model-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/model-validation into .cursor/skills/bio-machine-learning-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-model-validation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path machine-learning/model-validation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-model-validation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/machine-learning/model-validation .gemini/skills/bio-machine-learning-model-validation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-machine-learning-model-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/model-validation into .gemini/skills/bio-machine-learning-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-model-validation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-machine-learning-model-validationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/machine-learning/model-validation .github/skills/bio-machine-learning-model-validation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-machine-learning-model-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/model-validation into .github/skills/bio-machine-learning-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-model-validation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-model-validation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/machine-learning/model-validation .opencode/skills/bio-machine-learning-model-validation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-machine-learning-model-validation" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/model-validation into .opencode/skills/bio-machine-learning-model-validation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-model-validation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-machine-learning-model-validationValidates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net…
Bio Machine Learning Model Validation is an agent skill from GPTomics/bioSkills. Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net benefit, optimism correction, sample-size planning, and TRIPOD+AI reporting. Use when estimating model performance honestly, choosing a CV scheme, detecting leakage, or judging whether reported discrimination means the model is actually useful. For feature selection itself see machine-learning/biomarker-discovery; for…
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/nested_cv_biomarker.py` and `usage-guide.md`).
It sits in Data & Analytics, covering Machine learning. It works with scikit-learn. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Machine Learning Model Validation loads about 5k tokens when it runs. Until then it costs about 157 tokens; SKILL.md has 2,210 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,210 words, ~5,000 tokens.
.claude/skills/bio-machine-learning-model-validation/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Reference examples tested with: numpy 1.26+, scikit-learn 1.4+ (note 1.6/1.8 API changes below).
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturesscikit-learn drift to watch: CalibratedClassifierCV(cv='prefit') was deprecated in 1.6 and removed in 1.8 (it now raises; wrap a fitted model in sklearn.frozen.FrozenEstimator instead); ensemble default became 'auto' in 1.6; method='temperature' was added in 1.8. If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.
"Validate my omics classifier honestly" -> Keep every data-dependent step inside the resampling loop, never use the same data to both choose and grade, and report calibration and net benefit, not just AUC.
GridSearchCV (inner) wrapped by cross_val_score (outer)StratifiedGroupKFold, TimeSeriesSplitcalibration_curve, brier_score_loss, CalibratedClassifierCVA reported performance number is a claim about a data-generating process that will never recur. Almost every inflated result in ML-for-biology traces to one of two root causes: information from the test distribution leaked into model construction, or the same data was used to both choose and grade a decision. A clean train/test split is necessary but nowhere near sufficient -- the leakage has usually already contaminated the test set (a scaler fit on all data, a duplicate patient, ComBat run across the split). Leakage causes a reproducibility crisis across ML-based science (Kapoor 2023), and the bias is largest exactly when the true signal is weakest -- the omics regime.
A second, equally load-bearing insight: discrimination (AUC/C) and calibration (do predicted probabilities match observed frequencies) are orthogonal. AUC is invariant to any monotone transform of the score, so it is blind to calibration. For any decision that uses the probability itself, calibration -- not AUC -- is the property that matters, and it is the one routinely ignored (Van Calster 2019, "the Achilles heel").
| Leakage type | How it happens in omics | Symptom | Prevention |
|---|---|---|---|
| Preprocessing (most common, most missed) | z-scoring, quantile/library normalization, ComBat/SVA, PCA, kNN/MICE imputation, VST fit on the full dataset before splitting | Test performance suspiciously close to train; collapses on external data | Fit every transform inside the CV fold via a Pipeline |
| Feature selection (severe special case) | top-k DE genes / highest-variance / univariate filter chosen on all samples, then CV only the classifier | Near-perfect CV from pure noise; unstable selected set | Selection lives in the CV fold (Ambroise 2002) |
| Target / label | a feature is a proxy for or downstream of the outcome (post-diagnosis labs, treatment-derived fields, a collection-site that tracks case/control) | One feature dominates implausibly; fails when removed | Audit temporal/causal admissibility; exclude post-outcome variables |
| Group / patient / replicate | same patient, tumor, organoid, or technical replicate in train and test; KFold scatters them | Inflated metrics that vanish under leave-one-group-out | Split by the highest independent unit (GroupKFold/StratifiedGroupKFold) |
| Batch | batch correlated with outcome and not respected in the split, or ComBat across the train/test boundary | Model discriminates batches not biology; external batch destroys it | Block the split by batch; never run unsupervised correction across the split |
| Temporal | random-splitting time-ordered data; future-period statistics standardize the past | Backtest beats prospective deployment | Time-based split (TimeSeriesSplit); never shuffle first |
| Duplicate / homolog | near-identical samples, augmented copies, public-dataset overlap, homologous sequences across the split | Memorization passes as generalization | Deduplicate / cluster-then-split before CV |
| Test-reuse / threshold | repeatedly peeking to pick features, thresholds, "best epoch"; choosing the classification threshold on the test set | Irreproducible SOTA; fragile config | One locked test set; all tuning + thresholds inside nested CV |
| Scenario / generalization question | Recommended scheme | Why |
|---|---|---|
| "A new sample like training" (and any tuning occurs) | Nested CV: inner GridSearchCV, outer cross_val_score, Pipeline inside | Tuning and grading on the same CV is optimistic (Cawley-Talbot 2010) |
| "A new patient" (repeated measures) | GroupKFold/StratifiedGroupKFold by patient/donor | The unit of independence is not the row |
| "A new hospital/site" (transportability) | Leave-one-site-out (internal-external CV) | Approximates external validation |
| "Next year" (time-ordered) | TimeSeriesSplit forward-chaining | Random folds leak the future |
| Small n (dozens), need a stable estimate | RepeatedStratifiedKFold (5x10) with an interval | A single CV is one high-variance draw |
| Probabilities will drive a decision | Add calibration + decision-curve net benefit | AUC is blind to calibration and utility |
| Final evidence for a clinical model | External/temporal validation + TRIPOD+AI report | Internal CV cannot detect a whole-dataset confound |
| Choosing the features themselves | -> machine-learning/biomarker-discovery | Selection is its own discipline (run it inside the fold) |
| Confirmatory trial inference (HR, p-value) | -> clinical-biostatistics/trial-reporting | Estimand is a treatment effect, not a prediction |
Goal: Estimate the performance of the whole procedure (tuning + fit) without optimistic bias.
Approach: The inner loop does all tuning, feature selection, and threshold choice; the outer loop grades the winning configuration once on a fold it never touched. The reported number is the aggregate over outer folds; it answers "if I run this pipeline on new data, what do I get?"
from sklearn.model_selection import cross_val_score, StratifiedKFold, GridSearchCV
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.feature_selection import SelectKBest, f_classif
from sklearn.linear_model import LogisticRegression
pipe = Pipeline([('scaler', StandardScaler()),
('select', SelectKBest(f_classif)), # re-fit per inner fold -> no leakage
('clf', LogisticRegression(max_iter=5000))])
grid = {'select__k': [10, 50, 200], 'clf__C': [0.01, 0.1, 1]}
inner = StratifiedKFold(5, shuffle=True, random_state=0)
outer = StratifiedKFold(5, shuffle=True, random_state=1)
search = GridSearchCV(pipe, grid, cv=inner, scoring='roc_auc')
scores = cross_val_score(search, X, y, cv=outer, scoring='roc_auc') # unbiased estimate
print(f'Nested AUC: {scores.mean():.3f} +/- {scores.std():.3f}')Nested CV is needed whenever model selection happens -- even informal "tried three options, kept the best." Flat CV with tuning is a known reviewer red flag (Varma-Simon 2006).
Goal: Match the CV scheme to the real unit of independence and get a variance-aware estimate.
Approach: Pass a grouping vector so no group spans folds; for tiny n, repeat stratified k-fold and report the spread, not a bare number. Standard KFold assumes i.i.d. rows, which biomedical data almost never satisfy.
from sklearn.model_selection import StratifiedGroupKFold, RepeatedStratifiedKFold, cross_val_score
groups = meta['patient_id'].values # multiple samples per patient
gcv = StratifiedGroupKFold(n_splits=5) # group-disjoint AND class-balanced
g_auc = cross_val_score(pipe, X, y, cv=gcv, groups=groups, scoring='roc_auc')
rcv = RepeatedStratifiedKFold(n_splits=5, n_repeats=10, random_state=0)
r_auc = cross_val_score(pipe, X, y, cv=rcv, scoring='roc_auc') # report an intervalLeave-one-out is high-variance and degenerate for ranking metrics (AUC is undefined on a size-1 test fold) -- prefer repeated stratified k-fold. The .632+ bootstrap (Efron-Tibshirani 1997) is a defensible alternative but is optimistic for zero-apparent-error learners; for internal validation of a single fixed model, bootstrap optimism correction is the cleaner choice.
Goal: Verify that predicted probabilities mean what they say, not just that they rank correctly.
Approach: Plot a reliability curve, score it with the proper Brier score, and recalibrate on a held-out fold if needed. AUC measures only ranking; the calibration slope (<1 signals overfitting) and the reliability curve localize the failure.
from sklearn.calibration import calibration_curve, CalibratedClassifierCV
from sklearn.metrics import brier_score_loss
from sklearn.frozen import FrozenEstimator # sklearn >=1.6
prob_true, prob_pred = calibration_curve(y_test, p_test, n_bins=10, strategy='quantile')
brier = brier_score_loss(y_test, p_test) # proper score: calibration + refinement
# Recalibrate a fitted model on a disjoint calibration fold (cv='prefit' deprecated in 1.6, removed in 1.8):
calibrated = CalibratedClassifierCV(FrozenEstimator(fitted_model), method='isotonic')
calibrated.fit(X_cal, y_cal) # X_cal disjoint from train and testCalibration cautions: use strategy='quantile' (equal-mass bins) under imbalance; do not report a single Expected Calibration Error as ground truth -- equal-width ECE is biased and reports error even for perfectly calibrated models (Roelofs 2022). Use Platt (method='sigmoid') for small calibration sets, isotonic for hundreds-plus points. Recalibrating on the test set is leakage.
Net benefit / Decision Curve Analysis (Vickers-Elkin 2006): net_benefit = TP/n - (FP/n)*(pt/(1-pt)), where the threshold probability pt encodes the relative harm of a false positive. Plot it against treat-all and treat-none references; a model is clinically useful only where it sits above both. DCA requires good calibration to be valid and is the bridge from statistical performance to clinical usefulness -- a model can have high AUC yet zero net benefit at every plausible threshold.
| Metric | Use | Trap |
|---|---|---|
| Accuracy | Almost never headline it under imbalance | At 5% prevalence, "always negative" scores 95% |
| AUC / C | Discrimination, prevalence-independent | Blind to calibration; not a usefulness measure |
| AUPRC (average precision) | Rare-positive problems | Baseline is the prevalence, not 0.5 -- state it (Saito 2015) |
| Brier / log-loss | When probabilities are used | Proper; not comparable across prevalences without scaling |
| MCC | Balanced single-threshold summary | Still threshold-dependent (Chicco 2020) |
| F1 | Retrieval-style problems | Ignores true negatives; assumes a cost ratio |
The multiple-threshold problem: reporting the best F1/accuracy over thresholds is optimistic, and choosing that threshold on the test set is leakage. Pick the operating point on a separate fold (or by net benefit), then report the locked-threshold metric once; prefer threshold-free curves (ROC, PR, calibration) plus one pre-specified operating point.
pmsampsize). For p>>n omics these formulas are out of regime, which is precisely why heavy penalization + nested validation, not unpenalized multivariable fits, are mandatory.StandardScaler().fit_transform(X) (or ComBat, PCA, imputation) on all data, then CV.fit only sees training folds.imblearn.pipeline.Pipeline (train-fold only).cross_val_predict for visuals, not the headline metric.| Threshold | Source | Rationale |
|---|---|---|
| Selection/preprocessing inside every fold; nested CV for tuning | Ambroise 2002; Varma-Simon 2006 | Same-data tune-and-grade is optimistic |
| Repeated 5-fold x ~10, report the spread | field standard | A single CV is one high-variance draw at small n |
| Calibration slope ~1; <1 means overfitting | Van Calster 2019 | Basis of shrinkage |
| Sample size from Riley framework (shrinkage >=0.9) | Riley 2019 | "10 EPV" is obsolete |
| Report per TRIPOD+AI | Collins 2024 | 2024+ standard: discrimination + calibration + fairness |
| Error / symptom | Cause | Solution |
|---|---|---|
cv='prefit' warns or errors | Deprecated in 1.6, removed in 1.8 (now raises) | Wrap the fitted model in FrozenEstimator |
| Calibration looks perfect on the test set | Calibrated on the evaluation data | Calibrate on a disjoint fold |
| Scaler/selector fit outside the Pipeline | Preprocessing leakage | Move into the Pipeline passed to cross_val_score/GridSearchCV |
groups= ignored | Not threaded to the splitter | Pass groups= to cross_validate/GridSearchCV.fit |
cross_val_predict used as the headline AUC | Non-decomposable metric over pooled OOF | Average per-fold scores instead |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files in machine-learning/model-validation of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Machine Learning Model Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Machine Learning Model Validation this skillGPTomics/bioSkills | 1.2k | 1 repos | ~5k | Automated safety check: Pass | MIT | |
| Scikit LearnzLanqing/codex-claude-academic-skills | 4.6k | 17 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Time Series Analytics Useropen-edge-platform/edge-ai-libraries | 168 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | |
| Estimate Online Covariancemicroprediction/precise | 336 | — | ~535 | Automated safety check: Pass | MIT | |
| Aeon Time Series Machine Learningdavila7/claude-code-templates | 32k | 14 repos | ~2.6k | Automated safety check: Pass | MIT |
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
open-edge-platform/edge-ai-libraries
Build a new time-series analytics use case on top of the deployed Time Series Analytics microservice — bring it up with Docker Compose (from a repo clone, or by fetching the compose files from…
microprediction/precise
Estimate a covariance / correlation / precision matrix incrementally with precise.
davila7/claude-code-templates
Guides time series machine learning with the aeon toolkit: classification, regression, clustering, forecasting, anomaly detection, segmentation and similarity search.
microprediction/precise
Online (incremental) covariance, correlation, and precision estimation in Python — the streaming complement to sklearn.covariance.
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Works with
Categories
Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net…. Bio Machine Learning Model Validation is an agent skill from GPTomics/bioSkills. Validates predictive models on omics and biomedical data with nested cross-validation, group/batch/temporal-aware splits, the full data-leakage taxonomy, probability calibration, decision-curve net benefit, optimism correction, sample-size planning, and TRIPOD+AI reporting.
Bio Machine Learning Model Validation fits situations like: estimating model performance honestly; choosing a CV scheme; detecting leakage; judging whether reported discrimination means the model is actually useful.
Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a claude-code`. Or copy the skill folder (machine-learning/model-validation in GPTomics/bioSkills) into .claude/skills/bio-machine-learning-model-validation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a codex`. Or copy the skill folder (machine-learning/model-validation in GPTomics/bioSkills) into .agents/skills/bio-machine-learning-model-validation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-model-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-machine-learning-model-validation, .gemini/skills/bio-machine-learning-model-validation, .github/skills/bio-machine-learning-model-validation and .opencode/skills/bio-machine-learning-model-validation in your project.
Going by SKILL.md and its folder, Bio Machine Learning Model Validation needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Machine Learning Model Validation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Machine Learning Model Validation: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Time Series Analytics User (open-edge-platform/edge-ai-libraries, 168 stars) and Estimate Online Covariance (microprediction/precise, 336 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.