Agent skill

Bio Machine Learning Omics Classifiers

by GPTomics in GPTomics/bioSkills

Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut…

MITAuto-check passedData & Analytics

Install Bio Machine Learning Omics Classifiers

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-machine-learning-omics-classifiers --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/machine-learning/omics-classifiers .claude/skills/bio-machine-learning-omics-classifiers && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-machine-learning-omics-classifiers
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.7k tokens
SKILL.md length
1,898 words
Files
4
Skills in repo
553
Repo updated
First seen
Licence
MIT

At a glance

Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut…

  • Building a classifier from expression
  • SKILL.md covers Version Compatibility, The Single Most Important…, Algorithm Choice for p>>n and Decision Tree by Scenario, plus 12 more sections
  • Runs Python scripts from its folder; calls pip
  • Choosing an algorithm for high-dimensional small-n data

What it does

Bio Machine Learning Omics Classifiers is an agent skill from GPTomics/bioSkills. Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut learning, class imbalance, and probability calibration. Use when building a classifier from expression, methylation, or variant data, choosing an algorithm for high-dimensional small-n data, or diagnosing a suspiciously perfect AUC. For unbiased evaluation see machine-learning/model-validation; for feature selection see…

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/logistic_regression.py`, `examples/rf_xgboost_classifier.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Machine learning. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Building a classifier from expression
  • Choosing an algorithm for high-dimensional small-n data
  • Diagnosing a suspiciously perfect AUC

Example prompts

  • “Use the bio-machine-learning-omics-classifiers skill to build diagnostic and prognostic classifiers on omics feature matrices with regularized…”
  • “/bio-machine-learning-omics-classifiers”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Machine Learning Omics Classifiers loads about 4.7k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 1,898 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~165
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,898 words, ~4,722 tokens.

Download SKILL.mdSave it as .claude/skills/bio-machine-learning-omics-classifiers/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-machine-learning-omics-classifiers
description
Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the p>>n regime, batch shortcut learning, class imbalance, and probability calibration. Use when building a classifier from expression, methylation, or variant data, choosing an algorithm for high-dimensional small-n data, or diagnosing a suspiciously perfect AUC. For unbiased evaluation see machine-learning/model-validation; for feature selection see machine-learning/biomarker-discovery; for time-to-event outcomes see machine-learning/survival-analysis.
tool_type
python
primary_tool
sklearn

Version Compatibility

Reference examples tested with: pandas 2.2+, scikit-learn 1.4+, xgboost 2.0+, imbalanced-learn 0.12+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

Two high-risk drifts: XGBoost moved early_stopping_rounds from fit() to the constructor (deprecated 1.6, removed from fit() in 2.1); scikit-learn deprecated LogisticRegression(penalty=) in 1.8 (use l1_ratio+C) and CalibratedClassifierCV(cv='prefit') in 1.6 (use FrozenEstimator). If code throws TypeError/FutureWarning, switch to the constructor / l1_ratio / FrozenEstimator form.

Classification Models for Omics Data

"Build a classifier from my expression data" -> Start with a regularized linear model (often the ceiling in p>>n), check for batch shortcuts, and treat the probability -- not the label -- as the product.

  • Linear (often best): LogisticRegression(penalty='elasticnet', solver='saga')
  • Trees when nonlinear/interaction signal: RandomForestClassifier, xgboost.XGBClassifier
  • Imbalance: class_weight='balanced' or threshold tuning, NOT SMOTE for risk models

The Single Most Important Modern Insight -- In p>>n, Simple Often Wins and the Probability, Not the Label, Is the Product

Omics classification almost always lives in p>>n (thousands of features, tens-to-hundreds of samples). Two counterintuitive consequences follow. First, more flexible is not better: with n in the dozens the variance of a flexible learner dominates, the full covariance is singular so QDA/full-LDA are undefined, and simple diagonal/linear methods match or beat elaborate ones (Dudoit 2002). "Random forest is the obvious choice for expression data" is a myth -- SVM/regularized logistic frequently win on microarray-style problems (Statnikov 2008), and gradient-boosted trees beat deep nets on tabular/omics data (Grinsztajn 2022). Regularization is the load-bearing wall, not a tuning nicety.

Second, in diagnostic/prognostic use the probability is the product, not the label -- which makes calibration, not accuracy, the thing that breaks silently (Van Calster 2019). A model can rank perfectly (AUC 0.9) and still output dishonest risks. And the most common cause of a beautiful AUC is not skill but a batch artifact: if batch correlates with the outcome, the classifier learns the cleaner technical signal and the performance collapses on any independent cohort.

Algorithm Choice for p>>n

ModelWins whenOverfits / fails whenCalibrationScaling
L1 logistic (lasso)Sparse signal, want a small signatureCorrelated features -> unstable selection; >n true signalsGood (proper loss); shrinks toward base rateStandardize
L2 / elastic-net logisticMany small correlated effects; omics defaultNeeds C (and l1_ratio) tuningGood; preferred when calibration mattersStandardize
DLDA / nearest-centroidTiny n, roughly linear (Dudoit 2002)Strong interactions; non-GaussianCrude; recalibrateVariance-scaled
Linear SVMHigh-dim linear separability (Statnikov 2008)Heavy overlap; needs CNo native probabilities -- Platt-scale decision_functionCritical
Random forestNonlinear/interaction signal; robust baselineSparse-linear signal; tiny n; OOB-as-test leakageBagged votes bounded away from 0 and 1Scale-invariant
GBDT (XGBoost/LightGBM)Best general tabular performerTiny n + deep/many rounds; needs early stoppingLog-loss overfitting tends to overconfident extremesScale-invariant
Tabular deep netsVery large n; multimodal/transferTypical omics n -> loses to GBDTVariable; often needs temperature scalingStandardize

Tree ensembles are often miscalibrated and the direction depends on the learner and loss: classic boosted ensembles push probabilities toward 0.5 (sigmoid distortion; Niculescu-Mizil 2005), bagged forests are comparatively well-calibrated but bound their votes away from 0 and 1, and modern gradient boosting trained to log-loss for many rounds tends to overfit toward overconfident extremes. Check a reliability curve and recalibrate rather than assuming a direction.

Decision Tree by Scenario

ScenarioRecommended approachWhy
Default omics classifier, want a signatureElastic-net logisticOften the ceiling in p>>n; sparse + grouping; well-calibrated
Suspected nonlinear/interaction (epistasis, thresholds)Random forest then XGBoost with early stoppingTrees capture interactions; benchmark vs the linear baseline
Probabilities will drive a clinical decisionLinear model + calibration check; recalibrate if neededProbability is the product; AUC is blind to calibration
Class imbalanceclass_weight='balanced' or threshold tuning; never SMOTE for riskResampling destroys calibration for no AUC gain
Mixed continuous + categorical featuresColumnTransformer (scale continuous, encode categorical)Different feature types need different handling
Missing values, especially below-detectionXGBoost/LightGBM native NaN handlingMissingness is often informative (MNAR)
Considering a deep netOnly at very large n or multimodal/raw inputsGBDT beats deep on engineered omics matrices (Grinsztajn 2022)
Need unbiased performance / nested CV / calibration metrics-> machine-learning/model-validationEvaluation is its own discipline
Time-to-event outcome-> machine-learning/survival-analysisCensoring needs survival models, not classifiers

Core Workflow: Regularized Logistic First

Goal: A calibrated, interpretable baseline that is often the best omics classifier.

Approach: Standardize inside a Pipeline and fit elastic-net logistic with cross-validated penalty; the L2 component keeps correlated genes together, the L1 component yields a sparse signature.

python
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegressionCV

# saga supports elasticnet; C = 1/lambda (small C = strong shrinkage). Standardize: the penalty is scale-sensitive.
clf = LogisticRegressionCV(penalty='elasticnet', solver='saga', l1_ratios=[0.1, 0.5, 0.9],
                           Cs=20, cv=5, max_iter=10000, class_weight='balanced')
pipe = Pipeline([('scaler', StandardScaler()), ('clf', clf)])
pipe.fit(X_train, y_train)

Tree Ensembles

Goal: Capture nonlinear and interaction structure when the linear baseline leaves signal on the table.

Approach: Random forest needs no scaling and is a robust baseline; XGBoost needs a low learning rate, shallow depth, and early stopping (set in the constructor in 2.x) to avoid overfitting tiny n.

python
from sklearn.ensemble import RandomForestClassifier
from xgboost import XGBClassifier

rf = RandomForestClassifier(n_estimators=500, max_features='sqrt', min_samples_leaf=3,
                            class_weight='balanced', n_jobs=-1, random_state=0)

# XGBoost 2.x: early_stopping_rounds and eval_metric go in the CONSTRUCTOR, not fit().
# scale_pos_weight is omitted on purpose: like resampling, it reweights the prior and
# distorts calibration -- use it only for hard-label problems, not risk models (see Class Imbalance).
xgb = XGBClassifier(n_estimators=2000, learning_rate=0.03, max_depth=4, subsample=0.8,
                    colsample_bytree=0.5, reg_lambda=1.0,
                    early_stopping_rounds=50, eval_metric='aucpr', n_jobs=-1, random_state=0)
xgb.fit(X_train, y_train, eval_set=[(X_val, y_val)])      # NaN handled natively (missing=np.nan)

Detecting Batch Shortcut Learning

Goal: Rule out that a high AUC is a batch artifact rather than biology.

Approach: Try to predict the batch from the features and use batch-aware splits; if batch is confounded with the outcome, no correction rescues the design (Soneson 2014) -- fix it at the design stage.

python
from sklearn.model_selection import cross_val_score, StratifiedGroupKFold
from scipy.stats import chi2_contingency
import pandas as pd

# 1. Can the classifier predict the BATCH? If yes, batch is a strong axis and the label model is suspect.
batch_auc = cross_val_score(pipe, X, batch_labels, cv=5, scoring='roc_auc')
print(f'Batch predictability AUC: {batch_auc.mean():.2f} (high = shortcut risk)')

# 2. Is the outcome associated with batch by design?
print('label vs batch p:', chi2_contingency(pd.crosstab(y, batch_labels))[1])

# 3. Leave-one-batch-out is the honest generalization estimate (usually << random-split CV).
gcv = StratifiedGroupKFold(n_splits=5)
honest = cross_val_score(pipe, X, y, cv=gcv, groups=batch_labels, scoring='roc_auc')
print(f'Batch-aware AUC: {honest.mean():.2f}')

Class Imbalance: What Works and What Fails

Goal: Handle a rare positive class without destroying the probabilities.

Approach: For a risk model, do not resample -- class-weight cautiously or tune the threshold on a validation fold; SMOTE/oversampling change the training prior, inflate minority probabilities, give no AUC gain, and the same sensitivity is recoverable by moving the threshold (van den Goorbergh 2022). When resampling is unavoidable (a hard-label problem), use an imblearn Pipeline so only training folds are resampled.

python
from imblearn.pipeline import Pipeline as ImbPipeline   # NOT sklearn's Pipeline
from imblearn.over_sampling import SMOTE
from sklearn.linear_model import LogisticRegression

# Correct placement: SMOTE's fit_resample runs only during fit on the train fold, no-op on transform.
imb = ImbPipeline([('smote', SMOTE(random_state=0)),
                   ('clf', LogisticRegression(max_iter=5000))])
# Prefer for risk models: no resampling, then pick the operating threshold by cost on a validation fold.

Probability Calibration

Goal: Ensure a "0.9" means a 90% risk, not just a high rank.

Approach: Tree ensembles are often miscalibrated -- bagged forests bound votes away from 0 and 1, classic boosting is sigmoid-distorted toward 0.5 (Niculescu-Mizil 2005), and log-loss GBDT can overfit to overconfident extremes -- so check a reliability curve and recalibrate on a disjoint fold. Logistic regression optimizes a proper scoring rule and is usually best-calibrated out of the box. See machine-learning/model-validation for reliability curves, Brier, and the full protocol.

python
from sklearn.calibration import CalibratedClassifierCV
from sklearn.frozen import FrozenEstimator               # sklearn >=1.6; cv='prefit' deprecated
calibrated = CalibratedClassifierCV(FrozenEstimator(rf.fit(X_tr, y_tr)), method='isotonic')
calibrated.fit(X_cal, y_cal)                             # X_cal disjoint from train and test

Preprocessing, Scaling, Encoding, Missing Data

  • Inside the CV fold (refit per fold via Pipeline): feature selection, scaling, log/VST/quantile normalization, ComBat batch correction, imputation, PCA. Fitting any of these on the full matrix is leakage (machine-learning/model-validation).
  • Scale-sensitive: regularized logistic, SVM, kNN, PCA, neural nets -- standardize. Scale-invariant: trees/RF/GBDT.
  • Variant features: additive 0/1/2 ordinal for trees/additive logistic; one-hot for non-additive (dominant/recessive) effects; CatBoost-style encoding for high-cardinality (HLA).
  • Missing data: XGBoost/LightGBM learn a default split direction for NaN; in omics missingness is often MNAR (below detection limit) and a missingness indicator can carry signal -- naive zero-fill conflates "absent" with "not measured."

Hyperparameters That Matter

ModelTune theseLeave default
LogisticC (log-spaced), l1_ratio, class_weightsolver (saga for elasticnet)
Random forestmax_features, min_samples_leaf, max_depth (cap for tiny n)n_estimators (more is safe; 500-1000)
XGBoostlearning_rate+n_estimators+early stopping, max_depth (3-6), subsample, colsample_bytree, reg_lambdamost others

Per-Method Failure Modes

Show full SKILL.md (763 more words)Show less
Beautiful AUC that is a batch artifact
  • Trigger: Cases and controls processed in different batches/sites/times.
  • Mechanism: The classifier exploits the cleaner technical signal; even nested CV is optimistic, and ComBat cannot rescue a confounded design (Soneson 2014).
  • Symptom: Near-perfect CV AUC; collapse on an independent cohort; the model predicts batch easily.
  • Fix: Detect with the batch-prediction check; use leave-one-batch-out; fix at design (balance batches across outcome).
RF/boosting probabilities trusted as risks
  • Trigger: Reading predict_proba from RF or XGBoost as a calibrated risk.
  • Mechanism: RF bounds votes away from 0 and 1; classic boosting is sigmoid-distorted (Niculescu-Mizil 2005) and log-loss GBDT can overfit to overconfident extremes.
  • Symptom: Good AUC, reliability curve far from diagonal.
  • Fix: Recalibrate on a disjoint fold; or prefer logistic when the probability matters.
SMOTE-before-split / SMOTE for a risk model
  • Trigger: Resampling before the CV split, or to "fix" imbalance for a probability model.
  • Mechanism: Synthetic points derived from test samples leak; resampling inflates minority risk and wrecks calibration (van den Goorbergh 2022; Carriero 2025).
  • Symptom: Inflated CV performance; predicted risks systematically too high.
  • Fix: imblearn Pipeline (train-fold only); for risk models, do not resample -- tune the threshold.
OOB error read as an unbiased test estimate
  • Trigger: Reporting RF out-of-bag error after selecting features on the full data.
  • Mechanism: Selection leaked; OOB then reflects the contaminated feature set.
  • Symptom: Optimistic OOB; external collapse.
  • Fix: Selection inside CV; estimate performance by nested CV.

Quantitative Thresholds

ThresholdSourceRationale
Try a regularized linear model firstDudoit 2002; Statnikov 2008Simple often beats complex in p>>n
GBDT over deep nets for tabular omicsGrinsztajn 2022Trees handle uninformative features and non-rotational data
Do not resample for risk modelsvan den Goorbergh 2022; Carriero 2025Resampling destroys calibration for no AUC gain
XGBoost: low LR + many rounds + early stoppingfield standardPrevents overfitting tiny n
Report AUPRC + MCC under imbalanceSaito 2015; Chicco 2020Accuracy and ROC-AUC mislead when positives are rare

Common Errors

Error / symptomCauseSolution
XGBoost early_stopping_rounds TypeError in fit()Moved to constructor in 2.xPass it (and eval_metric) in XGBClassifier(...)
penalty='l1' FutureWarningDeprecated in sklearn 1.8Use l1_ratio=1 + C (1.8+) or keep penalty on 1.4-1.7
elasticnet solver errorOnly saga supports itsolver='saga' + l1_ratio
CalibratedClassifierCV(cv='prefit') deprecatedsklearn 1.6Wrap in FrozenEstimator
95% accuracy but useless modelImbalance + accuracy metricReport AUPRC/MCC; check the confusion matrix

References

  • Dudoit S, Fridlyand J, Speed TP. 2002. Comparison of discrimination methods for the classification of tumors using gene expression data. J Am Stat Assoc 97:77-87.
  • Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. 2002. SMOTE: Synthetic Minority Over-sampling Technique. J Artif Intell Res 16:321-357.
  • Zou H, Hastie T. 2005. Regularization and variable selection via the elastic net. J R Stat Soc B 67:301-320.
  • Niculescu-Mizil A, Caruana R. 2005. Predicting good probabilities with supervised learning. Proc 22nd ICML 625-632.
  • Diaz-Uriarte R, Alvarez de Andres S. 2006. Gene selection and classification of microarray data using random forest. BMC Bioinformatics 7:3.
  • Statnikov A, Wang L, Aliferis CF. 2008. A comprehensive comparison of random forests and support vector machines for microarray-based cancer classification. BMC Bioinformatics 9:319.
  • Soneson C, Gerster S, Delorenzi M. 2014. Batch effect confounding leads to strong bias in performance estimates obtained by cross-validation. PLoS ONE 9:e100335.
  • Saito T, Rehmsmeier M. 2015. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PLoS ONE 10:e0118432.
  • Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. 2019. Calibration: the Achilles heel of predictive analytics. BMC Med 17:230.
  • Chicco D, Jurman G. 2020. The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics 21:6.
  • Shwartz-Ziv R, Armon A. 2022. Tabular data: deep learning is not all you need. Inf Fusion 81:84-90.
  • Grinsztajn L, Oyallon E, Varoquaux G. 2022. Why do tree-based models still outperform deep learning on typical tabular data? NeurIPS Datasets and Benchmarks.
  • van den Goorbergh R, van Smeden M, Timmerman D, Van Calster B. 2022. The harm of class imbalance corrections for risk prediction models. J Am Med Inform Assoc 29:1525-1534.
  • Carriero A, Luijken K, de Hond A, Moons KGM, van Calster B, van Smeden M. 2025. The harms of class imbalance corrections for machine learning based prediction models. Stat Med 44:e10320.
  • machine-learning/model-validation - Nested CV, calibration, and net benefit for the trained classifier
  • machine-learning/biomarker-discovery - Select features before modeling (inside the CV fold)
  • machine-learning/prediction-explanation - Interpret the classifier and detect shortcuts with SHAP
  • machine-learning/survival-analysis - Time-to-event outcomes that classifiers cannot handle
  • differential-expression/batch-correction - Batch correction done design-aware, not across the split
  • expression-matrix/normalization - Per-sample normalization that is safe outside the CV fold

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in machine-learning/omics-classifiers of GPTomics/bioSkills.

  • SKILL.md
  • examples/logistic_regression.py
  • examples/rf_xgboost_classifier.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Machine Learning Omics Classifiers next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Machine Learning Omics Classifiers compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Machine Learning Omics Classifiers this skillGPTomics/bioSkills1.2k1 repos~4.7kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.6k17 repos~3.9kAutomated safety check: PassBSD-3-Clause
Senior Data ScientistRaidriar7170/hermes-skilleval1256 repos~1.4kAutomated safety check: PassMIT
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Retention Analysisliangdabiao/claude-data-analysis-ultra-main2901 repos~1.3kAutomated safety check: NotesNone
Geomlitalo-goncalves/geoML108—~4.2kAutomated safety check: PassGPL-3.0

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 17 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 6 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Retention Analysis

    liangdabiao/claude-data-analysis-ultra-main

    Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.

    290 GitHub starsUsed in 1 repo~1.3k tokens
    Data & AnalyticsAuto-check: notes
  • Geoml

    italo-goncalves/geoML

    Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…

    108 GitHub stars~4.2k tokensUpdated 5 days ago
    Data & AnalyticsAuto-check passed
  • Radiomics ML

    Aperivue/medsci-skills

    A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).

    329 GitHub starsUsed in 1 repo~2.7k tokens
    Data & AnalyticsAuto-check passed

More from GPTomics/bioSkills

All 553 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed
  • Bio Alignment Sorting

    GPTomics/bioSkills

    Sort alignment files by coordinate or read name using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.6k tokens
    Auto-check passed

Questions about Bio Machine Learning Omics Classifiers

What does Bio Machine Learning Omics Classifiers do?

Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut…. Bio Machine Learning Omics Classifiers is an agent skill from GPTomics/bioSkills. Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut learning, class imbalance, and probability calibration.

When should I use Bio Machine Learning Omics Classifiers?

Bio Machine Learning Omics Classifiers fits situations like: building a classifier from expression; choosing an algorithm for high-dimensional small-n data; diagnosing a suspiciously perfect AUC.

How do I install Bio Machine Learning Omics Classifiers in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a claude-code`. Or copy the skill folder (machine-learning/omics-classifiers in GPTomics/bioSkills) into .claude/skills/bio-machine-learning-omics-classifiers in your project. Claude Code loads it when a task matches its description.

How do I install Bio Machine Learning Omics Classifiers in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a codex`. Or copy the skill folder (machine-learning/omics-classifiers in GPTomics/bioSkills) into .agents/skills/bio-machine-learning-omics-classifiers in your project. Codex loads it when a task matches its description.

Can I use Bio Machine Learning Omics Classifiers in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-machine-learning-omics-classifiers, .gemini/skills/bio-machine-learning-omics-classifiers, .github/skills/bio-machine-learning-omics-classifiers and .opencode/skills/bio-machine-learning-omics-classifiers in your project.

What does Bio Machine Learning Omics Classifiers need to run?

Going by SKILL.md and its folder, Bio Machine Learning Omics Classifiers needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Machine Learning Omics Classifiers access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Machine Learning Omics Classifiers safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Machine Learning Omics Classifiers use?

Bio Machine Learning Omics Classifiers is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Machine Learning Omics Classifiers use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Machine Learning Omics Classifiers?

Skills that share tags, products or a category with Bio Machine Learning Omics Classifiers: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars) and Retention Analysis (liangdabiao/claude-data-analysis-ultra-main, 290 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Machine Learning Omics Classifiers?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.