Scikit Learn
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut…
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-omics-classifiers --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/machine-learning/omics-classifiers .claude/skills/bio-machine-learning-omics-classifiers && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "bio-machine-learning-omics-classifiers" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/omics-classifiers into .claude/skills/bio-machine-learning-omics-classifiers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-omics-classifiers", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/GPTomics/bioSkills/tree/main/machine-learning/omics-classifiersType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-omics-classifiers --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/machine-learning/omics-classifiers .agents/skills/bio-machine-learning-omics-classifiers && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "bio-machine-learning-omics-classifiers" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/omics-classifiers into .agents/skills/bio-machine-learning-omics-classifiers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-omics-classifiers", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-omics-classifiers --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/machine-learning/omics-classifiers .cursor/skills/bio-machine-learning-omics-classifiers && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "bio-machine-learning-omics-classifiers" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/omics-classifiers into .cursor/skills/bio-machine-learning-omics-classifiers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-omics-classifiers", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/GPTomics/bioSkills.git --path machine-learning/omics-classifiers--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-omics-classifiers --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/machine-learning/omics-classifiers .gemini/skills/bio-machine-learning-omics-classifiers && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "bio-machine-learning-omics-classifiers" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/omics-classifiers into .gemini/skills/bio-machine-learning-omics-classifiers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-omics-classifiers", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install GPTomics/bioSkills bio-machine-learning-omics-classifiersInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .github/skills && cp -r skills-src/machine-learning/omics-classifiers .github/skills/bio-machine-learning-omics-classifiers && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "bio-machine-learning-omics-classifiers" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/omics-classifiers into .github/skills/bio-machine-learning-omics-classifiers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-omics-classifiers", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install GPTomics/bioSkills bio-machine-learning-omics-classifiers --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/machine-learning/omics-classifiers .opencode/skills/bio-machine-learning-omics-classifiers && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "bio-machine-learning-omics-classifiers" agent skill from https://github.com/GPTomics/bioSkills/tree/main/machine-learning/omics-classifiers into .opencode/skills/bio-machine-learning-omics-classifiers/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "bio-machine-learning-omics-classifiers", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
bio-machine-learning-omics-classifiersBuilds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut…
Bio Machine Learning Omics Classifiers is an agent skill from GPTomics/bioSkills. Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut learning, class imbalance, and probability calibration. Use when building a classifier from expression, methylation, or variant data, choosing an algorithm for high-dimensional small-n data, or diagnosing a suspiciously perfect AUC. For unbiased evaluation see machine-learning/model-validation; for feature selection see…
Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/logistic_regression.py`, `examples/rf_xgboost_classifier.py` and `usage-guide.md`).
It sits in Data & Analytics, covering Machine learning. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.
Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python), which the agent can run.
Shell commands in SKILL.md call:
pipFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Bio Machine Learning Omics Classifiers loads about 4.7k tokens when it runs. Until then it costs about 165 tokens; SKILL.md has 1,898 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,898 words, ~4,722 tokens.
.claude/skills/bio-machine-learning-omics-classifiers/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.Reference examples tested with: pandas 2.2+, scikit-learn 1.4+, xgboost 2.0+, imbalanced-learn 0.12+.
Before using code patterns, verify installed versions match. If versions differ:
pip show <package> then help(module.function) to check signaturesTwo high-risk drifts: XGBoost moved early_stopping_rounds from fit() to the constructor (deprecated 1.6, removed from fit() in 2.1); scikit-learn deprecated LogisticRegression(penalty=) in 1.8 (use l1_ratio+C) and CalibratedClassifierCV(cv='prefit') in 1.6 (use FrozenEstimator). If code throws TypeError/FutureWarning, switch to the constructor / l1_ratio / FrozenEstimator form.
"Build a classifier from my expression data" -> Start with a regularized linear model (often the ceiling in p>>n), check for batch shortcuts, and treat the probability -- not the label -- as the product.
LogisticRegression(penalty='elasticnet', solver='saga')RandomForestClassifier, xgboost.XGBClassifierclass_weight='balanced' or threshold tuning, NOT SMOTE for risk modelsOmics classification almost always lives in p>>n (thousands of features, tens-to-hundreds of samples). Two counterintuitive consequences follow. First, more flexible is not better: with n in the dozens the variance of a flexible learner dominates, the full covariance is singular so QDA/full-LDA are undefined, and simple diagonal/linear methods match or beat elaborate ones (Dudoit 2002). "Random forest is the obvious choice for expression data" is a myth -- SVM/regularized logistic frequently win on microarray-style problems (Statnikov 2008), and gradient-boosted trees beat deep nets on tabular/omics data (Grinsztajn 2022). Regularization is the load-bearing wall, not a tuning nicety.
Second, in diagnostic/prognostic use the probability is the product, not the label -- which makes calibration, not accuracy, the thing that breaks silently (Van Calster 2019). A model can rank perfectly (AUC 0.9) and still output dishonest risks. And the most common cause of a beautiful AUC is not skill but a batch artifact: if batch correlates with the outcome, the classifier learns the cleaner technical signal and the performance collapses on any independent cohort.
| Model | Wins when | Overfits / fails when | Calibration | Scaling |
|---|---|---|---|---|
| L1 logistic (lasso) | Sparse signal, want a small signature | Correlated features -> unstable selection; >n true signals | Good (proper loss); shrinks toward base rate | Standardize |
| L2 / elastic-net logistic | Many small correlated effects; omics default | Needs C (and l1_ratio) tuning | Good; preferred when calibration matters | Standardize |
| DLDA / nearest-centroid | Tiny n, roughly linear (Dudoit 2002) | Strong interactions; non-Gaussian | Crude; recalibrate | Variance-scaled |
| Linear SVM | High-dim linear separability (Statnikov 2008) | Heavy overlap; needs C | No native probabilities -- Platt-scale decision_function | Critical |
| Random forest | Nonlinear/interaction signal; robust baseline | Sparse-linear signal; tiny n; OOB-as-test leakage | Bagged votes bounded away from 0 and 1 | Scale-invariant |
| GBDT (XGBoost/LightGBM) | Best general tabular performer | Tiny n + deep/many rounds; needs early stopping | Log-loss overfitting tends to overconfident extremes | Scale-invariant |
| Tabular deep nets | Very large n; multimodal/transfer | Typical omics n -> loses to GBDT | Variable; often needs temperature scaling | Standardize |
Tree ensembles are often miscalibrated and the direction depends on the learner and loss: classic boosted ensembles push probabilities toward 0.5 (sigmoid distortion; Niculescu-Mizil 2005), bagged forests are comparatively well-calibrated but bound their votes away from 0 and 1, and modern gradient boosting trained to log-loss for many rounds tends to overfit toward overconfident extremes. Check a reliability curve and recalibrate rather than assuming a direction.
| Scenario | Recommended approach | Why |
|---|---|---|
| Default omics classifier, want a signature | Elastic-net logistic | Often the ceiling in p>>n; sparse + grouping; well-calibrated |
| Suspected nonlinear/interaction (epistasis, thresholds) | Random forest then XGBoost with early stopping | Trees capture interactions; benchmark vs the linear baseline |
| Probabilities will drive a clinical decision | Linear model + calibration check; recalibrate if needed | Probability is the product; AUC is blind to calibration |
| Class imbalance | class_weight='balanced' or threshold tuning; never SMOTE for risk | Resampling destroys calibration for no AUC gain |
| Mixed continuous + categorical features | ColumnTransformer (scale continuous, encode categorical) | Different feature types need different handling |
| Missing values, especially below-detection | XGBoost/LightGBM native NaN handling | Missingness is often informative (MNAR) |
| Considering a deep net | Only at very large n or multimodal/raw inputs | GBDT beats deep on engineered omics matrices (Grinsztajn 2022) |
| Need unbiased performance / nested CV / calibration metrics | -> machine-learning/model-validation | Evaluation is its own discipline |
| Time-to-event outcome | -> machine-learning/survival-analysis | Censoring needs survival models, not classifiers |
Goal: A calibrated, interpretable baseline that is often the best omics classifier.
Approach: Standardize inside a Pipeline and fit elastic-net logistic with cross-validated penalty; the L2 component keeps correlated genes together, the L1 component yields a sparse signature.
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegressionCV
# saga supports elasticnet; C = 1/lambda (small C = strong shrinkage). Standardize: the penalty is scale-sensitive.
clf = LogisticRegressionCV(penalty='elasticnet', solver='saga', l1_ratios=[0.1, 0.5, 0.9],
Cs=20, cv=5, max_iter=10000, class_weight='balanced')
pipe = Pipeline([('scaler', StandardScaler()), ('clf', clf)])
pipe.fit(X_train, y_train)Goal: Capture nonlinear and interaction structure when the linear baseline leaves signal on the table.
Approach: Random forest needs no scaling and is a robust baseline; XGBoost needs a low learning rate, shallow depth, and early stopping (set in the constructor in 2.x) to avoid overfitting tiny n.
from sklearn.ensemble import RandomForestClassifier
from xgboost import XGBClassifier
rf = RandomForestClassifier(n_estimators=500, max_features='sqrt', min_samples_leaf=3,
class_weight='balanced', n_jobs=-1, random_state=0)
# XGBoost 2.x: early_stopping_rounds and eval_metric go in the CONSTRUCTOR, not fit().
# scale_pos_weight is omitted on purpose: like resampling, it reweights the prior and
# distorts calibration -- use it only for hard-label problems, not risk models (see Class Imbalance).
xgb = XGBClassifier(n_estimators=2000, learning_rate=0.03, max_depth=4, subsample=0.8,
colsample_bytree=0.5, reg_lambda=1.0,
early_stopping_rounds=50, eval_metric='aucpr', n_jobs=-1, random_state=0)
xgb.fit(X_train, y_train, eval_set=[(X_val, y_val)]) # NaN handled natively (missing=np.nan)Goal: Rule out that a high AUC is a batch artifact rather than biology.
Approach: Try to predict the batch from the features and use batch-aware splits; if batch is confounded with the outcome, no correction rescues the design (Soneson 2014) -- fix it at the design stage.
from sklearn.model_selection import cross_val_score, StratifiedGroupKFold
from scipy.stats import chi2_contingency
import pandas as pd
# 1. Can the classifier predict the BATCH? If yes, batch is a strong axis and the label model is suspect.
batch_auc = cross_val_score(pipe, X, batch_labels, cv=5, scoring='roc_auc')
print(f'Batch predictability AUC: {batch_auc.mean():.2f} (high = shortcut risk)')
# 2. Is the outcome associated with batch by design?
print('label vs batch p:', chi2_contingency(pd.crosstab(y, batch_labels))[1])
# 3. Leave-one-batch-out is the honest generalization estimate (usually << random-split CV).
gcv = StratifiedGroupKFold(n_splits=5)
honest = cross_val_score(pipe, X, y, cv=gcv, groups=batch_labels, scoring='roc_auc')
print(f'Batch-aware AUC: {honest.mean():.2f}')Goal: Handle a rare positive class without destroying the probabilities.
Approach: For a risk model, do not resample -- class-weight cautiously or tune the threshold on a validation fold; SMOTE/oversampling change the training prior, inflate minority probabilities, give no AUC gain, and the same sensitivity is recoverable by moving the threshold (van den Goorbergh 2022). When resampling is unavoidable (a hard-label problem), use an imblearn Pipeline so only training folds are resampled.
from imblearn.pipeline import Pipeline as ImbPipeline # NOT sklearn's Pipeline
from imblearn.over_sampling import SMOTE
from sklearn.linear_model import LogisticRegression
# Correct placement: SMOTE's fit_resample runs only during fit on the train fold, no-op on transform.
imb = ImbPipeline([('smote', SMOTE(random_state=0)),
('clf', LogisticRegression(max_iter=5000))])
# Prefer for risk models: no resampling, then pick the operating threshold by cost on a validation fold.Goal: Ensure a "0.9" means a 90% risk, not just a high rank.
Approach: Tree ensembles are often miscalibrated -- bagged forests bound votes away from 0 and 1, classic boosting is sigmoid-distorted toward 0.5 (Niculescu-Mizil 2005), and log-loss GBDT can overfit to overconfident extremes -- so check a reliability curve and recalibrate on a disjoint fold. Logistic regression optimizes a proper scoring rule and is usually best-calibrated out of the box. See machine-learning/model-validation for reliability curves, Brier, and the full protocol.
from sklearn.calibration import CalibratedClassifierCV
from sklearn.frozen import FrozenEstimator # sklearn >=1.6; cv='prefit' deprecated
calibrated = CalibratedClassifierCV(FrozenEstimator(rf.fit(X_tr, y_tr)), method='isotonic')
calibrated.fit(X_cal, y_cal) # X_cal disjoint from train and test| Model | Tune these | Leave default |
|---|---|---|
| Logistic | C (log-spaced), l1_ratio, class_weight | solver (saga for elasticnet) |
| Random forest | max_features, min_samples_leaf, max_depth (cap for tiny n) | n_estimators (more is safe; 500-1000) |
| XGBoost | learning_rate+n_estimators+early stopping, max_depth (3-6), subsample, colsample_bytree, reg_lambda | most others |
predict_proba from RF or XGBoost as a calibrated risk.imblearn Pipeline (train-fold only); for risk models, do not resample -- tune the threshold.| Threshold | Source | Rationale |
|---|---|---|
| Try a regularized linear model first | Dudoit 2002; Statnikov 2008 | Simple often beats complex in p>>n |
| GBDT over deep nets for tabular omics | Grinsztajn 2022 | Trees handle uninformative features and non-rotational data |
| Do not resample for risk models | van den Goorbergh 2022; Carriero 2025 | Resampling destroys calibration for no AUC gain |
| XGBoost: low LR + many rounds + early stopping | field standard | Prevents overfitting tiny n |
| Report AUPRC + MCC under imbalance | Saito 2015; Chicco 2020 | Accuracy and ROC-AUC mislead when positives are rare |
| Error / symptom | Cause | Solution |
|---|---|---|
XGBoost early_stopping_rounds TypeError in fit() | Moved to constructor in 2.x | Pass it (and eval_metric) in XGBClassifier(...) |
penalty='l1' FutureWarning | Deprecated in sklearn 1.8 | Use l1_ratio=1 + C (1.8+) or keep penalty on 1.4-1.7 |
elasticnet solver error | Only saga supports it | solver='saga' + l1_ratio |
CalibratedClassifierCV(cv='prefit') deprecated | sklearn 1.6 | Wrap in FrozenEstimator |
| 95% accuracy but useless model | Imbalance + accuracy metric | Report AUPRC/MCC; check the confusion matrix |
© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files in machine-learning/omics-classifiers of GPTomics/bioSkills.
Open the folder on GitHubat commit d91ed3d
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.
Bio Machine Learning Omics Classifiers next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Bio Machine Learning Omics Classifiers this skillGPTomics/bioSkills | 1.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Scikit LearnzLanqing/codex-claude-academic-skills | 4.6k | 17 repos | ~3.9k | Automated safety check: Pass | BSD-3-Clause | |
| Senior Data ScientistRaidriar7170/hermes-skilleval | 125 | 6 repos | ~1.4k | Automated safety check: Pass | MIT | |
| Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill | 188 | — | ~4k | Automated safety check: Pass | MIT | |
| Retention Analysisliangdabiao/claude-data-analysis-ultra-main | 290 | 1 repos | ~1.3k | Automated safety check: Notes | None | |
| Geomlitalo-goncalves/geoML | 108 | — | ~4.2k | Automated safety check: Pass | GPL-3.0 |
zLanqing/codex-claude-academic-skills
Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.
Raidriar7170/hermes-skilleval
World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.
FrankS-IntelLab/agentic-kaggle-skill
Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.
liangdabiao/claude-data-analysis-ultra-main
Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.
italo-goncalves/geoML
Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…
Aperivue/medsci-skills
A skill your agent uses when building or auditing a radiomics or tabular clinical-ML prediction model with a classical learner (LASSO, SVM, random forest, XGBoost and similar).
GPTomics/bioSkills
Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.
GPTomics/bioSkills
Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.
GPTomics/bioSkills
Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.
GPTomics/bioSkills
Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.
GPTomics/bioSkills
Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.
GPTomics/bioSkills
Sort alignment files by coordinate or read name using samtools and pysam.
Categories
Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut…. Bio Machine Learning Omics Classifiers is an agent skill from GPTomics/bioSkills. Builds diagnostic and prognostic classifiers on omics feature matrices with regularized logistic regression, random forest, and gradient-boosted trees, handling the pn regime, batch shortcut learning, class imbalance, and probability calibration.
Bio Machine Learning Omics Classifiers fits situations like: building a classifier from expression; choosing an algorithm for high-dimensional small-n data; diagnosing a suspiciously perfect AUC.
Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a claude-code`. Or copy the skill folder (machine-learning/omics-classifiers in GPTomics/bioSkills) into .claude/skills/bio-machine-learning-omics-classifiers in your project. Claude Code loads it when a task matches its description.
Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a codex`. Or copy the skill folder (machine-learning/omics-classifiers in GPTomics/bioSkills) into .agents/skills/bio-machine-learning-omics-classifiers in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-omics-classifiers -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-machine-learning-omics-classifiers, .gemini/skills/bio-machine-learning-omics-classifiers, .github/skills/bio-machine-learning-omics-classifiers and .opencode/skills/bio-machine-learning-omics-classifiers in your project.
Going by SKILL.md and its folder, Bio Machine Learning Omics Classifiers needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Bio Machine Learning Omics Classifiers is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Bio Machine Learning Omics Classifiers: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.6k stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars) and Retention Analysis (liangdabiao/claude-data-analysis-ultra-main, 290 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,215 GitHub stars. The repository holds 553 skills in this directory. The repository was last updated on August 15, 2026.
Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.