Agent skill

Bio Machine Learning Prediction Explanation

by GPTomics in GPTomics/bioSkills

Explains ML predictions on omics data with SHAP, LIME, and permutation importance, handling the correlated-feature trap, the conditional-vs-interventional Shapley choice, and the…

MITAuto-check passedData & Analytics

Install Bio Machine Learning Prediction Explanation

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-machine-learning-prediction-explanation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-machine-learning-prediction-explanation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/machine-learning/prediction-explanation .claude/skills/bio-machine-learning-prediction-explanation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-machine-learning-prediction-explanation
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.3k tokens
SKILL.md length
1,922 words
Files
4
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Explains ML predictions on omics data with SHAP, LIME, and permutation importance, handling the correlated-feature trap, the conditional-vs-interventional Shapley choice, and the…

  • Interpreting an omics classifier
  • SKILL.md covers Version Compatibility, The Single Most Important…, Method Taxonomy and The Correlated-Feature /…, plus 11 more sections
  • Runs Python scripts from its folder; calls pip
  • Debugging shortcut/batch learning

What it does

Bio Machine Learning Prediction Explanation is an agent skill from GPTomics/bioSkills. Explains ML predictions on omics data with SHAP, LIME, and permutation importance, handling the correlated-feature trap, the conditional-vs-interventional Shapley choice, and the attribution-is-not-causation boundary. Use when interpreting an omics classifier, debugging shortcut/batch learning, or deciding whether an attribution ranking can be trusted as biology. For validated feature selection see machine-learning/biomarker-discovery; explanations are not a selection method.

Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files (for example `examples/lime_explanation.py`, `examples/shap_omics_classifier.py` and `usage-guide.md`).

It sits in Data & Analytics, covering Machine learning. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Interpreting an omics classifier
  • Debugging shortcut/batch learning
  • Deciding whether an attribution ranking can be trusted as biology

Example prompts

  • “Use the bio-machine-learning-prediction-explanation skill to explain ML predictions on omics data with SHAP, LIME, and permutation importance…”
  • “/bio-machine-learning-prediction-explanation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Machine Learning Prediction Explanation loads about 4.3k tokens when it runs. Until then it costs about 131 tokens; SKILL.md has 1,922 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~131
When it runs · the whole SKILL.md, loaded when a task matches
~4.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,922 words, ~4,291 tokens.

Download SKILL.mdSave it as .claude/skills/bio-machine-learning-prediction-explanation/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
bio-machine-learning-prediction-explanation
description
Explains ML predictions on omics data with SHAP, LIME, and permutation importance, handling the correlated-feature trap, the conditional-vs-interventional Shapley choice, and the attribution-is-not-causation boundary. Use when interpreting an omics classifier, debugging shortcut/batch learning, or deciding whether an attribution ranking can be trusted as biology. For validated feature selection see machine-learning/biomarker-discovery; explanations are not a selection method.
tool_type
python
primary_tool
shap

Version Compatibility

Reference examples tested with: numpy 1.26+, pandas 2.2+, scikit-learn 1.4+, shap 0.44+, lime 0.2+.

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures

Two shap drifts: the Explanation-object plotting API arrived ~0.36 (new-style shap.plots.* take an Explanation, legacy shap.summary_plot/dependence_plot take numpy arrays -- mixing them is the most common runtime error); and TreeExplainer(..., feature_perturbation='auto') became the default in 0.47 (was interventional), so providing or omitting data= silently changes the estimand. Always set feature_perturbation explicitly. If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt rather than retrying.

Model Interpretation for Omics Classifiers

"Which genes drive my classifier?" -> Compute attributions, but treat them as a description of the model (not biology), choose the Shapley conditioning deliberately, and aggregate over correlated gene modules before ranking.

  • Tree models: shap.TreeExplainer(model, data=background, feature_perturbation='interventional')
  • Model-agnostic local: lime.lime_tabular.LimeTabularExplainer
  • Model-reliance screen: sklearn.inspection.permutation_importance

The Single Most Important Modern Insight -- Attributions Explain the Model, Not Biology, and Under Correlation the Algorithm Chooses How to Split Credit

A feature attribution describes the function the model learned on this training distribution; it is not a measurement of biology. A high-SHAP gene can be a pure correlate of a batch, scanner, or library-prep signal the model exploited (DeGrave 2021 is the canonical proof). And because genes co-express in tight modules, the attribution algorithm has genuine freedom in how it splits credit among correlated genes -- the choice of conditioning (tree_path_dependent/conditional vs interventional/marginal) is not a cosmetic knob, it changes which genes get credit, and it is a live methodological controversy with no universally correct answer (Janzing 2020). The operational rule: SHAP/LIME rankings are a debugging and hypothesis-generation tool, never a validated biomarker-selection criterion, and within a co-expression module the ordering is not a finding.

Method Taxonomy

MethodWhat it estimatesCorrelated-feature behaviorCostBest use
TreeSHAP tree_path_dependentConditional Shapley via tree coverage; approximates E[f | x_S]Can give nonzero credit to a feature the model never uses (correlation leak); no background neededFast, exact for this estimandFast cohort summaries when conditional semantics are acceptable
TreeSHAP interventionalMarginal/do-operator Shapley; features replaced from a backgroundZero credit to unused features even if correlatedScales with background size (~100-1000)"What the model actually uses"; most defensible default
KernelSHAPModel-agnostic Shapley via masking; assumes independenceMasking lands off-manifold under correlation; corruptedExpensiveLast resort for non-tree/non-net models
DeepSHAP / GradientSHAPSHAP for nets via backprop relative to a backgroundBackground-dependent under correlationModerateNeural omics models
LinearExplainerExact Shapley for linear modelsinterventional vs correlation_dependent give different valuesCheapPenalized linear models; choose the mode
LIMELocal sparse linear surrogate on perturbed samplesOff-manifold perturbations; unstable across seedsModerateEyeballing one local prediction, never global ranking
Permutation importanceDrop in score when a feature is shuffledShuffling A while correlated B intact zeros BOTH; extrapolatesn_repeats x n_featuresGlobal screen on decorrelated features
Conditional permutation (Strobl)Importance permuting within correlated strataFairer among correlated predictorsHigherRF importance under correlation

The Correlated-Feature / Conditional-vs-Marginal Problem (the core)

When Shapley values "drop" a feature subset, they replace it by some distribution, and two incompatible choices exist:

  • Conditional / observational (tree_path_dependent): dropped features drawn from p(x_dropped | x_S). A feature the model never uses can still receive nonzero attribution purely because it is correlated with a used feature. So high SHAP does not mean the model relies on this gene.
  • Marginal / interventional (interventional): dropped features drawn from the marginal p(x_dropped), i.e. do(x_dropped = background). Features the model genuinely ignores get exactly zero, even if correlated. Janzing 2020 argues this is the principled "drop" for attribution; it is why modern SHAP added and (pre-0.47) defaulted to it.

There is no free lunch: every method either extrapolates off-manifold (interventional SHAP, unrestricted permutation -- Hooker 2021's "no free variable importance") or leaks credit through correlation (conditional SHAP). Decide which pathology the question can tolerate, and aggregate attributions over co-expression modules before ranking -- "gene A ranked above gene B" within a correlated module is governed by off-manifold value-function behavior, not biology.

Decision Tree by Scenario

ScenarioRecommended approachWhy
"Which genes is my model actually keying on" (debug shortcuts)TreeSHAP interventional with a representative backgroundGives unused genes zero; reveals true reliance
"Which genes are informative about the outcome here" (descriptive)TreeSHAP tree_path_dependent, but never call it model relianceConditional semantics answer the descriptive question
Ranking importance across correlated genesAggregate |SHAP| within co-expression clusters firstWithin-module order is arbitrary
Penalized linear modelLinearExplainer (choose interventional vs correlation_dependent)Or just read the coefficients -- the model is its own explanation
One local prediction, communicationLIME or SHAP waterfall, pinned seed, labeled model-internalLocal surrogate; not global, not reproducible across seeds
"I want to pick a biomarker panel"-> machine-learning/biomarker-discoverySHAP ranking is not validated selection (no FDR, no replication)
High-stakes clinical decisionPrefer an inherently interpretable modelSparse linear/rule list is exact; avoids the conditional-vs-marginal ambiguity (Rudin 2019)

SHAP TreeExplainer (set the conditioning explicitly)

Goal: Attribute a tree model's predictions to features with a chosen, stated estimand.

Approach: Pass a background and feature_perturbation='interventional' for "what the model uses," or omit the background and use tree_path_dependent for the conditional/descriptive view. The default 'auto' flips between them based on whether data= is given.

python
import shap
import numpy as np

# Interventional ('what the model uses'): needs a background (~100-1000 rows).
background = shap.utils.sample(X_train, 200)
explainer = shap.TreeExplainer(model, data=background, feature_perturbation='interventional')
sv = explainer(X_test)                                  # modern Explanation object

# Aggregate over correlated modules BEFORE ranking (clusters = a precomputed gene->module map).
mean_abs = np.abs(sv.values).mean(axis=0)
module_importance = {}
for gene, m in zip(X_test.columns, mean_abs):
    module_importance[clusters[gene]] = module_importance.get(clusters[gene], 0) + m

Attribution Is Not Causation, Mechanism, or a Validated Biomarker

Three layers of "not": (1) not biology -- the attribution describes the model, which may have exploited a batch/confounder shortcut; (2) not causation -- predictive features conflate direct effects, confounders, mediators, and colliders, and turning attribution into a causal claim needs an explicit causal model SHAP does not contain; (3) not a validated biomarker -- taking "top-20 SHAP genes" as a panel is same-data feature selection with no FDR and no replication (winner's curse). The strongest legitimate use runs the other way: because attribution exposes what the model used, it is one of the best tools to catch shortcut/batch learning -- if top features are batch indicators or depth-tracking housekeeping genes, the model is cheating (DeGrave 2021). That is where attribution is most trustworthy.

LIME and Explanation Instability

LIME fits a sparse linear surrogate to predictions on perturbed samples around one instance. It is non-reproducible by construction: different seeds, different kernel_width, and discretize_continuous=True (the default) each flip the top features, and the per-feature perturbations land off-manifold for correlated genes. Worse, perturbation-based explainers can be deliberately fooled -- a biased model can be wrapped to look innocuous on the out-of-distribution points LIME/KernelSHAP probe (Slack 2020). Use LIME only to eyeball a single prediction's local logic with a pinned seed, never for global ranking, and never as evidence a model is unbiased.

python
from lime.lime_tabular import LimeTabularExplainer

explainer = LimeTabularExplainer(X_train.values, feature_names=list(X_train.columns),
                                 mode='classification', discretize_continuous=True,
                                 random_state=0)         # pin the seed; still only conditional stability
exp = explainer.explain_instance(X_test.values[0], model.predict_proba, num_features=10, num_samples=5000)

Background / Baseline Choice (the silent attribution-changer)

SHAP explains the deviation from E[f(X)] over the background dataset, so the background defines what "absence of a feature" means and changes every attribution. A tumor sample explained against a tumor-heavy vs a healthy-tissue background yields different "important genes" -- only one matches the scientific question. Use a background of real samples representative of the contrast of interest (a single global mean across a heterogeneous cohort is no real sample). check_additivity=True (default) raises when the SHAP values plus the base value do not sum to the model output (a local-accuracy violation) -- often a probability-vs-raw-margin or implementation mismatch; investigate rather than disabling it. Attributions in log-odds (model_output='raw') differ from probability space -- report the scale.

Show full SKILL.md (716 more words)Show less

Permutation Importance Also Breaks Under Correlation

A common error is to "fix" SHAP's correlation problem by switching to permutation importance, but it has the same root pathology: shuffling gene A while correlated gene B is intact lets the model recover the signal through B, so both look unimportant (scikit-learn's own multicollinearity example). Unrestricted permutation also forces the model to predict on points that cannot occur (Hooker 2021). For correlated omics predictors, cluster features and keep one per cluster, or use Strobl's conditional permutation (R party::cforest, varimp(conditional=TRUE)) -- there is no conditional= flag in sklearn's permutation_importance. Always evaluate permutation importance on held-out data, not training data.

Per-Method Failure Modes

Reading high tree_path_dependent SHAP as model reliance
  • Trigger: Using path-dependent SHAP (no background) and concluding the model depends on a top gene.
  • Mechanism: Conditional Shapley leaks credit to unused-but-correlated features.
  • Symptom: A gene the model never splits on ranks high.
  • Fix: Use interventional with a background for reliance questions; state the estimand.
Within-module ranking treated as a finding
  • Trigger: Reporting "gene A more important than gene B" for co-expressed A, B.
  • Mechanism: Shapley fairly splits credit, but the split is governed by off-manifold behavior, not biology.
  • Symptom: The credit split (and sometimes the order) changes with the conditioning mode or when the other module member is dropped.
  • Fix: Aggregate |SHAP| over co-expression modules before ranking.
SHAP ranking used as feature selection
  • Trigger: Taking top-k SHAP genes as a biomarker panel.
  • Mechanism: Same-data selection with no FDR/replication; winner's curse.
  • Symptom: The panel fails to replicate in an independent cohort.
  • Fix: Generate hypotheses with SHAP, validate with biomarker-discovery + independent data.
LIME/KernelSHAP global aggregates trusted
  • Trigger: Averaging LIME or KernelSHAP across samples for a global ranking.
  • Mechanism: Seed/kernel instability + off-manifold perturbation; adversarially foolable (Slack 2020).
  • Symptom: Ranking changes across runs.
  • Fix: Use TreeSHAP/LinearExplainer for global; keep LIME local.

Quantitative Thresholds

ThresholdSourceRationale
Background ~100-1000 rowsshap docsInterventional/Kernel SHAP cost scales with background; shap warns above ~1000
Set feature_perturbation explicitlyshap 0.47 changelog'auto' default silently flips estimand on data= presence
Aggregate over co-expression modules before rankingJanzing 2020; Aas 2021Within-module order is not identifiable from attributions
Validate SHAP-derived genes in independent cohortsRudin 2019Attribution rankings are model-internal, not replicated associations

Common Errors

Error / symptomCauseSolution
shap.plots.beeswarm errors on a numpy arrayExplanation-object API since ~0.36Pass explainer(X) (Explanation); use legacy summary_plot for arrays
Attribution estimand changed silentlyfeature_perturbation='auto' (0.47+)Set 'interventional' or 'tree_path_dependent' explicitly
check_additivity raisesSHAP values + base do not sum to output (local-accuracy)Fix the config (raw vs probability); do not just disable
interventional errorsNo data= background suppliedProvide a background sample
Permutation importance zeros real featuresCorrelation dilutionCluster features or use conditional permutation

References

  • Ribeiro MT, Singh S, Guestrin C. 2016. "Why Should I Trust You?": Explaining the Predictions of Any Classifier. Proc KDD 1135-1144.
  • Lundberg SM, Lee S-I. 2017. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst 30:4765-4774.
  • Strobl C, Boulesteix A-L, Kneib T, Augustin T, Zeileis A. 2008. Conditional variable importance for random forests. BMC Bioinformatics 9:307.
  • Rudin C. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nat Mach Intell 1:206-215.
  • Lundberg SM, Erion G, Chen H, et al. 2020. From local explanations to global understanding with explainable AI for trees. Nat Mach Intell 2:56-67.
  • Janzing D, Minorics L, Blobaum P. 2020. Feature relevance quantification in explainable AI: a causal problem. Proc AISTATS PMLR 108:2907-2916.
  • Kumar IE, Venkatasubramanian S, Scheidegger C, Friedler S. 2020. Problems with Shapley-value-based explanations as feature importance measures. Proc ICML PMLR 119:5491-5500.
  • Slack D, Hilgard S, Jia E, Singh S, Lakkaraju H. 2020. Fooling LIME and SHAP: adversarial attacks on post hoc explanation methods. Proc AIES.
  • Aas K, Jullum M, Loland A. 2021. Explaining individual predictions when features are dependent: more accurate approximations to Shapley values. Artif Intell 298:103502.
  • DeGrave AJ, Janizek JD, Lee S-I. 2021. AI for radiographic COVID-19 detection selects shortcuts over signal. Nat Mach Intell 3:610-619.
  • Hooker G, Mentch L, Zhou S. 2021. Unrestricted permutation forces extrapolation: variable importance requires at least one more model. Stat Comput 31:82.
  • machine-learning/omics-classifiers - Train the model being explained; debug batch shortcuts
  • machine-learning/biomarker-discovery - Validated feature selection (SHAP ranking is not selection)
  • machine-learning/model-validation - Confirm the model generalizes before interpreting it
  • data-visualization/heatmaps-clustering - Visualize module-aggregated attributions

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files in machine-learning/prediction-explanation of GPTomics/bioSkills.

  • SKILL.md
  • examples/lime_explanation.py
  • examples/shap_omics_classifier.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Machine Learning Prediction Explanation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Machine Learning Prediction Explanation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Machine Learning Prediction Explanation this skillGPTomics/bioSkills1.2k1 repos~4.3kAutomated safety check: PassMIT
Scikit LearnzLanqing/codex-claude-academic-skills4.7k16 repos~3.9kAutomated safety check: PassBSD-3-Clause
Agentic Kaggle WorkflowFrankS-IntelLab/agentic-kaggle-skill188—~4kAutomated safety check: PassMIT
Senior Data ScientistRaidriar7170/hermes-skilleval1255 repos~1.4kAutomated safety check: PassMIT
Geomlitalo-goncalves/geoML109—~4.6kAutomated safety check: PassGPL-3.0
QuantMind Training Config Generatorqusong0627/QuantMind1.7k—~1.5kAutomated safety check: PassAGPL-3.0

Similar skills

  • Scikit Learn

    zLanqing/codex-claude-academic-skills

    Machine learning in Python with scikit-learn. An agent skill from zLanqing/codex-claude-academic-skills.

    4.7k GitHub starsUsed in 16 repos~3.9k tokens
    Data & AnalyticsAuto-check passed
  • Agentic Kaggle Workflow

    FrankS-IntelLab/agentic-kaggle-skill

    Takes a Kaggle competition from rules and validation design through baselines, ensembling and notebook architecture to a scored submission.

    188 GitHub stars~4k tokensUpdated 3 mo ago
    Data & AnalyticsAuto-check passed
  • Senior Data Scientist

    Raidriar7170/hermes-skilleval

    World-class data science skill for statistical modeling, experimentation, causal inference, and advanced analytics.

    125 GitHub starsUsed in 5 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Geoml

    italo-goncalves/geoML

    Working knowledge of the geoML Python package (github.com/italo-goncalves/geoML): variational Gaussian processes for spatial data, implicit geological modelling, block models, drillhole data…

    109 GitHub stars~4.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Turns a plain-language model training request into a validated QuantMind training config file that can be imported from the Model Training page.

    1.7k GitHub stars~1.5k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Retention Analysis

    liangdabiao/claude-data-analysis-ultra-main

    Analyze user retention and churn using survival analysis, cohort analysis, and machine learning.

    290 GitHub stars~1.3k tokensUpdated 5 mo ago
    Data & AnalyticsAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Machine Learning Prediction Explanation

What does Bio Machine Learning Prediction Explanation do?

Explains ML predictions on omics data with SHAP, LIME, and permutation importance, handling the correlated-feature trap, the conditional-vs-interventional Shapley choice, and the…. Bio Machine Learning Prediction Explanation is an agent skill from GPTomics/bioSkills. Explains ML predictions on omics data with SHAP, LIME, and permutation importance, handling the correlated-feature trap, the conditional-vs-interventional Shapley choice, and the attribution-is-not-causation boundary.

When should I use Bio Machine Learning Prediction Explanation?

Bio Machine Learning Prediction Explanation fits situations like: interpreting an omics classifier; debugging shortcut/batch learning; deciding whether an attribution ranking can be trusted as biology.

How do I install Bio Machine Learning Prediction Explanation in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-prediction-explanation -a claude-code`. Or copy the skill folder (machine-learning/prediction-explanation in GPTomics/bioSkills) into .claude/skills/bio-machine-learning-prediction-explanation in your project. Claude Code loads it when a task matches its description.

How do I install Bio Machine Learning Prediction Explanation in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-prediction-explanation -a codex`. Or copy the skill folder (machine-learning/prediction-explanation in GPTomics/bioSkills) into .agents/skills/bio-machine-learning-prediction-explanation in your project. Codex loads it when a task matches its description.

Can I use Bio Machine Learning Prediction Explanation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-machine-learning-prediction-explanation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-machine-learning-prediction-explanation, .gemini/skills/bio-machine-learning-prediction-explanation, .github/skills/bio-machine-learning-prediction-explanation and .opencode/skills/bio-machine-learning-prediction-explanation in your project.

What does Bio Machine Learning Prediction Explanation need to run?

Going by SKILL.md and its folder, Bio Machine Learning Prediction Explanation needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Machine Learning Prediction Explanation access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Machine Learning Prediction Explanation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Machine Learning Prediction Explanation use?

Bio Machine Learning Prediction Explanation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Machine Learning Prediction Explanation use?

About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Machine Learning Prediction Explanation?

Skills that share tags, products or a category with Bio Machine Learning Prediction Explanation: Scikit Learn (zLanqing/codex-claude-academic-skills, 4.7k stars), Agentic Kaggle Workflow (FrankS-IntelLab/agentic-kaggle-skill, 188 stars), Senior Data Scientist (Raidriar7170/hermes-skilleval, 125 stars) and Geoml (italo-goncalves/geoML, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Machine Learning Prediction Explanation?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,217 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.