Agent skill

Bio Proteomics Protein Inference

by GPTomics in GPTomics/bioSkills

Groups proteins from peptide identifications and controls protein-level FDR, framing inference as a chosen explanation (parsimony or a probability model) of underdetermined peptide evidence rather…

MITAuto-check passedResearch & Science

Install Bio Proteomics Protein Inference

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-protein-inference -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-proteomics-protein-inference --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/protein-inference .claude/skills/bio-proteomics-protein-inference && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-proteomics-protein-inference
GitHub stars
1.2k
Used in
1 other repo
Token cost
~4.7k tokens
SKILL.md length
1,955 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Groups proteins from peptide identifications and controls protein-level FDR, framing inference as a chosen explanation (parsimony or a probability model) of underdetermined peptide evidence rather…

  • Works in 3 steps: The protein set is not uniquely… → Protein FDR is its own estimation… → The two-peptide rule is wrong -- it…
  • Resolving which proteins are present from a peptide list
  • SKILL.md covers Version Compatibility, The Single Most Important…, Vocabulary the Rest of This… and Tool Taxonomy, plus 6 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Proteomics Protein Inference is an agent skill from GPTomics/bioSkills. Groups proteins from peptide identifications and controls protein-level FDR, framing inference as a chosen explanation (parsimony or a probability model) of underdetermined peptide evidence rather than a measurement. Reports protein GROUPS (proteins indistinguishable by observed peptides) with a leading protein, not flat lists. Covers shared-vs-unique peptides, indistinguishable/subsumable proteins, parsimony vs probabilistic (ProteinProphet, EPIFANY) vs razor inference, picked-protein and picked-group FDR, and…

Its SKILL.md is about 4.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/protein_groups.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics. It works with Python. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Resolving which proteins are present from a peptide list
  • Building protein groups
  • Estimating protein-level FDR

Example prompts

  • “Use the bio-proteomics-protein-inference skill to group proteins from peptide identifications and controls protein-level FDR, framing inference as a…”
  • “/bio-proteomics-protein-inference”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. The protein set is not uniquely recoverable from peptides, so a protein group -- not a flat protein list -- is the only honest reporting…
  2. Protein FDR is its own estimation problem that INFLATES on large data; the fix is PICKED FDR, not the PSM formula reused. Controlling…
  3. The two-peptide rule is wrong -- it increases protein FDR and discards real proteins. Requiring >=2 peptides per protein (Gupta & Pevzner…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Proteomics Protein Inference loads about 4.7k tokens when it runs. Until then it costs about 222 tokens; SKILL.md has 1,955 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~222
When it runs · the whole SKILL.md, loaded when a task matches
~4.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 1,955 words, ~4,740 tokens.

Download SKILL.mdSave it as .claude/skills/bio-proteomics-protein-inference/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-proteomics-protein-inference
description
Groups proteins from peptide identifications and controls protein-level FDR, framing inference as a chosen explanation (parsimony or a probability model) of underdetermined peptide evidence rather than a measurement. Reports protein GROUPS (proteins indistinguishable by observed peptides) with a leading protein, not flat lists. Covers shared-vs-unique peptides, indistinguishable/subsumable proteins, parsimony vs probabilistic (ProteinProphet, EPIFANY) vs razor inference, picked-protein and picked-group FDR, and why the two-peptide rule is wrong. Use when resolving which proteins are present from a peptide list, building protein groups, or estimating protein-level FDR. PSM/peptide FDR and search engines are peptide-identification; razor-vs-unique quant consequences are quantification; isoform/proteoform resolution is top-down and out of scope.
tool_type
mixed
primary_tool
pyOpenMS

Version Compatibility

Reference examples tested with: pyOpenMS 3.1+, pandas 2.2+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

The pyOpenMS protein-inference class names have varied across releases. Confirm the exact spelling at the installed version with help(pyopenms.EpifanyAlgorithm) and help(pyopenms.BasicProteinInferenceAlgorithm) before relying on the reference code.

Protein Inference -- A Chosen Explanation of Peptide Evidence, Reported as Groups

"Tell me which proteins are present from my identified peptides" -> Assign the observed peptides to a minimal or probability-weighted set of proteins, reported as groups of indistinguishable proteins with a leading accession -- because bottom-up MS measures peptides, and the protein set behind them is inferred, not observed.

  • Python: pyopenms.BasicProteinInferenceAlgorithm().run(peptide_ids, protein_ids) for parsimony grouping
  • Python: pyopenms.EpifanyAlgorithm (TOPP tool Epifany) for Bayesian belief-propagation inference
  • CLI: ProteinProphet (TPP) for EM-based probabilistic inference; Philosopher filter for FragPipe FDR

Scope: this skill OWNS peptide-to-protein grouping, the indistinguishable/subsumable distinction, the leading-protein convention, inference-method choice, and protein/protein-group FDR. PSM-level and peptide-level FDR plus the search engines that produce the peptide list -> peptide-identification. The quantitative fallout of razor vs unique peptides on protein abundance -> quantification. OUT OF SCOPE: resolving splice isoforms, single-AA variants, or PTM-defined proteoforms (bottom-up groups cannot separate them; that is top-down / proteoform work).

The Single Most Important Modern Insight -- Protein Inference Is Underdetermined, So the Honest Unit Is a Group, Not a List

  1. The protein set is not uniquely recoverable from peptides, so a protein group -- not a flat protein list -- is the only honest reporting unit. Many peptides are shared across paralogs, gene families, and isoforms, so distinct protein sets can explain the same peptide evidence equally well. The inference picks ONE explanation under an assumption (parsimony, or a probability model); proteins that the observed peptides cannot tell apart (indistinguishable) MUST be reported as one group with a designated leading protein. A flat list double-counts indistinguishable proteins and breaks target/decoy symmetry at the protein level, silently corrupting FDR.

  2. Protein FDR is its own estimation problem that INFLATES on large data; the fix is PICKED FDR, not the PSM formula reused. Controlling PSM-FDR at 1% does not give 1% protein-FDR. A deep run has many false PSMs in absolute terms, and each can nucleate a one-hit-wonder false protein; because true proteins accumulate many peptides while false proteins are hit once, the naive protein-FDR balloons to 10-30% on deep datasets. Savitski 2015 picked-protein FDR pairs each target protein with its decoy and keeps only the higher-scoring of the pair before counting, removing the target/decoy asymmetry; The & Kall 2016 extends this to the group level (picked-group FDR), which is required because parsimony grouping is anticonservative otherwise.

  3. The two-peptide rule is wrong -- it increases protein FDR and discards real proteins. Requiring >=2 peptides per protein (Gupta & Pevzner 2009, "A strike against the two-peptide rule") removes MORE target proteins than decoy proteins, so it raises protein-level FDR rather than lowering it, while throwing away legitimate low-abundance single-peptide IDs. Replace the blanket rule with: control protein-level (picked) FDR, then judge single-peptide IDs by their score, not their peptide count.

Vocabulary the Rest of This Depends On

  • Shared (degenerate) peptide: maps to >1 protein in the searched database. Cannot, alone, distinguish which protein is present.
  • Unique peptide: maps to exactly one protein -- the only direct evidence for a specific protein. "Unique" is DATABASE-RELATIVE: a peptide unique against SwissProt may be shared against TrEMBL+isoforms+contaminants. Always document the exact database (isoforms, contaminants, decoys included).
  • Indistinguishable proteins: explained by the SAME set of observed peptides -> one group, never two confident IDs.
  • Subset / subsumable protein: its observed peptides are a subset of another protein's -> parsimony drops it (the larger protein explains everything it would).
  • Leading / representative protein: the group's reported accession. Convention: most peptides, then highest score, then SwissProt canonical over TrEMBL. Downstream tables key on this accession but must retain group membership -- "protein P12345" usually means "the group led by P12345".
  • Protein group vs proteoform: a group is an inference artifact (proteins lumped because peptides cannot separate them); a proteoform is a real molecular species (one gene product with a specific sequence + PTM + cleavage state). Bottom-up groups DO NOT resolve proteoforms -- claiming "isoform X present" from a shared-peptide group is overreach.

Tool Taxonomy

Tool / methodCitationMechanism / roleWhen
Parsimony (Occam)--Greedy minimal protein set explaining all peptidesFast default; ties broken arbitrarily; anticonservative group-FDR on large data unless picked
ProteinProphetNesvizhskii 2003EM APPORTIONS shared peptides across candidate proteins, weighted by other evidenceTPP / FragPipe pipelines; the classic probabilistic standard
EPIFANYPfeuffer 2020Bayesian network over the peptide-protein graph, loopy belief propagation + convolution treesOpenMS-recommended modern inference; strong at controlled protein-group FDR
Fido--Bayesian generative model (Percolator --protein)Percolator pipelines; superseded by picked-protein for FDR
Razor peptide--Shared peptide assigned winner-take-all to the group with most evidence (MaxQuant)MaxQuant default; ID-fine but distorts QUANT (route to quantification)
Picked-protein FDRSavitski 2015Pair target with its decoy, keep the higher-scoring of the pair, then count decoysProtein-level FDR on any non-trivial dataset
Picked-group FDRThe & Kall 2016Picking applied at the protein-GROUP levelWhen the inference unit is the group (the correct unit on deep data)
All-proteins / inclusive--Report every protein any peptide could come fromAlmost never; massive false-positive protein inflation

Decision Tree by Scenario

ScenarioRecommendedWhy
Standard DDA run, OpenMS-based pipelineEpifanyAlgorithm (or BasicProteinInferenceAlgorithm for parsimony) + picked-group FDRModern, group-FDR aware; well-calibrated on benchmarks
MaxQuant output (proteinGroups.txt)Parse groups as-is; quantify on UNIQUE peptidesGroups already inferred; razor quant is the trap, not the inference
FragPipe / TPP pipelineProteinProphet inference + Philosopher/Philosopher-style FDR filteringNative EM apportionment + 2-level FDR
Deep dataset (many thousands of proteins)Picked-GROUP FDR, NOT naive decoy/targetNaive protein-FDR inflates to 10-30% from one-hit-wonders
Sensitive differential abundance downstreamQuantify on unique peptides only -> quantificationRazor assignment can flip between conditions and fake DE
Want isoform-level answersStop -- route to top-down / proteoform methodsBottom-up groups cannot resolve proteoforms
Few PSMs (single-protein pulldown)Report evidence, do not trust a "0% protein FDR"Target-decoy FDR is meaningless at tiny counts

Default when uncertain: run parsimony grouping (BasicProteinInferenceAlgorithm with annotate_indistinguishable_groups), report protein GROUPS with a leading accession, and control protein-GROUP FDR with picked-group FDR at 1%. Do NOT impose a two-peptide rule.

Group Proteins by Parsimony with pyOpenMS

Goal: Turn an FDR-filtered peptide identification list into protein groups with a leading protein, resolving shared-peptide ambiguity.

Approach: Load the idXML from peptide identification, run the parsimony algorithm with indistinguishable-group annotation on, then read the inferred groups off the protein identification run.

python
from pyopenms import IdXMLFile, BasicProteinInferenceAlgorithm

protein_ids = []
peptide_ids = []
# protein_ids is FIRST in both load() and store() for IdXMLFile
IdXMLFile().load('peptides_1pct_fdr.idXML', protein_ids, peptide_ids)

inference = BasicProteinInferenceAlgorithm()
params = inference.getParameters()
# annotate_indistinguishable_groups reports indistinguishable proteins as ONE group
params.setValue('annotate_indistinguishable_groups', 'true')
inference.setParameters(params)
inference.run(peptide_ids, protein_ids)

# indistinguishable groups live on the protein identification run
for prot_id in protein_ids:
    for group in prot_id.getIndistinguishableProteins():
        leading = group.accessions[0]  # convention: highest-evidence accession first
        print(leading, group.probability, list(group.accessions))
Bayesian Inference + Group FDR with EPIFANY

Goal: Assign calibrated protein/group posteriors and control protein-group FDR with a probability model rather than greedy parsimony.

Approach: EPIFANY consumes idXML whose PSMs already carry posterior error probabilities (from Percolator or IDPosteriorErrorProbability), then propagates belief over the peptide-protein graph. The TOPP tool is reliably named Epifany; the pyOpenMS class spelling has varied across releases, so introspect first.

python
import pyopenms
from pyopenms import IdXMLFile

# CONFIRM the class name at the installed version before use:
#   help(pyopenms.EpifanyAlgorithm)
algo_cls = getattr(pyopenms, 'EpifanyAlgorithm')

protein_ids = []
peptide_ids = []
IdXMLFile().load('peptides_with_pep.idXML', protein_ids, peptide_ids)

algo = algo_cls()
# EPIFANY expects PSM posteriors as input; greedy_group_resolution controls
# whether shared peptides are razor-resolved after inference
algo.inferPosteriorProbabilities(protein_ids, peptide_ids, False)

for prot_id in protein_ids:
    for group in prot_id.getIndistinguishableProteins():
        print(group.accessions[0], group.probability)
Show full SKILL.md (784 more words)Show less
Picked Protein-Group FDR

Goal: Estimate protein-group FDR without the inflation that the reused PSM formula causes on large data.

Approach: For each target group, find its decoy counterpart (same accessions with the decoy prefix); keep only the higher-scoring member of each target/decoy PAIR; rank the picked set and count decoys as the FDR estimate. This is the operation the reference example demonstrates end to end.

python
def picked_group_fdr(groups, decoy_prefix='DECOY_'):
    # groups: list of dicts with 'accessions', 'score', 'is_decoy'
    by_base = {}
    for g in groups:
        base = frozenset(a.replace(decoy_prefix, '') for a in g['accessions'])
        # keep only the higher-scoring of the target/decoy pair (the 'pick')
        if base not in by_base or g['score'] > by_base[base]['score']:
            by_base[base] = g
    picked = sorted(by_base.values(), key=lambda g: g['score'], reverse=True)

    targets = decoys = 0
    for g in picked:
        if g['is_decoy']:
            decoys += 1
        else:
            targets += 1
        g['fdr'] = decoys / targets if targets else 1.0
    running_min = 1.0
    for g in reversed(picked):  # monotone q-values from the bottom up
        running_min = min(running_min, g['fdr'])
        g['qvalue'] = running_min
    return [g for g in picked if not g['is_decoy'] and g['qvalue'] <= 0.01]

Per-Method Failure Modes

Naive (non-picked) protein/group FDR

Trigger: Reusing the PSM-level decoys/targets formula at the protein level on a deep dataset. Mechanism: False target proteins (one-hit-wonders) and decoy proteins are not symmetric once peptides are mapped to proteins; true proteins absorb many peptides, false ones do not. Symptom: Reported 1% protein FDR, actual 10-30%; reviewer or entrapment check exposes it. Fix: Picked-protein FDR (Savitski 2015) or picked-group FDR (The & Kall 2016); validate with a two-species or entrapment search.

Two-peptide rule

Trigger: Filtering to proteins with >=2 (unique) peptides "for confidence". Mechanism: The rule removes more target proteins than decoy proteins, inverting the FDR effect, and deletes real low-abundance single-peptide proteins. Symptom: Fewer proteins AND higher true FDR than picked FDR at the same nominal cutoff. Fix: Drop the rule; control picked protein-level FDR and score single-peptide IDs individually.

Razor-peptide quantification

Trigger: Quantifying on MaxQuant's default unique+razor peptides for a sensitive comparison. Mechanism: A shared peptide's full intensity is credited to one group; that razor assignment can flip between conditions when peptide counts shift, so a protein's quantity changes for inference reasons, not biology. Symptom: Spurious differential abundance concentrated on proteins sharing peptides with paralogs. Fix: Quantify on unique peptides only for sensitive comparisons -> quantification.

Parsimony tie-breaking

Trigger: Multiple minimal protein sets explain the peptides equally well. Mechanism: Greedy parsimony breaks ties arbitrarily; minimality is a heuristic, not truth, and a real protein with only shared peptides is silently dropped. Symptom: Reported lead protein differs run-to-run or pipeline-to-pipeline on the same data. Fix: Prefer a probabilistic method (EPIFANY/ProteinProphet) that apportions shared evidence; retain group membership.

Proteoform overreach

Trigger: Reporting "isoform X is present" from a group whose evidence is shared peptides. Mechanism: Splice isoforms, variants, and PTM forms collapse into groups in bottom-up data; the group cannot separate them. Symptom: Isoform-specific claim with no isoform-unique peptide behind it. Fix: Require an isoform-unique peptide for any isoform claim, or use top-down / proteoform methods.

Quantitative Thresholds

ThresholdSourceRationale
Protein / protein-group FDR 1% (sometimes 5% for discovery)community standardSEPARATE estimation from PSM FDR; never assume 1% PSM implies 1% protein
Picked FDR (target/decoy pairing)Savitski 2015; The & Kall 2016Removes target/decoy asymmetry; dataset-size-independent, unlike naive decoy/target
Decoy:target ratio 1:1community standardStandard null; unequal ratios require formula correction
Min PSMs for trustworthy protein FDRhundreds+Below ~100s of items decoy counts are too noisy; "0% FDR" from zero decoys is luck, not control
Two-peptide ruleDO NOT USE (Gupta & Pevzner 2009)Increases protein FDR and drops real proteins; replaced by picked FDR + per-ID score
Single-peptide IDsjudge by score, not countA high-confidence unique peptide can be a legitimate ID

Common Errors

Error / symptomCauseSolution
Protein FDR much higher than nominal on deep dataNaive decoy/target reused from PSM levelPicked-protein or picked-group FDR
Real low-abundance proteins missingTwo-peptide rule appliedRemove the rule; control picked FDR
AttributeError on EpifanyAlgorithm / infer_proteinsClass name varies by version; the R ProteinInference::infer_proteins could not be confirmed to existhelp(pyopenms.EpifanyAlgorithm) to find the real name; use pyOpenMS, not an unverified R package
Indistinguishable proteins reported as separate IDsFlat protein list instead of groupsEnable annotate_indistinguishable_groups; report groups with a leading protein
Spurious DE on paralog-sharing proteinsRazor-peptide quant flipped between conditionsQuantify on unique peptides -> quantification
"Unique" peptide count changed when DB changedUniqueness is database-relativeFix and document the database (isoforms, contaminants, decoys)

References

  • Nesvizhskii, A.I., Keller, A., Kolker, E. & Aebersold, R. (2003). A statistical model for identifying proteins by tandem mass spectrometry. Analytical Chemistry 75(17):4646-4658.
  • Gupta, N. & Pevzner, P.A. (2009). False discovery rates of protein identifications: a strike against the two-peptide rule. Journal of Proteome Research 8(9):4173-4181.
  • Savitski, M.M., Wilhelm, M., Hahne, H., Kuster, B. & Bantscheff, M. (2015). A scalable approach for protein false discovery rate estimation in large proteomic data sets. Molecular & Cellular Proteomics 14(9):2394-2404.
  • The, M., Tasnim, A. & Kall, L. (2016). How to talk about protein-level false discovery rates in shotgun proteomics. Proteomics 16(18):2461-2469.
  • Pfeuffer, J., Sachsenberg, T., Dijkstra, T.M.H., Serang, O., Reinert, K. & Kohlbacher, O. (2020). EPIFANY: a method for efficient high-confidence protein inference. Journal of Proteome Research 19(3):1060-1072.
  • peptide-identification - Produces the FDR-filtered peptide list that feeds inference and shares the target-decoy machinery
  • quantification - Consumes inferred groups; razor-vs-unique peptide choice lives here
  • data-import - Loads idXML/mzML identification files
  • database-access/uniprot-access - Canonical-vs-isoform databases drive uniqueness and the leading-protein convention

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in proteomics/protein-inference of GPTomics/bioSkills.

  • SKILL.md
  • examples/protein_groups.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Proteomics Protein Inference next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Proteomics Protein Inference compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Proteomics Protein Inference this skillGPTomics/bioSkills1.2k1 repos~4.7kAutomated safety check: PassMIT
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Singlecell Qcxuzhougeng/wisp-science1k—~1.6kAutomated safety check: PassAGPL-3.0
Trackplotygidtu/trackplot109—~1.9kAutomated safety check: PassBSD-3-Clause
UniProt Database Accessdavila7/claude-code-templates33k14 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Singlecell Qc

    xuzhougeng/wisp-science

    A skill your agent uses when designing, reviewing, or implementing single-cell RNA-seq QC in Python or R with a human-in-the-loop, data-driven approach.

    1k GitHub stars~1.6k tokensUpdated today
    Research & ScienceAuto-check passed
  • Trackplot

    ygidtu/trackplot

    Generate sashimi-style genome visualization plots (coverage, line, heatmap, IGV read-by-read, HiC, circRNA, motif) from BAM/bigWig/depth/HiC inputs.

    109 GitHub stars~1.9k tokensUpdated 14 days ago
    Research & ScienceAuto-check passed
  • UniProt Database Access

    davila7/claude-code-templates

    Queries the UniProt REST API directly to search proteins, fetch FASTA sequences, map IDs between databases and read Swiss-Prot and TrEMBL entries.

    33k GitHub starsUsed in 14 repos~1.7k tokens
    Research & ScienceAuto-check passed
  • End-to-end 10x Visium spatial transcriptomics analysis workflow with staged execution and human review gates.

    101 GitHub stars~1.4k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Works with

Questions about Bio Proteomics Protein Inference

What does Bio Proteomics Protein Inference do?

Groups proteins from peptide identifications and controls protein-level FDR, framing inference as a chosen explanation (parsimony or a probability model) of underdetermined peptide evidence rather…. Bio Proteomics Protein Inference is an agent skill from GPTomics/bioSkills. Groups proteins from peptide identifications and controls protein-level FDR, framing inference as a chosen explanation (parsimony or a probability model) of underdetermined peptide evidence rather than a measurement.

When should I use Bio Proteomics Protein Inference?

Bio Proteomics Protein Inference fits situations like: resolving which proteins are present from a peptide list; building protein groups; estimating protein-level FDR.

How do I install Bio Proteomics Protein Inference in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-protein-inference -a claude-code`. Or copy the skill folder (proteomics/protein-inference in GPTomics/bioSkills) into .claude/skills/bio-proteomics-protein-inference in your project. Claude Code loads it when a task matches its description.

How do I install Bio Proteomics Protein Inference in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-protein-inference -a codex`. Or copy the skill folder (proteomics/protein-inference in GPTomics/bioSkills) into .agents/skills/bio-proteomics-protein-inference in your project. Codex loads it when a task matches its description.

Can I use Bio Proteomics Protein Inference in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-protein-inference -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-protein-inference, .gemini/skills/bio-proteomics-protein-inference, .github/skills/bio-proteomics-protein-inference and .opencode/skills/bio-proteomics-protein-inference in your project.

What does Bio Proteomics Protein Inference need to run?

Going by SKILL.md and its folder, Bio Proteomics Protein Inference needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Proteomics Protein Inference access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Proteomics Protein Inference safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Proteomics Protein Inference use?

Bio Proteomics Protein Inference is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Proteomics Protein Inference use?

About 4.7k tokens (SKILL.md is roughly 19k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Proteomics Protein Inference?

Skills that share tags, products or a category with Bio Proteomics Protein Inference: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Singlecell Qc (xuzhougeng/wisp-science, 1k stars) and Trackplot (ygidtu/trackplot, 109 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Proteomics Protein Inference?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.