Agent skill

Bio Proteomics Peptide Identification

by GPTomics in GPTomics/bioSkills

Peptide-spectrum matching from MS/MS with target-decoy FDR control, framing identification confidence as a property of a ranked list (q-value/PEP) rather than a raw engine score (XCorr, hyperscore…

MITAuto-check passedResearch & Science

Install Bio Proteomics Peptide Identification

skills CLI
$ npx skills add GPTomics/bioSkills --skill bio-proteomics-peptide-identification -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install GPTomics/bioSkills bio-proteomics-peptide-identification --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/GPTomics/bioSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/proteomics/peptide-identification .claude/skills/bio-proteomics-peptide-identification && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
bio-proteomics-peptide-identification
GitHub stars
1.2k
Used in
1 other repo
Token cost
~5.3k tokens
SKILL.md length
2,363 words
Files
3
Skills in repo
559
Repo updated
First seen
Licence
MIT

At a glance

Peptide-spectrum matching from MS/MS with target-decoy FDR control, framing identification confidence as a property of a ranked list (q-value/PEP) rather than a raw engine score (XCorr, hyperscore…

  • Works in 3 steps: Identification confidence is a property… → A q-value is valid only if (a) the decoy… → PEP and q-value answer different…
  • Identifying peptides from tandem mass spectra and deciding what FDR threshold to act on
  • SKILL.md covers Version Compatibility, The Single Most Important…, The FDR Vocabulary, Precisely and Tool Taxonomy, plus 6 more sections
  • Runs Python scripts from its folder; calls pip

What it does

Bio Proteomics Peptide Identification is an agent skill from GPTomics/bioSkills. Peptide-spectrum matching from MS/MS with target-decoy FDR control, framing identification confidence as a property of a ranked list (q-value/PEP) rather than a raw engine score (XCorr, hyperscore, Andromeda, SpecEValue). Covers sequence-database search engines (Comet, MS-GF+, MSFragger, Sage, MaxQuant, MetaMorpheus), concatenated vs separate target-decoy competition, PEP vs q-value, the multi-level FDR cascade, open/mass-tolerant search, rescoring (Percolator, mokapot, MS2Rescore), and pyOpenMS…

Its SKILL.md is about 5.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `examples/fdr_filtering.py` and `usage-guide.md`).

It sits in Research & Science, covering Bioinformatics and Internationalization. The repository describes itself as: a set of SKILLS.md for doing bioinformatics with agents like claude code. The licence is MIT.

When your agent uses it

  • Identifying peptides from tandem mass spectra and deciding what FDR threshold to act on
  • Tasks that involve Bioinformatics
  • Tasks that involve Internationalization

Example prompts

  • “/bio-proteomics-peptide-identification”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Identification confidence is a property of a ranked LIST controlled by target-decoy competition, never a property of one PSM. The number…
  2. A q-value is valid only if (a) the decoy DB is a faithful null, (b) targets and decoys competed in ONE concatenated search, and (c) there…
  3. PEP and q-value answer different questions; filtering at "PEP <= 0.01" is far stricter than "q <= 0.01." PEP (posterior error probability…

What it can do on your machine

Read from SKILL.md and the folder at commit d91ed3d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Bio Proteomics Peptide Identification loads about 5.3k tokens when it runs. Until then it costs about 217 tokens; SKILL.md has 2,363 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~217
When it runs · the whole SKILL.md, loaded when a task matches
~5.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from GPTomics/bioSkills at commit d91ed3d, republished under its MIT licence (© GPTomics). 2,363 words, ~5,349 tokens.

Download SKILL.mdSave it as .claude/skills/bio-proteomics-peptide-identification/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
bio-proteomics-peptide-identification
description
Peptide-spectrum matching from MS/MS with target-decoy FDR control, framing identification confidence as a property of a ranked list (q-value/PEP) rather than a raw engine score (XCorr, hyperscore, Andromeda, SpecEValue). Covers sequence-database search engines (Comet, MS-GF+, MSFragger, Sage, MaxQuant, MetaMorpheus), concatenated vs separate target-decoy competition, PEP vs q-value, the multi-level FDR cascade, open/mass-tolerant search, rescoring (Percolator, mokapot, MS2Rescore), and pyOpenMS SimpleSearchEngineAlgorithm + FalseDiscoveryRate. Use when identifying peptides from tandem mass spectra and deciding what FDR threshold to act on. Protein grouping and protein-level FDR are protein-inference; PTM site localization is ptm-analysis; DIA peptide-centric scoring is dia-analysis; intensity quant is quantification.
tool_type
mixed
primary_tool
pyOpenMS

Version Compatibility

Reference examples tested with: pyOpenMS 3.1+, pandas 2.2+, numpy 1.26+

Before using code patterns, verify installed versions match. If versions differ:

  • Python: pip show <package> then help(module.function) to check signatures
  • R: packageVersion('<pkg>') then ?function_name to verify parameters

If code throws ImportError, AttributeError, or TypeError, introspect the installed package and adapt the example to match the actual API rather than retrying.

Peptide Identification -- Confidence Is a Property of a Ranked List, Not a Single PSM

"Identify peptides from my MS/MS spectra" -> Match tandem mass spectra against a protein database, then control false discovery rate by target-decoy competition and act on a q-value -- because a raw match score is meaningless in isolation; only the list-level error rate is interpretable.

  • Python: pyopenms.SimpleSearchEngineAlgorithm().search(...) for in-process database search, FalseDiscoveryRate for q-values
  • CLI: comet, msfragger, sage, MSGFPlus for high-throughput database searching, percolator/mokapot for rescoring
  • R: mzID::mzID() + flatten() or mzR::openIDfile() + psms() to read mzIdentML search results

Scope: this skill owns spectrum-to-peptide matching and PSM/peptide-level FDR. Protein grouping and protein-level (picked) FDR -> protein-inference. PTM site localization and open-search mod discovery follow-up -> ptm-analysis. DIA peptide-centric extraction and scoring -> dia-analysis. FDR-filtered IDs feeding intensities -> quantification. mzML/raw loading -> data-import. OUT OF SCOPE: protein inference, PTM localization scoring, DIA peptide-centric pipelines, label-free/TMT quantification.

The Single Most Important Modern Insight -- A q-value Is a Verdict on the List, a Raw Score Is Not Even Comparable

  1. Identification confidence is a property of a ranked LIST controlled by target-decoy competition, never a property of one PSM. The number to act on is a q-value (list-level) or PEP (per-PSM), NOT the engine's raw score. XCorr (Comet), hyperscore (MSFragger/X!Tandem), Andromeda score (MaxQuant), and SpecEValue (MS-GF+) live on different scales, are charge- and length-dependent, and are frequently not even monotone in true probability within a single engine -- which is exactly why rescoring (Percolator/mokapot) exists. "1% FDR" answers "what fraction of the list I keep is wrong," NOT "I am 99% sure of this one ID." The catastrophic error is thresholding on a raw score, or comparing scores across engines.

  2. A q-value is valid only if (a) the decoy DB is a faithful null, (b) targets and decoys competed in ONE concatenated search, and (c) there are enough PSMs for the decoy count to be stable. Generate decoys at the PROTEIN level then digest (so decoy peptides obey the same enzyme rules), matching the target in size and composition. Concatenated competition gives FDR = (#decoys above threshold) / (#targets above threshold) -- one decoy above threshold estimates one false target. Separate target/decoy searches instead need either the simple Elias-Gygi 2x-decoy estimator FDR = 2 * #decoy / (#target + #decoy) or the more refined mix-max estimator (Keich, Kertesz-Farkas & Noble 2015) -- two distinct options for the separate-search setting, NOT the same formula. Mixing the concatenated and separate forms up is the most common silent FDR error.

  3. PEP and q-value answer different questions; filtering at "PEP <= 0.01" is far stricter than "q <= 0.01." PEP (posterior error probability, local FDR) is the probability that THIS PSM is wrong; q-value is the FDR of the list cut at this PSM. FDR is the average of PEP over the accepted set (Kall 2008). The worst PSM in a 1%-FDR list typically has a PEP of 10-50%. Use q-value for list cutoffs; use PEP only for per-ID decisions (e.g. picking one PTM site). And PSM-FDR at 1% does NOT give 1% peptide-FDR or 1% protein-FDR -- each level needs its own estimation; hand protein-level control to protein-inference.

The FDR Vocabulary, Precisely

  • FDR: the expected proportion of false positives among ALL accepted items at a threshold -- a property of the whole list.
  • q-value: the minimum FDR at which a given PSM is still accepted; monotone after taking the running minimum from the bottom of the ranked list. Filter on q <= 0.01.
  • PEP (local FDR): the probability that THIS PSM is wrong given its score. Local, per-PSM; FDR is the integral of PEP over the accepted set (Kall 2008, "two sides of the same coin").
  • The estimator must match the search mode. Concatenated target-decoy competition (TDC): FDR = #decoy / #target (no factor 2 -- one best hit per spectrum already resolves the competition). Separate target and decoy searches: either the simple Elias-Gygi 2x-decoy estimator FDR = 2 * #decoy / (#target + #decoy), or the more refined mix-max estimator (Keich, Kertesz-Farkas & Noble 2015). Mix-max is a distinct, calibrated-score procedure for the separate-search setting -- it is NOT a rename of the 2x formula.

Tool Taxonomy

Tool / methodCitationMechanism / roleWhen
CometEng 2013XCorr + E-value; SEQUEST lineage, open-sourceRobust default, TPP pipelines; pairs with Percolator
X!Tandem--hyperscore + refinement passesLegacy/free; semi-tryptic refinement niche
MS-GF+Kim & Pevzner 2014SpecEValue via generating-function DPCalibrated cross-instrument E-value; ETD/CID, low-res, non-standard enzymes
MaxQuant / AndromedaCox 2011binomial probability score; integrated MBR/LFQ/TMTAll-in-one quant pipeline (LFQ, TMT, SILAC); GUI
MSFraggerKong 2017hyperscore via fragment-ion indexing (~100x faster)Open/mass-tolerant search, PTM discovery, huge datasets; core of FragPipe
SageLazear 2023hyperscore-style, Rust, rescoring-nativeModern scalable open-source pipelines; emits Percolator-ready features
MetaMorpheusSolntsev 2018calibration + G-PTM-D multinotchPTM discovery with built-in calibration; proteoform-aware
pFind 3Chi 2018open-search engineMaximal unrestricted-PTM/mutation discovery
PercolatorKall 2007semi-supervised SVM re-rank on decoy negativesBoost IDs at fixed FDR; non-tryptic/PTM/large search spaces
mokapotFondrie & Noble 2021Percolator in Python; swappable XGBoost classifierPython pipelines, Sage output, custom features
MS2Rescore + DeepLC + MS2PIPDeclercq 2022; Bouwmeester 2021; Gabriels 2019predicted-RT + predicted-intensity rescoring featuresSharpen target/decoy separation; immunopeptidomics
Spectral-library search--match empirical reference spectra (intensity + RT)Faster/more specific for known peptides -> spectral-libraries
Protein grouping / protein FDRSavitski 2015; The 2016picked / picked-group FDRroute OUT -> protein-inference
PTM site localization--per-site PEP, localization scoringroute OUT -> ptm-analysis

Decision Tree by Scenario

ScenarioRecommendedWhy
Standard DDA, clean FDR, scriptableComet or Sage + Percolator/mokapot at q <= 0.01well-validated; rescoring boosts IDs at fixed FDR
Cross-instrument / varied fragmentation / odd enzymeMS-GF+SpecEValue is calibrated so a threshold means the same everywhere
Discover unknown PTMs / mass shiftsMSFragger open search (-150..+500 Da)fragment indexing makes wide-window search feasible; then closed search on discovered mods -> ptm-analysis
Huge dataset, reproducible, cloud-scaleSage (rescoring-native)Rust speed; emits Percolator features directly
All-in-one with quant in the same toolMaxQuant/Andromedaintegrated LFQ/TMT/SILAC and MBR
Non-tryptic (immunopeptidomics, degradomics)any engine + Percolator/MS2Rescorerescoring gains are largest where search space explodes
Few PSMs (single-protein pulldown)do NOT trust decoy FDR; inspect spectra manuallydecoy counts too noisy below ~hundreds of PSMs
Need per-site / per-ID confidenceact on PEP, not q-valueq-value is list-level; PEP is local

Default when uncertain: concatenated target-decoy search with Comet or Sage, rescore with Percolator/mokapot, filter at q <= 0.01, and hand protein-level FDR to protein-inference.

Database Search with pyOpenMS

Goal: Match tandem mass spectra in an mzML file against a protein FASTA and produce scored PSMs as idXML.

Approach: SimpleSearchEngineAlgorithm actually scores spectra (the hand-rolled ProteaseDigestion loop only digests, it never matches a spectrum). The FASTA must already contain target + decoy sequences concatenated for downstream FDR; decoys carry a recognizable prefix.

python
from pyopenms import SimpleSearchEngineAlgorithm, IdXMLFile

protein_ids = []
peptide_ids = []
search = SimpleSearchEngineAlgorithm()
# spectra are scored against in-silico fragment ions of every candidate peptide
search.search('sample.mzML', 'human_target_decoy.fasta', protein_ids, peptide_ids)

# protein_ids FIRST in load/store -- the OpenMS argument order is fixed
IdXMLFile().store('search_results.idXML', protein_ids, peptide_ids)
Annotate Target/Decoy and Estimate FDR with pyOpenMS

Goal: Convert raw PSM scores into q-values and keep only PSMs at 1% FDR.

Approach: PeptideIndexing maps each PSM back to proteins and flags target vs decoy from the decoy prefix; FalseDiscoveryRate.apply runs the concatenated competition; IDFilter keeps q <= 0.01. This is the real pyOpenMS path -- not a hand-rolled decoy/target ratio of unknown provenance.

python
from pyopenms import PeptideIndexing, FalseDiscoveryRate, IDFilter, FASTAFile

fasta = []
FASTAFile().load('human_target_decoy.fasta', fasta)
indexer = PeptideIndexing()
params = indexer.getParameters()
params.setValue('decoy_string', 'DECOY_')      # must match the decoy prefix in the FASTA
params.setValue('decoy_string_position', 'prefix')
indexer.setParameters(params)
indexer.run(fasta, protein_ids, peptide_ids)   # sets target/decoy flags on every hit

FalseDiscoveryRate().apply(peptide_ids)         # concatenated competition -> per-PSM q-value as the new score
IDFilter().filterHitsByScore(peptide_ids, 0.01) # 0.01 = 1% FDR, the community list-level standard
IDFilter().removeDecoyHits(peptide_ids)
FDR from a Results Table (concatenated competition, made explicit)

Goal: Compute q-values from any engine's PSM table when the search was a single concatenated target-decoy search.

Approach: Rank by score, walk down accumulating target and decoy counts, FDR = decoys/targets, then take the running minimum from the bottom to get monotone q-values. The decoy/target form is correct ONLY for concatenated competition; separate searches need either the Elias-Gygi 2x-decoy form or the mix-max estimator (Keich, Kertesz-Farkas & Noble 2015).

python
import pandas as pd

psms = pd.read_csv('search_results.tsv', sep='\t')
psms['is_decoy'] = psms['protein'].str.startswith(('DECOY_', 'REV_', 'XXX_'))
psms = psms.sort_values('score', ascending=False).reset_index(drop=True)

# concatenated target-decoy competition: each decoy above threshold estimates one false target
targets = (~psms['is_decoy']).cumsum()
decoys = psms['is_decoy'].cumsum()
psms['fdr'] = decoys / targets
psms['qvalue'] = psms['fdr'][::-1].cummin()[::-1]   # running min from the bottom -> monotone q-values

kept = psms[(psms['qvalue'] <= 0.01) & (~psms['is_decoy'])]   # 1% list-level FDR

Per-Method Failure Modes

Concatenated vs separate FDR formula mismatch

Trigger: applying #decoy/#target to separately-searched targets and decoys, or 2*decoy/(target+decoy) to concatenated competition. Mechanism: the factor of 2 accounts for false hits that could land in either independent database; concatenated competition already resolves that by a single best hit per spectrum. Symptom: systematically under- or over-estimated FDR; irreproducible ID counts. Fix: confirm the search mode; concatenated -> #decoy/#target; separate -> Elias-Gygi 2x-decoy or the mix-max estimator (Keich, Kertesz-Farkas & Noble 2015). In Percolator, mix-max is the default for separate-search input and -Y/--post-processing-tdc selects target-decoy competition instead; concatenated input forces TDC automatically.

Thresholding on raw engine score

Trigger: filtering on XCorr/hyperscore/Andromeda score, or comparing scores from two engines. Mechanism: scores are uncalibrated, charge/length-dependent, and not monotone in true probability. Symptom: different cutoffs admit different real FDRs; cross-engine merges nonsensical. Fix: always convert to q-value (or SpecEValue/PEP) first; rescore with Percolator/mokapot.

Show full SKILL.md (936 more words)Show less
Decoy FDR on too few PSMs

Trigger: reporting "0% FDR" from a single-protein pulldown or tiny PSM list. Mechanism: the decoy count is a noisy Poisson-like estimate; zero observed decoys does not mean zero false targets. Symptom: spuriously confident IDs from small experiments. Fix: below ~hundreds of PSMs, inspect spectra manually; do not act on the decoy q-value.

Open-search results used for clean FDR or quant

Trigger: taking IDs from a wide-window (-150..+500 Da) search as final, FDR-controlled results. Mechanism: wide windows admit "free" mass shifts that inflate random matches; the target-decoy null differs per mass-shift bin. Symptom: inflated, unreliable FDR on open-search output. Fix: treat open search as discovery; follow with a closed search restricted to the discovered mods -> ptm-analysis.

Rescoring overfitting

Trigger: custom features that leak label information, or training without proper cross-validation. Mechanism: the model learns the decoys, making rescored FDR optimistic. Symptom: ID counts jump but downstream validation fails. Fix: use Percolator/mokapot default cross-validation; predicted-feature rescoring (DeepLC/MS2PIP) is safer; validate with entrapment for high-stakes claims (Wen 2025).

DIA tool FDR taken at face value

Trigger: trusting a DIA tool's reported 1% peptide/protein FDR. Mechanism: entrapment shows several DIA tools do not reliably control FDR (Wen 2025). Symptom: real error rate exceeds the reported FDR. Fix: validate with entrapment for high-stakes DIA claims -> dia-analysis.

Quantitative Thresholds

ThresholdSourceRationale
Precursor tolerance 10-20 ppm (high-res Orbitrap)--matches FT mass accuracy; tighter = fewer random candidates at fixed FDR
Precursor tolerance -150..+500 Da (open search)Kong 2017captures arbitrary PTM/mutation shifts; feasible only with fragment indexing
Fragment tolerance 0.02 Da (HCD Orbitrap) / 0.6 Da (ion-trap CID)--instrument-dependent; 0.6 Da on Orbitrap discards resolving power
Missed cleavages 2--covers incomplete trypsin digestion without exploding search space
PSM/peptide FDR 1% (q <= 0.01)Elias & Gygi 2007community standard; list-level error, not per-PSM
Decoy:target ratio 1:1Elias & Gygi 2007standard; unequal ratios need formula correction
Min PSMs for trustworthy decoy FDR: hundreds+--below this the decoy count is too noisy
Variable mods per peptide <= 2-3--each variable mod multiplies search space and random-match rate

Common Errors

Error / symptomCauseSolution
pyOpenMS "search" returns peptides but never scores spectraused ProteaseDigestion, which only digests a FASTAuse SimpleSearchEngineAlgorithm().search(mzML, fasta, protein_ids, peptide_ids)
IdXMLFile().load/store argument errorwrong orderprotein_ids FIRST: IdXMLFile().load(path, protein_ids, peptide_ids)
FDR ignores decoys / all q-values 0decoys not annotated before FalseDiscoveryRaterun PeptideIndexing with matching decoy_string first
R: MSnbase::readMzIdData not foundthat function name does not existuse mzID::mzID(file) + flatten(), or mzR::openIDfile() + psms() (PSMatch/Spectra is the modern path)
Percolator q-method mismatched to search modemix-max is the default for separate-search inputfor separate searches, mix-max (default) or -Y/--post-processing-tdc for target-decoy competition; concatenated input forces TDC automatically; use --picked-protein for protein FDR
1% PSM FDR assumed to give 1% protein FDReach level needs its own estimationestimate protein-level (picked) FDR -> protein-inference
"PEP <= 0.01" returns far fewer IDs than expectedPEP is per-PSM and far stricter than q-valuefilter list cutoffs on q-value; reserve PEP for per-ID decisions

References

  • Elias, J.E. & Gygi, S.P. 2007. Target-decoy search strategy for increased confidence in large-scale protein identifications by mass spectrometry. Nature Methods 4(3):207-214.
  • Keich, U., Kertesz-Farkas, A. & Noble, W.S. 2015. Improved false discovery rate estimation procedure for shotgun proteomics. Journal of Proteome Research 14(8):3148-3161.
  • Kall, L., Canterbury, J.D., Weston, J., Noble, W.S. & MacCoss, M.J. 2007. Semi-supervised learning for peptide identification from shotgun proteomics datasets. Nature Methods 4(11):923-925.
  • Kall, L., Storey, J.D., MacCoss, M.J. & Noble, W.S. 2008. Posterior error probabilities and false discovery rates: two sides of the same coin. Journal of Proteome Research 7(1):40-44.
  • Eng, J.K., Jahan, T.A. & Hoopmann, M.R. 2013. Comet: an open-source MS/MS sequence database search tool. Proteomics 13(1):22-24.
  • Kim, S. & Pevzner, P.A. 2014. MS-GF+ makes progress towards a universal database search tool for proteomics. Nature Communications 5:5277.
  • Cox, J., Neuhauser, N., Michalski, A., Scheltema, R.A., Olsen, J.V. & Mann, M. 2011. Andromeda: a peptide search engine integrated into the MaxQuant environment. Journal of Proteome Research 10(4):1794-1805.
  • Kong, A.T., Leprevost, F.V., Avtonomov, D.M., Mellacheruvu, D. & Nesvizhskii, A.I. 2017. MSFragger: ultrafast and comprehensive peptide identification in mass spectrometry-based proteomics. Nature Methods 14(5):513-520.
  • Lazear, M.R. 2023. Sage: an open-source tool for fast proteomics searching and quantification at scale. Journal of Proteome Research 22(11):3652-3659.
  • Solntsev, S.K., Shortreed, M.R., Frey, B.L. & Smith, L.M. 2018. Enhanced global post-translational modification discovery with MetaMorpheus. Journal of Proteome Research 17(5):1844-1851.
  • Chi, H., Liu, C., Yang, H. et al. 2018. Comprehensive identification of peptides in tandem mass spectra using an efficient open search engine. Nature Biotechnology 36:1059-1061.
  • Fondrie, W.E. & Noble, W.S. 2021. mokapot: fast and flexible semisupervised learning for peptide detection. Journal of Proteome Research 20(4):1966-1971.
  • Bouwmeester, R., Gabriels, R., Hulstaert, N., Martens, L. & Degroeve, S. 2021. DeepLC can predict retention times for peptides that carry as-yet unseen modifications. Nature Methods 18:1363-1369.
  • Gabriels, R., Martens, L. & Degroeve, S. 2019. Updated MS2PIP web server delivers fast and accurate MS2 peak intensity prediction for multiple fragmentation methods, instruments and labeling techniques. Nucleic Acids Research 47(W1):W295-W299.
  • Declercq, A., Bouwmeester, R., Hirschler, A., Carapito, C., Degroeve, S., Martens, L. & Gabriels, R. 2022. MS2Rescore: data-driven rescoring dramatically boosts immunopeptide identification rates. Molecular & Cellular Proteomics 21(8):100266.
  • Wen, B., Freestone, J., Riffle, M., MacCoss, M.J., Noble, W.S. & Keich, U. 2025. Assessment of false discovery rate control in tandem mass spectrometry analysis using entrapment. Nature Methods 22:1454-1463.
  • protein-inference - Group peptides to protein groups and control protein-level (picked) FDR
  • ptm-analysis - Open/variable-mod search follow-up and per-site PTM localization
  • dia-analysis - DIA peptide-centric extraction and scoring; entrapment FDR validation
  • quantification - FDR-filtered IDs feed label-free/TMT intensity quantification
  • spectral-libraries - Empirical and predicted spectral-library search as an ID alternative
  • data-import - Load mzML/raw MS data before identification
  • database-access/uniprot-access - Build the target FASTA (canonical vs isoform, contaminants)

© GPTomics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in proteomics/peptide-identification of GPTomics/bioSkills.

  • SKILL.md
  • examples/fdr_filtering.py
  • usage-guide.md

Open the folder on GitHubat commit d91ed3d

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in GPTomics/bioSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Bio Proteomics Peptide Identification next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Bio Proteomics Peptide Identification compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Bio Proteomics Peptide Identification this skillGPTomics/bioSkills1.2k1 repos~5.3kAutomated safety check: PassMIT
Spatial S5 DownstreamQING1105/ezST101—~513Automated safety check: PassMIT
Bio Proteomics Ptm AnalysisFreedomIntelligence/OpenClaw-Medical-Skills3.1k1 repos~1.2kAutomated safety check: PassNone
Proteomics PtmTianGzlab/OmicsClaw161—~989Automated safety check: PassApache-2.0
Bioconductor BandlebioMate-AI/biomate-bioconductor-kb804—~1.3kAutomated safety check: PassCustom licence
Bioconductor StatialbioMate-AI/biomate-bioconductor-kb804—~1.5kAutomated safety check: PassCustom licence

Similar skills

  • Stage 5 of the spatial transcriptomics workflow — neighborhood enrichment and cell-cell communication analysis.

    101 GitHub stars~513 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Bio Proteomics Ptm Analysis

    FreedomIntelligence/OpenClaw-Medical-Skills

    Post-translational modification analysis including phosphorylation, acetylation, and ubiquitination.

    3.1k GitHub starsUsed in 1 repo~1.2k tokens
    Research & ScienceAuto-check passed
  • Proteomics Ptm

    TianGzlab/OmicsClaw

    Load when summarising PTM sites (phosphorylation, acetylation, ubiquitination, etc.) from a per-site CSV — site-class assignment (Olsen et al.

    161 GitHub stars~989 tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Bioconductor Bandle

    bioMate-AI/biomate-bioconductor-kb

    The Bandle package enables the analysis and visualisation of differential localisation experiments using mass-spectrometry data.

    804 GitHub stars~1.3k tokensUpdated 3 mo ago
    Frontend & DesignAuto-check passed
  • Bioconductor Statial

    bioMate-AI/biomate-bioconductor-kb

    Statial is a suite of functions for identifying changes in cell state.

    804 GitHub stars~1.5k tokensUpdated 3 mo ago
    Frontend & DesignAuto-check passed
  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes

More from GPTomics/bioSkills

All 559 skills in this repo
  • Bio Alignment Io

    GPTomics/bioSkills

    Read, write, and convert multiple sequence alignment files using Biopython Bio.AlignIO.

    1.2k GitHub starsUsed in 3 repos~4.9k tokens
    Auto-check passed
  • bioSkills Installer

    GPTomics/bioSkills

    Installs the bioSkills collection of 425 bioinformatics skills in one step, or only chosen categories, so sequencing, RNA-seq, single-cell and variant tasks get specialized help.

    1.2k GitHub starsUsed in 1 repo~789 tokens
    Auto-check passed
  • Bio Write Sequences

    GPTomics/bioSkills

    Write biological sequences to files (FASTA, FASTQ, GenBank, EMBL) using Biopython Bio.SeqIO.

    1.2k GitHub starsUsed in 3 repos~2.1k tokens
    Auto-check passed
  • Amplicon Primer Clipping

    GPTomics/bioSkills

    Soft- or hard-clips PCR primer footprints from aligned amplicon BAMs so primer bases stop masquerading as confirmed reference sequence.

    1.2k GitHub starsUsed in 2 repos~2.2k tokens
    Auto-check passed
  • Filters BAM alignments by FLAG bits, mapping quality and regions with samtools view or pysam, with recipes for common keep and drop cases.

    1.2k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Bio Alignment Indexing

    GPTomics/bioSkills

    Create and use BAI/CSI indices for BAM/CRAM files using samtools and pysam.

    1.2k GitHub starsUsed in 2 repos~2.4k tokens
    Auto-check passed

Questions about Bio Proteomics Peptide Identification

What does Bio Proteomics Peptide Identification do?

Peptide-spectrum matching from MS/MS with target-decoy FDR control, framing identification confidence as a property of a ranked list (q-value/PEP) rather than a raw engine score (XCorr, hyperscore…. Bio Proteomics Peptide Identification is an agent skill from GPTomics/bioSkills. Peptide-spectrum matching from MS/MS with target-decoy FDR control, framing identification confidence as a property of a ranked list (q-value/PEP) rather than a raw engine score (XCorr, hyperscore, Andromeda, SpecEValue).

When should I use Bio Proteomics Peptide Identification?

Bio Proteomics Peptide Identification fits situations like: identifying peptides from tandem mass spectra and deciding what FDR threshold to act on; tasks that involve Bioinformatics; tasks that involve Internationalization.

How do I install Bio Proteomics Peptide Identification in Claude Code?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-peptide-identification -a claude-code`. Or copy the skill folder (proteomics/peptide-identification in GPTomics/bioSkills) into .claude/skills/bio-proteomics-peptide-identification in your project. Claude Code loads it when a task matches its description.

How do I install Bio Proteomics Peptide Identification in Codex?

Run `npx skills add GPTomics/bioSkills --skill bio-proteomics-peptide-identification -a codex`. Or copy the skill folder (proteomics/peptide-identification in GPTomics/bioSkills) into .agents/skills/bio-proteomics-peptide-identification in your project. Codex loads it when a task matches its description.

Can I use Bio Proteomics Peptide Identification in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add GPTomics/bioSkills --skill bio-proteomics-peptide-identification -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/bio-proteomics-peptide-identification, .gemini/skills/bio-proteomics-peptide-identification, .github/skills/bio-proteomics-peptide-identification and .opencode/skills/bio-proteomics-peptide-identification in your project.

What does Bio Proteomics Peptide Identification need to run?

Going by SKILL.md and its folder, Bio Proteomics Peptide Identification needs Python for the scripts in its folder and the command-line tools its instructions call (pip). Our summary lists: Python 3.

Does Bio Proteomics Peptide Identification access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Bio Proteomics Peptide Identification safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Bio Proteomics Peptide Identification use?

Bio Proteomics Peptide Identification is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Bio Proteomics Peptide Identification use?

About 5.3k tokens (SKILL.md is roughly 21k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Bio Proteomics Peptide Identification?

Skills that share tags, products or a category with Bio Proteomics Peptide Identification: Spatial S5 Downstream (QING1105/ezST, 101 stars), Bio Proteomics Ptm Analysis (FreedomIntelligence/OpenClaw-Medical-Skills, 3.1k stars), Proteomics Ptm (TianGzlab/OmicsClaw, 161 stars) and Bioconductor Bandle (bioMate-AI/biomate-bioconductor-kb, 804 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Bio Proteomics Peptide Identification?

GPTomics (a GitHub organization) maintains it in GPTomics/bioSkills, which has 1,218 GitHub stars. The repository holds 559 skills in this directory. The repository was last updated on August 15, 2026.

Source: GPTomics/bioSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.