Agent skill

Pyopenms Mass Spectrometry

by jaechang-hits in jaechang-hits/SciAgent-Skills

MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID…

BSD-3-ClauseAuto-check passedResearch & Science

Install Pyopenms Mass Spectrometry

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills pyopenms-mass-spectrometry --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry .claude/skills/pyopenms-mass-spectrometry && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
pyopenms-mass-spectrometry
GitHub stars
374
Used in
1 other repo
Token cost
~5.9k tokens
SKILL.md length
1,133 words
Files
4 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
BSD-3-Clause

At a glance

MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID…

  • Works in 6 steps: Load mzML files for all samples (Core… → Centroid each sample (Core API Module 2… → Detect features per sample (Core API… → …
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 11 more sections
  • Calls uv

What it does

Pyopenms Mass Spectrometry is an agent skill from jaechang-hits/SciAgent-Skills. MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID with FDR, untargeted metabolomics. Use matchms for simple spectral matching.

Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/data_structures_reference.md`, `references/identification_metabolomics.md` and `references/signal_feature_processing.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/pyopenms-mass-spectrometry”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Load mzML files for all samples (Core API Module 1)
  2. Centroid each sample (Core API Module 2 — PeakPickerHiRes)
  3. Detect features per sample (Core API Module 3 — FeatureFinder)
  4. Align retention times (Core API Module 3 — MapAlignmentAlgorithmPoseClustering)
  5. Link features to consensus map (Core API Module 3 — FeatureGroupingAlgorithmQT)
  6. Export consensus table to pandas and filter by CV across replicates

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • pyopenms.readthedocs.io
    • openms.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Pyopenms Mass Spectrometry loads about 5.9k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 1,133 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~75
When it runs · the whole SKILL.md, loaded when a task matches
~5.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,133 words, ~5,945 tokens.

Download SKILL.mdSave it as .claude/skills/pyopenms-mass-spectrometry/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
pyopenms-mass-spectrometry
description
MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID with FDR, untargeted metabolomics. Use matchms for simple spectral matching.
license
BSD-3-Clause

PyOpenMS — Mass Spectrometry Analysis

Overview

PyOpenMS provides Python bindings to the OpenMS C++ library for computational mass spectrometry. It supports proteomics and metabolomics data processing including file I/O for 10+ MS formats, signal processing, feature detection, peptide/protein identification, and quantitative analysis across samples.

When to Use

  • Processing raw LC-MS/MS data (mzML, mzXML) for proteomics or metabolomics
  • Detecting chromatographic features and linking them across multiple samples
  • Identifying peptides and proteins from MS/MS search engine results with FDR control
  • Running untargeted metabolomics workflows (peak picking → feature detection → alignment → annotation)
  • Converting between mass spectrometry file formats (mzML, mzXML, featureXML, idXML)
  • Smoothing, filtering, and centroiding raw spectral data
  • For simple spectral library matching and metabolite identification, use matchms instead
  • For protein sequence analysis (not mass spec), use biopython instead

Prerequisites

bash
uv pip install pyopenms numpy pandas matplotlib
  • Python 3.8+; NumPy for peak array operations
  • Input data: mzML files (standard MS format), FASTA databases (for identification)
  • All algorithms follow a consistent pattern: algo = Algorithm(); params = algo.getParameters(); params.setValue(...); algo.setParameters(params)

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: instrumentResolution
    kind: required
    source: data
    ask: "Is this high-resolution data, and are the spectra already centroided or still in profile mode?"
    default: null

  - id: D2
    param: signalToNoiseThreshold
    kind: required
    source: user
    depends_on: [D1]
    ask: "How far above noise must a peak rise to be kept during centroiding?"
    default: 1.0
    skip_if: "spectra already centroided by the acquisition software"

  - id: D3
    param: massTolerance
    kind: required
    source: data
    depends_on: [D1]
    ask: "What mass accuracy should feature detection and linking assume?"
    default: "10 ppm"

  - id: D4
    param: retentionTimeTolerance
    kind: required
    source: data
    ask: "How far apart in retention time may the same feature appear across runs?"
    default: "100 s"

  - id: D5
    param: chargeStateRange
    kind: required
    source: user
    ask: "Which charge states should features be searched for?"
    default: "1 to 3 - widen for intact protein or metabolite work"

  - id: D6
    param: intensityNormalization
    kind: optional
    source: user
    ask: "Should spectra be normalized before comparison, and to total ion current or to the base peak?"
    default: "not normalized"

  - id: D7
    param: smoothingFilter
    kind: optional_conditional
    source: data
    ask: "Does the signal need smoothing before peak picking, and with which filter width?"
    default: "no smoothing"

D1 governs almost everything after it: profile data that skips centroiding produces one feature per scan point, and high-resolution tolerances applied to low-resolution data link features that are not the same compound. Both yield a full feature table.

Quick Start

python
import pyopenms as ms

# Load mzML file
exp = ms.MSExperiment()
ms.MzMLFile().load("sample.mzML", exp)
print(f"Spectra: {exp.getNrSpectra()}, Chromatograms: {exp.getNrChromatograms()}")

# Examine first spectrum
spec = exp.getSpectrum(0)
mz, intensity = spec.get_peaks()
print(f"MS level: {spec.getMSLevel()}, RT: {spec.getRT():.2f}s, Peaks: {len(mz)}")

# Quick preprocessing: smooth + centroid
gauss = ms.GaussFilter()
p = gauss.getParameters(); p.setValue("gaussian_width", 0.1); gauss.setParameters(p)
gauss.filterExperiment(exp)

picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
print(f"Centroided spectra: {centroided.getNrSpectra()}")

Core API

Module 1: File I/O & Data Access

Read and write mass spectrometry data in multiple formats.

python
import pyopenms as ms

# Read mzML (standard MS format)
exp = ms.MSExperiment()
ms.MzMLFile().load("data.mzML", exp)

# Indexed access for large files (memory-efficient)
loader = ms.IndexedMzMLFileLoader()
indexed_file = ms.OnDiscMSExperiment()
loader.load("large_data.mzML", indexed_file)
spec = indexed_file.getSpectrum(0)  # Load single spectrum on demand
print(f"Total spectra: {indexed_file.getNrSpectra()}")

# Read identification results (idXML)
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("results.idXML", protein_ids, peptide_ids)
print(f"Peptide IDs: {len(peptide_ids)}, Protein IDs: {len(protein_ids)}")

# Read feature map
fm = ms.FeatureMap()
ms.FeatureXMLFile().load("features.featureXML", fm)
print(f"Features: {fm.size()}")
python
# Write mzML with compression
exp_out = ms.MSExperiment()
# ... populate experiment ...
ms.MzMLFile().store("output.mzML", exp_out)

# Read FASTA database
entries = []
ms.FASTAFile().load("database.fasta", entries)
print(f"Proteins in DB: {len(entries)}")
for e in entries[:3]:
    print(f"  {e.identifier}: {e.sequence[:30]}...")

Supported formats: mzML, mzXML, mzData (spectra); featureXML, consensusXML (features); idXML, mzIdentML, pepXML (identifications); TraML (transitions); mzTab (results); FASTA (sequences)

Module 2: Signal Processing & Peak Picking

Preprocess raw spectral data for downstream analysis.

python
import pyopenms as ms

exp = ms.MSExperiment()
ms.MzMLFile().load("raw.mzML", exp)

# Gaussian smoothing
gauss = ms.GaussFilter()
p = gauss.getParameters()
p.setValue("gaussian_width", 0.15)  # m/z width
gauss.setParameters(p)
gauss.filterExperiment(exp)

# Savitzky-Golay smoothing (alternative)
sg = ms.SavitzkyGolayFilter()
p = sg.getParameters()
p.setValue("frame_length", 15)  # Must be odd
sg.setParameters(p)
# sg.filterExperiment(exp)  # Use one smoother, not both

# Peak picking (centroiding) — required before feature detection
picker = ms.PeakPickerHiRes()
p = picker.getParameters()
p.setValue("signal_to_noise", 1.0)
picker.setParameters(p)
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
print(f"Raw peaks in spec 0: {exp.getSpectrum(0).size()}")
print(f"Centroided peaks: {centroided.getSpectrum(0).size()}")
python
# Normalization
normalizer = ms.Normalizer()
p = normalizer.getParameters()
p.setValue("method", "to_one")  # "to_one" or "to_TIC"
normalizer.setParameters(p)
normalizer.filterPeakMap(centroided)

# Peak filtering — remove low-intensity noise
mower = ms.ThresholdMower()
p = mower.getParameters()
p.setValue("threshold", 100.0)  # Minimum intensity
mower.setParameters(p)
mower.filterPeakMap(centroided)

# Baseline reduction
morph = ms.MorphologicalFilter()
p = morph.getParameters()
p.setValue("struc_elem_length", 3.0)  # m/z window
morph.setParameters(p)
morph.filterExperiment(exp)
Module 3: Feature Detection & Linking

Detect chromatographic features and link them across samples.

python
import pyopenms as ms

# Load centroided data
exp = ms.MSExperiment()
ms.MzMLFile().load("centroided.mzML", exp)

# Feature detection (proteomics — centroided data)
ff = ms.FeatureFinder()
features = ms.FeatureMap()
seeds = ms.FeatureMap()
params = ms.FeatureFinder().getParameters("centroided")
ff.run("centroided", exp, features, params, seeds)
print(f"Detected {features.size()} features")

# Access feature properties
for f in features[:5]:
    print(f"  RT: {f.getRT():.1f}s, m/z: {f.getMZ():.4f}, "
          f"intensity: {f.getIntensity():.0f}, quality: {f.getOverallQuality():.3f}")
python
# Feature linking across samples — align retention times first
aligner = ms.MapAlignmentAlgorithmPoseClustering()
p = aligner.getParameters()
p.setValue("max_num_peaks_considered", 1000)
aligner.setParameters(p)

# Link features into consensus map
linker = ms.FeatureGroupingAlgorithmQT()
p = linker.getParameters()
p.setValue("distance_RT:max_difference", 60.0)  # seconds
p.setValue("distance_MZ:max_difference", 10.0)   # ppm
linker.setParameters(p)

consensus = ms.ConsensusMap()
linker.group([features_sample1, features_sample2, features_sample3], consensus)
print(f"Consensus features: {consensus.size()}")

# Export to pandas for downstream analysis
import pandas as pd
df = consensus.get_df()
print(f"Consensus table: {df.shape}")
Module 4: Peptide & Protein Identification

Process search engine results with FDR control and protein inference.

python
import pyopenms as ms

# Load search engine results
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("search_results.idXML", protein_ids, peptide_ids)

# Examine peptide hits
for pep_id in peptide_ids[:3]:
    print(f"Spectrum: RT={pep_id.getRT():.1f}, MZ={pep_id.getMZ():.4f}")
    for hit in pep_id.getHits():
        seq = hit.getSequence()
        print(f"  {seq} score={hit.getScore():.4f} charge={hit.getCharge()}")

# FDR filtering (target-decoy approach)
fdr = ms.FalseDiscoveryRate()
fdr.apply(peptide_ids)

# Filter at 1% FDR
filtered = []
for pep_id in peptide_ids:
    hits = [h for h in pep_id.getHits() if h.getScore() <= 0.01]
    if hits:
        pep_id.setHits(hits)
        filtered.append(pep_id)
print(f"Peptide IDs at 1% FDR: {len(filtered)}")
python
# Protein inference
inference = ms.BasicProteinInferenceAlgorithm()
inference.run(peptide_ids, protein_ids)

for prot_id in protein_ids:
    for hit in prot_id.getHits()[:5]:
        print(f"Protein: {hit.getAccession()}, score: {hit.getScore():.4f}")

# Peptide sequence handling
seq = ms.AASequence.fromString("PEPTIDER")
print(f"Molecular weight: {seq.getMonoWeight():.4f}")
print(f"Formula: {seq.getFormula()}")

# Modified sequence
mod_seq = ms.AASequence.fromString("PEPTM(Oxidation)DER")
print(f"Modified weight: {mod_seq.getMonoWeight():.4f}")

# Enzymatic digestion
digestor = ms.ProteaseDigestion()
digestor.setEnzyme("Trypsin")
digest = []
digestor.digest(ms.AASequence.fromString("MKWVTFISLLLLFSSAYSRGVFRR"), digest)
print(f"Tryptic peptides: {len(digest)}")
Module 5: Metabolomics Pipeline

Complete untargeted metabolomics workflow from raw data to feature table.

python
import pyopenms as ms

# Step 1: Load and centroid raw data
exp = ms.MSExperiment()
ms.MzMLFile().load("metabolomics_sample.mzML", exp)

picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)

# Step 2: Feature detection for metabolomics (small molecules)
ff = ms.FeatureFinder()
features = ms.FeatureMap()
seeds = ms.FeatureMap()
params = ms.FeatureFinder().getParameters("centroided")
params.setValue("isotopic_pattern:charge_low", 1)
params.setValue("isotopic_pattern:charge_high", 3)
ff.run("centroided", centroided, features, params, seeds)
print(f"Detected features: {features.size()}")

# Step 3: Adduct detection (group related adducts)
decharger = ms.MetaboliteAdductDecharger()
p = decharger.getParameters()
p.setValue("potential_adducts", "H:+:0.6;Na:+:0.3;K:+:0.1")  # Positive mode
decharger.setParameters(p)
# decharger.compute(features, feature_map_out, consensus_map_out)
python
# Step 4: RT alignment across samples
import pyopenms as ms
import pandas as pd

# Assuming feature maps from multiple samples
sample_files = ["sample1.featureXML", "sample2.featureXML", "sample3.featureXML"]
feature_maps = []
for f in sample_files:
    fm = ms.FeatureMap()
    ms.FeatureXMLFile().load(f, fm)
    feature_maps.append(fm)

# Align retention times
aligner = ms.MapAlignmentAlgorithmPoseClustering()
aligner.setReference(0)  # Use first sample as reference

# Step 5: Link features to consensus map
linker = ms.FeatureGroupingAlgorithmQT()
p = linker.getParameters()
p.setValue("distance_RT:max_difference", 30.0)
p.setValue("distance_MZ:max_difference", 10.0)
linker.setParameters(p)

consensus = ms.ConsensusMap()
linker.group(feature_maps, consensus)

# Step 6: Export metabolite table
df = consensus.get_df()
print(f"Feature table: {df.shape[0]} features × {df.shape[1]} columns")
print(df[["RT", "mz", "intensity_0", "intensity_1", "intensity_2"]].head())

Key Concepts

Algorithm Parameter Pattern

All PyOpenMS algorithms follow the same 3-step pattern:

python
algo = ms.AlgorithmClass()           # 1. Instantiate
params = algo.getParameters()         # 2. Get parameters
params.setValue("param_name", value)   # 3. Set values
algo.setParameters(params)            # 4. Apply
# algo.process(input, output)         # 5. Execute

To discover available parameters:

python
for key in params.keys():
    print(f"{key}: {params.getValue(key)} (type: {params.getDescription(key)})")
Core Data Structures
ObjectDescriptionKey Methods
MSExperimentCollection of spectra + chromatogramsgetNrSpectra(), getSpectrum(i), addSpectrum()
MSSpectrumSingle mass spectrum (m/z, intensity)get_peaks(), getMSLevel(), getRT(), size()
MSChromatogramChromatographic traceget_peaks(), getNativeID()
FeatureDetected chromatographic peakgetRT(), getMZ(), getIntensity(), getOverallQuality()
FeatureMapCollection of featuressize(), get_df(), iteration
ConsensusMapFeatures linked across samplessize(), get_df()
PeptideIdentificationSearch results for one spectrumgetHits(), getRT(), getMZ()
AASequenceAmino acid sequence with modificationsfromString(), getMonoWeight(), getFormula()
Supported File Formats
FormatReader ClassContent
mzMLMzMLFileRaw spectra (standard)
mzXMLMzXMLFileRaw spectra (legacy)
featureXMLFeatureXMLFileDetected features
consensusXMLConsensusXMLFileLinked features
idXMLIdXMLFilePeptide/protein IDs
mzIdentMLMzIdentMLFilePeptide/protein IDs (standard)
FASTAFASTAFileProtein sequences
TraMLTraMLFileMRM transitions
mzTabMzTabFileQuantification results

Common Workflows

Workflow 1: Proteomics Feature Detection Pipeline
python
import pyopenms as ms

# Load → Smooth → Centroid → Detect → Export
exp = ms.MSExperiment()
ms.MzMLFile().load("proteomics.mzML", exp)

# Preprocessing
gauss = ms.GaussFilter()
p = gauss.getParameters(); p.setValue("gaussian_width", 0.1); gauss.setParameters(p)
gauss.filterExperiment(exp)

picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)

# Feature detection
ff = ms.FeatureFinder()
features = ms.FeatureMap()
params = ff.getParameters("centroided")
ff.run("centroided", centroided, features, params, ms.FeatureMap())

# Export to pandas
df = features.get_df()
print(f"Features: {df.shape}")
df.to_csv("proteomics_features.csv", index=False)
Workflow 2: Multi-Sample Metabolomics Comparison
  1. Load mzML files for all samples (Core API Module 1)
  2. Centroid each sample (Core API Module 2 — PeakPickerHiRes)
  3. Detect features per sample (Core API Module 3 — FeatureFinder)
  4. Align retention times (Core API Module 3 — MapAlignmentAlgorithmPoseClustering)
  5. Link features to consensus map (Core API Module 3 — FeatureGroupingAlgorithmQT)
  6. Export consensus table to pandas and filter by CV across replicates
Workflow 3: Identification Results Analysis
python
import pyopenms as ms
import pandas as pd

# Load and filter identifications
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("search.idXML", protein_ids, peptide_ids)

# FDR filtering
fdr = ms.FalseDiscoveryRate()
fdr.apply(peptide_ids)

# Build results table
rows = []
for pep_id in peptide_ids:
    for hit in pep_id.getHits():
        if hit.getScore() <= 0.01:  # 1% FDR
            rows.append({
                "sequence": str(hit.getSequence()),
                "score": hit.getScore(),
                "charge": hit.getCharge(),
                "rt": pep_id.getRT(),
                "mz": pep_id.getMZ()
            })

df = pd.DataFrame(rows)
print(f"Identified peptides at 1% FDR: {len(df)}")
print(f"Unique sequences: {df['sequence'].nunique()}")
df.to_csv("peptide_identifications.csv", index=False)

Key Parameters

ParameterModuleDefaultRange/OptionsEffect
gaussian_widthGaussFilter0.20.05–1.0m/z smoothing window width
frame_lengthSavitzkyGolayFilter115–31 (odd)Smoothing window points
signal_to_noisePeakPickerHiRes1.00.5–10.0Minimum S/N for centroiding
isotopic_pattern:charge_lowFeatureFinder11–6Minimum charge state
isotopic_pattern:charge_highFeatureFinder31–10Maximum charge state
distance_RT:max_differenceFeatureGroupingAlgorithmQT10010–300 sRT tolerance for feature linking
distance_MZ:max_differenceFeatureGroupingAlgorithmQT101–50 ppmm/z tolerance for feature linking
thresholdThresholdMower0.00–10000Minimum peak intensity
methodNormalizer"to_one""to_one"/"to_TIC"Normalization method

Best Practices

  1. Always centroid before feature detection — FeatureFinder requires centroided data. Use PeakPickerHiRes first
  2. Use indexed loading for large files — OnDiscMSExperiment loads spectra on demand, avoiding memory issues for files >1 GB
  3. Preserve original data — Store raw experiment before processing: orig = ms.MSExperiment(exp). Processing is destructive
  4. Profile vs centroid awareness — Check data type: spec.getType() returns 1 (profile) or 2 (centroid). Some algorithms require specific types
  5. Parameter discovery — Print parameters before tuning: for k in params.keys(): print(k, params.getValue(k))
  6. Export to pandas early — Use features.get_df() or consensus.get_df() to leverage pandas/numpy for statistical analysis
Show full SKILL.md (447 more words)Show less

Common Recipes

Recipe 1: Spectrum Statistics and Quality Check
python
import pyopenms as ms
import numpy as np

exp = ms.MSExperiment()
ms.MzMLFile().load("sample.mzML", exp)

ms1_count = sum(1 for s in exp if s.getMSLevel() == 1)
ms2_count = sum(1 for s in exp if s.getMSLevel() == 2)
rt_range = (exp.getSpectrum(0).getRT(), exp.getSpectrum(exp.getNrSpectra()-1).getRT())
peak_counts = [s.size() for s in exp]
print(f"MS1: {ms1_count}, MS2: {ms2_count}")
print(f"RT range: {rt_range[0]:.1f}–{rt_range[1]:.1f}s")
print(f"Peaks per spectrum: mean={np.mean(peak_counts):.0f}, "
      f"median={np.median(peak_counts):.0f}")
Recipe 2: Theoretical Spectrum Generation
python
import pyopenms as ms

# Generate theoretical fragment spectrum for a peptide
tsg = ms.TheoreticalSpectrumGenerator()
p = tsg.getParameters()
p.setValue("add_b_ions", "true")
p.setValue("add_y_ions", "true")
p.setValue("add_metainfo", "true")
tsg.setParameters(p)

spec = ms.MSSpectrum()
seq = ms.AASequence.fromString("PEPTIDER")
tsg.getSpectrum(spec, seq, 1, 2)  # charge 1 to 2
mz, intensity = spec.get_peaks()
print(f"Theoretical fragments for PEPTIDER: {len(mz)} ions")
Recipe 3: Format Conversion (mzXML → mzML)
python
import pyopenms as ms

exp = ms.MSExperiment()
ms.MzXMLFile().load("old_data.mzXML", exp)
ms.MzMLFile().store("converted.mzML", exp)
print(f"Converted {exp.getNrSpectra()} spectra to mzML format")

Troubleshooting

ProblemCauseSolution
FileNotFound on loadWrong path or formatVerify file exists; use correct reader class (MzMLFile for .mzML)
Empty feature map after detectionProfile data instead of centroidedRun PeakPickerHiRes before FeatureFinder
Very few features detectedToo-strict S/N thresholdLower signal_to_noise to 0.5–1.0; adjust isotopic_pattern charge range
Memory error on large filesLoading entire file at onceUse OnDiscMSExperiment for indexed access
setValue type errorWrong value type for parameterCheck expected type: params.getDescription(key). Use float for numeric, string for enum
RT alignment failsToo few common featuresIncrease max_num_peaks_considered; verify samples are from same experiment
FDR values all 1.0No decoy hits in search resultsEnsure search was run with target-decoy database; check score orientation
Feature linking too aggressiveLarge RT/m/z tolerancesReduce distance_RT:max_difference and distance_MZ:max_difference
Slow processingProcessing all spectraFilter by MS level first: [s for s in exp if s.getMSLevel() == 1]

Bundled Resources

  • references/data_structures_reference.md — Comprehensive documentation of PyOpenMS core objects (MSExperiment, MSSpectrum, Feature, FeatureMap, ConsensusMap, PeptideIdentification, AASequence, Param). Covers: object creation, attribute access, iteration patterns, memory management, type conversions.

    • Covers: all 13+ core data structure classes with creation/access/iteration examples
    • Relocated inline: Core Data Structures summary table in Key Concepts, AASequence usage in Core API Module 4
    • Omitted: exhaustive attribute listings duplicating PyOpenMS API docs
  • references/signal_feature_processing.md — Signal processing algorithms and feature detection/linking workflows consolidated from two original references.

    • Covers: smoothing (Gaussian, Savitzky-Golay), peak picking (HiRes, CWT), normalization, peak filtering (threshold, window, N-largest), baseline reduction, deconvolution, spectrum merging, RT alignment, mass calibration, feature finding (centroided, metabolomics), feature linking, adduct detection, quality control
    • Relocated inline: GaussFilter, PeakPickerHiRes, Normalizer, ThresholdMower in Core API Module 2; FeatureFinder, MapAlignment, FeatureGroupingAlgorithmQT in Core API Module 3
    • Omitted: CWT peak picker detailed parameters — rarely used vs HiRes; isotope wavelet transform — specialized use case
  • references/identification_metabolomics.md — Peptide/protein identification and untargeted metabolomics pipelines consolidated from two original references.

    • Covers: search engine integration, FDR calculation, protein inference, enzymatic digestion, spectral library search, theoretical spectrum generation, untargeted metabolomics pipeline (peak picking → feature detection → adduct detection → alignment → linking → gap filling → annotation → export), compound identification (mass-based, MS/MS-based), normalization (TIC), QC (CV filtering, blank filtering), MetaboAnalyst export
    • Relocated inline: FDR filtering, protein inference, AASequence handling in Core API Module 4; metabolomics pipeline in Core API Module 5
    • Omitted: detailed Comet/Mascot parameter configuration — search engine-specific; MetaboAnalyst export format details — external tool
  • matchms-spectral-matching — spectral library matching and metabolite identification from MS/MS
  • biopython-molecular-biology — protein sequence analysis (FASTA parsing, BLAST)

References

© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in skills/proteomics-protein-engineering/pyopenms-mass-spectrometry of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/data_structures_reference.md
  • references/identification_metabolomics.md
  • references/signal_feature_processing.md

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Pyopenms Mass Spectrometry next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Pyopenms Mass Spectrometry compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Pyopenms Mass Spectrometry this skilljaechang-hits/SciAgent-Skills3741 repos~5.9kAutomated safety check: PassBSD-3-Clause
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Pyopenms Mass Spectrometry

What does Pyopenms Mass Spectrometry do?

MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID…. Pyopenms Mass Spectrometry is an agent skill from jaechang-hits/SciAgent-Skills. MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID with FDR, untargeted metabolomics.

When should I use Pyopenms Mass Spectrometry?

Pyopenms Mass Spectrometry fits situations like: tasks that involve Bioinformatics.

How do I install Pyopenms Mass Spectrometry in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/pyopenms-mass-spectrometry in jaechang-hits/SciAgent-Skills) into .claude/skills/pyopenms-mass-spectrometry in your project. Claude Code loads it when a task matches its description.

How do I install Pyopenms Mass Spectrometry in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/pyopenms-mass-spectrometry in jaechang-hits/SciAgent-Skills) into .agents/skills/pyopenms-mass-spectrometry in your project. Codex loads it when a task matches its description.

Can I use Pyopenms Mass Spectrometry in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pyopenms-mass-spectrometry, .gemini/skills/pyopenms-mass-spectrometry, .github/skills/pyopenms-mass-spectrometry and .opencode/skills/pyopenms-mass-spectrometry in your project.

What does Pyopenms Mass Spectrometry need to run?

Going by SKILL.md and its folder, Pyopenms Mass Spectrometry needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Pyopenms Mass Spectrometry access the network?

SKILL.md names 3 domains. As links in the text: pyopenms.readthedocs.io, openms.org and github.com. This is read from the text; nothing was executed.

Is Pyopenms Mass Spectrometry safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Pyopenms Mass Spectrometry use?

Pyopenms Mass Spectrometry is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Pyopenms Mass Spectrometry use?

About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.

What are the alternatives to Pyopenms Mass Spectrometry?

Skills that share tags, products or a category with Pyopenms Mass Spectrometry: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Pyopenms Mass Spectrometry?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.