Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID…
$ npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pyopenms-mass-spectrometry --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry .claude/skills/pyopenms-mass-spectrometry && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "pyopenms-mass-spectrometry" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry into .claude/skills/pyopenms-mass-spectrometry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyopenms-mass-spectrometry", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/pyopenms-mass-spectrometryType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pyopenms-mass-spectrometry --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry .agents/skills/pyopenms-mass-spectrometry && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "pyopenms-mass-spectrometry" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry into .agents/skills/pyopenms-mass-spectrometry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyopenms-mass-spectrometry", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pyopenms-mass-spectrometry --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry .cursor/skills/pyopenms-mass-spectrometry && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "pyopenms-mass-spectrometry" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry into .cursor/skills/pyopenms-mass-spectrometry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyopenms-mass-spectrometry", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/proteomics-protein-engineering/pyopenms-mass-spectrometry--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pyopenms-mass-spectrometry --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry .gemini/skills/pyopenms-mass-spectrometry && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "pyopenms-mass-spectrometry" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry into .gemini/skills/pyopenms-mass-spectrometry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyopenms-mass-spectrometry", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills pyopenms-mass-spectrometryInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry .github/skills/pyopenms-mass-spectrometry && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "pyopenms-mass-spectrometry" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry into .github/skills/pyopenms-mass-spectrometry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyopenms-mass-spectrometry", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills pyopenms-mass-spectrometry --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry .opencode/skills/pyopenms-mass-spectrometry && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "pyopenms-mass-spectrometry" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/pyopenms-mass-spectrometry into .opencode/skills/pyopenms-mass-spectrometry/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "pyopenms-mass-spectrometry", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
pyopenms-mass-spectrometryMS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID…
Pyopenms Mass Spectrometry is an agent skill from jaechang-hits/SciAgent-Skills. MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID with FDR, untargeted metabolomics. Use matchms for simple spectral matching.
Its SKILL.md is about 5.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `references/data_structures_reference.md`, `references/identification_metabolomics.md` and `references/signal_feature_processing.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is BSD-3-Clause.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
pyopenms.readthedocs.ioopenms.orggithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Pyopenms Mass Spectrometry loads about 5.9k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 75 tokens; SKILL.md has 1,133 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its BSD-3-Clause licence (© jaechang-hits). 1,133 words, ~5,945 tokens.
.claude/skills/pyopenms-mass-spectrometry/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.PyOpenMS provides Python bindings to the OpenMS C++ library for computational mass spectrometry. It supports proteomics and metabolomics data processing including file I/O for 10+ MS formats, signal processing, feature detection, peptide/protein identification, and quantitative analysis across samples.
uv pip install pyopenms numpy pandas matplotlibalgo = Algorithm(); params = algo.getParameters(); params.setValue(...); algo.setParameters(params)Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: instrumentResolution
kind: required
source: data
ask: "Is this high-resolution data, and are the spectra already centroided or still in profile mode?"
default: null
- id: D2
param: signalToNoiseThreshold
kind: required
source: user
depends_on: [D1]
ask: "How far above noise must a peak rise to be kept during centroiding?"
default: 1.0
skip_if: "spectra already centroided by the acquisition software"
- id: D3
param: massTolerance
kind: required
source: data
depends_on: [D1]
ask: "What mass accuracy should feature detection and linking assume?"
default: "10 ppm"
- id: D4
param: retentionTimeTolerance
kind: required
source: data
ask: "How far apart in retention time may the same feature appear across runs?"
default: "100 s"
- id: D5
param: chargeStateRange
kind: required
source: user
ask: "Which charge states should features be searched for?"
default: "1 to 3 - widen for intact protein or metabolite work"
- id: D6
param: intensityNormalization
kind: optional
source: user
ask: "Should spectra be normalized before comparison, and to total ion current or to the base peak?"
default: "not normalized"
- id: D7
param: smoothingFilter
kind: optional_conditional
source: data
ask: "Does the signal need smoothing before peak picking, and with which filter width?"
default: "no smoothing"D1 governs almost everything after it: profile data that skips centroiding produces one feature per scan point, and high-resolution tolerances applied to low-resolution data link features that are not the same compound. Both yield a full feature table.
import pyopenms as ms
# Load mzML file
exp = ms.MSExperiment()
ms.MzMLFile().load("sample.mzML", exp)
print(f"Spectra: {exp.getNrSpectra()}, Chromatograms: {exp.getNrChromatograms()}")
# Examine first spectrum
spec = exp.getSpectrum(0)
mz, intensity = spec.get_peaks()
print(f"MS level: {spec.getMSLevel()}, RT: {spec.getRT():.2f}s, Peaks: {len(mz)}")
# Quick preprocessing: smooth + centroid
gauss = ms.GaussFilter()
p = gauss.getParameters(); p.setValue("gaussian_width", 0.1); gauss.setParameters(p)
gauss.filterExperiment(exp)
picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
print(f"Centroided spectra: {centroided.getNrSpectra()}")Read and write mass spectrometry data in multiple formats.
import pyopenms as ms
# Read mzML (standard MS format)
exp = ms.MSExperiment()
ms.MzMLFile().load("data.mzML", exp)
# Indexed access for large files (memory-efficient)
loader = ms.IndexedMzMLFileLoader()
indexed_file = ms.OnDiscMSExperiment()
loader.load("large_data.mzML", indexed_file)
spec = indexed_file.getSpectrum(0) # Load single spectrum on demand
print(f"Total spectra: {indexed_file.getNrSpectra()}")
# Read identification results (idXML)
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("results.idXML", protein_ids, peptide_ids)
print(f"Peptide IDs: {len(peptide_ids)}, Protein IDs: {len(protein_ids)}")
# Read feature map
fm = ms.FeatureMap()
ms.FeatureXMLFile().load("features.featureXML", fm)
print(f"Features: {fm.size()}")# Write mzML with compression
exp_out = ms.MSExperiment()
# ... populate experiment ...
ms.MzMLFile().store("output.mzML", exp_out)
# Read FASTA database
entries = []
ms.FASTAFile().load("database.fasta", entries)
print(f"Proteins in DB: {len(entries)}")
for e in entries[:3]:
print(f" {e.identifier}: {e.sequence[:30]}...")Supported formats: mzML, mzXML, mzData (spectra); featureXML, consensusXML (features); idXML, mzIdentML, pepXML (identifications); TraML (transitions); mzTab (results); FASTA (sequences)
Preprocess raw spectral data for downstream analysis.
import pyopenms as ms
exp = ms.MSExperiment()
ms.MzMLFile().load("raw.mzML", exp)
# Gaussian smoothing
gauss = ms.GaussFilter()
p = gauss.getParameters()
p.setValue("gaussian_width", 0.15) # m/z width
gauss.setParameters(p)
gauss.filterExperiment(exp)
# Savitzky-Golay smoothing (alternative)
sg = ms.SavitzkyGolayFilter()
p = sg.getParameters()
p.setValue("frame_length", 15) # Must be odd
sg.setParameters(p)
# sg.filterExperiment(exp) # Use one smoother, not both
# Peak picking (centroiding) — required before feature detection
picker = ms.PeakPickerHiRes()
p = picker.getParameters()
p.setValue("signal_to_noise", 1.0)
picker.setParameters(p)
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
print(f"Raw peaks in spec 0: {exp.getSpectrum(0).size()}")
print(f"Centroided peaks: {centroided.getSpectrum(0).size()}")# Normalization
normalizer = ms.Normalizer()
p = normalizer.getParameters()
p.setValue("method", "to_one") # "to_one" or "to_TIC"
normalizer.setParameters(p)
normalizer.filterPeakMap(centroided)
# Peak filtering — remove low-intensity noise
mower = ms.ThresholdMower()
p = mower.getParameters()
p.setValue("threshold", 100.0) # Minimum intensity
mower.setParameters(p)
mower.filterPeakMap(centroided)
# Baseline reduction
morph = ms.MorphologicalFilter()
p = morph.getParameters()
p.setValue("struc_elem_length", 3.0) # m/z window
morph.setParameters(p)
morph.filterExperiment(exp)Detect chromatographic features and link them across samples.
import pyopenms as ms
# Load centroided data
exp = ms.MSExperiment()
ms.MzMLFile().load("centroided.mzML", exp)
# Feature detection (proteomics — centroided data)
ff = ms.FeatureFinder()
features = ms.FeatureMap()
seeds = ms.FeatureMap()
params = ms.FeatureFinder().getParameters("centroided")
ff.run("centroided", exp, features, params, seeds)
print(f"Detected {features.size()} features")
# Access feature properties
for f in features[:5]:
print(f" RT: {f.getRT():.1f}s, m/z: {f.getMZ():.4f}, "
f"intensity: {f.getIntensity():.0f}, quality: {f.getOverallQuality():.3f}")# Feature linking across samples — align retention times first
aligner = ms.MapAlignmentAlgorithmPoseClustering()
p = aligner.getParameters()
p.setValue("max_num_peaks_considered", 1000)
aligner.setParameters(p)
# Link features into consensus map
linker = ms.FeatureGroupingAlgorithmQT()
p = linker.getParameters()
p.setValue("distance_RT:max_difference", 60.0) # seconds
p.setValue("distance_MZ:max_difference", 10.0) # ppm
linker.setParameters(p)
consensus = ms.ConsensusMap()
linker.group([features_sample1, features_sample2, features_sample3], consensus)
print(f"Consensus features: {consensus.size()}")
# Export to pandas for downstream analysis
import pandas as pd
df = consensus.get_df()
print(f"Consensus table: {df.shape}")Process search engine results with FDR control and protein inference.
import pyopenms as ms
# Load search engine results
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("search_results.idXML", protein_ids, peptide_ids)
# Examine peptide hits
for pep_id in peptide_ids[:3]:
print(f"Spectrum: RT={pep_id.getRT():.1f}, MZ={pep_id.getMZ():.4f}")
for hit in pep_id.getHits():
seq = hit.getSequence()
print(f" {seq} score={hit.getScore():.4f} charge={hit.getCharge()}")
# FDR filtering (target-decoy approach)
fdr = ms.FalseDiscoveryRate()
fdr.apply(peptide_ids)
# Filter at 1% FDR
filtered = []
for pep_id in peptide_ids:
hits = [h for h in pep_id.getHits() if h.getScore() <= 0.01]
if hits:
pep_id.setHits(hits)
filtered.append(pep_id)
print(f"Peptide IDs at 1% FDR: {len(filtered)}")# Protein inference
inference = ms.BasicProteinInferenceAlgorithm()
inference.run(peptide_ids, protein_ids)
for prot_id in protein_ids:
for hit in prot_id.getHits()[:5]:
print(f"Protein: {hit.getAccession()}, score: {hit.getScore():.4f}")
# Peptide sequence handling
seq = ms.AASequence.fromString("PEPTIDER")
print(f"Molecular weight: {seq.getMonoWeight():.4f}")
print(f"Formula: {seq.getFormula()}")
# Modified sequence
mod_seq = ms.AASequence.fromString("PEPTM(Oxidation)DER")
print(f"Modified weight: {mod_seq.getMonoWeight():.4f}")
# Enzymatic digestion
digestor = ms.ProteaseDigestion()
digestor.setEnzyme("Trypsin")
digest = []
digestor.digest(ms.AASequence.fromString("MKWVTFISLLLLFSSAYSRGVFRR"), digest)
print(f"Tryptic peptides: {len(digest)}")Complete untargeted metabolomics workflow from raw data to feature table.
import pyopenms as ms
# Step 1: Load and centroid raw data
exp = ms.MSExperiment()
ms.MzMLFile().load("metabolomics_sample.mzML", exp)
picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
# Step 2: Feature detection for metabolomics (small molecules)
ff = ms.FeatureFinder()
features = ms.FeatureMap()
seeds = ms.FeatureMap()
params = ms.FeatureFinder().getParameters("centroided")
params.setValue("isotopic_pattern:charge_low", 1)
params.setValue("isotopic_pattern:charge_high", 3)
ff.run("centroided", centroided, features, params, seeds)
print(f"Detected features: {features.size()}")
# Step 3: Adduct detection (group related adducts)
decharger = ms.MetaboliteAdductDecharger()
p = decharger.getParameters()
p.setValue("potential_adducts", "H:+:0.6;Na:+:0.3;K:+:0.1") # Positive mode
decharger.setParameters(p)
# decharger.compute(features, feature_map_out, consensus_map_out)# Step 4: RT alignment across samples
import pyopenms as ms
import pandas as pd
# Assuming feature maps from multiple samples
sample_files = ["sample1.featureXML", "sample2.featureXML", "sample3.featureXML"]
feature_maps = []
for f in sample_files:
fm = ms.FeatureMap()
ms.FeatureXMLFile().load(f, fm)
feature_maps.append(fm)
# Align retention times
aligner = ms.MapAlignmentAlgorithmPoseClustering()
aligner.setReference(0) # Use first sample as reference
# Step 5: Link features to consensus map
linker = ms.FeatureGroupingAlgorithmQT()
p = linker.getParameters()
p.setValue("distance_RT:max_difference", 30.0)
p.setValue("distance_MZ:max_difference", 10.0)
linker.setParameters(p)
consensus = ms.ConsensusMap()
linker.group(feature_maps, consensus)
# Step 6: Export metabolite table
df = consensus.get_df()
print(f"Feature table: {df.shape[0]} features × {df.shape[1]} columns")
print(df[["RT", "mz", "intensity_0", "intensity_1", "intensity_2"]].head())All PyOpenMS algorithms follow the same 3-step pattern:
algo = ms.AlgorithmClass() # 1. Instantiate
params = algo.getParameters() # 2. Get parameters
params.setValue("param_name", value) # 3. Set values
algo.setParameters(params) # 4. Apply
# algo.process(input, output) # 5. ExecuteTo discover available parameters:
for key in params.keys():
print(f"{key}: {params.getValue(key)} (type: {params.getDescription(key)})")| Object | Description | Key Methods |
|---|---|---|
MSExperiment | Collection of spectra + chromatograms | getNrSpectra(), getSpectrum(i), addSpectrum() |
MSSpectrum | Single mass spectrum (m/z, intensity) | get_peaks(), getMSLevel(), getRT(), size() |
MSChromatogram | Chromatographic trace | get_peaks(), getNativeID() |
Feature | Detected chromatographic peak | getRT(), getMZ(), getIntensity(), getOverallQuality() |
FeatureMap | Collection of features | size(), get_df(), iteration |
ConsensusMap | Features linked across samples | size(), get_df() |
PeptideIdentification | Search results for one spectrum | getHits(), getRT(), getMZ() |
AASequence | Amino acid sequence with modifications | fromString(), getMonoWeight(), getFormula() |
| Format | Reader Class | Content |
|---|---|---|
| mzML | MzMLFile | Raw spectra (standard) |
| mzXML | MzXMLFile | Raw spectra (legacy) |
| featureXML | FeatureXMLFile | Detected features |
| consensusXML | ConsensusXMLFile | Linked features |
| idXML | IdXMLFile | Peptide/protein IDs |
| mzIdentML | MzIdentMLFile | Peptide/protein IDs (standard) |
| FASTA | FASTAFile | Protein sequences |
| TraML | TraMLFile | MRM transitions |
| mzTab | MzTabFile | Quantification results |
import pyopenms as ms
# Load → Smooth → Centroid → Detect → Export
exp = ms.MSExperiment()
ms.MzMLFile().load("proteomics.mzML", exp)
# Preprocessing
gauss = ms.GaussFilter()
p = gauss.getParameters(); p.setValue("gaussian_width", 0.1); gauss.setParameters(p)
gauss.filterExperiment(exp)
picker = ms.PeakPickerHiRes()
centroided = ms.MSExperiment()
picker.pickExperiment(exp, centroided)
# Feature detection
ff = ms.FeatureFinder()
features = ms.FeatureMap()
params = ff.getParameters("centroided")
ff.run("centroided", centroided, features, params, ms.FeatureMap())
# Export to pandas
df = features.get_df()
print(f"Features: {df.shape}")
df.to_csv("proteomics_features.csv", index=False)PeakPickerHiRes)FeatureFinder)MapAlignmentAlgorithmPoseClustering)FeatureGroupingAlgorithmQT)import pyopenms as ms
import pandas as pd
# Load and filter identifications
protein_ids, peptide_ids = [], []
ms.IdXMLFile().load("search.idXML", protein_ids, peptide_ids)
# FDR filtering
fdr = ms.FalseDiscoveryRate()
fdr.apply(peptide_ids)
# Build results table
rows = []
for pep_id in peptide_ids:
for hit in pep_id.getHits():
if hit.getScore() <= 0.01: # 1% FDR
rows.append({
"sequence": str(hit.getSequence()),
"score": hit.getScore(),
"charge": hit.getCharge(),
"rt": pep_id.getRT(),
"mz": pep_id.getMZ()
})
df = pd.DataFrame(rows)
print(f"Identified peptides at 1% FDR: {len(df)}")
print(f"Unique sequences: {df['sequence'].nunique()}")
df.to_csv("peptide_identifications.csv", index=False)| Parameter | Module | Default | Range/Options | Effect |
|---|---|---|---|---|
gaussian_width | GaussFilter | 0.2 | 0.05–1.0 | m/z smoothing window width |
frame_length | SavitzkyGolayFilter | 11 | 5–31 (odd) | Smoothing window points |
signal_to_noise | PeakPickerHiRes | 1.0 | 0.5–10.0 | Minimum S/N for centroiding |
isotopic_pattern:charge_low | FeatureFinder | 1 | 1–6 | Minimum charge state |
isotopic_pattern:charge_high | FeatureFinder | 3 | 1–10 | Maximum charge state |
distance_RT:max_difference | FeatureGroupingAlgorithmQT | 100 | 10–300 s | RT tolerance for feature linking |
distance_MZ:max_difference | FeatureGroupingAlgorithmQT | 10 | 1–50 ppm | m/z tolerance for feature linking |
threshold | ThresholdMower | 0.0 | 0–10000 | Minimum peak intensity |
method | Normalizer | "to_one" | "to_one"/"to_TIC" | Normalization method |
PeakPickerHiRes firstOnDiscMSExperiment loads spectra on demand, avoiding memory issues for files >1 GBorig = ms.MSExperiment(exp). Processing is destructivespec.getType() returns 1 (profile) or 2 (centroid). Some algorithms require specific typesfor k in params.keys(): print(k, params.getValue(k))features.get_df() or consensus.get_df() to leverage pandas/numpy for statistical analysisimport pyopenms as ms
import numpy as np
exp = ms.MSExperiment()
ms.MzMLFile().load("sample.mzML", exp)
ms1_count = sum(1 for s in exp if s.getMSLevel() == 1)
ms2_count = sum(1 for s in exp if s.getMSLevel() == 2)
rt_range = (exp.getSpectrum(0).getRT(), exp.getSpectrum(exp.getNrSpectra()-1).getRT())
peak_counts = [s.size() for s in exp]
print(f"MS1: {ms1_count}, MS2: {ms2_count}")
print(f"RT range: {rt_range[0]:.1f}–{rt_range[1]:.1f}s")
print(f"Peaks per spectrum: mean={np.mean(peak_counts):.0f}, "
f"median={np.median(peak_counts):.0f}")import pyopenms as ms
# Generate theoretical fragment spectrum for a peptide
tsg = ms.TheoreticalSpectrumGenerator()
p = tsg.getParameters()
p.setValue("add_b_ions", "true")
p.setValue("add_y_ions", "true")
p.setValue("add_metainfo", "true")
tsg.setParameters(p)
spec = ms.MSSpectrum()
seq = ms.AASequence.fromString("PEPTIDER")
tsg.getSpectrum(spec, seq, 1, 2) # charge 1 to 2
mz, intensity = spec.get_peaks()
print(f"Theoretical fragments for PEPTIDER: {len(mz)} ions")import pyopenms as ms
exp = ms.MSExperiment()
ms.MzXMLFile().load("old_data.mzXML", exp)
ms.MzMLFile().store("converted.mzML", exp)
print(f"Converted {exp.getNrSpectra()} spectra to mzML format")| Problem | Cause | Solution |
|---|---|---|
FileNotFound on load | Wrong path or format | Verify file exists; use correct reader class (MzMLFile for .mzML) |
| Empty feature map after detection | Profile data instead of centroided | Run PeakPickerHiRes before FeatureFinder |
| Very few features detected | Too-strict S/N threshold | Lower signal_to_noise to 0.5–1.0; adjust isotopic_pattern charge range |
| Memory error on large files | Loading entire file at once | Use OnDiscMSExperiment for indexed access |
setValue type error | Wrong value type for parameter | Check expected type: params.getDescription(key). Use float for numeric, string for enum |
| RT alignment fails | Too few common features | Increase max_num_peaks_considered; verify samples are from same experiment |
| FDR values all 1.0 | No decoy hits in search results | Ensure search was run with target-decoy database; check score orientation |
| Feature linking too aggressive | Large RT/m/z tolerances | Reduce distance_RT:max_difference and distance_MZ:max_difference |
| Slow processing | Processing all spectra | Filter by MS level first: [s for s in exp if s.getMSLevel() == 1] |
references/data_structures_reference.md — Comprehensive documentation of PyOpenMS core objects (MSExperiment, MSSpectrum, Feature, FeatureMap, ConsensusMap, PeptideIdentification, AASequence, Param). Covers: object creation, attribute access, iteration patterns, memory management, type conversions.
references/signal_feature_processing.md — Signal processing algorithms and feature detection/linking workflows consolidated from two original references.
references/identification_metabolomics.md — Peptide/protein identification and untargeted metabolomics pipelines consolidated from two original references.
© jaechang-hits, BSD-3-Clause. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (references) in skills/proteomics-protein-engineering/pyopenms-mass-spectrometry of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Pyopenms Mass Spectrometry next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Pyopenms Mass Spectrometry this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~5.9k | Automated safety check: Pass | BSD-3-Clause | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID…. Pyopenms Mass Spectrometry is an agent skill from jaechang-hits/SciAgent-Skills. MS data processing with PyOpenMS for LC-MS/MS proteomics and metabolomics — mzML/mzXML I/O, signal processing (smoothing, peak picking, centroiding), feature detection/linking, peptide/protein ID with FDR, untargeted metabolomics.
Pyopenms Mass Spectrometry fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/pyopenms-mass-spectrometry in jaechang-hits/SciAgent-Skills) into .claude/skills/pyopenms-mass-spectrometry in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/pyopenms-mass-spectrometry in jaechang-hits/SciAgent-Skills) into .agents/skills/pyopenms-mass-spectrometry in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill pyopenms-mass-spectrometry -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/pyopenms-mass-spectrometry, .gemini/skills/pyopenms-mass-spectrometry, .github/skills/pyopenms-mass-spectrometry and .opencode/skills/pyopenms-mass-spectrometry in your project.
Going by SKILL.md and its folder, Pyopenms Mass Spectrometry needs the command-line tools its instructions call (uv). Our summary lists: Python 3.
SKILL.md names 3 domains. As links in the text: pyopenms.readthedocs.io, openms.org and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Pyopenms Mass Spectrometry is published under the BSD-3-Clause licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.9k tokens (SKILL.md is roughly 24k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Pyopenms Mass Spectrometry: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.