Agent skill

Matchms Spectral Matching

by jaechang-hits in jaechang-hits/SciAgent-Skills

MS spectral matching and metabolite ID with matchms. An agent skill from jaechang-hits/SciAgent-Skills.

Apache-2.0Auto-check passedResearch & Science

Install Matchms Spectral Matching

skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jaechang-hits/SciAgent-Skills matchms-spectral-matching --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/matchms-spectral-matching .claude/skills/matchms-spectral-matching && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
matchms-spectral-matching
GitHub stars
374
Used in
1 other repo
Token cost
~6.3k tokens
SKILL.md length
1,123 words
Files
3 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
Apache-2.0

At a glance

MS spectral matching and metabolite ID with matchms. An agent skill from jaechang-hits/SciAgent-Skills.

  • Works in 3 steps: Load spectra from source format (Core… → Apply default_filters for metadata… → Export to target format (Core API Module…
  • Tasks that involve Bioinformatics
  • SKILL.md covers Overview, When to Use, Prerequisites and Pre-flight Interview, plus 11 more sections
  • Calls uv and pip

What it does

Matchms Spectral Matching is an agent skill from jaechang-hits/SciAgent-Skills. MS spectral matching and metabolite ID with matchms. Import spectra (mzML, MGF, MSP, JSON), filter/normalize peaks, score similarity (cosine, modified cosine, fingerprint), build reproducible pipelines, identify unknowns vs spectral libraries. Use pyopenms for full LC-MS/MS proteomics.

Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/filtering_catalog.md` and `references/workflows_similarity.md`).

It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Bioinformatics

Example prompts

  • “/matchms-spectral-matching”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Load spectra from source format (Core API Module 1 — load_from_mzml)
  2. Apply default_filters for metadata harmonization (Core API Module 2)
  3. Export to target format (Core API Module 1 — save_as_mgf)

What it can do on your machine

Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • matchms.readthedocs.io
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Matchms Spectral Matching loads about 6.3k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 1,123 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~6.3k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~13k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 1,123 words, ~6,301 tokens.

Download SKILL.mdSave it as .claude/skills/matchms-spectral-matching/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
matchms-spectral-matching
description
MS spectral matching and metabolite ID with matchms. Import spectra (mzML, MGF, MSP, JSON), filter/normalize peaks, score similarity (cosine, modified cosine, fingerprint), build reproducible pipelines, identify unknowns vs spectral libraries. Use pyopenms for full LC-MS/MS proteomics.
license
Apache-2.0

Matchms — Spectral Matching & Metabolite Identification

Overview

Matchms is a Python library for mass spectrometry data processing focused on spectral similarity calculation and compound identification. It provides multi-format I/O, 50+ spectrum filters for metadata harmonization and peak processing, 8 similarity scoring functions, and a pipeline framework for reproducible analytical workflows.

When to Use

  • Identifying unknown metabolites by matching MS/MS spectra against reference libraries
  • Computing spectral similarity scores (cosine, modified cosine, fingerprint-based)
  • Processing and standardizing mass spectral data from multiple formats (mzML, MGF, MSP, JSON)
  • Building reproducible spectral processing pipelines for quality control
  • Harmonizing metadata across spectral databases (compound names, SMILES, InChI, adducts)
  • Large-scale spectral library comparisons and duplicate detection
  • For full LC-MS/MS proteomics workflows (feature detection, protein ID), use pyopenms instead
  • For chemical structure similarity without mass spectra, use rdkit fingerprint comparison

Prerequisites

bash
uv pip install matchms numpy pandas
# For chemical structure processing (SMILES, InChI, fingerprints):
uv pip install matchms[chemistry]
  • Python 3.8+; NumPy for peak array operations
  • Input: spectral data in MGF, MSP, mzML, mzXML, JSON, or pickle format
  • Reference library in any supported format for matching

Pre-flight Interview

Settle these with the user before writing any analysis code.

yaml
decisions:
  - id: D1
    param: referenceLibrary
    kind: required
    source: user
    ask: "Which spectral library should queries be matched against?"
    default: null

  - id: D2
    param: similarityMeasure
    kind: required
    source: user
    ask: "Should matching require the same precursor mass, or allow a mass shift so modified analogues still match?"
    default: "cosine on matched peaks, precursor mass fixed"

  - id: D3
    param: peakMatchTolerance
    kind: required
    source: data
    ask: "How close must two peaks be in m/z to count as the same fragment?"
    default: "0.1 Da - tighten toward 0.005 for high-resolution data"

  - id: D4
    param: spectrumPreprocessing
    kind: required
    source: user
    ask: "Which peaks should be discarded before scoring - low-intensity noise, peaks near the precursor, or spectra with too few peaks to be informative?"
    default: "drop peaks under 1% relative intensity, require at least 10 peaks"

  - id: D5
    param: scoreCutoff
    kind: required
    source: user
    ask: "What similarity score, and how many matched peaks, should a hit need before it is reported as an identification?"
    default: null

  - id: D6
    param: peakWeighting
    kind: optional
    source: user
    ask: "Should the score weight fragment m/z as well as intensity, favouring informative high-mass fragments?"
    default: "intensity only"

  - id: D7
    param: fingerprintSettings
    kind: optional_conditional
    source: user
    ask: "Which molecular fingerprint should structural similarity use, when comparing hits to known structures?"
    default: "not computed"

D3 and D5 have no safe shared default because they trade off directly: a loose tolerance with a low score cutoff returns an identification for every query spectrum, all of them plausible-looking. Set the tolerance from the instrument's actual accuracy and the cutoff from what the library supports.

Quick Start

python
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters, normalize_intensities
from matchms.filtering import select_by_relative_intensity, require_minimum_number_of_peaks
from matchms import calculate_scores
from matchms.similarity import CosineGreedy

# Load and process query spectra
queries = list(load_from_mgf("queries.mgf"))
queries = [default_filters(s) for s in queries]
queries = [normalize_intensities(s) for s in queries if s is not None]
queries = [require_minimum_number_of_peaks(s, n_required=5) for s in queries if s is not None]

# Load reference library
refs = list(load_from_mgf("library.mgf"))
refs = [default_filters(s) for s in refs]
refs = [normalize_intensities(s) for s in refs if s is not None]

# Calculate similarity scores
scores = calculate_scores(references=refs, queries=queries,
                          similarity_function=CosineGreedy(tolerance=0.1))
# Get best matches for first query
best = scores.scores_by_query(queries[0], sort=True)[:5]
for match, score_tuple in best:
    print(f"Score: {score_tuple['score']:.3f}, Matches: {score_tuple['matches']}")

Core API

Module 1: Spectrum I/O

Import spectra from multiple file formats and export processed data.

python
from matchms.importing import (load_from_mgf, load_from_mzml, load_from_msp,
                                load_from_json, load_from_mzxml, load_from_pickle,
                                load_from_usi)
from matchms.exporting import save_as_mgf, save_as_msp, save_as_json, save_as_pickle

# Import from various formats (returns generators)
spectra_mgf = list(load_from_mgf("library.mgf"))
spectra_mzml = list(load_from_mzml("data.mzML"))
spectra_msp = list(load_from_msp("nist_library.msp"))
spectra_json = list(load_from_json("gnps_spectra.json"))
print(f"Loaded: MGF={len(spectra_mgf)}, mzML={len(spectra_mzml)}")

# Export processed spectra
save_as_mgf(spectra_mgf, "processed.mgf")
save_as_json(spectra_mgf, "processed.json")
save_as_pickle(spectra_mgf, "spectra.pickle")  # Fast for intermediate results

# Pickle for large datasets (fastest I/O)
from matchms.importing import load_from_pickle
spectra = list(load_from_pickle("spectra.pickle"))
python
from matchms import Spectrum
import numpy as np

# Create spectrum manually
mz = np.array([100.0, 150.0, 200.0, 250.0, 300.0])
intensities = np.array([0.1, 0.5, 0.9, 0.3, 0.7])
metadata = {
    "precursor_mz": 325.5,
    "ionmode": "positive",
    "compound_name": "Caffeine",
    "smiles": "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"
}
spectrum = Spectrum(mz=mz, intensities=intensities, metadata=metadata)

# Access spectrum data
print(f"Peaks: {spectrum.peaks.mz}")
print(f"Precursor: {spectrum.get('precursor_mz')}")
print(f"Name: {spectrum.get('compound_name')}")

# Visualize
spectrum.plot()
Module 2: Spectrum Filtering & Processing

Apply metadata harmonization and peak processing filters. Matchms provides 50+ filters.

python
from matchms.filtering import (
    default_filters, normalize_intensities,
    select_by_relative_intensity, select_by_mz,
    require_minimum_number_of_peaks, reduce_to_number_of_peaks,
    remove_peaks_around_precursor_mz, add_losses
)

# default_filters applies: metadata cleanup, charge correction, adduct parsing
spectrum = default_filters(spectrum)

# Peak normalization (max intensity → 1.0)
spectrum = normalize_intensities(spectrum)

# Filter peaks by relative intensity (remove noise below 1%)
spectrum = select_by_relative_intensity(spectrum, intensity_from=0.01, intensity_to=1.0)

# Filter peaks by m/z range
spectrum = select_by_mz(spectrum, mz_from=50.0, mz_to=500.0)

# Keep top N peaks only
spectrum = reduce_to_number_of_peaks(spectrum, n_max=50)

# Remove peaks near precursor (common contaminants)
spectrum = remove_peaks_around_precursor_mz(spectrum, mz_tolerance=17.0)

# Require minimum peaks for matching
spectrum = require_minimum_number_of_peaks(spectrum, n_required=5)

# Add neutral losses (useful for NeutralLossesCosine)
spectrum = add_losses(spectrum)

if spectrum is not None:
    print(f"After filtering: {len(spectrum.peaks.mz)} peaks")
python
# Chemical annotation filters (require matchms[chemistry])
from matchms.filtering import (
    derive_inchi_from_smiles, derive_inchikey_from_inchi,
    derive_smiles_from_inchi, add_fingerprint,
    repair_inchi_inchikey_smiles, require_valid_annotation
)

# Derive chemical identifiers from SMILES
spectrum = derive_inchi_from_smiles(spectrum)
spectrum = derive_inchikey_from_inchi(spectrum)

# Add molecular fingerprint for structural similarity
spectrum = add_fingerprint(spectrum, fingerprint_type="morgan", nbits=2048)

# Validate annotations
spectrum = require_valid_annotation(spectrum)
if spectrum is not None:
    print(f"InChIKey: {spectrum.get('inchikey')}")
Module 3: Similarity Scoring

Compare spectra using multiple similarity metrics.

python
from matchms import calculate_scores
from matchms.similarity import (
    CosineGreedy, CosineHungarian, ModifiedCosine,
    NeutralLossesCosine, FingerprintSimilarity,
    MetadataMatch, PrecursorMzMatch
)

# CosineGreedy — fast peak matching (greedy algorithm)
scores = calculate_scores(references=library, queries=unknowns,
                          similarity_function=CosineGreedy(tolerance=0.1))

# ModifiedCosine — accounts for precursor mass differences (best for analog search)
scores = calculate_scores(references=library, queries=unknowns,
                          similarity_function=ModifiedCosine(tolerance=0.1))

# CosineHungarian — optimal peak matching (slower but more accurate)
scores = calculate_scores(references=library, queries=unknowns,
                          similarity_function=CosineHungarian(tolerance=0.1))

# NeutralLossesCosine — similarity based on neutral loss patterns
scores = calculate_scores(references=library, queries=unknowns,
                          similarity_function=NeutralLossesCosine(tolerance=0.1))

# Access results
for i, query in enumerate(unknowns[:3]):
    best_matches = scores.scores_by_query(query, sort=True)[:3]
    print(f"\nQuery {i}: precursor_mz={query.get('precursor_mz')}")
    for ref, score_tuple in best_matches:
        print(f"  {ref.get('compound_name', 'Unknown')}: "
              f"score={score_tuple['score']:.3f}, matches={score_tuple['matches']}")
python
# FingerprintSimilarity — structural similarity (requires fingerprints)
from matchms.similarity import FingerprintSimilarity
scores = calculate_scores(references=library, queries=unknowns,
                          similarity_function=FingerprintSimilarity(
                              similarity_measure="jaccard"))

# PrecursorMzMatch — fast mass-based pre-filtering
from matchms.similarity import PrecursorMzMatch
scores = calculate_scores(references=library, queries=unknowns,
                          similarity_function=PrecursorMzMatch(tolerance=0.1))

# Multi-metric scoring: combine peak + structural similarity
cosine_scores = calculate_scores(references=library, queries=unknowns,
                                  similarity_function=CosineGreedy(tolerance=0.1))
fp_scores = calculate_scores(references=library, queries=unknowns,
                              similarity_function=FingerprintSimilarity())
Module 4: Processing Pipelines

Build reusable, reproducible multi-step processing workflows.

python
from matchms import SpectrumProcessor
from matchms.filtering import (
    default_filters, normalize_intensities,
    select_by_relative_intensity, require_minimum_number_of_peaks,
    remove_peaks_around_precursor_mz, add_losses
)

# Define reusable pipeline
pipeline = SpectrumProcessor([
    default_filters,
    normalize_intensities,
    lambda s: select_by_relative_intensity(s, intensity_from=0.01),
    lambda s: remove_peaks_around_precursor_mz(s, mz_tolerance=17.0),
    lambda s: require_minimum_number_of_peaks(s, n_required=5),
    add_losses
])

# Apply to all spectra (filters returning None remove the spectrum)
processed = [pipeline(s) for s in raw_spectra]
processed = [s for s in processed if s is not None]
print(f"Processed: {len(processed)}/{len(raw_spectra)} spectra retained")

Key Concepts

Similarity Function Comparison
FunctionSpeedAccuracyBest For
CosineGreedyFastGoodGeneral library matching
CosineHungarianSlowBestSmall comparisons, validation
ModifiedCosineFastGoodAnalog search (different precursors)
NeutralLossesCosineMediumGoodStructural class identification
FingerprintSimilarityFastModerateStructure-based pre-filtering
PrecursorMzMatchFastestN/AMass-based pre-filtering
Filter Categories
CategoryExamplesPurpose
Metadata cleanupdefault_filters, clean_compound_name, clean_adductStandardize metadata fields
Chemical derivationderive_inchi_from_smiles, add_fingerprintCompute chemical identifiers
Mass/chargeadd_precursor_mz, correct_charge, add_parent_massFix and validate mass info
Peak normalizationnormalize_intensities, select_by_relative_intensityScale and filter peaks
Peak reductionreduce_to_number_of_peaks, remove_peaks_around_precursor_mzRemove noise/artifacts
Quality controlrequire_minimum_number_of_peaks, require_precursor_mzEnforce minimum quality
Neutral lossesadd_lossesCompute precursor-fragment losses
Score Tuple Structure

All similarity functions return (score, matches):

  • score: float 0.0–1.0 (cosine similarity value)
  • matches: int (number of matched peaks between query and reference)

Higher scores and more matched peaks indicate better matches. Typical thresholds: score > 0.7 and matches > 6 for confident identifications.

Common Workflows

Workflow 1: Library Matching for Metabolite Identification
python
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters, normalize_intensities
from matchms.filtering import select_by_relative_intensity, require_minimum_number_of_peaks
from matchms import calculate_scores
from matchms.similarity import ModifiedCosine
import pandas as pd

# Load and process both queries and library identically
def process_spectra(spectra):
    processed = []
    for s in spectra:
        s = default_filters(s)
        if s is None: continue
        s = normalize_intensities(s)
        s = select_by_relative_intensity(s, intensity_from=0.01)
        s = require_minimum_number_of_peaks(s, n_required=5)
        if s is not None:
            processed.append(s)
    return processed

queries = process_spectra(load_from_mgf("unknowns.mgf"))
library = process_spectra(load_from_mgf("reference_library.mgf"))
print(f"Queries: {len(queries)}, Library: {len(library)}")

# Score all query-reference pairs
scores = calculate_scores(references=library, queries=queries,
                          similarity_function=ModifiedCosine(tolerance=0.1))

# Extract best matches
results = []
for query in queries:
    best = scores.scores_by_query(query, sort=True)[:1]
    if best:
        ref, score_tuple = best[0]
        results.append({
            "query_precursor_mz": query.get("precursor_mz"),
            "match_name": ref.get("compound_name", "Unknown"),
            "match_smiles": ref.get("smiles", ""),
            "score": score_tuple["score"],
            "matched_peaks": score_tuple["matches"]
        })

df = pd.DataFrame(results)
confident = df[df["score"] > 0.7]
print(f"Confident matches (score>0.7): {len(confident)}/{len(df)}")
df.to_csv("identification_results.csv", index=False)
Workflow 2: Quality Control and Data Cleaning
python
from matchms.importing import load_from_msp
from matchms.exporting import save_as_mgf
from matchms import SpectrumProcessor
from matchms.filtering import (
    default_filters, normalize_intensities,
    select_by_relative_intensity, require_minimum_number_of_peaks,
    require_precursor_mz, add_parent_mass
)

# Define QC pipeline
qc_pipeline = SpectrumProcessor([
    default_filters,
    require_precursor_mz,
    add_parent_mass,
    normalize_intensities,
    lambda s: select_by_relative_intensity(s, intensity_from=0.001),
    lambda s: require_minimum_number_of_peaks(s, n_required=3)
])

# Process and filter
raw = list(load_from_msp("raw_library.msp"))
cleaned = [qc_pipeline(s) for s in raw]
cleaned = [s for s in cleaned if s is not None]

print(f"Input: {len(raw)}, Output: {len(cleaned)} ({len(cleaned)/len(raw)*100:.1f}% retained)")
save_as_mgf(cleaned, "cleaned_library.mgf")
Workflow 3: Format Conversion
  1. Load spectra from source format (Core API Module 1 — load_from_mzml)
  2. Apply default_filters for metadata harmonization (Core API Module 2)
  3. Export to target format (Core API Module 1 — save_as_mgf)

Key Parameters

ParameterFunction/ModuleDefaultRange/OptionsEffect
toleranceCosineGreedy/ModifiedCosine0.10.005–0.5 Dam/z tolerance for peak matching
mz_powerCosineGreedy0.00.0–2.0Weight of m/z in scoring (0=ignore)
intensity_powerCosineGreedy1.00.0–2.0Weight of intensity in scoring
intensity_fromselect_by_relative_intensity0.00.0–1.0Minimum relative intensity to keep
n_requiredrequire_minimum_number_of_peaks101–100Minimum peaks to retain spectrum
n_maxreduce_to_number_of_peaks10010–500Maximum peaks to retain
mz_toleranceremove_peaks_around_precursor_mz17.00.5–50 DaWindow around precursor to remove
fingerprint_typeadd_fingerprint"daylight""daylight"/"morgan"/"maccs"Molecular fingerprint type
nbitsadd_fingerprint2048256–4096Fingerprint bit vector length

Best Practices

  1. Always process queries and references identically — apply the same filtering pipeline to both sets to avoid systematic bias in similarity scores
  2. Save intermediate results in pickle — pickle format is fastest for re-loading; use MGF/MSP for sharing with other tools
  3. Pre-filter by precursor mass — for large libraries, use PrecursorMzMatch first to reduce the comparison space, then score with CosineGreedy
  4. Combine multiple metrics — use both CosineGreedy (peak matching) and FingerprintSimilarity (structure) for more robust identification
  5. Check for None after filtering — filters return None when a spectrum fails quality requirements. Always filter: [s for s in processed if s is not None]
  6. Use ModifiedCosine for analog search — when querying against libraries that may not have the exact compound, ModifiedCosine handles precursor mass differences
Show full SKILL.md (439 more words)Show less

Common Recipes

Recipe 1: Precursor-Filtered Library Search (Efficient)
python
from matchms import calculate_scores
from matchms.similarity import PrecursorMzMatch, CosineGreedy

# Step 1: Fast mass filter
mass_scores = calculate_scores(references=library, queries=unknowns,
                                similarity_function=PrecursorMzMatch(tolerance=0.5))

# Step 2: Detailed scoring only for mass-matched pairs
cosine = CosineGreedy(tolerance=0.1)
for query in unknowns:
    candidates = mass_scores.scores_by_query(query, sort=True)
    mass_matched = [ref for ref, score in candidates if score["score"] > 0]
    if mass_matched:
        detailed = calculate_scores(references=mass_matched, queries=[query],
                                     similarity_function=cosine)
        best = detailed.scores_by_query(query, sort=True)[:3]
        for ref, s in best:
            print(f"{ref.get('compound_name')}: {s['score']:.3f}")
Recipe 2: Ion Mode-Specific Processing
python
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters, normalize_intensities

spectra = list(load_from_mgf("mixed_library.mgf"))
spectra = [default_filters(s) for s in spectra]
spectra = [s for s in spectra if s is not None]

# Separate by ion mode
positive = [s for s in spectra if s.get("ionmode") == "positive"]
negative = [s for s in spectra if s.get("ionmode") == "negative"]
print(f"Positive: {len(positive)}, Negative: {len(negative)}")

# Process each mode with mode-specific filtering
positive = [normalize_intensities(s) for s in positive]
negative = [normalize_intensities(s) for s in negative]
Recipe 3: Metadata Enrichment Report
python
import pandas as pd
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters

spectra = [default_filters(s) for s in load_from_mgf("library.mgf")]
spectra = [s for s in spectra if s is not None]

# Extract metadata summary
rows = []
for s in spectra:
    rows.append({
        "compound_name": s.get("compound_name", ""),
        "precursor_mz": s.get("precursor_mz"),
        "ionmode": s.get("ionmode", ""),
        "smiles": s.get("smiles", ""),
        "inchikey": s.get("inchikey", ""),
        "num_peaks": len(s.peaks.mz)
    })

df = pd.DataFrame(rows)
print(f"Library: {len(df)} spectra")
print(f"Named: {(df.compound_name != '').sum()}")
print(f"With SMILES: {(df.smiles != '').sum()}")
print(f"Ion modes: {df.ionmode.value_counts().to_dict()}")

Troubleshooting

ProblemCauseSolution
All scores are 0.0No matching peaks within toleranceIncrease tolerance (try 0.2–0.5 Da); verify both spectra have peaks
Low scores despite same compoundDifferent fragmentation conditionsUse ModifiedCosine instead of CosineGreedy; check ion mode consistency
Many spectra filtered to NoneToo strict quality filtersLower n_required in require_minimum_number_of_peaks; relax intensity thresholds
KeyError on metadata fieldField name not harmonizedApply default_filters first to harmonize metadata keys
Memory error with large libraryAll-vs-all comparisonPre-filter by precursor mass (PrecursorMzMatch) before detailed scoring
add_fingerprint failsRDKit not installedInstall chemistry extras: pip install matchms[chemistry]
Import returns empty listWrong file format or pathVerify format matches loader (MGF for .mgf, MSP for .msp); check file is not empty
Inconsistent scores between runsDifferent processing pipelinesUse SpectrumProcessor to ensure identical processing for queries and references

Bundled Resources

  • references/filtering_catalog.md — Complete catalog of 50+ matchms filter functions organized by category (metadata processing, chemical structure, mass/charge, peak processing, quality control).

    • Covers: all filter function signatures with parameters and brief descriptions, common filter combinations
    • Relocated inline: key filters (default_filters, normalize_intensities, select_by_relative_intensity, reduce_to_number_of_peaks, remove_peaks_around_precursor_mz, add_fingerprint) in Core API Module 2; filter category summary in Key Concepts
    • Omitted: individual filter examples duplicating Core API patterns
  • references/workflows_similarity.md — Extended workflows and detailed similarity function documentation consolidated from two original references.

    • Covers: all 8 similarity functions with parameter tables, large-scale comparison strategies, multi-metric scoring patterns, 7 extended workflows (library matching, QC, multi-metric, format conversion, metadata enrichment, large-scale comparison, automated identification report), performance considerations
    • Relocated inline: CosineGreedy/ModifiedCosine/CosineHungarian/FingerprintSimilarity usage in Core API Module 3; similarity comparison table in Key Concepts; library matching + QC workflows in Common Workflows
    • Omitted: network-based spectral clustering workflow — requires external tool (spec2vec); spectra visualization tutorial — covered by spectrum.plot() in Core API Module 1

Disposition of original reference files:

  • importing_exporting.md (417 lines) → fully consolidated into Core API Module 1 (I/O functions, format list, Spectrum creation, pickle usage). Retained: all 7 import functions, 4 export functions, Spectrum class creation, format selection guidance. Omitted: USI detailed examples — niche use case
  • filtering.md (289 lines) → migrated as references/filtering_catalog.md with key filters relocated to Core API Module 2
  • similarity.md (381 lines) → consolidated into references/workflows_similarity.md with core functions in Core API Module 3
  • workflows.md (648 lines) → consolidated into references/workflows_similarity.md with top workflows in Common Workflows
  • pyopenms-mass-spectrometry — full LC-MS/MS proteomics and metabolomics pipelines (feature detection, protein ID)
  • rdkit-cheminformatics — molecular fingerprint generation and chemical structure similarity

References

© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/proteomics-protein-engineering/matchms-spectral-matching of jaechang-hits/SciAgent-Skills.

  • SKILL.md
  • references/filtering_catalog.md
  • references/workflows_similarity.md

Open the folder on GitHubat commit 82c862c

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Matchms Spectral Matching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Matchms Spectral Matching compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Matchms Spectral Matching this skilljaechang-hits/SciAgent-Skills3741 repos~6.3kAutomated safety check: PassApache-2.0
Alphagenome Single Variant Analysisgoogle-deepmind/science-skills3.2k2 repos~3kAutomated safety check: NotesApache-2.0
13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills48k1 repos~3.2kAutomated safety check: PassMIT
Clinvar Databasegoogle-deepmind/science-skills3.2k2 repos~3.9kAutomated safety check: NotesApache-2.0
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT
Dbsnp Databasegoogle-deepmind/science-skills3.2k2 repos~3.4kAutomated safety check: NotesApache-2.0

Similar skills

  • Alphagenome Single Variant Analysis

    google-deepmind/science-skills

    Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.

    3.2k GitHub starsUsed in 2 repos~3k tokens
    Research & ScienceAuto-check: notes
  • 13C Metabolic Flux Analysis

    K-Dense-AI/scientific-agent-skills

    Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.

    48k GitHub starsUsed in 1 repo~3.2k tokens
    Research & ScienceAuto-check passed
  • Clinvar Database

    google-deepmind/science-skills

    A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…

    3.2k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Dbsnp Database

    google-deepmind/science-skills

    A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.

    3.2k GitHub starsUsed in 2 repos~3.4k tokens
    Research & ScienceAuto-check: notes
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed

More from jaechang-hits/SciAgent-Skills

All 169 skills in this repo
  • Neb Irc Activation Energy

    jaechang-hits/SciAgent-Skills

    NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.

    374 GitHub stars~4k tokensUpdated 12 days ago
    Auto-check passed
  • Molecular Visualization 3dmol

    jaechang-hits/SciAgent-Skills

    3Dmol.js WebGL molecular visualization emitted as self-contained HTML.

    374 GitHub stars~3.2k tokensUpdated 12 days ago
    Auto-check passed
  • Cobrapy Metabolic Modeling

    jaechang-hits/SciAgent-Skills

    Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.

    374 GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed
  • Rdkit Chemdraw Cdxml

    jaechang-hits/SciAgent-Skills

    Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.

    374 GitHub stars~6.9k tokensUpdated 12 days ago
    Auto-check passed
  • Pubmed Database

    jaechang-hits/SciAgent-Skills

    Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Sciagent Skill Creator

    jaechang-hits/SciAgent-Skills

    Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.

    374 GitHub stars~2.3k tokensUpdated 12 days ago
    Auto-check passed

Questions about Matchms Spectral Matching

What does Matchms Spectral Matching do?

MS spectral matching and metabolite ID with matchms. An agent skill from jaechang-hits/SciAgent-Skills. Matchms Spectral Matching is an agent skill from jaechang-hits/SciAgent-Skills. MS spectral matching and metabolite ID with matchms.

When should I use Matchms Spectral Matching?

Matchms Spectral Matching fits situations like: tasks that involve Bioinformatics.

How do I install Matchms Spectral Matching in Claude Code?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/matchms-spectral-matching in jaechang-hits/SciAgent-Skills) into .claude/skills/matchms-spectral-matching in your project. Claude Code loads it when a task matches its description.

How do I install Matchms Spectral Matching in Codex?

Run `npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/matchms-spectral-matching in jaechang-hits/SciAgent-Skills) into .agents/skills/matchms-spectral-matching in your project. Codex loads it when a task matches its description.

Can I use Matchms Spectral Matching in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/matchms-spectral-matching, .gemini/skills/matchms-spectral-matching, .github/skills/matchms-spectral-matching and .opencode/skills/matchms-spectral-matching in your project.

What does Matchms Spectral Matching need to run?

Going by SKILL.md and its folder, Matchms Spectral Matching needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.

Does Matchms Spectral Matching access the network?

SKILL.md names 2 domains. As links in the text: matchms.readthedocs.io and github.com. This is read from the text; nothing was executed.

Is Matchms Spectral Matching safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Matchms Spectral Matching use?

Matchms Spectral Matching is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Matchms Spectral Matching use?

About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.3k tokens, read only when the agent opens those files.

What are the alternatives to Matchms Spectral Matching?

Skills that share tags, products or a category with Matchms Spectral Matching: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Matchms Spectral Matching?

jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.

Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.