Alphagenome Single Variant Analysis
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
MS spectral matching and metabolite ID with matchms. An agent skill from jaechang-hits/SciAgent-Skills.
$ npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills matchms-spectral-matching --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/proteomics-protein-engineering/matchms-spectral-matching .claude/skills/matchms-spectral-matching && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "matchms-spectral-matching" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/matchms-spectral-matching into .claude/skills/matchms-spectral-matching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "matchms-spectral-matching", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/matchms-spectral-matchingType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills matchms-spectral-matching --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/proteomics-protein-engineering/matchms-spectral-matching .agents/skills/matchms-spectral-matching && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "matchms-spectral-matching" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/matchms-spectral-matching into .agents/skills/matchms-spectral-matching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "matchms-spectral-matching", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills matchms-spectral-matching --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/proteomics-protein-engineering/matchms-spectral-matching .cursor/skills/matchms-spectral-matching && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "matchms-spectral-matching" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/matchms-spectral-matching into .cursor/skills/matchms-spectral-matching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "matchms-spectral-matching", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/jaechang-hits/SciAgent-Skills.git --path skills/proteomics-protein-engineering/matchms-spectral-matching--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills matchms-spectral-matching --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/proteomics-protein-engineering/matchms-spectral-matching .gemini/skills/matchms-spectral-matching && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "matchms-spectral-matching" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/matchms-spectral-matching into .gemini/skills/matchms-spectral-matching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "matchms-spectral-matching", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install jaechang-hits/SciAgent-Skills matchms-spectral-matchingInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/proteomics-protein-engineering/matchms-spectral-matching .github/skills/matchms-spectral-matching && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "matchms-spectral-matching" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/matchms-spectral-matching into .github/skills/matchms-spectral-matching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "matchms-spectral-matching", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install jaechang-hits/SciAgent-Skills matchms-spectral-matching --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/jaechang-hits/SciAgent-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/proteomics-protein-engineering/matchms-spectral-matching .opencode/skills/matchms-spectral-matching && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "matchms-spectral-matching" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/proteomics-protein-engineering/matchms-spectral-matching into .opencode/skills/matchms-spectral-matching/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "matchms-spectral-matching", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
matchms-spectral-matchingMS spectral matching and metabolite ID with matchms. An agent skill from jaechang-hits/SciAgent-Skills.
Matchms Spectral Matching is an agent skill from jaechang-hits/SciAgent-Skills. MS spectral matching and metabolite ID with matchms. Import spectra (mzML, MGF, MSP, JSON), filter/normalize peaks, score similarity (cosine, modified cosine, fingerprint), build reproducible pipelines, identify unknowns vs spectral libraries. Use pyopenms for full LC-MS/MS proteomics.
Its SKILL.md is about 6.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/filtering_catalog.md` and `references/workflows_similarity.md`).
It sits in Research & Science, covering Bioinformatics. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
3 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvpipFrom the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
matchms.readthedocs.iogithub.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Matchms Spectral Matching loads about 6.3k tokens when it runs, and up to ~13k if it reads all its reference files. Until then it costs about 78 tokens; SKILL.md has 1,123 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from jaechang-hits/SciAgent-Skills at commit 82c862c, republished under its Apache-2.0 licence (© jaechang-hits). 1,123 words, ~6,301 tokens.
.claude/skills/matchms-spectral-matching/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Matchms is a Python library for mass spectrometry data processing focused on spectral similarity calculation and compound identification. It provides multi-format I/O, 50+ spectrum filters for metadata harmonization and peak processing, 8 similarity scoring functions, and a pipeline framework for reproducible analytical workflows.
uv pip install matchms numpy pandas
# For chemical structure processing (SMILES, InChI, fingerprints):
uv pip install matchms[chemistry]Settle these with the user before writing any analysis code.
decisions:
- id: D1
param: referenceLibrary
kind: required
source: user
ask: "Which spectral library should queries be matched against?"
default: null
- id: D2
param: similarityMeasure
kind: required
source: user
ask: "Should matching require the same precursor mass, or allow a mass shift so modified analogues still match?"
default: "cosine on matched peaks, precursor mass fixed"
- id: D3
param: peakMatchTolerance
kind: required
source: data
ask: "How close must two peaks be in m/z to count as the same fragment?"
default: "0.1 Da - tighten toward 0.005 for high-resolution data"
- id: D4
param: spectrumPreprocessing
kind: required
source: user
ask: "Which peaks should be discarded before scoring - low-intensity noise, peaks near the precursor, or spectra with too few peaks to be informative?"
default: "drop peaks under 1% relative intensity, require at least 10 peaks"
- id: D5
param: scoreCutoff
kind: required
source: user
ask: "What similarity score, and how many matched peaks, should a hit need before it is reported as an identification?"
default: null
- id: D6
param: peakWeighting
kind: optional
source: user
ask: "Should the score weight fragment m/z as well as intensity, favouring informative high-mass fragments?"
default: "intensity only"
- id: D7
param: fingerprintSettings
kind: optional_conditional
source: user
ask: "Which molecular fingerprint should structural similarity use, when comparing hits to known structures?"
default: "not computed"D3 and D5 have no safe shared default because they trade off directly: a loose tolerance with a low score cutoff returns an identification for every query spectrum, all of them plausible-looking. Set the tolerance from the instrument's actual accuracy and the cutoff from what the library supports.
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters, normalize_intensities
from matchms.filtering import select_by_relative_intensity, require_minimum_number_of_peaks
from matchms import calculate_scores
from matchms.similarity import CosineGreedy
# Load and process query spectra
queries = list(load_from_mgf("queries.mgf"))
queries = [default_filters(s) for s in queries]
queries = [normalize_intensities(s) for s in queries if s is not None]
queries = [require_minimum_number_of_peaks(s, n_required=5) for s in queries if s is not None]
# Load reference library
refs = list(load_from_mgf("library.mgf"))
refs = [default_filters(s) for s in refs]
refs = [normalize_intensities(s) for s in refs if s is not None]
# Calculate similarity scores
scores = calculate_scores(references=refs, queries=queries,
similarity_function=CosineGreedy(tolerance=0.1))
# Get best matches for first query
best = scores.scores_by_query(queries[0], sort=True)[:5]
for match, score_tuple in best:
print(f"Score: {score_tuple['score']:.3f}, Matches: {score_tuple['matches']}")Import spectra from multiple file formats and export processed data.
from matchms.importing import (load_from_mgf, load_from_mzml, load_from_msp,
load_from_json, load_from_mzxml, load_from_pickle,
load_from_usi)
from matchms.exporting import save_as_mgf, save_as_msp, save_as_json, save_as_pickle
# Import from various formats (returns generators)
spectra_mgf = list(load_from_mgf("library.mgf"))
spectra_mzml = list(load_from_mzml("data.mzML"))
spectra_msp = list(load_from_msp("nist_library.msp"))
spectra_json = list(load_from_json("gnps_spectra.json"))
print(f"Loaded: MGF={len(spectra_mgf)}, mzML={len(spectra_mzml)}")
# Export processed spectra
save_as_mgf(spectra_mgf, "processed.mgf")
save_as_json(spectra_mgf, "processed.json")
save_as_pickle(spectra_mgf, "spectra.pickle") # Fast for intermediate results
# Pickle for large datasets (fastest I/O)
from matchms.importing import load_from_pickle
spectra = list(load_from_pickle("spectra.pickle"))from matchms import Spectrum
import numpy as np
# Create spectrum manually
mz = np.array([100.0, 150.0, 200.0, 250.0, 300.0])
intensities = np.array([0.1, 0.5, 0.9, 0.3, 0.7])
metadata = {
"precursor_mz": 325.5,
"ionmode": "positive",
"compound_name": "Caffeine",
"smiles": "CN1C=NC2=C1C(=O)N(C(=O)N2C)C"
}
spectrum = Spectrum(mz=mz, intensities=intensities, metadata=metadata)
# Access spectrum data
print(f"Peaks: {spectrum.peaks.mz}")
print(f"Precursor: {spectrum.get('precursor_mz')}")
print(f"Name: {spectrum.get('compound_name')}")
# Visualize
spectrum.plot()Apply metadata harmonization and peak processing filters. Matchms provides 50+ filters.
from matchms.filtering import (
default_filters, normalize_intensities,
select_by_relative_intensity, select_by_mz,
require_minimum_number_of_peaks, reduce_to_number_of_peaks,
remove_peaks_around_precursor_mz, add_losses
)
# default_filters applies: metadata cleanup, charge correction, adduct parsing
spectrum = default_filters(spectrum)
# Peak normalization (max intensity → 1.0)
spectrum = normalize_intensities(spectrum)
# Filter peaks by relative intensity (remove noise below 1%)
spectrum = select_by_relative_intensity(spectrum, intensity_from=0.01, intensity_to=1.0)
# Filter peaks by m/z range
spectrum = select_by_mz(spectrum, mz_from=50.0, mz_to=500.0)
# Keep top N peaks only
spectrum = reduce_to_number_of_peaks(spectrum, n_max=50)
# Remove peaks near precursor (common contaminants)
spectrum = remove_peaks_around_precursor_mz(spectrum, mz_tolerance=17.0)
# Require minimum peaks for matching
spectrum = require_minimum_number_of_peaks(spectrum, n_required=5)
# Add neutral losses (useful for NeutralLossesCosine)
spectrum = add_losses(spectrum)
if spectrum is not None:
print(f"After filtering: {len(spectrum.peaks.mz)} peaks")# Chemical annotation filters (require matchms[chemistry])
from matchms.filtering import (
derive_inchi_from_smiles, derive_inchikey_from_inchi,
derive_smiles_from_inchi, add_fingerprint,
repair_inchi_inchikey_smiles, require_valid_annotation
)
# Derive chemical identifiers from SMILES
spectrum = derive_inchi_from_smiles(spectrum)
spectrum = derive_inchikey_from_inchi(spectrum)
# Add molecular fingerprint for structural similarity
spectrum = add_fingerprint(spectrum, fingerprint_type="morgan", nbits=2048)
# Validate annotations
spectrum = require_valid_annotation(spectrum)
if spectrum is not None:
print(f"InChIKey: {spectrum.get('inchikey')}")Compare spectra using multiple similarity metrics.
from matchms import calculate_scores
from matchms.similarity import (
CosineGreedy, CosineHungarian, ModifiedCosine,
NeutralLossesCosine, FingerprintSimilarity,
MetadataMatch, PrecursorMzMatch
)
# CosineGreedy — fast peak matching (greedy algorithm)
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=CosineGreedy(tolerance=0.1))
# ModifiedCosine — accounts for precursor mass differences (best for analog search)
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=ModifiedCosine(tolerance=0.1))
# CosineHungarian — optimal peak matching (slower but more accurate)
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=CosineHungarian(tolerance=0.1))
# NeutralLossesCosine — similarity based on neutral loss patterns
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=NeutralLossesCosine(tolerance=0.1))
# Access results
for i, query in enumerate(unknowns[:3]):
best_matches = scores.scores_by_query(query, sort=True)[:3]
print(f"\nQuery {i}: precursor_mz={query.get('precursor_mz')}")
for ref, score_tuple in best_matches:
print(f" {ref.get('compound_name', 'Unknown')}: "
f"score={score_tuple['score']:.3f}, matches={score_tuple['matches']}")# FingerprintSimilarity — structural similarity (requires fingerprints)
from matchms.similarity import FingerprintSimilarity
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=FingerprintSimilarity(
similarity_measure="jaccard"))
# PrecursorMzMatch — fast mass-based pre-filtering
from matchms.similarity import PrecursorMzMatch
scores = calculate_scores(references=library, queries=unknowns,
similarity_function=PrecursorMzMatch(tolerance=0.1))
# Multi-metric scoring: combine peak + structural similarity
cosine_scores = calculate_scores(references=library, queries=unknowns,
similarity_function=CosineGreedy(tolerance=0.1))
fp_scores = calculate_scores(references=library, queries=unknowns,
similarity_function=FingerprintSimilarity())Build reusable, reproducible multi-step processing workflows.
from matchms import SpectrumProcessor
from matchms.filtering import (
default_filters, normalize_intensities,
select_by_relative_intensity, require_minimum_number_of_peaks,
remove_peaks_around_precursor_mz, add_losses
)
# Define reusable pipeline
pipeline = SpectrumProcessor([
default_filters,
normalize_intensities,
lambda s: select_by_relative_intensity(s, intensity_from=0.01),
lambda s: remove_peaks_around_precursor_mz(s, mz_tolerance=17.0),
lambda s: require_minimum_number_of_peaks(s, n_required=5),
add_losses
])
# Apply to all spectra (filters returning None remove the spectrum)
processed = [pipeline(s) for s in raw_spectra]
processed = [s for s in processed if s is not None]
print(f"Processed: {len(processed)}/{len(raw_spectra)} spectra retained")| Function | Speed | Accuracy | Best For |
|---|---|---|---|
CosineGreedy | Fast | Good | General library matching |
CosineHungarian | Slow | Best | Small comparisons, validation |
ModifiedCosine | Fast | Good | Analog search (different precursors) |
NeutralLossesCosine | Medium | Good | Structural class identification |
FingerprintSimilarity | Fast | Moderate | Structure-based pre-filtering |
PrecursorMzMatch | Fastest | N/A | Mass-based pre-filtering |
| Category | Examples | Purpose |
|---|---|---|
| Metadata cleanup | default_filters, clean_compound_name, clean_adduct | Standardize metadata fields |
| Chemical derivation | derive_inchi_from_smiles, add_fingerprint | Compute chemical identifiers |
| Mass/charge | add_precursor_mz, correct_charge, add_parent_mass | Fix and validate mass info |
| Peak normalization | normalize_intensities, select_by_relative_intensity | Scale and filter peaks |
| Peak reduction | reduce_to_number_of_peaks, remove_peaks_around_precursor_mz | Remove noise/artifacts |
| Quality control | require_minimum_number_of_peaks, require_precursor_mz | Enforce minimum quality |
| Neutral losses | add_losses | Compute precursor-fragment losses |
All similarity functions return (score, matches):
score: float 0.0–1.0 (cosine similarity value)matches: int (number of matched peaks between query and reference)Higher scores and more matched peaks indicate better matches. Typical thresholds: score > 0.7 and matches > 6 for confident identifications.
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters, normalize_intensities
from matchms.filtering import select_by_relative_intensity, require_minimum_number_of_peaks
from matchms import calculate_scores
from matchms.similarity import ModifiedCosine
import pandas as pd
# Load and process both queries and library identically
def process_spectra(spectra):
processed = []
for s in spectra:
s = default_filters(s)
if s is None: continue
s = normalize_intensities(s)
s = select_by_relative_intensity(s, intensity_from=0.01)
s = require_minimum_number_of_peaks(s, n_required=5)
if s is not None:
processed.append(s)
return processed
queries = process_spectra(load_from_mgf("unknowns.mgf"))
library = process_spectra(load_from_mgf("reference_library.mgf"))
print(f"Queries: {len(queries)}, Library: {len(library)}")
# Score all query-reference pairs
scores = calculate_scores(references=library, queries=queries,
similarity_function=ModifiedCosine(tolerance=0.1))
# Extract best matches
results = []
for query in queries:
best = scores.scores_by_query(query, sort=True)[:1]
if best:
ref, score_tuple = best[0]
results.append({
"query_precursor_mz": query.get("precursor_mz"),
"match_name": ref.get("compound_name", "Unknown"),
"match_smiles": ref.get("smiles", ""),
"score": score_tuple["score"],
"matched_peaks": score_tuple["matches"]
})
df = pd.DataFrame(results)
confident = df[df["score"] > 0.7]
print(f"Confident matches (score>0.7): {len(confident)}/{len(df)}")
df.to_csv("identification_results.csv", index=False)from matchms.importing import load_from_msp
from matchms.exporting import save_as_mgf
from matchms import SpectrumProcessor
from matchms.filtering import (
default_filters, normalize_intensities,
select_by_relative_intensity, require_minimum_number_of_peaks,
require_precursor_mz, add_parent_mass
)
# Define QC pipeline
qc_pipeline = SpectrumProcessor([
default_filters,
require_precursor_mz,
add_parent_mass,
normalize_intensities,
lambda s: select_by_relative_intensity(s, intensity_from=0.001),
lambda s: require_minimum_number_of_peaks(s, n_required=3)
])
# Process and filter
raw = list(load_from_msp("raw_library.msp"))
cleaned = [qc_pipeline(s) for s in raw]
cleaned = [s for s in cleaned if s is not None]
print(f"Input: {len(raw)}, Output: {len(cleaned)} ({len(cleaned)/len(raw)*100:.1f}% retained)")
save_as_mgf(cleaned, "cleaned_library.mgf")load_from_mzml)default_filters for metadata harmonization (Core API Module 2)save_as_mgf)| Parameter | Function/Module | Default | Range/Options | Effect |
|---|---|---|---|---|
tolerance | CosineGreedy/ModifiedCosine | 0.1 | 0.005–0.5 Da | m/z tolerance for peak matching |
mz_power | CosineGreedy | 0.0 | 0.0–2.0 | Weight of m/z in scoring (0=ignore) |
intensity_power | CosineGreedy | 1.0 | 0.0–2.0 | Weight of intensity in scoring |
intensity_from | select_by_relative_intensity | 0.0 | 0.0–1.0 | Minimum relative intensity to keep |
n_required | require_minimum_number_of_peaks | 10 | 1–100 | Minimum peaks to retain spectrum |
n_max | reduce_to_number_of_peaks | 100 | 10–500 | Maximum peaks to retain |
mz_tolerance | remove_peaks_around_precursor_mz | 17.0 | 0.5–50 Da | Window around precursor to remove |
fingerprint_type | add_fingerprint | "daylight" | "daylight"/"morgan"/"maccs" | Molecular fingerprint type |
nbits | add_fingerprint | 2048 | 256–4096 | Fingerprint bit vector length |
PrecursorMzMatch first to reduce the comparison space, then score with CosineGreedyNone when a spectrum fails quality requirements. Always filter: [s for s in processed if s is not None]from matchms import calculate_scores
from matchms.similarity import PrecursorMzMatch, CosineGreedy
# Step 1: Fast mass filter
mass_scores = calculate_scores(references=library, queries=unknowns,
similarity_function=PrecursorMzMatch(tolerance=0.5))
# Step 2: Detailed scoring only for mass-matched pairs
cosine = CosineGreedy(tolerance=0.1)
for query in unknowns:
candidates = mass_scores.scores_by_query(query, sort=True)
mass_matched = [ref for ref, score in candidates if score["score"] > 0]
if mass_matched:
detailed = calculate_scores(references=mass_matched, queries=[query],
similarity_function=cosine)
best = detailed.scores_by_query(query, sort=True)[:3]
for ref, s in best:
print(f"{ref.get('compound_name')}: {s['score']:.3f}")from matchms.importing import load_from_mgf
from matchms.filtering import default_filters, normalize_intensities
spectra = list(load_from_mgf("mixed_library.mgf"))
spectra = [default_filters(s) for s in spectra]
spectra = [s for s in spectra if s is not None]
# Separate by ion mode
positive = [s for s in spectra if s.get("ionmode") == "positive"]
negative = [s for s in spectra if s.get("ionmode") == "negative"]
print(f"Positive: {len(positive)}, Negative: {len(negative)}")
# Process each mode with mode-specific filtering
positive = [normalize_intensities(s) for s in positive]
negative = [normalize_intensities(s) for s in negative]import pandas as pd
from matchms.importing import load_from_mgf
from matchms.filtering import default_filters
spectra = [default_filters(s) for s in load_from_mgf("library.mgf")]
spectra = [s for s in spectra if s is not None]
# Extract metadata summary
rows = []
for s in spectra:
rows.append({
"compound_name": s.get("compound_name", ""),
"precursor_mz": s.get("precursor_mz"),
"ionmode": s.get("ionmode", ""),
"smiles": s.get("smiles", ""),
"inchikey": s.get("inchikey", ""),
"num_peaks": len(s.peaks.mz)
})
df = pd.DataFrame(rows)
print(f"Library: {len(df)} spectra")
print(f"Named: {(df.compound_name != '').sum()}")
print(f"With SMILES: {(df.smiles != '').sum()}")
print(f"Ion modes: {df.ionmode.value_counts().to_dict()}")| Problem | Cause | Solution |
|---|---|---|
| All scores are 0.0 | No matching peaks within tolerance | Increase tolerance (try 0.2–0.5 Da); verify both spectra have peaks |
| Low scores despite same compound | Different fragmentation conditions | Use ModifiedCosine instead of CosineGreedy; check ion mode consistency |
| Many spectra filtered to None | Too strict quality filters | Lower n_required in require_minimum_number_of_peaks; relax intensity thresholds |
KeyError on metadata field | Field name not harmonized | Apply default_filters first to harmonize metadata keys |
| Memory error with large library | All-vs-all comparison | Pre-filter by precursor mass (PrecursorMzMatch) before detailed scoring |
add_fingerprint fails | RDKit not installed | Install chemistry extras: pip install matchms[chemistry] |
| Import returns empty list | Wrong file format or path | Verify format matches loader (MGF for .mgf, MSP for .msp); check file is not empty |
| Inconsistent scores between runs | Different processing pipelines | Use SpectrumProcessor to ensure identical processing for queries and references |
references/filtering_catalog.md — Complete catalog of 50+ matchms filter functions organized by category (metadata processing, chemical structure, mass/charge, peak processing, quality control).
references/workflows_similarity.md — Extended workflows and detailed similarity function documentation consolidated from two original references.
spectrum.plot() in Core API Module 1Disposition of original reference files:
© jaechang-hits, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/proteomics-protein-engineering/matchms-spectral-matching of jaechang-hits/SciAgent-Skills.
Open the folder on GitHubat commit 82c862c
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Matchms Spectral Matching next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Matchms Spectral Matching this skilljaechang-hits/SciAgent-Skills | 374 | 1 repos | ~6.3k | Automated safety check: Pass | Apache-2.0 | |
| Alphagenome Single Variant Analysisgoogle-deepmind/science-skills | 3.2k | 2 repos | ~3k | Automated safety check: Notes | Apache-2.0 | |
| 13C Metabolic Flux AnalysisK-Dense-AI/scientific-agent-skills | 48k | 1 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Clinvar Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.9k | Automated safety check: Notes | Apache-2.0 | |
| Metabolic Study Planneraiming-lab/AutoResearchClaw | 15k | — | ~1.9k | Automated safety check: Pass | MIT | |
| Dbsnp Databasegoogle-deepmind/science-skills | 3.2k | 2 repos | ~3.4k | Automated safety check: Notes | Apache-2.0 |
google-deepmind/science-skills
Analyzes genetic variant effects on gene expression (RNA-seq), chromatin accessibility (DNASE), histone marks (ChIP), and transcription factors using the AlphaGenome API.
K-Dense-AI/scientific-agent-skills
Estimates reaction fluxes inside cells from steady-state carbon-13 labeling data with a bundled mfapy-based solver, and reports which fluxes the data pin down.
google-deepmind/science-skills
A skill your agent uses when needing clinical significance, pathogenicity classifications (e.g., Pathogenic, Benign, VUS), clinical evidence rationales, or finding "hard positive" benchmark controls…
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
google-deepmind/science-skills
A skill your agent uses when you want to look up, map, and search for short genetic variants (SNPs, indels) in NCBI's dbSNP database.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
jaechang-hits/SciAgent-Skills
NEB-IRC activation energy pipeline for reaction barriers using GFN2-xTB and pysisyphus.
jaechang-hits/SciAgent-Skills
3Dmol.js WebGL molecular visualization emitted as self-contained HTML.
jaechang-hits/SciAgent-Skills
Constraint-based (COBRA) analysis of genome-scale metabolic models: FBA, FVA, knockouts, flux sampling, production envelopes, gapfilling, media optimization.
jaechang-hits/SciAgent-Skills
Read, write, and edit ChemDraw CDX/CDXML files with RDKit's rdkit.Chem.rdChemDraw plus direct XML editing, always paired with a rendered PNG.
jaechang-hits/SciAgent-Skills
Programmatic PubMed access via NCBI E-utilities REST API. An agent skill from jaechang-hits/SciAgent-Skills.
jaechang-hits/SciAgent-Skills
Scaffold a new SciAgent-Skills entry. An agent skill from jaechang-hits/SciAgent-Skills.
Categories
MS spectral matching and metabolite ID with matchms. An agent skill from jaechang-hits/SciAgent-Skills. Matchms Spectral Matching is an agent skill from jaechang-hits/SciAgent-Skills. MS spectral matching and metabolite ID with matchms.
Matchms Spectral Matching fits situations like: tasks that involve Bioinformatics.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a claude-code`. Or copy the skill folder (skills/proteomics-protein-engineering/matchms-spectral-matching in jaechang-hits/SciAgent-Skills) into .claude/skills/matchms-spectral-matching in your project. Claude Code loads it when a task matches its description.
Run `npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a codex`. Or copy the skill folder (skills/proteomics-protein-engineering/matchms-spectral-matching in jaechang-hits/SciAgent-Skills) into .agents/skills/matchms-spectral-matching in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill matchms-spectral-matching -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/matchms-spectral-matching, .gemini/skills/matchms-spectral-matching, .github/skills/matchms-spectral-matching and .opencode/skills/matchms-spectral-matching in your project.
Going by SKILL.md and its folder, Matchms Spectral Matching needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: matchms.readthedocs.io and github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Matchms Spectral Matching is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 6.3k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Matchms Spectral Matching: Alphagenome Single Variant Analysis (google-deepmind/science-skills, 3.2k stars), 13C Metabolic Flux Analysis (K-Dense-AI/scientific-agent-skills, 48k stars), Clinvar Database (google-deepmind/science-skills, 3.2k stars) and Metabolic Study Planner (aiming-lab/AutoResearchClaw, 15k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 374 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.