Agent skill

Chem Spectrum Matcher

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Match an experimental spectrum (1H NMR, 13C NMR, IR) against predicted or database reference spectra for candidate ranking and structure confirmation.

MITAuto-check passed

Install Chem Spectrum Matcher

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill chem-spectrum-matcher -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills chem-spectrum-matcher --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/chem-spectrum-matcher .claude/skills/chem-spectrum-matcher && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
chem-spectrum-matcher
GitHub stars
176
Token cost
~2.5k tokens
SKILL.md length
838 words
Files
4 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Match an experimental spectrum (1H NMR, 13C NMR, IR) against predicted or database reference spectra for candidate ranking and structure confirmation.

  • Works in 3 steps: Prepare Candidate Reference Spectra → Match and Rank → Interpret Results
  • SKILL.md covers Goal, When to Use This Skill, When NOT to Use This Skill and Architecture, plus 6 more sections
  • Runs Python scripts from its folder

What it does

Chem Spectrum Matcher is an agent skill from learningmatter-mit/AtomisticSkills. Match an experimental spectrum (1H NMR, 13C NMR, IR) against predicted or database reference spectra for candidate ranking and structure confirmation. Supports local catalog lookup, public database fallback, and pluggable similarity metrics.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts (for example `examples/nmr-ethanol/README.md`, `scripts/match_spectrum.py` and `scripts/register_spectrum.py`).

The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

Example prompts

  • “/chem-spectrum-matcher”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Prepare Candidate Reference Spectra
  2. Match and Rank
  3. Interpret Results

What it can do on your machine

Read from SKILL.md and the folder at commit 6257444. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • webbook.nist.gov
    • rdkit.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Chem Spectrum Matcher loads about 2.5k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 838 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 6257444, republished under its MIT licence (© learningmatter-mit). 838 words, ~2,515 tokens.

Download SKILL.mdSave it as .claude/skills/chem-spectrum-matcher/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
chem-spectrum-matcher
description
Match an experimental spectrum (1H NMR, 13C NMR, IR) against predicted or database reference spectra for candidate ranking and structure confirmation. Supports local catalog lookup, public database fallback, and pluggable similarity metrics.
metadata.category
chemistry, drug-discovery
metadata.venv
cpu

Spectrum Matcher

Goal

To retrieve or generate reference spectra for a set of candidate molecules and rank them by similarity to an experimental query spectrum. The skill abstracts a common three-component pattern:

  1. Prediction — generate reference spectra from structure (SMILES) via empirical predictors or QM.
  2. Reference DB — cache computed spectra locally; fall back to public databases for known compounds.
  3. Similarity metric — score each candidate against the query and rank.

This pattern applies to any spectral modality: 1H NMR, 13C NMR, IR, mass spectrometry, UV-Vis, Raman. The concrete implementation here covers 1H NMR and IR, with the NMR path fully implemented and IR sketched for extension.

When to Use This Skill

  • Confirm a proposed structure against an experimental spectrum.
  • Screen a shortlist of candidates and rank by spectral similarity.
  • Avoid re-running expensive predictions by retrieving cached spectra from the local catalog.
  • Fetch experimental reference spectra from public databases (NMRShiftDB2, NIST WebBook) before committing to a prediction run.

When NOT to Use This Skill

  • Unknown structure elucidation from scratch — this skill requires a candidate list. For open-ended structure identification, use general-query-literature-database first.
  • Mixture deconvolution — use chem-nmr-analysis (Wasserstein deconvolution) for quantifying component ratios.
  • 13C, 19F, 31P NMR prediction — chem-nmr-predict (SPINUS) covers 1H only. Extension needed.
  • Mass spectrometry — not yet implemented. Scaffold is in place; add a predictor and similarity metric.

Architecture

SMILES  ──► [Predictor]  ──► predicted spectrum (.xy / .jdx)
                                        │
                                        ▼
                              [Local Catalog]  ◄──  register_spectrum.py
                                        │
        [Public DB fallback]  ──────────┤  (NMRShiftDB2, NIST WebBook)
                                        │
                                        ▼
            Experimental query  ──► [match_spectrum.py]  ──► ranked candidates
Modality–Predictor–Metric table
ModalityPredictor skillPublic DBSimilarity metric
1H NMRchem-nmr-predict (SPINUS + nmrsim)NMRShiftDB2L2 / Wasserstein
IRchem-db-spectra (NIST) or ORCA DFTNIST WebBookCosine
13C NMR(not yet implemented)NMRShiftDB2L2
Mass spec(not yet implemented)NIST WebBookDot product

Workflow

Step 1 — Prepare Candidate Reference Spectra
Option A: Retrieve from catalog or public DB (fast, no prediction)
bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/match_spectrum.py \
  --query experimental_spectrum.xy \
  --smiles "CCO" \
  --names "ethanol" \
  --modality nmr_1h \
  --catalog_dir research/spectrum_catalog/ \
  --output_dir <research_dir>/spectrum_match/ \
  --fallback_public_db
Option B: Predict first, then match

Run the appropriate predictor for the modality, then register outputs into the catalog, then match.

1H NMR:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/../chem-nmr-predict/scripts/predict_nmr.py \
  --smiles "CCO" \
  --names "ethanol" \
  --field_mhz 400 \
  --output_dir <research_dir>/nmr_predictions/

${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/register_spectrum.py \
  --source_dir <research_dir>/nmr_predictions/ \
  --modality nmr_1h \
  --catalog_dir research/spectrum_catalog/

IR (from NIST WebBook):

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/../chem-db-spectra/scripts/query_spectra.py \
  C10H18O <research_dir>/ir_references/ --type IR

${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/register_spectrum.py \
  --source_dir <research_dir>/ir_references/ \
  --modality ir \
  --catalog_dir research/spectrum_catalog/

IR (QM-backed, high accuracy):

bash
# Run ORCA frequency calculation → extract IR spectrum → register
# ORCA setup (ORCA_BINARY_PATH, x86_64 only): see the chem-dft-orca-singlepoint skill.
# After ORCA run, convert output with ${CLAUDE_SKILL_DIR}/../../src/utils/dft/orca_utils.py
# then call register_spectrum.py --modality ir
Step 2 — Match and Rank

match_spectrum.py retrieves reference spectra (catalog → public DB fallback) and computes similarity scores between the query and each candidate.

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/match_spectrum.py \
  --query experimental_spectrum.xy \
  --smiles "CCO" \
  --names "ethanol" \
  --modality nmr_1h \
  --metric l2 \
  --catalog_dir research/spectrum_catalog/ \
  --output_dir <research_dir>/spectrum_match/ \
  --plot

Arguments:

  • --query: experimental spectrum file (two-column, ppm/wavenumber vs intensity; .xy, .csv, .jdx).
  • --smiles: candidate SMILES strings.
  • --names: human-readable labels matching SMILES order.
  • --modality: nmr_1h, nmr_13c, ir. Controls which catalog partition and public DB to query.
  • --metric: similarity metric — l2 (default), cosine, wasserstein. Choose based on modality (see table above).
  • --catalog_dir: local spectrum catalog directory.
  • --fallback_public_db: query NMRShiftDB2 or NIST WebBook for any candidate not in catalog.
  • --field_mhz: spectrometer field (NMR only, default 400). Must match experimental spectrum.
  • --plot: emit overlay plot (match_plot.png) with query and top-3 candidates.
  • --output_dir: directory for outputs.

Outputs:

  • match_results.json — ranked candidates with similarity scores, source (catalog/public_db/predicted), and spectrum paths.
  • match_plot.png — overlay of query vs ranked references (if --plot).
  • input_configs.yaml — all parameters for reproducibility.
Show full SKILL.md (387 more words)Show less
Step 3 — Interpret Results

Read match_results.json. Candidates are ranked by descending similarity score (1.0 = perfect match, 0.0 = no overlap).

If/Then rules:

Score (L2 / cosine)InterpretationAgent action
> 0.90Strong matchReport top candidate with confidence.
0.70–0.90Plausible matchReport with caveat; check overlay plot for unmatched peaks.
< 0.70Poor matchLikely wrong candidate or missing structure. Expand candidate list or re-examine experimental spectrum.

After reading scores, the agent must:

  1. Inspect match_plot.png — verify visual agreement, check for systematic shifts.
  2. Cross-check with signal table (NMR) — compare predicted multiplicity/coupling to experimental assignments.
  3. If top candidate has score < 0.70 — consider running QM-backed prediction (ORCA IR, or higher-level NMR) rather than empirical.

Failure Modes

FailureSymptomAgent action
Catalog miss, no public DB hitmatch_results.json candidate marked missedRun appropriate predictor then register_spectrum.py.
Ppm/wavenumber axis mismatchSimilarity scores all near 0Query and reference use different x-axis. Check --field_mhz or unit convention.
SMILES canonicalization failsRDKit errorSMILES invalid. Verify with RDKit before retry.
NMRShiftDB2 / NIST timeoutHTTP error during public DB queryRetry once; if persistent, disable --fallback_public_db and predict locally.
ORCA IR prediction unavailableORCA_BINARY_PATH not setSet it as described in the chem-dft-orca-singlepoint skill, or fall back to NIST WebBook IR.

Relationship to Other Skills

drug-db-pubchem          → resolve compound name to SMILES
chem-nmr-predict         → 1H NMR prediction (SPINUS + nmrsim)
chem-db-spectra          → experimental IR/MS from NIST WebBook
chem-nmr-analysis        → mixture deconvolution (Wasserstein)
chem-spectrum-matcher    → this skill: catalog + retrieval + similarity ranking

Environment

Primary (NMR matching):

bash

Required packages: numpy, scipy, rdkit, requests, matplotlib.

IR prediction via QM (optional):

bash

Requires the ORCA_BINARY_PATH environment variable (see the chem-dft-orca-singlepoint skill); x86_64 only, since SCINE has no aarch64 wheels.


Constraints

  • Catalog format: catalog.json keyed by (canonical_smiles, modality). Do not edit manually.
  • Spectrum format: .xy files are two-column tab-separated (x-axis descending, intensity). .jdx files are parsed via the jcamp package.
  • Modality isolation: NMR and IR spectra are stored in separate catalog partitions; cross-modality lookup is not supported.
  • Field strength: NMR catalog entries store the field (MHz) used during prediction. Retrieval warns if query field differs.
  • Stereochemistry: Diastereomers stored separately by canonical SMILES. Enantiomers produce identical achiral NMR/IR spectra but are stored separately for traceability.

References

  • Steinbeck, C. et al., "NMRShiftDB — constructing a free chemical information system with open-source components", J. Chem. Inf. Comput. Sci., 2003. DOI
  • Linstrom, P.J. and Mallard, W.G., Eds., NIST Chemistry WebBook, NIST Standard Reference Database 69. URL
  • Banfi, D. & Patiny, L., "www.nmrdb.org: Resurrecting and processing NMR spectra on-line", Chimia, 2008.
  • Landrum, G. et al., RDKit: Open-Source Cheminformatics. URL

Author: Magdalena Lederbauer Contact: GitHub @mlederbauer

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/chem-spectrum-matcher of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/nmr-ethanol/README.md
  • scripts/match_spectrum.py
  • scripts/register_spectrum.py

Open the folder on GitHubat commit 6257444

Compare with similar skills

Chem Spectrum Matcher next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Chem Spectrum Matcher compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Chem Spectrum Matcher this skilllearningmatter-mit/AtomisticSkills176—~2.5kAutomated safety check: PassMIT
Experimental Designaiming-lab/AutoResearchClaw15k—~286Automated safety check: PassMIT
Matchersonsi/gomega2.4k—~2.1kAutomated safety check: PassMIT
Find Matching Tenderssickn33/agentic-awesome-skills47k1 repos~1kAutomated safety check: PassApache-2.0
Custom Matchersonsi/gomega2.4k—~2.4kAutomated safety check: PassMIT
Composing Matchersonsi/gomega2.4k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Experimental Design

    aiming-lab/AutoResearchClaw

    Best practices for designing reproducible ML experiments. An agent skill from aiming-lab/AutoResearchClaw.

    15k GitHub stars~286 tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Matchers

    onsi/gomega

    The complete catalog of Gomega's built-in matchers, grouped by category — equivalence (Equal/BeEquivalentTo/BeComparableTo/BeIdenticalTo/BeAssignableToTypeOf), presence (BeNil/BeZero/BeEmpty)…

    2.4k GitHub stars~2.1k tokensUpdated 14 days ago
    DevelopmentAuto-check passed
  • Find Matching Tenders

    sickn33/agentic-awesome-skills

    Find open AU/NZ government tenders matching what a company does, ranked by fit with why and gap analysis.

    47k GitHub starsUsed in 1 repo~1k tokens
    Sales & SupportAuto-check passed
  • Custom Matchers

    onsi/gomega

    Writing your own Gomega matchers — the GomegaMatcher interface (Match/FailureMessage/NegatedFailureMessage), gcustom.MakeMatcher with message templates and template data, the format package helpers…

    2.4k GitHub stars~2.4k tokensUpdated 14 days ago
    DevelopmentAuto-check passed
  • Build compound Gomega assertions by combining matchers — And/SatisfyAll (all pass), Or/SatisfyAny (any pass), Not (negate), WithTransform to map the actual before matching, Satisfy for an ad-hoc…

    2.4k GitHub stars~1.5k tokensUpdated 14 days ago
    DevelopmentAuto-check passed
  • Experimentation

    cbrock84/headcount

    Designs, runs, and reads A/B tests and growth experiments — hypothesis, sample size, duration, and honest interpretation.

    2k GitHub stars~971 tokensUpdated 20 days ago
    Marketing & SEOAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    176 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    176 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    176 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    176 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    176 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    176 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about Chem Spectrum Matcher

What does Chem Spectrum Matcher do?

Match an experimental spectrum (1H NMR, 13C NMR, IR) against predicted or database reference spectra for candidate ranking and structure confirmation. Chem Spectrum Matcher is an agent skill from learningmatter-mit/AtomisticSkills. Match an experimental spectrum (1H NMR, 13C NMR, IR) against predicted or database reference spectra for candidate ranking and structure confirmation.

How do I install Chem Spectrum Matcher in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill chem-spectrum-matcher -a claude-code`. Or copy the skill folder (skills/chem-spectrum-matcher in learningmatter-mit/AtomisticSkills) into .claude/skills/chem-spectrum-matcher in your project. Claude Code loads it when a task matches its description.

How do I install Chem Spectrum Matcher in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill chem-spectrum-matcher -a codex`. Or copy the skill folder (skills/chem-spectrum-matcher in learningmatter-mit/AtomisticSkills) into .agents/skills/chem-spectrum-matcher in your project. Codex loads it when a task matches its description.

Can I use Chem Spectrum Matcher in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill chem-spectrum-matcher -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/chem-spectrum-matcher, .gemini/skills/chem-spectrum-matcher, .github/skills/chem-spectrum-matcher and .opencode/skills/chem-spectrum-matcher in your project.

What does Chem Spectrum Matcher need to run?

Going by SKILL.md and its folder, Chem Spectrum Matcher needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Chem Spectrum Matcher access the network?

SKILL.md names 4 domains. As links in the text: doi.org, webbook.nist.gov, rdkit.org and github.com. This is read from the text; nothing was executed.

Is Chem Spectrum Matcher safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Chem Spectrum Matcher use?

Chem Spectrum Matcher is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Chem Spectrum Matcher use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Chem Spectrum Matcher?

Skills that share tags, products or a category with Chem Spectrum Matcher: Experimental Design (aiming-lab/AutoResearchClaw, 15k stars), Matchers (onsi/gomega, 2.4k stars), Find Matching Tenders (sickn33/agentic-awesome-skills, 47k stars) and Custom Matchers (onsi/gomega, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Chem Spectrum Matcher?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 176 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 7, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.