---
name: chem-msms-predict
description: Predict LC-MS/MS (MS2, tandem mass spectra) from SMILES via ICEBERG, a two-stage deep neural network. Outputs predicted m/z vs intensity spectrum, fragment ion SMILES, and a spectrum plot.
metadata:
  category: [chemistry, drug-discovery]
  venv: [msms]
---

# LC-MS/MS Spectrum Prediction

## Goal

Predict the LC-MS/MS (tandem mass) spectrum of a molecule given its SMILES string using ICEBERG — a two-stage GNN that first generates a fragmentation DAG (fragment ions) and then predicts their intensities. Output is a predicted spectrum (m/z, intensity) with optional fragment SMILES assignments per peak.

## When to Use This Skill

- A SMILES string is known and a predicted LC-MS/MS spectrum (m/z vs intensity) is needed.
- Fragment ion assignments (SMILES per peak) are required.
- No reference spectrum exists, or comparison to a predicted spectrum is desired.
- Companion skill `chem-spectrum-matcher` can compare predicted vs experimental spectra.

## When NOT to Use This Skill

- **Experimental spectrum already available** — use it directly; no prediction needed.
- **Only compound name known** — first resolve to SMILES via `drug-db-pubchem`, then call this skill.
- **GC-MS or other MS types** — ICEBERG is trained on LC-MS/MS only; flag a warning before proceeding.
- **Organometallics or MW > 1000** — predictions may be unreliable or fail due to unsupported element types.

## Prerequisites

Scripts run in the `msms` environment, created on first use by `venv/run`. It
installs ICEBERG 2.1 (`ms-pred`, pinned to a commit) on a CPU torch build, and
is x86_64 Linux only: DGL publishes no aarch64 wheels.

### Download the ICEBERG 2.1 checkpoints

```bash
${CLAUDE_SKILL_DIR}/../../venv/run msms python ${CLAUDE_SKILL_DIR}/scripts/download_weights.py
```

This fetches the public weights trained on MassSpecGym (`msg_all`, ~80 MB) from
the ms-pred authors:

```
downloads/iceberg_msg_all/
├── gen/best.ckpt           # generator (stage 1)
└── inten_contr/best.ckpt   # intensity predictor (stage 2)
```

Weights trained on NIST are available from the ms-pred authors on proof of a
NIST license. ICEBERG 2.0 checkpoints do not load: 2.1 adds instrument types.
**Flag error and stop** if either checkpoint is missing.

## Instructions

### Step 1 — Run inference and generate spectrum

```bash
${CLAUDE_SKILL_DIR}/../../venv/run msms python ${CLAUDE_SKILL_DIR}/scripts/predict_msms.py \
    --smiles "c1ccccc1C(=O)OCCN" \
    --gen_ckpt downloads/iceberg_msg_all/gen/best.ckpt \
    --inten_ckpt downloads/iceberg_msg_all/inten_contr/best.ckpt \
    --collision_energies 20 40 \
    --adduct "[M+H]+" \
    --instrument "Orbitrap" \
    --output_dir results/msms_prediction
```

**Key parameters:**
- `--smiles` — input molecule as SMILES string
- `--gen_ckpt` / `--inten_ckpt` — paths to ICEBERG checkpoints
- `--collision_energies` — one or more collision energies in eV (e.g. `20 40 60`); model was trained on absolute eV values
- `--adduct` — supported adducts: `[M+H]+`, `[M-H]-`, `[M+Na]+`, `[M+NH4]+`, and others from `ms_pred.common.ion2mass`
- `--instrument` — instrument type for intensity prediction (e.g. `"Orbitrap"`, `"QTOF"`)
- `--threshold` — confidence cutoff for DAG fragment generator (default `0.1`; lower = more fragments)
- `--sparse_k` — maximum number of peaks returned (default `100`)
- `--num_workers` — parallel CPU workers (default `0`: serial); inference runs on the CPU

**Outputs written to `--output_dir`:**
| File | Description |
|------|-------------|
| `spectrum.png` | Stem plot of predicted spectrum, one panel per collision energy |
| `fragments.json` | JSON list per CE: `{mz, intensity, fragment_smiles}` sorted by intensity |
| `input_configs.yaml` | All run parameters for reproducibility |

### Step 2 — Inspect fragment assignments (optional)

`fragments.json` maps each predicted peak to the fragment ICEBERG assigns it,
as the Kekulé SMILES of the heavy-atom substructure (hydrogen shifts are not
shown). There is one entry per fragment, so fragments with the same formula
repeat an m/z; `spectrum.png` sums them. For 2-aminoethyl benzoate at 20 eV:

```json
{
  "20": [
    {"mz": 149.0597, "intensity": 1.0, "fragment_smiles": "CCOC(=O)C1=CC=CC=C1"},
    {"mz": 105.0335, "intensity": 0.944, "fragment_smiles": "O=CC1=CC=CC=C1"},
    ...
  ]
}
```

Use this to rationalize which bonds fragment at which energy.

### Step 3 — Compare with experimental spectrum (optional)

If an experimental spectrum is available, use the companion skill:

→ [`chem-spectrum-matcher`](../chem-spectrum-matcher/SKILL.md)

## Examples

### 2-Aminoethyl benzoate (`c1ccccc1C(=O)OCCN`)

```bash
${CLAUDE_SKILL_DIR}/../../venv/run msms python ${CLAUDE_SKILL_DIR}/examples/predict_smiles.py \
    --gen_ckpt downloads/iceberg_msg_all/gen/best.ckpt \
    --inten_ckpt downloads/iceberg_msg_all/inten_contr/best.ckpt \
    --output_dir .agents/test/msms_example
```

Expected output:
- `spectrum.png` — two-panel spectrum (20 eV + 40 eV)
- `fragments.json` — fragment assignments for both energies
- Precursor `[M+H]+` ≈ 166.086 Da

## Constraints

- **Environment**: All scripts run in the `msms` environment (`venv/run msms`, x86_64 Linux only), on the CPU.
- **Checkpoints required**: Script raises `FileNotFoundError` if `--gen_ckpt` or `--inten_ckpt` are missing.
- **Collision energy units**: Use absolute eV values. To convert NCE → eV, set `nce=True` in `iceberg_prediction()` directly.
- **Non-binned output only**: This skill uses `binned_out=False` (high-precision m/z). Binned output disables fragment assignment.
- **Single-compound inference**: Provide one SMILES per call. For batch prediction, loop over SMILES and use separate output dirs.
- **Unsupported elements**: Molecules containing metals, lanthanides, or rare main-group elements may fail or produce low-quality predictions.
- **MW limit**: ICEBERG is unreliable for MW > 1000 Da.

## References

- Alberts, M. et al., "Artificial intelligence for context-aware mass spectrometry", *Nature Methods*, 2025. [DOI:10.1038/s41592-025-02658-z](https://doi.org/10.1038/s41592-025-02658-z)
- ICEBERG source code: [github.com/coleygroup/ms-pred](https://github.com/coleygroup/ms-pred)

---

**Author:** Magdalena Lederbauer
**Contact:** [GitHub @mlederbauer](https://github.com/mlederbauer)
