---
name: metabolomics-annotation
description: Load when matching LC-MS m/z features to an explicit local metabolite reference within a ppm
  tolerance; bundled HMDB entries are for explicit demonstrations only. Skip pathway ORA (use
  metabolomics-pathway-enrichment) and spectral matching or online searches (use external SIRIUS / GNPS).
trigger: metabolite annotation, SIRIUS, GNPS, MetFrag, spectral matching, metabolite ID, ppm tolerance
tags:
- metabolomics
- annotation
- hmdb
- demo
- mz-match
---

# metabolomics-annotation

## When to use

Match m/z to adduct masses in an explicit reference or the bundled 15-metabolite demo. Use external SIRIUS/GNPS for spectral or database-scale identification.

## Use from a step

```python
import pandas as pd
from skills._sdk.notebook import load_skill, read_input, write_output
library = load_skill("metabolomics-annotation")
data = read_input('features.csv', reader=pd.read_csv)
reference = read_input('reference.csv', reader=pd.read_csv)
result = library.annotate(data, reference=reference)
write_output(result, 'tables/result.csv')
```

[examples/example_step.py](examples/example_step.py) runs a seeded synthetic
example through the step runner and writes a table and Figure. Computations
return new DataFrames, leave the input unchanged and expose diagnostics through
`run_info(result)`. Plotting functions write no files.

## API

<!-- api:begin generated from _api.py; regenerate with run.py api <skill dir> --write -->

### `annotate(data, *, database='hmdb', ppm=10.0, adducts=None, reference=None)`

Match every observed m/z to all reference adducts within tolerance.

:param data: Feature DataFrame with a numeric mz column.
:param database: CLI default hmdb; label for the supplied reference, not a database fetch.
:param ppm: CLI default 10; nonnegative mass error tolerance in parts per million.
:param adducts: CLI default None resolves to [M+H]+ and [M-H]-.
:param reference: Required DataFrame with name, neutral_mass, database_id and formula; demo_reference() is for demonstrations only.
:returns: A new annotations DataFrame; attrs['run_info'] names the reference scope.
:raises ValueError: Reference, observed masses, tolerance or adducts are invalid.

### `demo_reference()`

Return the 15 bundled metabolites for explicit demonstrations.

:returns: An independent reference DataFrame marked as demo in its attrs.

### `run_info(data, *, keep=True)`

Read diagnostics attached to a returned table.

:param data: DataFrame returned by this library.
:param keep: Default True; use False in the CLI to remove diagnostics.
:returns: An independent dictionary describing the run.
:raises ValueError: The table carries no run_info.

### `mass_error_figure(data)`

Plot the ppm error of matched metabolite candidates.

:param data: Annotation table returned by annotate.
:returns: A matplotlib Figure.
:raises KeyError: ppm_error is absent.

<!-- api:end -->

## Methods and parameters

Pass `reference=` with name, neutral_mass, database_id and formula for local reference mass matching. The CLI requires `--reference-file reference.csv` for real input. No network lookup runs; `database` labels the supplied reference. `demo_reference()` explicitly selects 15 illustrative HMDB entries, also used by `--demo`.

## Gotchas

- `annotate` rejects missing reference data; demo references cannot be relabelled as another database. Each query can have multiple candidate rows in `tables/annotations.csv`; Unknown rows retain unmatched queries. Confidence labels describe ppm bins, not identification probability.
- `result.json` preserves the reference scope in `data.run_info.reference_scope`; demo matches are not biological identification evidence.

## Inputs and outputs

CSV input; `tables/annotations.csv`, `report.md` and `result.json`. The CLI writes `reproducibility/commands.sh`.
The function library returns objects; the CLI and step own file writes.

## CLI

```bash
python skills/metabolomics/metabolomics-annotation/metabolomics_annotation.py --demo --output /tmp/metabolomics_annotation
python skills/metabolomics/metabolomics-annotation/metabolomics_annotation.py --input features.csv --reference-file reference.csv --output /tmp/metabolomics_annotation_real
```

## See also

- [Parameters](references/parameters.md)
- [Methodology](references/methodology.md)
- [Output contract](references/output_contract.md)

## Dependencies

`numpy`, `pandas`, `matplotlib`
