Agent skill

General Plot Digitizer

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Extract continuous X-Y data from experimental spectrum images (Raman, XRD, UV-Vis, IR, etc.) via hybrid VLM + CV pipeline and agent-in-the-loop workflow.

MITAuto-check passed

Install General Plot Digitizer

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill general-plot-digitizer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills general-plot-digitizer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/general-plot-digitizer .claude/skills/general-plot-digitizer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
general-plot-digitizer
GitHub stars
176
Token cost
~2.8k tokens
SKILL.md length
1,130 words
Files
59 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Extract continuous X-Y data from experimental spectrum images (Raman, XRD, UV-Vis, IR, etc.) via hybrid VLM + CV pipeline and agent-in-the-loop workflow.

  • Works in 4 steps: Visual Inspection (VLM) → Metadata Construction (Coding Agent) → Pipeline Execution → …
  • SKILL.md covers Goal, Instructions, Agent Rules and Key Parameters, plus 5 more sections
  • Calls python

What it does

General Plot Digitizer is an agent skill from learningmatter-mit/AtomisticSkills. Extract continuous X-Y data from experimental spectrum images (Raman, XRD, UV-Vis, IR, etc.) via hybrid VLM + CV pipeline and agent-in-the-loop workflow.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 61 other files, including scripts (for example `examples/01-single-curve/README.md`, `examples/01-single-curve/metadata.json` and `examples/01-single-curve/source_digitized.md`).

The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

Example prompts

  • “/general-plot-digitizer”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Visual Inspection (VLM)
  2. Metadata Construction (Coding Agent)
  3. Pipeline Execution
  4. Overlay Inspection & Iteration

What it can do on your machine

Read from SKILL.md and the folder at commit 6257444. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

General Plot Digitizer loads about 2.8k tokens when it runs. Until then it costs about 44 tokens; SKILL.md has 1,130 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~44
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 6257444, republished under its MIT licence (© learningmatter-mit). 1,130 words, ~2,829 tokens.

Download SKILL.mdSave it as .claude/skills/general-plot-digitizer/SKILL.md (or your agent's skills folder). This skill also uses 58 other files; get the full folder from GitHub.
name
general-plot-digitizer
description
Extract continuous X-Y data from experimental spectrum images (Raman, XRD, UV-Vis, IR, etc.) via hybrid VLM + CV pipeline and agent-in-the-loop workflow.
metadata.category
general, machine-learning
metadata.venv
cpu

General Plot Digitizer

Goal

Extract calibrated numeric X-Y data from images of experimental spectra (Raman, XRD, UV-Vis, IR, NMR, etc.) using a deterministic "Agent-in-the-Loop" workflow.

The labor is divided between two models:

  1. Vision-Language Model (Visual Sensor): Reads the image and returns a rich, unstructured narrative description of axes, colors, and visual obstacles. It does not produce JSON.
  2. Coding Agent (Translator & Executor): Translates the VLM narrative into a precise metadata.json, runs the CV pipeline, inspects the overlay, and iterates until the curve is correctly isolated.

Instructions

Phase 1: Visual Inspection (VLM)

Do not attempt to generate JSON with the VLM. It acts only as a visual sensor.

  1. Generate grid overlay:
bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/plot_utils.py plot.png --draw-grid

This produces plot_grid.png with a labeled pixel grid for precise coordinate reading.

  1. Prompt the VLM to analyze plot_grid.png (not the raw image). Use the built-in vision capabilities or the notify_user VLM inspection tool. Provide the prompt guidelines from scripts/vlm_prompt_template.txt.

  2. Expected VLM output — a natural-language report covering:

    • Axis labels, numeric ranges, and directions (is X reversed?).
    • Bounding box of the data region in pixels (read from the grid).
    • Color and style of each target curve (hex guess from the pixels on the line itself).
    • Pixel bounding boxes of all visual obstacles (legends, text annotations, gridlines, tick marks) that overlap the data curves.
    • Trace quality hints: thin/needle-like, thick/noisy, anti-aliased, JPEG artifacts.
Phase 2: Metadata Construction (Coding Agent)

Read the VLM narrative and construct metadata.json. Schema: resources/metadata_schema.json.

Required fields:

json
{
  "plot_title": "",
  "x_axis_label": "Wavelength (nm)",
  "y_axis_label": "Absorbance",
  "x_tick_min": 400, "x_tick_max": 800,
  "y_tick_min": 0, "y_tick_max": 1,
  "x_calibration_points": [
    { "pixel": 70, "value": 400 },
    { "pixel": 450, "value": 800 }
  ],
  "x_scale": "linear", "y_scale": "linear",
  "bounding_box": {"x_min": 72, "y_min": 28, "x_max": 452, "y_max": 318},
  "x_reversed": false, "y_reversed": false,
  "spectrum_type": "UV-Vis",
  "curves": [{"label": "sample", "color_hint": "#1f77b4"}],
  "text_regions": [{"x_min": 300, "y_min": 50, "x_max": 400, "y_max": 80, "label": "legend"}]
}

x_calibration_points (strongly recommended): anchor the X-axis transform to exact pixel→value pairs read from the grid, rather than assuming axis ticks align perfectly with bbox edges. Pick two well-separated ticks visible on the grid. If provided, these override x_tick_min/max for pixel-to-data mapping.

Translation rules (VLM narrative → metadata fields):

  • Obstacles → text_regions[] and/or mask_regions[] with pixel bounding boxes.
  • Curve colors → curves[].color_hint (from pixels on the plotted line, not from legend swatches).
  • "Curve is black / same as axes" → "cli_hints": {"curve_is_black": true}.
  • "Noisy / jagged trace" → "cli_hints": {"smooth": true}.
  • "Thin, needle-like peaks" → do not set smooth; set upscale in Phase 3 instead.

If the VLM color guess is uncertain, run:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/suggest_colors.py plot.png \
  --bounding-box x_min,y_min,x_max,y_max

This reports dominant non-background colors in the cropped region. Use the top result as color_hint.

Phase 3: Pipeline Execution

Select CLI flags based on VLM visual cues:

VLM describes...Required CLI flagsAvoid
Thin, needle-like peaks (XRD, FTIR)--crop-upscale 4.0--smooth, --cluster-centroid
Fuzzy / anti-aliased / JPEG artifacts--curve-tolerance 75 (up to 85)—
Thick, noisy trace / scatter points--smooth --smooth-window 5 --smooth-deviation 15.0—
Black curve on black axes--allow-black (auto-enables --spatial-filter --cluster-centroid)—
Thin anti-aliased colored line--extraction-method edge+color--morph-open

Run the pipeline:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/digitize_pipeline.py \
  plot.png \
  --full \
  --metadata metadata.json \
  --output-dir ./output \
  --overlay \
  --format both

Append the VLM-dictated flags from the table above. The pipeline also reads cli_hints from metadata and auto-applies safe flags (--allow-black, --smooth, --all-curves).

Phase 4: Overlay Inspection & Iteration
  1. Inspect *_digitized.overlay.png visually.
  2. If extraction is wrong, diagnose using this table and re-run Phase 3:
SymptomFix (metadata or CLI)
Wrong curve extracted (e.g. legend ink)Set curves[].color_hint from actual line pixels; or run suggest_colors.py on a tight crop
Text / labels contaminating traceAdd bounding boxes to text_regions[] in metadata; add --smooth
Black curve picks up axis lines--allow-black + add axis regions to mask_regions[]
Trace too sparse / broken gaps--curve-tolerance 55 or --preset lowres
Thin line lost entirely--extraction-method edge+color; or --crop-upscale 4.0 for needle peaks
Same-color text blobs on thick curve--morph-open (caution: destroys thin <3px curves)
Low-res image, everything pixelated--upscale-strategy force (pre-upscales image, scales metadata bbox)
X/Y values shifted or invertedFix x_tick_min/max, x_reversed, y_reversed, or x_calibration_points in metadata
  1. Prefer metadata edits (adjusting color_hint, text_regions, mask_regions, bounding_box) over adding CLI flags. Re-run the same pipeline command after editing metadata.json.

Outputs:

  • *_digitized.csv — comma-separated with x,y header
  • *_digitized.xy — space-separated, no header (if --format xy or both)
  • *_digitized.overlay.png — visual QC
  • *_digitized.md — summary

Agent Rules

  • Do not write ad-hoc NumPy/OpenCV pixel-scanning code. The pipeline already handles HSV masking, spatial filtering, and smoothing internally.
  • Do not skip the grid overlay step. The VLM is significantly more accurate on plot_grid.png than on raw images.
  • Do not guess CLI flags without VLM visual evidence. The flag table above is the complete decision tree.
  • Always inspect the overlay image after each run before declaring success.
Show full SKILL.md (446 more words)Show less

Key Parameters

FlagUse when...Default
--curve-color HEXVLM identified a specific curve colorauto-detect
--curve-tolerance NTrace is sparse or image has JPEG artifacts (increase); or mask bleeds into nearby colors (decrease)40
--overlayAlways recommended for QCoff
--format {csv,xy,both}Downstream tool needs specific formatcsv
--all-curvesMultiple curves in metadata curves[] (auto-enabled when >1 curve)off
--allow-blackData curve is black (same color as axes/frame)off
--smoothNoisy, jagged, or thick trace with outlier pointsoff
--smooth-window NTune smoothing aggressiveness (larger = more smoothing)5
--smooth-deviation PXPixel distance from local median beyond which a point is rejected as outlier15.0
--crop-upscale FACTORThin peaks (XRD, FTIR) need more pixel width to register; or generally low-res crop1.0
--upscale-strategy {none,auto,force}auto (default) pre-upscales when metadata or heuristic says low-res; force always pre-upscales; none skipsauto
--upscale-factor FACTORControls the multiplier used by --upscale-strategy auto|force2.0
--vlm-metadata-on-upscaleAfter pre-upscale, re-run VLM metadata on upscaled image instead of just scaling coordinates. Only if API keys are setoff
--extraction-method {color,edge,edge+color}color (default) fails on thin anti-aliased lines; try edge+colorcolor
--morph-openSame-color text blobs touching a thick curve; erodes then dilates to remove small blobs. Destroys thin (<3px) curvesoff
--spatial-filterCurve matches frame/axis color — keeps only the largest connected line-like componentoff
--cluster-centroidText or axes bleed into mask — uses largest-cluster centroid instead of full-column median. Auto-enabled with --allow-blackoff
--preset {lowres,thin-red}Quick combos: lowres = upscale 2 + tolerance 55 + edge+color; thin-red = tolerance 50, no morph-opennone
--debugSave intermediate crops/masks for diagnosing failuresoff
--json-summaryEmit machine-readable JSON summary to stdout after completionoff

Full flag list: python ${CLAUDE_SKILL_DIR}/scripts/digitize_pipeline.py --help

Examples

ScenarioDirectoryKey Flags
Single colored curve01-single-curve/--curve-color
Multiple curves by color02-multi-curve-color/--all-curves
Black curve + text masking03-black-curve-text-mask/--allow-black --smooth, text_regions
Stacked spectra04-stacked-spectra/--all-curves, per_curve_normalized

Constraints

  • Bounding box must enclose only the inner data region (inside axis lines, excluding labels/legend/title). VLMs often get this wrong — verify against the grid overlay.
  • Multi-curve: Each curve in curves[] must have a color_hint. Auto-detect is unreliable with multiple traces.
  • Per-curve regions: For stacked/offset plots, curves[].region.y_min/y_max must include generous padding (10-20px) above tallest peaks and below baseline.
  • Environments: All scripts require the cpu environment.

Resources

References

  • Gonzalez & Woods, Digital Image Processing, Pearson, 4th ed., 2018. Standard reference for HSV color-space segmentation, morphological operations, and connected-component analysis used in the extraction pipeline.
  • Bradski, G., "The OpenCV Library", Dr. Dobb's Journal of Software Tools, 2000. Core CV library underlying all image manipulation in this skill.

Author: Jesus Diaz Sanchez Contact: GitHub @jdsanc

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 58 other files (scripts) in skills/general-plot-digitizer of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/01-single-curve/README.md
  • examples/01-single-curve/metadata.json
  • examples/01-single-curve/source.png
  • examples/01-single-curve/source_digitized.csv
  • examples/01-single-curve/source_digitized.md
  • examples/01-single-curve/source_digitized.overlay.png
  • examples/01-single-curve/source_digitized.xy
  • examples/02-multi-curve-color/README.md
  • examples/02-multi-curve-color/metadata.json
  • examples/02-multi-curve-color/source.png
  • examples/02-multi-curve-color/source_citric_acid_aqueous_solution_digitized.csv
  • examples/02-multi-curve-color/source_citric_acid_aqueous_solution_digitized.md
  • examples/02-multi-curve-color/source_citric_acid_aqueous_solution_digitized.overlay.png
  • examples/02-multi-curve-color/source_citric_acid_aqueous_solution_digitized.xy
  • examples/02-multi-curve-color/source_citric_acid_solid_digitized.csv
  • examples/02-multi-curve-color/source_citric_acid_solid_digitized.md
  • examples/02-multi-curve-color/source_citric_acid_solid_digitized.overlay.png
  • … and 41 more

Open the folder on GitHubat commit 6257444

Compare with similar skills

General Plot Digitizer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

General Plot Digitizer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
General Plot Digitizer this skilllearningmatter-mit/AtomisticSkills176—~2.8kAutomated safety check: PassMIT
Meta Forest Continuous Plotaipoch/medical-research-skills2k—~1.8kAutomated safety check: PassMIT
Continuetelegramdesktop/tdesktop33k2 repos~9.4kAutomated safety check: PassGPL-3.0
Extractalirezarezvani/claude-skills28k—~1.4kAutomated safety check: PassMIT
PDF Extract Experimental Materialsaipoch/medical-research-skills2k—~2kAutomated safety check: PassMIT
Digital Forensicssickn33/agentic-awesome-skills47k1 repos~495Automated safety check: PassMIT

Similar skills

  • Meta Forest Continuous Plot

    aipoch/medical-research-skills

    Generate forest plots for meta-analysis of continuous data. An agent skill from aipoch/medical-research-skills.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Documents & OfficeAuto-check passed
  • Continue

    telegramdesktop/tdesktop

    Continue autonomous Telegram Desktop development from the shared ai-tdesktop repository.

    33k GitHub starsUsed in 2 repos~9.4k tokens
    Productivity & AutomationAuto-check passed
  • Extract

    alirezarezvani/claude-skills

    Turn a proven pattern or debugging solution into a standalone reusable skill with SKILL.md, reference docs, and examples.

    28k GitHub stars~1.4k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • PDF Extract Experimental Materials

    aipoch/medical-research-skills

    Extract experimental materials and instrument information from PDFs (or PDF-derived text/Markdown) into three CSV tables; use when a paper/report contains sections like Materials and Methods, Key…

    2k GitHub stars~2k tokensUpdated 21 days ago
    Documents & OfficeAuto-check passed
  • Digital Forensics

    sickn33/agentic-awesome-skills

    Authorized digital forensics: memory dumps, disk timelines, PCAP investigation, artifact triage, and incident-response evidence preservation.

    47k GitHub starsUsed in 1 repo~495 tokens
    SecurityAuto-check passed
  • Digital Forensics

    zhaoxuya520/reverse-skill

    A skill your agent uses for authorized digital forensics including memory dumps, disk timelines, PCAP investigation, artifact triage, and IR evidence preservation.

    40k GitHub starsUsed in 2 repos~389 tokens
    SecurityAuto-check: warnings

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    176 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    176 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    176 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    176 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    176 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    176 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about General Plot Digitizer

What does General Plot Digitizer do?

Extract continuous X-Y data from experimental spectrum images (Raman, XRD, UV-Vis, IR, etc.) via hybrid VLM + CV pipeline and agent-in-the-loop workflow. General Plot Digitizer is an agent skill from learningmatter-mit/AtomisticSkills.) via hybrid VLM + CV pipeline and agent-in-the-loop workflow.

How do I install General Plot Digitizer in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill general-plot-digitizer -a claude-code`. Or copy the skill folder (skills/general-plot-digitizer in learningmatter-mit/AtomisticSkills) into .claude/skills/general-plot-digitizer in your project. Claude Code loads it when a task matches its description.

How do I install General Plot Digitizer in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill general-plot-digitizer -a codex`. Or copy the skill folder (skills/general-plot-digitizer in learningmatter-mit/AtomisticSkills) into .agents/skills/general-plot-digitizer in your project. Codex loads it when a task matches its description.

Can I use General Plot Digitizer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill general-plot-digitizer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/general-plot-digitizer, .gemini/skills/general-plot-digitizer, .github/skills/general-plot-digitizer and .opencode/skills/general-plot-digitizer in your project.

What does General Plot Digitizer need to run?

Going by SKILL.md and its folder, General Plot Digitizer needs the command-line tools its instructions call (python). Our summary lists: Python 3.

Does General Plot Digitizer access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is General Plot Digitizer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does General Plot Digitizer use?

General Plot Digitizer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does General Plot Digitizer use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to General Plot Digitizer?

Skills that share tags, products or a category with General Plot Digitizer: Meta Forest Continuous Plot (aipoch/medical-research-skills, 2k stars), Continue (telegramdesktop/tdesktop, 33k stars), Extract (alirezarezvani/claude-skills, 28k stars) and PDF Extract Experimental Materials (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains General Plot Digitizer?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 176 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 7, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.