Agent skill

Mat Synthesis Extraction

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Extract structured synthesis procedures from a folder of PDFs using the LeMat-Synth GeneralSynthesisOntology schema, producing one JSON file per paper with per-material synthesis records.

MITAuto-check passedDocuments & Office

Install Mat Synthesis Extraction

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill mat-synthesis-extraction -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills mat-synthesis-extraction --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mat-synthesis-extraction .claude/skills/mat-synthesis-extraction && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mat-synthesis-extraction
GitHub stars
175
Token cost
~2.2k tokens
SKILL.md length
502 words
Files
3 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Extract structured synthesis procedures from a folder of PDFs using the LeMat-Synth GeneralSynthesisOntology schema, producing one JSON file per paper with per-material synthesis records.

  • Works in 5 steps: Parse PDFs to text → Extract synthesized material names → Extract GeneralSynthesisOntology per… → …
  • Documents & Office work in your project
  • SKILL.md covers Goal, Instructions, Examples and Constraints, plus 1 more section
  • Runs Python scripts from its folder

What it does

Mat Synthesis Extraction is an agent skill from learningmatter-mit/AtomisticSkills. Extract structured synthesis procedures from a folder of PDFs using the LeMat-Synth GeneralSynthesisOntology schema, producing one JSON file per paper with per-material synthesis records.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `examples/pt-cu-alloy-co-oxidation/README.md` and `scripts/parse_pdfs.py`).

It sits in Documents & Office. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

When your agent uses it

  • Documents & Office work in your project

Example prompts

  • “/mat-synthesis-extraction”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Parse PDFs to text
  2. Extract synthesized material names
  3. Extract GeneralSynthesisOntology per material
  4. Collect and write output JSON
  5. Review outputs

What it can do on your machine

Read from SKILL.md and the folder at commit 7f2d86d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • arxiv.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mat Synthesis Extraction loads about 2.2k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 502 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 7f2d86d, republished under its MIT licence (© learningmatter-mit). 502 words, ~2,208 tokens.

Download SKILL.mdSave it as .claude/skills/mat-synthesis-extraction/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
mat-synthesis-extraction
description
Extract structured synthesis procedures from a folder of PDFs using the LeMat-Synth GeneralSynthesisOntology schema, producing one JSON file per paper with per-material synthesis records.
metadata.category
materials
metadata.venv
cpu

mat-synthesis-extraction

Goal

Given a folder of scientific paper PDFs, extract all synthesis procedures described in each paper and structure them according to the GeneralSynthesisOntology developed in LeMat-Synth [1]. Output is one JSON file per paper containing a list of per-material synthesis records.

The ontology captures: target compound, compound type, synthesis method, starting materials (with amounts/units/purity), sequential process steps (with actions, conditions, equipment), and overall equipment list.


Instructions

Step 1 — Parse PDFs to text

Run the PDF parser to extract plain text from all PDFs in the input folder.

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/parse_pdfs.py \
    --pdf-dir /path/to/pdf_folder \
    --output-dir /path/to/output/texts

This produces:

  • One .txt file per PDF (named <paper_stem>.txt)
  • A parse_summary.json listing extraction status and character counts

Inspect parse_summary.json to confirm all PDFs extracted successfully. Papers with "status": "empty" are likely scanned images — skip them or obtain a text-layer PDF.


Step 2 — Extract synthesized material names

For each .txt file produced in Step 1, identify which materials are synthesized in the paper. Read the paper text and extract a comma-separated list of synthesized compound names.

System prompt to use:

You are a materials science expert. Given the full text of a scientific paper, identify ALL distinct materials that are synthesized (not just characterized or used as reagents). Return ONLY a comma-separated list of chemical names or formulas (e.g. "NiCo2O4, CoFe2O4, Fe3O4"). If no synthesis is described, return an empty string.

Input: full paper text from Step 1 Output: comma-separated string of material names → split into a Python list


Step 3 — Extract GeneralSynthesisOntology per material

For each (paper_text, material_name) pair from Step 2, extract the structured synthesis ontology. Use the system prompt and JSON schema below.

System prompt:

You are a helpful assistant that extracts the structured synthesis for a specific material from the paper text.

Focus ONLY on the synthesis procedure for the specified material. Search through the entire paper text to find the synthesis procedure that describes how this specific material is made.

IMPORTANT: You must output ONLY a valid JSON object with a "structured_synthesis" field. Do not include any reasoning, explanations, or markdown formatting.

If you cannot find a synthesis procedure for the specified material, return a minimal structure with the material name and an empty synthesis.

User message template:

Paper text:
<PAPER_TEXT>

Extract the synthesis procedure for: <MATERIAL_NAME>

Required JSON output schema (GeneralSynthesisOntology):

json
{
  "structured_synthesis": {
    "target_compound": "string (required) — composition and description of the target",
    "target_compound_type": "one of: 'metals & alloys' | 'ceramics & glasses' | 'polymers & soft matter' | 'composites' | 'semiconductors & electronic' | 'nanomaterials' | 'two-dimensional materials' | 'framework & porous materials' | 'biomaterials & biological' | 'liquid materials' | 'hybrid & organic-inorganic' | 'functional materials & catalysts' | 'energy & sustainability' | 'smart & responsive materials' | 'emerging & quantum materials' | 'other'",
    "synthesis_method": "one of: 'PVD' | 'CVD' | 'arc discharge' | 'ball milling' | 'spray pyrolysis' | 'electrospinning' | 'sol-gel' | 'hydrothermal' | 'solvothermal' | 'precipitation' | 'coprecipitation' | 'combustion' | 'microwave-assisted' | 'sonochemical' | 'template-directed' | 'solid-state' | 'flux growth' | 'float zone & Bridgman' | 'arc melting & induction melting' | 'spark plasma sintering' | 'electrochemical deposition' | 'chemical bath deposition' | 'liquid-phase epitaxy' | 'self-assembly' | 'atomic layer deposition' | 'molecular beam epitaxy' | 'pulsed laser deposition' | 'ion implantation' | 'lithographic patterning' | 'wet impregnation' | 'incipient wetness impregnation' | 'mechanical mixing' | 'solution-based' | 'mechanochemical' | 'other'",
    "starting_materials": [
      {
        "name": "string",
        "amount": "number or null",
        "unit": "string or null — e.g. 'g', 'mL', 'mmol', 'wt%'",
        "purity": "string or null — e.g. '99.9%', 'ACS grade'",
        "vendor": "string or null"
      }
    ],
    "steps": [
      {
        "step_number": "integer",
        "action": "one of: 'add' | 'mix' | 'heat' | 'cool' | 'reflux' | 'age' | 'filter' | 'wash' | 'dry' | 'reduce' | 'calcine' | 'dissolve' | 'precipitate' | 'centrifuge' | 'sonicate' | 'anneal' | 'ion exchange' | 'impregnate'",
        "description": "string or null",
        "materials": [{"name": "string", "amount": "number or null", "unit": "string or null", "purity": "string or null", "vendor": "string or null"}],
        "equipment": [{"name": "string", "instrument_vendor": "string or null", "settings": "string or null"}],
        "conditions": {
          "temperature": "number or null",
          "temp_unit": "string or null — 'C', 'K', or 'F'",
          "duration": "number or null",
          "time_unit": "string or null — 'h', 'min', 's', 'days'",
          "pressure": "number or null",
          "pressure_unit": "string or null",
          "atmosphere": "string or null — e.g. 'air', 'N2', 'Ar'",
          "stirring": "boolean or null",
          "stirring_speed": "number or null",
          "ph": "number or null"
        }
      }
    ],
    "equipment": [{"name": "string", "instrument_vendor": "string or null", "settings": "string or null"}],
    "notes": "string or null"
  }
}

Retry strategy: If extraction fails or JSON is invalid, retry with slightly increased temperature (0.3, then 0.5). If all retries fail, write a minimal record:

json
{"target_compound": "<material_name>", "target_compound_type": "other", "synthesis_method": "other", "starting_materials": [], "steps": [], "equipment": [], "notes": "Extraction failed."}

Step 4 — Collect and write output JSON

After extracting all materials for a paper, write one JSON output file per paper.

Output file: <output_dir>/<paper_stem>_synthesis.json

Output format:

json
{
  "paper": "<paper_stem>",
  "source_pdf": "<original_pdf_filename>",
  "materials_found": ["Material A", "Material B"],
  "syntheses": [
    {
      "material": "Material A",
      "synthesis": { ... }
    },
    {
      "material": "Material B",
      "synthesis": { ... }
    }
  ]
}

Write all output JSONs to the same --output-dir used in Step 1 (or a dedicated subdirectory).


Show full SKILL.md (213 more words)Show less
Step 5 — Review outputs

Inspect the extracted JSONs. Key things to verify:

  • target_compound matches the material identified in Step 2
  • synthesis_method is not "other" unless genuinely ambiguous
  • steps list is non-empty for papers that clearly describe synthesis
  • starting_materials includes amounts/units where reported in the paper

Flag papers where steps is empty and notes contains "Extraction failed" for manual review.


Examples

See examples/pt-cu-alloy-co-oxidation/ for a worked example using a catalysis paper from ChemRxiv.


Constraints

  • Environment: Step 1 (parse_pdfs.py) requires cpu (pymupdf installed).
  • Scanned PDFs: PyMuPDF extracts embedded text only. Scanned-image PDFs produce empty output — use a PDF with a text layer.
  • Figure text: PyMuPDF may extract figure captions and table text. The LLM extraction prompt instructs to focus on synthesis procedure sections only.
  • One JSON per paper: Output aggregates all materials for a paper into a single file.
  • No DSPy / no separate LLM env: LLM extraction is done by the agent directly using the schema and prompts above. No additional environment needed beyond cpu for the parsing script.
  • Input configs: Save any run-specific settings (model used, temperatures, pdf_dir, output_dir) to input_configs.yaml in the output directory for reproducibility.

References

[1] Lederbauer et al., "LeMat-Synth: a multi-modal toolbox to curate broad synthesis procedure databases from scientific literature", arXiv, 2025. arXiv:2510.26824


Author: Magdalena Lederbauer Contact: GitHub @mlederbauer

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in skills/mat-synthesis-extraction of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/pt-cu-alloy-co-oxidation/README.md
  • scripts/parse_pdfs.py

Open the folder on GitHubat commit 7f2d86d

Compare with similar skills

Mat Synthesis Extraction next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mat Synthesis Extraction compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mat Synthesis Extraction this skilllearningmatter-mit/AtomisticSkills175—~2.2kAutomated safety check: PassMIT
Markdown Article FormatterJimLiu/baoyu-skills26k6 repos~3.5kAutomated safety check: PassMIT
MarkitdownImCa0/just-laws78114 repos~3.2kAutomated safety check: NotesMIT
Obsidian MarkdownAtmosphere/atmosphere3.8k20 repos~1.3kAutomated safety check: PassApache-2.0
DOCXrvdbreemen/OTGW-firmware20733 repos~4.3kAutomated safety check: PassProprietary
Gzh Designisjiamu/gzh-design-skill3.9k1 repos~2.2kAutomated safety check: PassAGPL-3.0

Similar skills

  • Markdown Article Formatter

    JimLiu/baoyu-skills

    Reformats plain text or Markdown articles with frontmatter, a title, a summary, headings, bold, lists and code blocks, and saves a separate formatted copy.

    26k GitHub starsUsed in 6 repos~3.5k tokens
    Documents & OfficeAuto-check passed
  • Markitdown

    ImCa0/just-laws

    Convert files and office documents to Markdown. An agent skill from ImCa0/just-laws.

    781 GitHub starsUsed in 14 repos~3.2k tokens
    Documents & OfficeAuto-check: notes
  • Obsidian Markdown

    Atmosphere/atmosphere

    Create and edit Obsidian Flavored Markdown with wikilinks, embeds, callouts, properties, and other Obsidian-specific syntax.

    3.8k GitHub starsUsed in 20 repos~1.3k tokens
    Documents & OfficeAuto-check passed
  • DOCX

    rvdbreemen/OTGW-firmware

    A skill your agent uses whenever the user wants to create, read, edit, or manipulate Word documents (.docx files).

    207 GitHub starsUsed in 33 repos~4.3k tokens
    Documents & OfficeAuto-check passed
  • Gzh Design

    isjiamu/gzh-design-skill

    微信公众号文章排版引擎,将 Markdown 转换为可直接粘贴到公众号编辑器的 HTML。主题风格从 references/theme-index.md 注册的自定义主题库中选取,自动章节编号、关键词下划线标记、引言卡片、目录导航、代码块、图片/GIF、作者签名。支持 Markdown / Word(.docx) / PDF / 纯文本输入(非 Markdown…

    3.9k GitHub starsUsed in 1 repo~2.2k tokens
    Documents & OfficeAuto-check passed
  • Reads, creates and edits Word .docx files with python-docx, and drops to raw OOXML for tracked changes, comments and byte-exact edits.

    41k GitHub stars~2.5k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    175 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    175 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    175 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    175 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    175 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    175 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Questions about Mat Synthesis Extraction

What does Mat Synthesis Extraction do?

Extract structured synthesis procedures from a folder of PDFs using the LeMat-Synth GeneralSynthesisOntology schema, producing one JSON file per paper with per-material synthesis records. Mat Synthesis Extraction is an agent skill from learningmatter-mit/AtomisticSkills. Extract structured synthesis procedures from a folder of PDFs using the LeMat-Synth GeneralSynthesisOntology schema, producing one JSON file per paper with per-material synthesis records.

When should I use Mat Synthesis Extraction?

Mat Synthesis Extraction fits situations like: documents & Office work in your project.

How do I install Mat Synthesis Extraction in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill mat-synthesis-extraction -a claude-code`. Or copy the skill folder (skills/mat-synthesis-extraction in learningmatter-mit/AtomisticSkills) into .claude/skills/mat-synthesis-extraction in your project. Claude Code loads it when a task matches its description.

How do I install Mat Synthesis Extraction in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill mat-synthesis-extraction -a codex`. Or copy the skill folder (skills/mat-synthesis-extraction in learningmatter-mit/AtomisticSkills) into .agents/skills/mat-synthesis-extraction in your project. Codex loads it when a task matches its description.

Can I use Mat Synthesis Extraction in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill mat-synthesis-extraction -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mat-synthesis-extraction, .gemini/skills/mat-synthesis-extraction, .github/skills/mat-synthesis-extraction and .opencode/skills/mat-synthesis-extraction in your project.

What does Mat Synthesis Extraction need to run?

Going by SKILL.md and its folder, Mat Synthesis Extraction needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Mat Synthesis Extraction access the network?

SKILL.md names 2 domains. As links in the text: arxiv.org and github.com. This is read from the text; nothing was executed.

Is Mat Synthesis Extraction safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Mat Synthesis Extraction use?

Mat Synthesis Extraction is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Mat Synthesis Extraction use?

About 2.2k tokens (SKILL.md is roughly 8.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mat Synthesis Extraction?

Skills that share tags, products or a category with Mat Synthesis Extraction: Markdown Article Formatter (JimLiu/baoyu-skills, 26k stars), Markitdown (ImCa0/just-laws, 781 stars), Obsidian Markdown (Atmosphere/atmosphere, 3.8k stars) and DOCX (rvdbreemen/OTGW-firmware, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mat Synthesis Extraction?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 175 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 6, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.