Agent skill

ML Generative Mattergen

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Generate inorganic material structures using MatterGen, a diffusion-based generative model.

MITAuto-check passedAI & LLM Engineering

Install ML Generative Mattergen

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-generative-mattergen -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills ml-generative-mattergen --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-generative-mattergen .claude/skills/ml-generative-mattergen && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-generative-mattergen
GitHub stars
175
Token cost
~1.8k tokens
SKILL.md length
664 words
Files
9 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Generate inorganic material structures using MatterGen, a diffusion-based generative model.

  • Works in 9 steps: Prerequisites → Available Models → MCP Tool Usage → …
  • Tasks that involve Fine-tuning
  • SKILL.md covers 1. Prerequisites, 2. Available Models, 3. MCP Tool Usage and 4. Fine-Tuning (Skill Scripts), plus 5 more sections
  • Runs Python and Shell scripts from its folder

What it does

ML Generative Mattergen is an agent skill from learningmatter-mit/AtomisticSkills. Generate inorganic material structures using MatterGen, a diffusion-based generative model.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 12 other files, including scripts (for example `examples/finetuning/README.md`, `examples/finetuning/example_finetuned_model/adapter_config.json` and `examples/finetuning/make_dummy.py`).

It sits in AI & LLM Engineering, covering Fine-tuning. It works with CUDA. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

When your agent uses it

  • Tasks that involve Fine-tuning

Example prompts

  • “/ml-generative-mattergen”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Prerequisites
  2. Available Models
  3. MCP Tool Usage
  4. Fine-Tuning (Skill Scripts)
  5. Output Files
  6. Limitations
  7. Best Practices
  8. Workflow Integration
  9. Examples

What it can do on your machine

Read from SKILL.md and the folder at commit 7f2d86d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python and Shell), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Generative Mattergen loads about 1.8k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 664 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 7f2d86d, republished under its MIT licence (© learningmatter-mit). 664 words, ~1,755 tokens.

Download SKILL.mdSave it as .claude/skills/ml-generative-mattergen/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
ml-generative-mattergen
description
Generate inorganic material structures using MatterGen, a diffusion-based generative model.
metadata.category
machine-learning, materials
metadata.venv
mattergen

MatterGen Structure Generation Skill

This skill provides tools for generating novel inorganic material structures using MatterGen, a state-of-the-art diffusion-based generative model for crystalline materials.

1. Prerequisites

[!IMPORTANT] GPU Required: MatterGen generation needs a CUDA GPU (NVIDIA driver ≥ 525).

  • Runs as the mattergen MCP server and its scripts run in the mattergen environment: on x86_64 a uv environment created on first use (CUDA 12.6 or 13 by driver), on aarch64 the generative container image.

  • A MatterGen checkout next to this project as ../mattergen, or anywhere with MATTERGEN_REPO pointing to it. MatterGen's PyPI distribution omits the data files it needs (sampling configs, GemNet scale factors), so it runs from the checkout, as upstream installs it. Fetch it with the setup script, which pins the verified commit, skips the LFS checkpoints (the weights come from Hugging Face) and replaces np.math, which NumPy 2 removed:

    bash
    ${CLAUDE_SKILL_DIR}/../../venv/run mattergen python ${CLAUDE_SKILL_DIR}/scripts/setup_mattergen.py

    venv/run mounts it into the container on aarch64.

  • On aarch64 (e.g. NVIDIA DGX Spark) the image carries PyG's extensions compiled for CUDA, which PyG publishes no aarch64 wheels for; nothing needs building on the host.

2. Available Models

MatterGen provides several pretrained models:

  • mattergen_base: Base unconditional generative model
  • mp_20_base: Materials Project base model
  • dft_mag_density: Model for magnetic density conditioning
  • chemical_system: Model for chemical system conditioning

3. MCP Tool Usage

The MCP tool automatically loads models when needed - no explicit load step required.

Unconditional Generation

Generate novel structures without conditioning:

python
from mcp_base import mattergen.generate_structures

result = mattergen.generate_structures(
    model_name="mattergen_base",
    num_structures=10,
    batch_size=10,
    output_dir="research/my_project/generated"
)
Chemical System Conditioning

Generate structures from a specific chemical system (controls which elements appear):

python
result = mattergen.generate_structures(
    chemical_system="Li-Fe-P-O",  # Automatically uses chemical_system model
    guidance_scale=1.0,  # Recommended for chemical system conditioning
    num_structures=20,
    batch_size=10,
    output_dir="research/cathode_materials/generated"
)

[!NOTE] Chemical system conditioning controls which elements appear, but NOT the exact stoichiometry. For example, chemical_system="Li-Zr-Cl" can generate Li3Cl5, LiZrCl4, Li2ZrCl5, etc., but you cannot specify exactly "Li2ZrCl6".

4. Fine-Tuning (Skill Scripts)

Fine-tune MatterGen on custom datasets using the skill scripts:

Step 1: Prepare Training Data
bash
# Convert structures and properties to CSV format
${CLAUDE_SKILL_DIR}/../../venv/run mattergen python ${CLAUDE_SKILL_DIR}/scripts/prepare_training_data.py \
  --structures-json training_structures.json \
  --property-name "formation_energy" \
  --output training_data.csv

Training data JSON format:

json
[
  {
    "structure": {<pymatgen Structure dict>},
    "properties": {"formation_energy": -2.5}
  },
  ...
]
Step 2: Run Fine-Tuning
bash
${CLAUDE_SKILL_DIR}/../../venv/run mattergen python ${CLAUDE_SKILL_DIR}/scripts/run_finetuning.py \
  --training-data training_data.csv \
  --property-name "formation_energy" \
  --base-model "mattergen_base" \
  --epochs 100 \
  --output-dir finetuned_formation_energy

Fine-tuning parameters:

  • --training-data: Path to CSV from Step 1
  • --property-name: Property to condition on (must match CSV column)
  • --base-model: Starting model (mattergen_base, chemical_system, etc.)
  • --epochs: Training epochs (100-200 recommended)
  • --learning-rate: Learning rate (default: 5e-6)
  • --batch-size: Batch size (default: 32)

[!TIP]

  • Start with 2 epochs for quick testing
  • Use 100-200 epochs for actual fine-tuning
  • GPU required (fine-tuning on CPU is extremely slow)

5. Output Files

Generation Output
  • structure_XXXX.cif: Generated structure files
  • generation_metadata.json: Metadata about generation parameters
Fine-Tuning Output
  • checkpoints/: Model checkpoint files (.ckpt)
  • config.yaml: Hydra configuration used
  • Training logs and metrics
Show full SKILL.md (276 more words)Show less

6. Limitations

[!WARNING] Chemical System vs. Stoichiometry

  • chemical_system parameter controls which elements are encouraged to be present.
  • It does NOT guarantee that all specified elements will be in the output structure.
  • It does NOT prevent other elements from occasionally appearing if guidance is too low.
  • Example: chemical_system="Li-Zr-Cl" might generate LiCl, ZrCl4, or even structures missing Li, alongside the desired ternaries (e.g., Li2ZrCl6).
  • Action Required: You MUST write a post-processing script to filter the output .cif files and keep only the ones that match your exact target elemental composition.

[!WARNING] CSP Mode Not Available

  • Target composition control (target_compositions parameter) requires CSP-trained models
  • CSP models are NOT publicly available - must be custom-trained
  • Public models (mattergen_base, chemical_system, etc.) are generation models only

7. Best Practices

[!IMPORTANT]

  • GPU Required: MatterGen requires a CUDA-compatible GPU. CPU is extremely slow.
  • Batch Size: Use larger batches (10-50) for efficient GPU utilization
  • Guidance Scale: Higher values (1.0-5.0) enforce stronger conditioning
  • Composition Filtering: Always filter the generated output CIFs using pymatgen to verify that the structures contain exactly the target elements.
  • Validation: Always validate generated structures via relaxation and stability analysis

[!TIP]

  • Start with unconditional generation to understand model behavior
  • Use chemical_system conditioning to explore specific element combinations
  • Fine-tune on domain-specific data for specialized applications (e.g., cathode materials)

8. Workflow Integration

MatterGen works well in combination with:

  • Structure relaxation: Use MCP MLIP tools to optimize generated structures
  • Stability analysis: Calculate E_hull to identify stable phases
  • Property prediction: Use MLIPs or DFT to calculate properties
  • High-throughput screening: Generate → Relax → Screen workflow

9. Examples

See examples/ for:

  • Unconditional generation workflow
  • Chemical system conditioning
  • Fine-tuning on custom datasets with complete example data


Author: Bowen Deng Contact: GitHub @learningmatter-mit

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts) in skills/ml-generative-mattergen of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/finetuning/README.md
  • examples/finetuning/example_finetuned_model/adapter_config.json
  • examples/finetuning/example_finetuned_model/adapter_model.bin
  • examples/finetuning/make_dummy.py
  • examples/finetuning/run_finetune_example.sh
  • scripts/prepare_training_data.py
  • scripts/run_finetuning.py
  • scripts/setup_mattergen.py

Open the folder on GitHubat commit 7f2d86d

Compare with similar skills

ML Generative Mattergen next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Generative Mattergen compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Generative Mattergen this skilllearningmatter-mit/AtomisticSkills175—~1.8kAutomated safety check: PassMIT
Setup Workshop Nemoclawbrevdev/workshop-build-an-agent143—~5.2kAutomated safety check: PassApache-2.0
Setup GuideRed-Hat-AI-Innovation-Team/training_hub100—~959Automated safety check: PassApache-2.0
Training Hub GuideRed-Hat-AI-Innovation-Team/training_hub100—~2.8kAutomated safety check: PassApache-2.0
Cosmos3 Post TrainingNVIDIA/cosmos-framework556—~2.7kAutomated safety check: PassCustom licence
Spark Environment Setupwshobson/agents40k—~2kAutomated safety check: PassMIT

Similar skills

  • Setup Workshop Nemoclaw

    brevdev/workshop-build-an-agent

    Set up the NVIDIA "Build an Agent" DevX workshop as a working JupyterLab environment from INSIDE a locked-down OpenShell/NemoClaw sandbox, and hand the user the token URL + access commands.

    143 GitHub stars~5.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Setup Guide

    Red-Hat-AI-Innovation-Team/training_hub

    A skill your agent uses when the user wants to set up LLM training for the first time, or when traininghub is not yet installed/configured in the current environment.

    100 GitHub stars~959 tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Training Hub Guide

    Red-Hat-AI-Innovation-Team/training_hub

    Guides users through LLM post-training with Training Hub, including installation, algorithm selection (SFT, OSFT, LoRA), hyperparameter tuning, troubleshooting OOM errors, interpreting loss curves…

    100 GitHub stars~2.8k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Cosmos3 Post Training

    NVIDIA/cosmos-framework

    Official

    Guide users through Cosmos3 supervised fine-tuning (SFT) post-training: preparing the example dataset and Wan2.2 VAE, converting the base checkpoint to DCP, launching distributed training (paired…

    556 GitHub stars~2.7k tokensUpdated 11 days ago
    AI & LLM EngineeringAuto-check passed
  • Set up a working ML training/inference environment on NVIDIA DGX Spark (GB10, aarch64, CUDA 13).

    40k GitHub stars~2k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Kermt Finetune

    NVIDIA/skills

    Official

    Finetune a pretrained KERMT encoder on a labeled CSV. An agent skill from NVIDIA/skills.

    3.5k GitHub starsUsed in 1 repo~4.1k tokens
    AI & LLM EngineeringAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    175 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    175 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    175 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    175 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    175 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    175 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Works with

Questions about ML Generative Mattergen

What does ML Generative Mattergen do?

Generate inorganic material structures using MatterGen, a diffusion-based generative model. ML Generative Mattergen is an agent skill from learningmatter-mit/AtomisticSkills. Generate inorganic material structures using MatterGen, a diffusion-based generative model.

When should I use ML Generative Mattergen?

ML Generative Mattergen fits situations like: tasks that involve Fine-tuning.

How do I install ML Generative Mattergen in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-generative-mattergen -a claude-code`. Or copy the skill folder (skills/ml-generative-mattergen in learningmatter-mit/AtomisticSkills) into .claude/skills/ml-generative-mattergen in your project. Claude Code loads it when a task matches its description.

How do I install ML Generative Mattergen in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-generative-mattergen -a codex`. Or copy the skill folder (skills/ml-generative-mattergen in learningmatter-mit/AtomisticSkills) into .agents/skills/ml-generative-mattergen in your project. Codex loads it when a task matches its description.

Can I use ML Generative Mattergen in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-generative-mattergen -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-generative-mattergen, .gemini/skills/ml-generative-mattergen, .github/skills/ml-generative-mattergen and .opencode/skills/ml-generative-mattergen in your project.

What does ML Generative Mattergen need to run?

Going by SKILL.md and its folder, ML Generative Mattergen needs Python and a shell for the scripts in its folder. Our summary lists: Python 3; A Bash shell.

Does ML Generative Mattergen access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is ML Generative Mattergen safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does ML Generative Mattergen use?

ML Generative Mattergen is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Generative Mattergen use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Generative Mattergen?

Skills that share tags, products or a category with ML Generative Mattergen: Setup Workshop Nemoclaw (brevdev/workshop-build-an-agent, 143 stars), Setup Guide (Red-Hat-AI-Innovation-Team/training_hub, 100 stars), Training Hub Guide (Red-Hat-AI-Innovation-Team/training_hub, 100 stars) and Cosmos3 Post Training (NVIDIA/cosmos-framework, 556 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Generative Mattergen?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 175 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 6, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.