Agent skill

ML Mlip Benchmark

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.

MITAuto-check passed

Install ML Mlip Benchmark

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills ml-mlip-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-mlip-benchmark .claude/skills/ml-mlip-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-mlip-benchmark
GitHub stars
175
Token cost
~1.7k tokens
SKILL.md length
655 words
Files
9 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.

  • Works in 4 steps: Run Benchmark metrics → Generate Parity Plots → Reconcile units before comparing anything → …
  • SKILL.md covers Prerequisites, Instructions, Examples and Typical Combinations
  • Runs Python scripts from its folder

What it does

ML Mlip Benchmark is an agent skill from learningmatter-mit/AtomisticSkills. Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `examples/README.md`, `examples/fetch_r2scan.py` and `scripts/plot_benchmark.py`).

It works with Model Context Protocol and Python. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

Example prompts

  • “/ml-mlip-benchmark”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Run Benchmark metrics
  2. Generate Parity Plots
  3. Reconcile units before comparing anything
  4. Interpret Results

What it can do on your machine

Read from SKILL.md and the folder at commit 7f2d86d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Mlip Benchmark loads about 1.7k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 655 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~35
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 7f2d86d, republished under its MIT licence (© learningmatter-mit). 655 words, ~1,737 tokens.

Download SKILL.mdSave it as .claude/skills/ml-mlip-benchmark/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
ml-mlip-benchmark
description
Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.
metadata.category
machine-learning, materials, chemistry
metadata.venv
cpu, fairchem, mlip

Benchmark Machine Learning Interatomic Potentials (MLIP)

<!-- mcp-tools-note -->

[!NOTE] Steps written server.tool are MCP tool calls: mace.load_model is the load_model tool of the mace server (mcp__mace__load_model, or mcp__plugin_atomistic-skills_mace__load_model when installed as a plugin). Without a connected server, run the same tools from the shell. Tools named in one command share a process, so a model loaded by load_model stays loaded:

bash
${CLAUDE_SKILL_DIR}/../../venv/run mlip python -m src.mcp_server.cli mace load_model key=value
${CLAUDE_SKILL_DIR}/../../venv/run fairchem python -m src.mcp_server.cli fairchem load_model key=value
${CLAUDE_SKILL_DIR}/../../venv/run mlip python -m src.mcp_server.cli matgl load_model key=value

This skill evaluates the accuracy of a given MLIP against an existing ground-truth dataset (e.g., DFT calculations or a higher-fidelity foundation potential). It computes the Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) for both energy (per atom) and atomic forces, and optionally stress. It also generates parity plots for visual inspection of the model's correlation.

Prerequisites

  1. Model Loaded: An MLIP must be currently active via a load_model MCP tool call (e.g., mace.load_model, fairchem.load_model, matgl.load_model).
  2. Labeled Data: A JSON dataset where each entry contains a structural dictionary under "structure", along with scalar/vector ground truth values for "energy", "forces", and optionally "stress". This is identical to the format used in ml-mlip-training. (Data can be generated using Atomate2 MongoDB queries or MD sampling + labeling).

Instructions

1. Run Benchmark metrics

Use the ${CLAUDE_SKILL_DIR}/scripts/run_benchmark.py script to perform inference across the dataset and compute global error metrics.

Environment requirement: This script instantiates the MLIP models directly, so it must run in the uv project that provides the backend: venv/mlip for MACE and MatGL, venv/fairchem for FairChem.

bash
# (or venv/fairchem when --backend fairchem)
${CLAUDE_SKILL_DIR}/../../venv/run mlip python ${CLAUDE_SKILL_DIR}/scripts/run_benchmark.py \
    --data_path <path_to_labeled_data.json> \
    --model <model_name_or_path> \
    --backend <mace|fairchem|matgl> \
    --output <path_to_save_benchmark_results.json>

Note: The script utilizes src.utils.mlips.loader.load_wrapper to abstract backend details.

2. Generate Parity Plots

Once run_benchmark.py finishes, it writes a comprehensive JSON file containing original targets alongside the model's predictions and numerical metrics. Visualize these using the plotting script.

Environment requirement: venv/cpu is enough for the plotting script.

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/plot_benchmark.py \
    --results <path_to_benchmark_results.json> \
    --output_dir <path_to_save_plots>

This generates energy_parity.png, forces_parity.png, and (if stress was present) stress_parity.png.

Show full SKILL.md (340 more words)Show less
3. Reconcile units before comparing anything

A benchmark subtracts two numbers that came from different software, so a unit or sign mismatch shows up as a large "model error" that is not a model error at all. Energy (eV) and forces (eV/Å) agree across every backend here; stress does not. Settle it before computing a single metric -- see general-property-units for the full tables.

The three traps, in order of how often they bite:

  1. MatGL returns GPa, not eV/ų. matgl.ext.ase.PESCalculator defaults to stress_unit="GPa", and Potential.forward returns GPa as well, so MatGL is not a drop-in ASE calculator. Pass PESCalculator(potential=model, stress_unit="eV/A3"). Getting this wrong is a factor of 160.21766208.
  2. Raw model output != ASE calculator output. CHGNetCalculator converts GPa to eV/ų on the way out (stress_weight, default 1/160.21766208); MACE and FAIRChem convert nothing because their models already emit eV/ų. Know which layer you are reading.
  3. DFT labels usually carry the opposite sign. VASP reports stress compressive-positive in kB; ASE and every MLIP here are tensile-positive in eV/ų. Converting VASP labels to ASE convention is eV/A3 = -kB / 1602.1766208.

Sanity check that costs nothing: take a structure you have compressed, and confirm the diagonal stress is negative in ASE convention. If it is positive, you have a sign convention crossed somewhere.

4. Interpret Results

When presenting the plotted benchmarks to the user, consult the following rough heuristics for MLIP performance:

  • Energy MAE: Excellent (< 5 meV/atom), Good (5-20 meV/atom), Poor (> 50 meV/atom)
  • Forces MAE: Excellent (< 20 meV/Å), Good (20-50 meV/Å), Poor (> 100 meV/Å)

If the model is performing poorly on the labeled data, suggest fine-tuning it utilizing the ml-mlip-training skill.

Examples

Evaluating state-of-the-art MatPES-r2SCAN Foundation Models directly against f-block filtered analytical DFT data from the Materials Project:

bash
# Fetch 100 random r2SCAN structures from MP API (excluding Lanthanides/Actinides) into ./r2scan_data.json
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/examples/fetch_r2scan.py

# Benchmark MACE foundation potential
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/run_benchmark.py \
    --data_path r2scan_data.json \
    --model MACE-MATPES-R2SCAN-0 \
    --backend mace \
    --output research/2026-03-03_r2SCAN_benchmark/mace_results.json

# Plot the evaluation statistics
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/plot_benchmark.py \
    --results research/2026-03-03_r2SCAN_benchmark/mace_results.json \
    --output_dir research/2026-03-03_r2SCAN_benchmark/plots_mace
Resulting Parity Plots (Filtered R2SCAN Data)

Image: MACE-MATPES-R2SCAN-0 Parity Plot Image: CHGNet-MatPES-r2SCAN-2025.2.10-2.7M Parity Plot Image: M3GNet-MatPES-r2SCAN-v2025.1 Parity Plot Image: TensorNet-MatPES-r2SCAN-v2025.1 Parity Plot

Typical Combinations

  • Use mat-sample-pes-by-md to generate un-labeled configurations.
  • Use atomate2 MCP tools or VASP to evaluate configurations and produce a labeled dataset JSON.
  • Use ml-mlip-training if the benchmark metric thresholds are unsatisfactory.

Author: Bowen Deng Contact: GitHub @learningmatter-mit

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts) in skills/ml-mlip-benchmark of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/README.md
  • examples/chgnet_parity.png
  • examples/fetch_r2scan.py
  • examples/m3gnet_parity.png
  • examples/mace_parity.png
  • examples/tensornet_parity.png
  • scripts/plot_benchmark.py
  • scripts/run_benchmark.py

Open the folder on GitHubat commit 7f2d86d

Compare with similar skills

ML Mlip Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Mlip Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Mlip Benchmark this skilllearningmatter-mit/AtomisticSkills175—~1.7kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k62 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Fastmcp Client CLIPrefectHQ/fastmcp28k1 repos~823Automated safety check: PassApache-2.0
MemPalace Setup and OperationMemPalace/mempalace59k—~2.2kAutomated safety check: PassMIT
LangBot Plugin Developmentlangbot-app/LangBot18k—~3.9kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 62 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Fastmcp Client CLI

    PrefectHQ/fastmcp

    Query and invoke tools on MCP servers using fastmcp list and fastmcp call.

    28k GitHub starsUsed in 1 repo~823 tokens
    Agent WorkflowsAuto-check passed
  • Installs and configures MemPalace as a private local palace, a shared-brain hub or a client of an existing hub, including MCP registration and version-correct initialization.

    59k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • LangBot Plugin Development

    langbot-app/LangBot

    Guides building, debugging and testing LangBot plugins: components, SDK calls, README and locale rules, SDK pitfalls and WebSocket-based testing.

    18k GitHub stars~3.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Vexor CLI

    scarletkc/vexor

    Semantic file discovery via vexor. An agent skill from scarletkc/vexor.

    243 GitHub starsUsed in 3 repos~1.2k tokens
    AI & LLM EngineeringAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    175 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    175 GitHub stars~2k tokensUpdated yesterday
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    175 GitHub stars~4k tokensUpdated yesterday
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    175 GitHub stars~2.5k tokensUpdated yesterday
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    175 GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    175 GitHub stars~1.9k tokensUpdated yesterday
    Auto-check passed

Questions about ML Mlip Benchmark

What does ML Mlip Benchmark do?

Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots. ML Mlip Benchmark is an agent skill from learningmatter-mit/AtomisticSkills. Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.

How do I install ML Mlip Benchmark in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a claude-code`. Or copy the skill folder (skills/ml-mlip-benchmark in learningmatter-mit/AtomisticSkills) into .claude/skills/ml-mlip-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install ML Mlip Benchmark in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a codex`. Or copy the skill folder (skills/ml-mlip-benchmark in learningmatter-mit/AtomisticSkills) into .agents/skills/ml-mlip-benchmark in your project. Codex loads it when a task matches its description.

Can I use ML Mlip Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-mlip-benchmark, .gemini/skills/ml-mlip-benchmark, .github/skills/ml-mlip-benchmark and .opencode/skills/ml-mlip-benchmark in your project.

What does ML Mlip Benchmark need to run?

Going by SKILL.md and its folder, ML Mlip Benchmark needs Python for the scripts in its folder. Our summary lists: Python 3.

Does ML Mlip Benchmark access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is ML Mlip Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does ML Mlip Benchmark use?

ML Mlip Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Mlip Benchmark use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Mlip Benchmark?

Skills that share tags, products or a category with ML Mlip Benchmark: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), Fastmcp Client CLI (PrefectHQ/fastmcp, 28k stars) and MemPalace Setup and Operation (MemPalace/mempalace, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Mlip Benchmark?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 175 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 6, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.