Install the "ml-mlip-benchmark" agent skill from https://github.com/learningmatter-mit/AtomisticSkills/tree/main/skills/ml-mlip-benchmark into .claude/skills/ml-mlip-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-mlip-benchmark", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "ml-mlip-benchmark" agent skill from https://github.com/learningmatter-mit/AtomisticSkills/tree/main/skills/ml-mlip-benchmark into .agents/skills/ml-mlip-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-mlip-benchmark", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "ml-mlip-benchmark" agent skill from https://github.com/learningmatter-mit/AtomisticSkills/tree/main/skills/ml-mlip-benchmark into .cursor/skills/ml-mlip-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-mlip-benchmark", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "ml-mlip-benchmark" agent skill from https://github.com/learningmatter-mit/AtomisticSkills/tree/main/skills/ml-mlip-benchmark into .gemini/skills/ml-mlip-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-mlip-benchmark", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "ml-mlip-benchmark" agent skill from https://github.com/learningmatter-mit/AtomisticSkills/tree/main/skills/ml-mlip-benchmark into .github/skills/ml-mlip-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-mlip-benchmark", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "ml-mlip-benchmark" agent skill from https://github.com/learningmatter-mit/AtomisticSkills/tree/main/skills/ml-mlip-benchmark into .opencode/skills/ml-mlip-benchmark/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-mlip-benchmark", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
ml-mlip-benchmark
GitHub stars
175
Token cost
~1.7k tokens
SKILL.md length
655 words
Files
9 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT
At a glance
Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.
Works in 4 steps: Run Benchmark metrics → Generate Parity Plots → Reconcile units before comparing anything → …
SKILL.md covers Prerequisites, Instructions, Examples and Typical Combinations
Runs Python scripts from its folder
What it does
ML Mlip Benchmark is an agent skill from learningmatter-mit/AtomisticSkills. Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 10 other files, including scripts (for example `examples/README.md`, `examples/fetch_r2scan.py` and `scripts/plot_benchmark.py`).
It works with Model Context Protocol and Python. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.
Example prompts
“/ml-mlip-benchmark”
Requirements
Python 3
Workflow steps
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7f2d86d. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Network
Links to these hosts (documentation or services it may open):
github.com
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
ML Mlip Benchmark loads about 1.7k tokens when it runs. Until then it costs about 35 tokens; SKILL.md has 655 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~35
When it runs· the whole SKILL.md, loaded when a task matches
~1.7k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
Download SKILL.mdSave it as .claude/skills/ml-mlip-benchmark/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
ml-mlip-benchmark
description
Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.
[!NOTE]
Steps written server.tool are MCP tool calls: mace.load_model is the load_model
tool of the mace server (mcp__mace__load_model, or
mcp__plugin_atomistic-skills_mace__load_model when installed as a plugin).
Without a connected server, run the same tools from the shell. Tools named in
one command share a process, so a model loaded by load_model stays loaded:
This skill evaluates the accuracy of a given MLIP against an existing ground-truth dataset (e.g., DFT calculations or a higher-fidelity foundation potential). It computes the Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) for both energy (per atom) and atomic forces, and optionally stress. It also generates parity plots for visual inspection of the model's correlation.
Prerequisites
Model Loaded: An MLIP must be currently active via a load_model MCP tool call (e.g., mace.load_model, fairchem.load_model, matgl.load_model).
Labeled Data: A JSON dataset where each entry contains a structural dictionary under "structure", along with scalar/vector ground truth values for "energy", "forces", and optionally "stress". This is identical to the format used in ml-mlip-training. (Data can be generated using Atomate2 MongoDB queries or MD sampling + labeling).
Instructions
1. Run Benchmark metrics
Use the ${CLAUDE_SKILL_DIR}/scripts/run_benchmark.py script to perform inference across the dataset and compute global error metrics.
Environment requirement: This script instantiates the MLIP models directly, so it must run in the uv project that provides the backend: venv/mlip for MACE and MatGL, venv/fairchem for FairChem.
Note: The script utilizes src.utils.mlips.loader.load_wrapper to abstract backend details.
2. Generate Parity Plots
Once run_benchmark.py finishes, it writes a comprehensive JSON file containing original targets alongside the model's predictions and numerical metrics. Visualize these using the plotting script.
Environment requirement: venv/cpu is enough for the plotting script.
bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/plot_benchmark.py \
--results <path_to_benchmark_results.json> \
--output_dir <path_to_save_plots>
This generates energy_parity.png, forces_parity.png, and (if stress was present) stress_parity.png.
Show full SKILL.md (340 more words)Show less
3. Reconcile units before comparing anything
A benchmark subtracts two numbers that came from different software, so a unit or
sign mismatch shows up as a large "model error" that is not a model error at all.
Energy (eV) and forces (eV/Å) agree across every backend here; stress does not.
Settle it before computing a single metric -- see
general-property-units for the full tables.
The three traps, in order of how often they bite:
MatGL returns GPa, not eV/ų.matgl.ext.ase.PESCalculator defaults to
stress_unit="GPa", and Potential.forward returns GPa as well, so MatGL is not
a drop-in ASE calculator. Pass PESCalculator(potential=model, stress_unit="eV/A3").
Getting this wrong is a factor of 160.21766208.
Raw model output != ASE calculator output.CHGNetCalculator converts GPa to
eV/ų on the way out (stress_weight, default 1/160.21766208); MACE and
FAIRChem convert nothing because their models already emit eV/ų. Know which
layer you are reading.
DFT labels usually carry the opposite sign. VASP reports stress
compressive-positive in kB; ASE and every MLIP here are tensile-positive in
eV/ų. Converting VASP labels to ASE convention is
eV/A3 = -kB / 1602.1766208.
Sanity check that costs nothing: take a structure you have compressed, and confirm
the diagonal stress is negative in ASE convention. If it is positive, you have a
sign convention crossed somewhere.
4. Interpret Results
When presenting the plotted benchmarks to the user, consult the following rough heuristics for MLIP performance:
Energy MAE: Excellent (< 5 meV/atom), Good (5-20 meV/atom), Poor (> 50 meV/atom)
ML Mlip Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
ML Mlip Benchmark compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
ML Mlip Benchmark this skilllearningmatter-mit/AtomisticSkills
Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.
Installs and configures MemPalace as a private local palace, a shared-brain hub or a client of an existing hub, including MCP registration and version-correct initialization.
Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.
Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots. ML Mlip Benchmark is an agent skill from learningmatter-mit/AtomisticSkills. Benchmark MLIP accuracy against a labeled dataset — compute MAE/RMSE for energy/atom and forces, and generate parity plots.
How do I install ML Mlip Benchmark in Claude Code?
Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a claude-code`. Or copy the skill folder (skills/ml-mlip-benchmark in learningmatter-mit/AtomisticSkills) into .claude/skills/ml-mlip-benchmark in your project. Claude Code loads it when a task matches its description.
How do I install ML Mlip Benchmark in Codex?
Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a codex`. Or copy the skill folder (skills/ml-mlip-benchmark in learningmatter-mit/AtomisticSkills) into .agents/skills/ml-mlip-benchmark in your project. Codex loads it when a task matches its description.
Can I use ML Mlip Benchmark in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-mlip-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-mlip-benchmark, .gemini/skills/ml-mlip-benchmark, .github/skills/ml-mlip-benchmark and .opencode/skills/ml-mlip-benchmark in your project.
What does ML Mlip Benchmark need to run?
Going by SKILL.md and its folder, ML Mlip Benchmark needs Python for the scripts in its folder. Our summary lists: Python 3.
Does ML Mlip Benchmark access the network?
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Is ML Mlip Benchmark safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
What licence does ML Mlip Benchmark use?
ML Mlip Benchmark is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does ML Mlip Benchmark use?
About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to ML Mlip Benchmark?
Skills that share tags, products or a category with ML Mlip Benchmark: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), Fastmcp Client CLI (PrefectHQ/fastmcp, 28k stars) and MemPalace Setup and Operation (MemPalace/mempalace, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains ML Mlip Benchmark?
learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 175 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 6, 2026.