Agent skill

ML Cluster Expansion

by learningmatter-mit in learningmatter-mit/AtomisticSkills

train a Cluster Expansion (CE) for lattice-based Monte Carlo simulation of disordered materials.

MITAuto-check passedAgent Workflows

Install ML Cluster Expansion

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-cluster-expansion -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills ml-cluster-expansion --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-cluster-expansion .claude/skills/ml-cluster-expansion && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-cluster-expansion
GitHub stars
176
Token cost
~2.1k tokens
SKILL.md length
760 words
Files
13 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

train a Cluster Expansion (CE) for lattice-based Monte Carlo simulation of disordered materials.

  • Works in 6 steps: Prepare the Primordial Structure → Iteration 0 (Ordered Structure Sampling) → Label Structures (Agent/MCP) → …
  • Tasks that involve MCP servers
  • SKILL.md covers Goal, Workflow Overview, Energy Format for Training Data and Constraints & Tips
  • Runs Python scripts from its folder

What it does

ML Cluster Expansion is an agent skill from learningmatter-mit/AtomisticSkills. train a Cluster Expansion (CE) for lattice-based Monte Carlo simulation of disordered materials.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including scripts (for example `examples/AgPd_DFT_CE/README.md`, `examples/CuAg_CE/README.md` and `examples/CuAg_CE/cluster_expansion.json`).

It sits in Agent Workflows, covering MCP servers. It works with Model Context Protocol. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

When your agent uses it

  • Tasks that involve MCP servers

Example prompts

  • “/ml-cluster-expansion”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Prepare the Primordial Structure
  2. Iteration 0 (Ordered Structure Sampling)
  3. Label Structures (Agent/MCP)
  4. Train CE
  5. Direct Feature Matrix Fitting (Optional)
  6. Active Learning Loop (Optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 6257444. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Cluster Expansion loads about 2.1k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 760 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 6257444, republished under its MIT licence (© learningmatter-mit). 760 words, ~2,063 tokens.

Download SKILL.mdSave it as .claude/skills/ml-cluster-expansion/SKILL.md (or your agent's skills folder). This skill also uses 12 other files; get the full folder from GitHub.
name
ml-cluster-expansion
description
train a Cluster Expansion (CE) for lattice-based Monte Carlo simulation of disordered materials.
metadata.category
machine-learning, materials
metadata.venv
cpu, mlip

Cluster Expansion

<!-- mcp-tools-note -->

[!NOTE] Steps written server.tool are MCP tool calls: smol.train_cluster_expansion is the train_cluster_expansion tool of the smol server (mcp__smol__train_cluster_expansion, or mcp__plugin_atomistic-skills_smol__train_cluster_expansion when installed as a plugin). Without a connected server, run the same tools from the shell. Tools named in one command share a process, so a model loaded by load_model stays loaded:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python -m src.mcp_server.cli smol train_cluster_expansion key=value run_monte_carlo key=value
${CLAUDE_SKILL_DIR}/../../venv/run mlip python -m src.mcp_server.cli mace load_model key=value relax_structure key=value

Goal

To automatically build and refine a Cluster Expansion (CE) model for a disordered material system using an Agent-driven iterative workflow that leverages MCP tools for efficient training, sampling, and labeling.

Workflow Overview

  1. Preparation: Generate a disordered primordial structure.
  2. Iteration 0: systematic enumeration to generate initial structures.
  3. Labeling: Relax structures with an MLIP (e.g., MACE, CHGNet) via MCP.
  4. Training: Train the CE model using smol.train_cluster_expansion.
  5. Sampling: Run MC with smol.run_monte_carlo to explore configuration space.
  6. Selection: Extract structures from MC, compute features, and select novel configurations.
  7. Loop: Repeat labeling, training, and sampling until convergence.

Step 1: Prepare the Primordial Structure

Use prepare_disordered.py to handle symmetry refinement and disorder creation. It is highly recommended to save the primordial structure as a JSON file to preserve exact occupancy and species information, avoiding "unrecognized species" errors during matching.

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/prepare_disordered.py \
    input_structure.cif \
    Li \
    0.5 \
    -o primordial.cif

Step 2: Iteration 0 (Ordered Structure Sampling)

Generate an initial set of structures using systematic enumeration and D-optimality via the MCP tool. It is recommended to generate around 1000 structures to ensure high coverage of the configuration space.

python
# MCP Tool: smol.sample_ordered_structures
result = smol.sample_ordered_structures(
    disordered_structure="primordial.cif",
    cutoffs={2: 5.0, 3: 4.0},
    num_structures=1000,
    target_num_sites=32,
    output_dir="./ce_project/iter_0/to_label"
)

[!NOTE] Training Workflow: You can perform a Simple Training by only using the initially sample_ordered_structures set. This is often sufficient for basic property predictions. For high-accuracy ground-state exploration, you should continue with the Active Learning loop (MC sampling and iterative refinement).

[!TIP] This tool works for arbitrary disorder, including multi-element alloys (e.g., Co-Ni-Cr) and multi-sublattice systems (e.g., (Li-Na)(Cl-Br)).


Step 3: Label Structures (Agent/MCP)

Use the appropriate MLIP MCP tool to relax the structures.

[!IMPORTANT] Fixed Cell: Always set relax_cell=False during MLIP relaxation for cluster expansion training. The cluster expansion model is built on a fixed lattice. Fixed Lattice: Ensure the supercell matrix used for enumeration/sampling matches your expectation.

Example (MACE):

python
mace.load_model(model_name="MACE-OMAT-0-small", device="cuda")
mace.relax_structure(
    structure_data="./ce_project/iter_0/to_label",
    fmax=0.02,
    relax_cell=False,
    output_dir="./ce_project/iter_0/results"
)

Step 4: Train CE

The smol.train_cluster_expansion tool can directly accept the directory containing relaxation results (from Step 3). It will automatically find and aggregate the training data.

python
result = smol.train_cluster_expansion(
    disordered_structure="primordial.cif",
    training_data="ce_project/iter_0/results",
    cutoffs={2: 5.0, 3: 4.0},
    ce_file="ce_project/cluster_expansion.json"
)

[!NOTE] Training Workflow: You can perform a Simple Training by only using the initially sample_ordered_structures set. This is often sufficient for basic property predictions. For high-accuracy ground-state exploration, you should continue with the Active Learning loop (MC sampling and iterative refinement).

Show full SKILL.md (328 more words)Show less
Step 5: Direct Feature Matrix Fitting (Optional)

In some scenarios, you may have pre-computed feature matrices and energies (e.g., from literature or external workflows). You can directly fit these without building a ClusterSubspace first, using the smol.fit_feature_matrix tool. This is also how you can utilize advanced techniques like Sparse Group Lasso (SGL).

python
# MCP Tool: smol.fit_feature_matrix
result = smol.fit_feature_matrix(
    feature_matrix_path="fm.npy",
    energies_path="e.npy",
    groups_path="groups.npy", # Required for sgl
    fit_method="sgl",         # 'ls', 'lasso', 'ridge', or 'sgl'
    alpha=0.002,
    lambda_mixing=0.5,
    test_size=0.2
)

[!TIP] Sparse Group Lasso (SGL) is highly recommended for complex, high-component systems as it naturally selects important cluster groups while maintaining sparsity. It requires a groups.npy file that assigns each feature column to a cluster orbit group.


Step 6: Active Learning Loop (Optional)
A. Run Monte Carlo Sampling
python
mc_result = smol.run_monte_carlo(
    supercell_matrix=[[2,0,0], [0,2,0], [0,0,2]],
    temperature=2000,
    steps=100000,
    ce_file="./ce_project/cluster_expansion.json",
    trajectory_file="./ce_project/iter_1/mc_trajectory.h5"
)
B. Extract & Select Candidates

Extract structures from the trajectory. The skill script handles dimension squeeze for single-sample trajectories.

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/extract_mc_structures.py \
    --trajectory_file ./ce_project/iter_1/mc_trajectory.h5 \
    --cluster_expansion ./ce_project/cluster_expansion.json \
    --output_dir ./ce_project/iter_1/to_label \
    --num_structures 20 \
    --strategy random

[!IMPORTANT] Loop Back: After extracting these new candidate structures, go back to Step 3 to relax and label them. Then, add these new results to your training set and repeat Step 4 to refine your Cluster Expansion.

Energy Format for Training Data

[!IMPORTANT] Use Total Energy (Extensive Property) Training structures should include the total energy for the entire supercell (not per-atom energy). Smol automatically normalizes energies by the primitive cell size during training via get_property_vector('energy', normalize=True).

Example: For a 32-atom supercell with total DFT energy of -160.0 eV:

  • ✅ Correct: {"structure": {...}, "energy": -160.0}
  • ❌ Incorrect: {"structure": {...}, "energy": -5.0} (per-atom)

Formation Energy: If you want to train on formation energies instead of total energies, calculate the formation energy for each structure first, then provide it as the total (extensive) energy value. Smol will still normalize by primitive cell size.

Constraints & Tips

  • Primordial Format: Use .cif for primordial structures. Ensure partial occupancies are correctly defined.
  • Supercell Matrices: If automatic matching fails (e.g. "Could not determine supercell matrix"), provide sc_matrix explicitly in the training data JSON.
  • Fixed Cell: Always set relax_cell=False to maintain the cluster expansion mapping.
  • Convergence: Target LOOCV < 10 meV/atom for high-accuracy models.
  • Simulations: For finite-temperature MC simulations using your trained model, refer to the mat-disorder skill.

Author: Bowen Deng Contact: GitHub @learningmatter-mit

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 12 other files (scripts) in skills/ml-cluster-expansion of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/AgPd_DFT_CE/README.md
  • examples/CuAg_CE/README.md
  • examples/CuAg_CE/cluster_expansion.json
  • examples/pdptag_ternary_surface/README.md
  • examples/pdptag_ternary_surface/analyze_mc.py
  • examples/pdptag_ternary_surface/mc_trajectory_final.cif
  • examples/pdptag_ternary_surface/mc_trajectory_initial.cif
  • examples/pdptag_ternary_surface/parse_atat.py
  • examples/pdptag_ternary_surface/parse_primordial.py
  • examples/pdptag_ternary_surface/primordial.cif
  • scripts/extract_mc_structures.py
  • scripts/prepare_disordered.py

Open the folder on GitHubat commit 6257444

Compare with similar skills

ML Cluster Expansion next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Cluster Expansion compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Cluster Expansion this skilllearningmatter-mit/AtomisticSkills176—~2.1kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
MCP Integration for Pluginsanthropics/claude-plugins-official38k11 repos~3.1kAutomated safety check: PassApache-2.0
Fastmcp Client CLIPrefectHQ/fastmcp28k1 repos~823Automated safety check: PassApache-2.0
Crush Configurationcharmbracelet/crush29k—~3.7kAutomated safety check: PassCustom licence

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • MCP Integration for Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.

    38k GitHub starsUsed in 11 repos~3.1k tokens
    Agent WorkflowsAuto-check passed
  • Fastmcp Client CLI

    PrefectHQ/fastmcp

    Query and invoke tools on MCP servers using fastmcp list and fastmcp call.

    28k GitHub starsUsed in 1 repo~823 tokens
    Agent WorkflowsAuto-check passed
  • Crush Configuration

    charmbracelet/crush

    Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.

    29k GitHub stars~3.7k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Context Mode Output Sandbox

    mksglu/context-mode

    Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.

    26k GitHub stars~4.1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    176 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    176 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    176 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    176 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    176 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    176 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Categories

Questions about ML Cluster Expansion

What does ML Cluster Expansion do?

train a Cluster Expansion (CE) for lattice-based Monte Carlo simulation of disordered materials. ML Cluster Expansion is an agent skill from learningmatter-mit/AtomisticSkills. train a Cluster Expansion (CE) for lattice-based Monte Carlo simulation of disordered materials.

When should I use ML Cluster Expansion?

ML Cluster Expansion fits situations like: tasks that involve MCP servers.

How do I install ML Cluster Expansion in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-cluster-expansion -a claude-code`. Or copy the skill folder (skills/ml-cluster-expansion in learningmatter-mit/AtomisticSkills) into .claude/skills/ml-cluster-expansion in your project. Claude Code loads it when a task matches its description.

How do I install ML Cluster Expansion in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-cluster-expansion -a codex`. Or copy the skill folder (skills/ml-cluster-expansion in learningmatter-mit/AtomisticSkills) into .agents/skills/ml-cluster-expansion in your project. Codex loads it when a task matches its description.

Can I use ML Cluster Expansion in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-cluster-expansion -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-cluster-expansion, .gemini/skills/ml-cluster-expansion, .github/skills/ml-cluster-expansion and .opencode/skills/ml-cluster-expansion in your project.

What does ML Cluster Expansion need to run?

Going by SKILL.md and its folder, ML Cluster Expansion needs Python for the scripts in its folder. Our summary lists: Python 3.

Does ML Cluster Expansion access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is ML Cluster Expansion safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does ML Cluster Expansion use?

ML Cluster Expansion is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Cluster Expansion use?

About 2.1k tokens (SKILL.md is roughly 8.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Cluster Expansion?

Skills that share tags, products or a category with ML Cluster Expansion: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and Fastmcp Client CLI (PrefectHQ/fastmcp, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Cluster Expansion?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 176 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 7, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.