Agent skill

ML Committee Uncertainty

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Quantify prediction uncertainty of MACE MLIPs using committee (ensemble) models; flag high-uncertainty structures for DFT verification.

MITAuto-check passed

Install ML Committee Uncertainty

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-committee-uncertainty -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills ml-committee-uncertainty --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-committee-uncertainty .claude/skills/ml-committee-uncertainty && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-committee-uncertainty
GitHub stars
176
Token cost
~2.3k tokens
SKILL.md length
865 words
Files
4 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Quantify prediction uncertainty of MACE MLIPs using committee (ensemble) models; flag high-uncertainty structures for DFT verification.

  • Works in 5 steps: Obtain Committee Models → Run Committee Inference → Interpret Results → …
  • Tasks that involve MLOps
  • SKILL.md covers Goal, Instructions, Constraints and References
  • Runs Python scripts from its folder

What it does

ML Committee Uncertainty is an agent skill from learningmatter-mit/AtomisticSkills. Quantify prediction uncertainty of MACE MLIPs using committee (ensemble) models; flag high-uncertainty structures for DFT verification.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts (for example `examples/mace-lipo4-committee/README.md`, `examples/mace-lipo4-committee/uncertainty_summary.json` and `scripts/run_committee_inference.py`).

It works with Model Context Protocol. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

When your agent uses it

  • Tasks that involve MLOps

Example prompts

  • “/ml-committee-uncertainty”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Obtain Committee Models
  2. Run Committee Inference
  3. Interpret Results
  4. Send High-Uncertainty Structures to DFT
  5. Register Threshold in the Model Registry

What it can do on your machine

Read from SKILL.md and the folder at commit 6257444. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Committee Uncertainty loads about 2.3k tokens when it runs. Until then it costs about 40 tokens; SKILL.md has 865 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~40
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 6257444, republished under its MIT licence (© learningmatter-mit). 865 words, ~2,279 tokens.

Download SKILL.mdSave it as .claude/skills/ml-committee-uncertainty/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
ml-committee-uncertainty
description
Quantify prediction uncertainty of MACE MLIPs using committee (ensemble) models; flag high-uncertainty structures for DFT verification.
metadata.category
machine-learning, materials, chemistry
metadata.venv
cpu, mlip

MACE Committee Model Uncertainty Quantification

<!-- mcp-tools-note -->

[!NOTE] Steps written server.tool are MCP tool calls: base.search_model_registry is the search_model_registry tool of the base server (mcp__base__search_model_registry, or mcp__plugin_atomistic-skills_base__search_model_registry when installed as a plugin). Without a connected server, run the same tools from the shell. Tools named in one command share a process, so a model loaded by load_model stays loaded:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python -m src.mcp_server.cli base search_model_registry key=value register_model key=value

Goal

To estimate the epistemic uncertainty of a MACE MLIP by running inference with a committee (ensemble) of independently trained models. Structures where the committee disagrees strongly (high energy or force variance) are flagged as candidates for DFT labelling, supporting active learning workflows and validating MLIP reliability in under-sampled regions of configuration space.

The uncertainty estimate is:

  • Energy uncertainty: standard deviation of predicted energies across committee members (meV/atom)
  • Force uncertainty: component-wise force RMSE of the across-committee standard deviations (meV/Å): take the sample standard deviation for every atom and Cartesian component, then take one root-mean-square over all 3N components. This is the conventional MLIP force-RMSE reduction used by MACE reporting.

[!WARNING] Energy std is only a valid disagreement signal when every committee member shares the same energy reference — the same training dataset, level of theory, and atomic reference energies (E0). If the committee instead pools several independently pretrained foundation models (e.g. different MACE-MP / MACE-OMAT / MACE-MATPES releases) rather than same-data/different-seed checkpoints, their absolute energies are not on a common scale: an energy-std ranking will mostly reflect per-model reference-energy offsets, not genuine epistemic disagreement. For this heterogeneous committee flavor, rank structures by force disagreement instead — forces are invariant to each model's arbitrary atomic reference energy, so they remain a reliable cross-model signal. See the heterogeneous-committee note under Constraints.

Instructions

1. Obtain Committee Models

A committee requires N ≥ 3 independently trained MACE checkpoints covering the same chemical system. There are two ways to obtain them:

Option A — Train with different random seeds (recommended for fine-tuned models)

Run ml-mace-finetune N times, varying only the random seed via the --seed flag in generate_mace_config.py. Save each checkpoint to a separate directory:

bash
# Run for seed=0, seed=1, seed=2 (at minimum)
for SEED in 0 1 2; do
    ${CLAUDE_SKILL_DIR}/../../venv/run mlip python ${CLAUDE_SKILL_DIR}/../ml-mace-finetune/scripts/generate_mace_config.py \
        --train-file ./mace_data/train.xyz \
        --valid-file ./mace_data/valid.xyz \
        --model MACE-MH-1 \
        --epochs 200 \
        --lr 1e-4 \
        --batch-size 4 \
        --freeze-backbone \
        --seed ${SEED} \
        --output-dir ./committee_models/seed_${SEED}
    ${CLAUDE_SKILL_DIR}/../../venv/run mlip mace_run_train \
        --config ./committee_models/seed_${SEED}/finetune_config.yaml
done

Option B — Use pre-existing models from the registry

bash
base.search_model_registry(
    chemical_system="Li-Fe-P-O",
    backend="mace",
)
# If multiple models exist for the same system, they can form a committee.
2. Run Committee Inference

Pass all checkpoint paths to the inference script. It will run each model independently and compute mean ± std across the committee.

bash
${CLAUDE_SKILL_DIR}/../../venv/run mlip python ${CLAUDE_SKILL_DIR}/scripts/run_committee_inference.py \
    --structures /path/to/structures_dir_or_file.cif \
    --models ./committee_models/seed_0/mace_finetuned.model \
             ./committee_models/seed_1/mace_finetuned.model \
             ./committee_models/seed_2/mace_finetuned.model \
    --output-dir ./uncertainty_results \
    --energy-threshold 10.0 \
    --force-threshold 200.0
# For MACE-MH foundation models, add:
#   --head omat_pbe

Key parameters:

ParameterDescriptionTypical value
--structuresPath to structure file, directory, or .xyz trajectory—
--modelsSpace-separated list of MACE checkpoint paths≥ 3 models
--energy-thresholdFlag if energy std > this value (meV/atom)5–20 meV/atom
--force-thresholdFlag if component-wise force RMSE > this value (meV/Å)100–300 meV/Å
--output-dirDirectory for results and plots—
--devicecuda or cpucuda
--headHead name for multi-head models (e.g. omat_pbe for MACE-MH-1)None
3. Interpret Results

The script produces:

  • uncertainty_summary.json — per-structure energy/force mean ± std
  • high_uncertainty_structures/ — .cif files of flagged structures requiring DFT
  • uncertainty_distribution.png — histogram of energy uncertainty across all structures

Thresholds for DFT flagging: There is no universal threshold. Empirically:

  • Energy std > 10 meV/atom is a conservative threshold suitable for phonon/stability work
  • Energy std > 5 meV/atom for high-accuracy MD (e.g., melting point, ionic conductivity)
  • Force RMSE > 200 meV/Å generally indicates the configuration is poorly represented in training data

[!TIP] Run this skill on your MD trajectory after a few nanoseconds to check whether the MLIP remains in-distribution throughout the simulation. High uncertainty at late simulation times suggests the trajectory has drifted into unexplored configuration space.

Show full SKILL.md (311 more words)Show less
4. Send High-Uncertainty Structures to DFT

Pass the flagged structures to the DFT labelling pipeline. See the mat-sample-pes-by-md skill for context on when and how to label new structures.

bash
# High-uncertainty structures needing DFT labels are in:
ls ./uncertainty_results/high_uncertainty_structures/
# → Pass these to atomate2 DFT workflow for labelling
5. Register Threshold in the Model Registry

After determining an appropriate uncertainty threshold for your system, update the registry so future tasks can apply the same criterion automatically:

bash
base.register_model(
    checkpoint_path="./committee_models/seed_0/mace_finetuned.model",
    chemical_system="Li-Fe-P-O",
    backend="mace",
    base_model="MACE-MH-1",
    notes="Committee of 3 models (seed 0/1/2). Use energy_std > 10 meV/atom as DFT flag threshold.",
    tags_json='["battery", "committee"]',
)

Constraints

  • Minimum committee size: Use at least 3 models. Two models can give misleading std estimates; 5+ models provide more robust uncertainty quantification.
  • Identical architecture: All committee members must share the same base model architecture and chemical elements (same --model flag during fine-tuning). Different architectures cannot be meaningfully ensembled.
  • Same data, different seeds (fine-tuned committees): When building a committee via Option A above, members should be trained on the same dataset with different random seeds. Mixing training datasets within this committee flavor introduces epistemic uncertainty from data mismatch, not model uncertainty, and confounds the estimate.
  • Heterogeneous / multi-foundation-model committees: A committee does not have to be same-data/different-seed — a fixed pool of independently pretrained foundation models (e.g. MACE-MP-0, MACE-MP-0b, MACE-MP-0b2, MACE-OMAT-0) is also a valid ensemble for epistemic UQ, and can surface genuine out-of-distribution structures (e.g. an element in an unusual coordination environment) that a same-data seed ensemble would miss. Because these models differ in training data and reference level of theory, however, their energy predictions are not on a common scale — always rank this committee flavor by force disagreement, never by energy std (see the warning under Goal).
  • Environment: This script requires the mlip environment.

References

  • Batatia et al., "MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force Fields", NeurIPS, 2022.
  • Musil et al., "Physics-Inspired Structural Representations for Molecules and Materials", Chem. Rev., 2021. (Committee model UQ)
  • Schran et al., "Committee Neural Network Potentials Control Generalization Errors and Enable Active Learning", J. Chem. Phys., 2020.

Author: Yu Yao Contact: GitHub @AI4SciDisc

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (scripts) in skills/ml-committee-uncertainty of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/mace-lipo4-committee/README.md
  • examples/mace-lipo4-committee/uncertainty_summary.json
  • scripts/run_committee_inference.py

Open the folder on GitHubat commit 6257444

Compare with similar skills

ML Committee Uncertainty next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Committee Uncertainty compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Committee Uncertainty this skilllearningmatter-mit/AtomisticSkills176—~2.3kAutomated safety check: PassMIT
Managing MCP IndexComfy-Org/workflow_templates1.3k—~1.9kAutomated safety check: NotesMIT
Model Registryartokun/comfyui-mcp795—~2.1kAutomated safety check: PassMIT
Validate Profileindranilbanerjee/digital-marketing-pro8551 repos~3.3kAutomated safety check: NotesMIT
Find AI Consultancyjeremylongshore/tons-of-skills-marketplace2.8k—~4kAutomated safety check: NotesMIT
Find Software Developerjeremylongshore/tons-of-skills-marketplace2.8k—~3.7kAutomated safety check: NotesMIT

Similar skills

  • Managing MCP Index

    Comfy-Org/workflow_templates

    Builds and maintains templates/index.mcp.json for Comfy Cloud MCP tools.

    1.3k GitHub stars~1.9k tokensUpdated today
    Frontend & DesignAuto-check: notes
  • Model Registry

    artokun/comfyui-mcp

    Curated download URLs and target directories, organized by family (Flux, WAN, LTX, Qwen, Z-Image, SD15/SDXL), for every model the comfyui-mcp skills reference, covering checkpoints, VAEs, text…

    795 GitHub stars~2.1k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Validate Profile

    indranilbanerjee/digital-marketing-pro

    Read-only health check that a brand profile is production-ready: required fields, voice and audience completeness, guardrails, compliance-jurisdiction coverage, connector configuration and MCP…

    855 GitHub starsUsed in 1 repo~3.3k tokens
    AI & LLM EngineeringAuto-check: notes
  • Find AI Consultancy

    jeremylongshore/tons-of-skills-marketplace

    A skill your agent uses whenever the user wants to find, shortlist, vet, or enrich US AI/ML/data consulting firms (consultancies) — AI/ML development, MLOps, generative AI / LLM apps (RAG, chatbots…

    2.8k GitHub stars~4k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Find Software Developer

    jeremylongshore/tons-of-skills-marketplace

    A skill your agent uses whenever the user wants to find, shortlist, vet, or enrich US software development firms — custom software, web development, mobile app development, backend/API development…

    2.8k GitHub stars~3.7k tokensUpdated today
    Frontend & DesignAuto-check: notes
  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    176 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    176 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    176 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    176 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    176 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    176 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Questions about ML Committee Uncertainty

What does ML Committee Uncertainty do?

Quantify prediction uncertainty of MACE MLIPs using committee (ensemble) models; flag high-uncertainty structures for DFT verification. ML Committee Uncertainty is an agent skill from learningmatter-mit/AtomisticSkills. Quantify prediction uncertainty of MACE MLIPs using committee (ensemble) models; flag high-uncertainty structures for DFT verification.

When should I use ML Committee Uncertainty?

ML Committee Uncertainty fits situations like: tasks that involve MLOps.

How do I install ML Committee Uncertainty in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-committee-uncertainty -a claude-code`. Or copy the skill folder (skills/ml-committee-uncertainty in learningmatter-mit/AtomisticSkills) into .claude/skills/ml-committee-uncertainty in your project. Claude Code loads it when a task matches its description.

How do I install ML Committee Uncertainty in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-committee-uncertainty -a codex`. Or copy the skill folder (skills/ml-committee-uncertainty in learningmatter-mit/AtomisticSkills) into .agents/skills/ml-committee-uncertainty in your project. Codex loads it when a task matches its description.

Can I use ML Committee Uncertainty in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-committee-uncertainty -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-committee-uncertainty, .gemini/skills/ml-committee-uncertainty, .github/skills/ml-committee-uncertainty and .opencode/skills/ml-committee-uncertainty in your project.

What does ML Committee Uncertainty need to run?

Going by SKILL.md and its folder, ML Committee Uncertainty needs Python for the scripts in its folder. Our summary lists: Python 3.

Does ML Committee Uncertainty access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is ML Committee Uncertainty safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does ML Committee Uncertainty use?

ML Committee Uncertainty is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Committee Uncertainty use?

About 2.3k tokens (SKILL.md is roughly 9.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Committee Uncertainty?

Skills that share tags, products or a category with ML Committee Uncertainty: Managing MCP Index (Comfy-Org/workflow_templates, 1.3k stars), Model Registry (artokun/comfyui-mcp, 795 stars), Validate Profile (indranilbanerjee/digital-marketing-pro, 855 stars) and Find AI Consultancy (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Committee Uncertainty?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 176 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 7, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.