Agent skill

ML Bayesian Optimization

by learningmatter-mit in learningmatter-mit/AtomisticSkills

Iteratively optimize expensive black-box objectives — such as materials properties, experimental yields, or simulation outputs — by learning from past evaluations to select the most promising next…

MITAuto-check passedAgent Workflows

Install ML Bayesian Optimization

skills CLI
$ npx skills add learningmatter-mit/AtomisticSkills --skill ml-bayesian-optimization -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install learningmatter-mit/AtomisticSkills ml-bayesian-optimization --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/learningmatter-mit/AtomisticSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ml-bayesian-optimization .claude/skills/ml-bayesian-optimization && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-bayesian-optimization
GitHub stars
176
Token cost
~2.3k tokens
SKILL.md length
846 words
Files
32 (incl. scripts)
Skills in repo
129
Repo updated
First seen
Licence
MIT

At a glance

Iteratively optimize expensive black-box objectives — such as materials properties, experimental yields, or simulation outputs — by learning from past evaluations to select the most promising next…

  • Works in 5 steps: Define the Search Space → Initialize with Quasi-Random Sobol Samples → Evaluate Candidates Using MCP Tools → …
  • Agent Workflows work in your project
  • SKILL.md covers Goal, Instructions, Examples and Constraints, plus 1 more section

What it does

ML Bayesian Optimization is an agent skill from learningmatter-mit/AtomisticSkills. Iteratively optimize expensive black-box objectives — such as materials properties, experimental yields, or simulation outputs — by learning from past evaluations to select the most promising next candidates.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 34 other files, including scripts (for example `examples/branin-function/README.md`, `examples/branin-function/campaign_state.json` and `examples/branin-function/search_space.yaml`).

It sits in Agent Workflows. It works with Model Context Protocol and Python. The repository describes itself as: Integrating AtomisticSkills into Agentic IDEs (Cursor, Claude Code, Codex, Google Antigravity, Hermes Agent, etc). The licence is MIT.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “/ml-bayesian-optimization”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Define the Search Space
  2. Initialize with Quasi-Random Sobol Samples
  3. Evaluate Candidates Using MCP Tools
  4. BO Loop — Suggest Next Candidates
  5. Convergence Check and Analysis

What it can do on your machine

Read from SKILL.md and the folder at commit 6257444. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • doi.org
    • arxiv.org
    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Bayesian Optimization loads about 2.3k tokens when it runs. Until then it costs about 58 tokens; SKILL.md has 846 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~58
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from learningmatter-mit/AtomisticSkills at commit 6257444, republished under its MIT licence (© learningmatter-mit). 846 words, ~2,324 tokens.

Download SKILL.mdSave it as .claude/skills/ml-bayesian-optimization/SKILL.md (or your agent's skills folder). This skill also uses 31 other files; get the full folder from GitHub.
name
ml-bayesian-optimization
description
Iteratively optimize expensive black-box objectives — such as materials properties, experimental yields, or simulation outputs — by learning from past evaluations to select the most promising next candidates.
metadata.category
machine-learning
metadata.venv
cpu, mlip

Bayesian Optimization

<!-- mcp-tools-note -->

[!NOTE] Steps written server.tool are MCP tool calls: mace.relax_structure is the relax_structure tool of the mace server (mcp__mace__relax_structure, or mcp__plugin_atomistic-skills_mace__relax_structure when installed as a plugin). Without a connected server, run the same tools from the shell. Tools named in one command share a process, so a model loaded by load_model stays loaded:

bash
${CLAUDE_SKILL_DIR}/../../venv/run mlip python -m src.mcp_server.cli mace relax_structure key=value predict_structure key=value
${CLAUDE_SKILL_DIR}/../../venv/run mlip python -m src.mcp_server.cli matgl relax_structure key=value predict_bandgap key=value
${CLAUDE_SKILL_DIR}/../../venv/run cpu python -m src.mcp_server.cli atomate2 run_atomate2_vasp_calculation key=value
${CLAUDE_SKILL_DIR}/../../venv/run cpu python -m src.mcp_server.cli base visualize_structure key=value

Goal

Efficiently find the optimal input parameters (e.g., alloy composition, simulation hyperparameters, process conditions) that minimize or maximize one or more expensive black-box objectives (e.g., formation energy, bandgap, elastic modulus) using Bayesian Optimization (BO). BO builds a probabilistic surrogate model (Gaussian Process) over the objective landscape and uses an acquisition function to intelligently select the next most informative experiments, minimizing the number of expensive evaluations required.

  • Single-objective: Expected Improvement (EI) maximized via multi-start L-BFGS-B.
  • Multi-objective: ParEGO — random Chebyshev scalarization with independent GPs, one weight vector per batch element, naturally steering candidates toward different Pareto-front regions.

Instructions

Step 1: Define the Search Space

Create a search_space.yaml in the research directory. Use the template at resources/search_space_template.yaml as a starting point:

yaml
# research_dir/search_space.yaml

parameters:
  # Continuous range parameter
  - name: x_Fe
    type: range
    bounds: [0.0, 1.0]
    value_type: float

  # Integer range parameter
  - name: supercell_size
    type: range
    bounds: [2, 6]
    value_type: int

objectives:
  # Single-objective: minimize formation energy
  - name: formation_energy_eV_atom
    minimize: true

  # Multi-objective: additionally maximize bandgap (uncomment to enable)
  # - name: bandgap_eV
  #   minimize: false

Guidelines:

  • Use type: range for continuous or integer parameters with known bounds. Only range parameters are passed to the GP surrogate.
  • choice and fixed types are recorded in output CSVs but not optimized. Run separate campaigns per discrete choice.
  • Choose bounds informed by domain knowledge; avoid unnecessarily wide ranges.
Step 2: Initialize with Quasi-Random Sobol Samples

Generate an initial space-filling design using Sobol sequences. Use a power-of-2 batch_size (4, 8, 16, …) for optimal Sobol balance:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/suggest_candidates.py \
    --config research_dir/search_space.yaml \
    --batch_size 8 \
    --output research_dir/candidates_round_0.csv \
    --output_dir research_dir/

With no --results provided (or fewer than 2 × N_range_params evaluated points), the script automatically uses Sobol initialization. The output CSV lists candidate parameter values only — it does not evaluate the objective.

Step 3: Evaluate Candidates Using MCP Tools

For each row in candidates_round_0.csv, call the appropriate MCP tool to evaluate the objective and record results. The script plays no role in this step.

Objective typeRecommended MCP tool
Energy / formation energymace.relax_structure or matgl.relax_structure
Bandgapmatgl.predict_bandgap
Arbitrary propertymace.predict_structure
DFT referenceatomate2.run_atomate2_vasp_calculation

Save all results to evaluated.csv — one row per candidate, parameter columns plus objective column(s):

x_Fe,formation_energy_eV_atom
0.12,-0.087
0.45,-0.143
0.73,-0.091
Step 4: BO Loop — Suggest Next Candidates

With evaluated results, fit the GP surrogate and suggest the next batch:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/suggest_candidates.py \
    --config research_dir/search_space.yaml \
    --results research_dir/evaluated.csv \
    --batch_size 4 \
    --output research_dir/candidates_round_1.csv \
    --output_dir research_dir/

The script will:

  1. Normalize all inputs to [0, 1] per dimension.
  2. Fit a GP with Matern-5/2 kernel (scikit-learn, normalize_y=True).
  3. Single-objective: maximize EI via multi-start L-BFGS-B with diversity enforcement.
  4. Multi-objective: fit one GP per batch element with a random Chebyshev weight vector (ParEGO).
  5. Write the next batch_size candidates to the output CSV.
  6. Append round metadata to campaign_state.json in --output_dir.

Repeat Steps 3–4 until the budget is exhausted or convergence is observed.

Show full SKILL.md (371 more words)Show less
Step 5: Convergence Check and Analysis
bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/plot_bo_results.py \
    --results research_dir/evaluated.csv \
    --config research_dir/search_space.yaml \
    --output_dir research_dir/

This generates:

  • convergence_curve.png — best observed value vs. evaluation number (flat line = converged)
  • parameter_importance.png — scatter of each parameter vs. objective across all evaluations
  • pareto_front.png — (multi-objective only) non-dominated Pareto front
  • gp_model_1d.png — (single range parameter only) GP mean ± 2σ and EI landscape
  • gp_model_2d.png — (exactly two range parameters) GP mean, uncertainty, and EI as contour maps

Convergence criteria (manual inspection):

  • Single-objective: best value improves by < 1% over the last 5 rounds.
  • Multi-objective: hypervolume of the Pareto front changes by < 2% over the last 5 rounds.

Visual Inspection: Use base.visualize_structure to inspect the best-found structure, and inspect all generated plots.


Examples

Synthetic 2D Test (Branin Function)

Validate the BO workflow on a known 2D benchmark with three global minima at $f^* = 0.3979$:

bash
${CLAUDE_SKILL_DIR}/../../venv/run cpu python ${CLAUDE_SKILL_DIR}/scripts/suggest_candidates.py \
    --config ${CLAUDE_SKILL_DIR}/examples/branin-function/search_space.yaml \
    --batch_size 8 \
    --output /tmp/bo_test/candidates_round_0.csv \
    --output_dir /tmp/bo_test/

See branin-function example for the full walkthrough including all 8 rounds and the GP model visualization.

Locate the most stable Li-Ag binary composition using Materials Project formation energies as the oracle:

See li-ag-phases example for the full walkthrough.


Constraints

  • Environment: cpu. Dependencies: scikit-learn, scipy, numpy, pandas, pyyaml. No BoTorch or ax-platform required.
  • GP Scaling: Training is O(n³). Practical upper limit is ~500 evaluated points before training time becomes noticeable.
  • Batch Size: Use a power-of-2 for initialization (batch_size = 4, 8, 16, …). For BO rounds, batch_size ≤ 8 is recommended; larger is fine when evaluations are embarrassingly parallel.
  • Minimum Data: The GP needs at least 2 × N_range_params evaluated points to fit reliably. The Sobol initialization (Step 2) should provide at least this many.
  • Parameter Bounds: Bounds in search_space.yaml must be physically meaningful. The GP has no information about the objective outside the specified domain.
  • Noise Handling: For deterministic objectives (analytic functions, noiseless ML models) set --noise_std 0. For stochastic objectives (MD-derived properties, noisy experiments) set --noise_std <estimated_std>.

References

  • Frazier, P.I., "A Tutorial on Bayesian Optimization", arXiv, 2018. arXiv:1807.02811
  • Knowles, J., "ParEGO: A Hybrid Algorithm with On-Line Landscape Approximation for Expensive Multiobjective Optimization Problems", IEEE Transactions on Evolutionary Computation, 2006. DOI:10.1109/TEVC.2005.851274
  • Balachandran, P.V. et al., "Adaptive strategies for materials design using uncertainties", Scientific Reports, 2016. DOI:10.1038/srep19660
  • Lookman, T. et al., "Active learning in materials science with emphasis on adaptive sampling using uncertainties for targeted design", npj Computational Materials, 2019. DOI:10.1038/s41524-019-0153-8

Author: Bowen Deng Contact: GitHub @bowen-bd

© learningmatter-mit, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 31 other files (scripts) in skills/ml-bayesian-optimization of learningmatter-mit/AtomisticSkills.

  • SKILL.md
  • examples/branin-function/README.md
  • examples/branin-function/campaign_state.json
  • examples/branin-function/candidates_round_0.csv
  • examples/branin-function/candidates_round_1.csv
  • examples/branin-function/candidates_round_2.csv
  • examples/branin-function/candidates_round_3.csv
  • examples/branin-function/candidates_round_4.csv
  • examples/branin-function/candidates_round_5.csv
  • examples/branin-function/candidates_round_6.csv
  • examples/branin-function/candidates_round_7.csv
  • examples/branin-function/convergence_curve.png
  • examples/branin-function/evaluated.csv
  • examples/branin-function/gp_model_2d.png
  • examples/branin-function/parameter_importance.png
  • examples/branin-function/search_space.yaml
  • examples/li-ag-phases/README.md
  • examples/li-ag-phases/campaign_state.json
  • … and 14 more

Open the folder on GitHubat commit 6257444

Compare with similar skills

ML Bayesian Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Bayesian Optimization compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Bayesian Optimization this skilllearningmatter-mit/AtomisticSkills176—~2.3kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
Fastmcp Client CLIPrefectHQ/fastmcp28k1 repos~823Automated safety check: PassApache-2.0
MemPalace Setup and OperationMemPalace/mempalace59k—~2.2kAutomated safety check: PassMIT
FastmcpTommy-yw/RunbookHermes5464 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Fastmcp Client CLI

    PrefectHQ/fastmcp

    Query and invoke tools on MCP servers using fastmcp list and fastmcp call.

    28k GitHub starsUsed in 1 repo~823 tokens
    Agent WorkflowsAuto-check passed
  • Installs and configures MemPalace as a private local palace, a shared-brain hub or a client of an existing hub, including MCP registration and version-correct initialization.

    59k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Fastmcp

    Tommy-yw/RunbookHermes

    Build, test, inspect, install, and deploy MCP servers with FastMCP in Python.

    546 GitHub starsUsed in 4 repos~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Google Antigravity SDK

    google-antigravity/antigravity-sdk-python

    Design, implement, and debug autonomous AI agents and multi-agent systems using the Google Antigravity (AGY) SDK.

    3.7k GitHub stars~2.1k tokensUpdated today
    Agent WorkflowsAuto-check: notes

More from learningmatter-mit/AtomisticSkills

All 129 skills in this repo
  • Drug Binding Site Definition

    learningmatter-mit/AtomisticSkills

    Define a docking search box (center coordinates + box dimensions in Angstroms) from a co-crystal ligand, binding-site residues, or a saved JSON specification.

    176 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Drug Complex System Builder

    learningmatter-mit/AtomisticSkills

    Build a solvated, charge-neutralized protein-ligand complex for OpenMM molecular dynamics simulation.

    176 GitHub stars~2k tokensUpdated today
    Auto-check passed
  • Drug Pocket Detection

    learningmatter-mit/AtomisticSkills

    Identify and rank ligandable pockets on a protein structure or model using geometry (fpocket) or an ML predictor (P2Rank).

    176 GitHub stars~4k tokensUpdated today
    Auto-check passed
  • Chem Bond Dissociation

    learningmatter-mit/AtomisticSkills

    Calculate homolytic and heterolytic bond dissociation energies (BDEs) for all single bonds in a molecule using MLIPs with RDKit fragmentation.

    176 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Chem Conformer Search

    learningmatter-mit/AtomisticSkills

    Generate molecular conformers with RDKit ETKDG, relax with MLIPs, and rank by energy with Boltzmann weighting.

    176 GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Chem DB Mof

    learningmatter-mit/AtomisticSkills

    Query multiple MOF databases (QMOF via MPContribs; ARC-MOF DB7/Majumdar et al.

    176 GitHub stars~1.9k tokensUpdated today
    Auto-check passed

Categories

Questions about ML Bayesian Optimization

What does ML Bayesian Optimization do?

Iteratively optimize expensive black-box objectives — such as materials properties, experimental yields, or simulation outputs — by learning from past evaluations to select the most promising next…. ML Bayesian Optimization is an agent skill from learningmatter-mit/AtomisticSkills. Iteratively optimize expensive black-box objectives — such as materials properties, experimental yields, or simulation outputs — by learning from past evaluations to select the most promising next candidates.

When should I use ML Bayesian Optimization?

ML Bayesian Optimization fits situations like: agent Workflows work in your project.

How do I install ML Bayesian Optimization in Claude Code?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-bayesian-optimization -a claude-code`. Or copy the skill folder (skills/ml-bayesian-optimization in learningmatter-mit/AtomisticSkills) into .claude/skills/ml-bayesian-optimization in your project. Claude Code loads it when a task matches its description.

How do I install ML Bayesian Optimization in Codex?

Run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-bayesian-optimization -a codex`. Or copy the skill folder (skills/ml-bayesian-optimization in learningmatter-mit/AtomisticSkills) into .agents/skills/ml-bayesian-optimization in your project. Codex loads it when a task matches its description.

Can I use ML Bayesian Optimization in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add learningmatter-mit/AtomisticSkills --skill ml-bayesian-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-bayesian-optimization, .gemini/skills/ml-bayesian-optimization, .github/skills/ml-bayesian-optimization and .opencode/skills/ml-bayesian-optimization in your project.

What does ML Bayesian Optimization need to run?

SKILL.md names no scripts, command-line tools or credentials: ML Bayesian Optimization is instructions for the agent only. Our summary lists: Python 3.

Does ML Bayesian Optimization access the network?

SKILL.md names 3 domains. As links in the text: doi.org, arxiv.org and github.com. This is read from the text; nothing was executed.

Is ML Bayesian Optimization safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does ML Bayesian Optimization use?

ML Bayesian Optimization is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Bayesian Optimization use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Bayesian Optimization?

Skills that share tags, products or a category with ML Bayesian Optimization: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), Fastmcp Client CLI (PrefectHQ/fastmcp, 28k stars) and MemPalace Setup and Operation (MemPalace/mempalace, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Bayesian Optimization?

learningmatter-mit (a GitHub organization) maintains it in learningmatter-mit/AtomisticSkills, which has 176 GitHub stars. The repository holds 129 skills in this directory. The repository was last updated on October 7, 2026.

Source: learningmatter-mit/AtomisticSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.