Agent skill

Alpha Evolve Experiment Design

by Google-Cloud-AI in Google-Cloud-AI/alphaevolve-on-googlecloud

Design AlphaEvolve experiments for the Cloud API. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud.

Apache-2.0Auto-check: warningsResearch & Science

Install Alpha Evolve Experiment Design

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-experiment-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-experiment-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/alpha_evolve_experiment_design .claude/skills/alpha-evolve-experiment-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
alpha-evolve-experiment-design
GitHub stars
118
Token cost
~2.5k tokens
SKILL.md length
970 words
Files
17 (incl. references)
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Design AlphaEvolve experiments for the Cloud API. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud.

  • Works in 2 steps: Clarify → Implement
  • : design an experiment
  • SKILL.md covers Preconditions, Postconditions, Phases and Critical Rules, plus 2 more sections
  • Runs Python scripts from its folder; calls uv and python3

What it does

Alpha Evolve Experiment Design is an agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud. Design AlphaEvolve experiments for the Cloud API. Takes a natural-language problem description and produces a complete, tested experiment directory ready for the experiment-runner skill. Triggers on: "design an experiment", "set up an AlphaEvolve experiment", "create an experiment for", "I want to evolve", "help me set up AlphaEvolve".

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 20 other files, including reference files (for example `README.md`, `examples/circle_packing/README.md` and `examples/circle_packing/evaluator.py`).

It sits in Research & Science, covering Experimental design. The licence is Apache-2.0.

When your agent uses it

  • : design an experiment
  • Set up an AlphaEvolve experiment
  • Create an experiment for
  • I want to evolve

Example prompts

  • “design an experiment”
  • “set up an AlphaEvolve experiment”
  • “create an experiment for”
  • “/alpha-evolve-experiment-design”

Requirements

  • Python 3

Workflow steps

2 steps, taken from the step headings in SKILL.md.

  1. Clarify
  2. Implement

What it can do on your machine

Read from SKILL.md and the folder at commit 674dd5e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Alpha Evolve Experiment Design loads about 2.5k tokens when it runs, and up to ~20k if it reads all its reference files. Until then it costs about 92 tokens; SKILL.md has 970 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~20k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningContains instruction-override wording (e.g. “without asking the user”)SKILL.md:183
    the evaluation metric or scoring logic without user consent.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Google-Cloud-AI/alphaevolve-on-googlecloud at commit 674dd5e, republished under its Apache-2.0 licence (© Google-Cloud-AI). 970 words, ~2,485 tokens.

Download SKILL.mdSave it as .claude/skills/alpha-evolve-experiment-design/SKILL.md (or your agent's skills folder). This skill also uses 16 other files; get the full folder from GitHub.
name
alpha-evolve-experiment-design
description
Design AlphaEvolve experiments for the Cloud API. Takes a natural-language problem description and produces a complete, tested experiment directory ready for the experiment-runner skill. Triggers on: "design an experiment", "set up an AlphaEvolve experiment", "create an experiment for", "I want to evolve", "help me set up AlphaEvolve".

Experiment Design Skill

You help users design AlphaEvolve experiments. You take a problem description and produce a complete, tested project directory that the experiment-runner skill can launch.

Preconditions

  • The user has a problem description (natural language). It may be rigorous or vague.
  • Optional: existing code to optimize.
  • Optional: a target directory path. If not provided, ask. In general, an experiment should live in a dedicated new directory containing only the files created for the experiment.

Postconditions

A project directory containing:

FilePurpose
.evolve/experiment_description.jsonComplete experiment specification
.evolve/source_map.jsonMaps code regions to original source
: : files (only when optimizing existing :
: : code; enables post-experiment :
: : integration) :
initial_program.pySeed program with EVOLVE-BLOCK
: : markers and ORIGIN comments :
evaluator.pyCLI-compatible evaluator script for
: : the ae CLI :
problem_description.mdDetailed technical problem
: : description (used in LLM prompts) :
example_evaluation.jsonSample evaluator output
test_program.pyPytest tests for the initial program
test_evaluator.pyPytest tests for the evaluator
pyproject.tomluv project configuration
README.mdExperiment documentation
*.py (multi-file only)Additional context files imported by
: : the initial program :

All pytest tests pass via uv run pytest.


Phases

The skill has exactly two phases. Complete Phase 1 before starting Phase 2.

Phase 1 — Clarify

Objective: Fill the ExperimentDescription data structure through conversation with the user.

Gate: Phase 1 is complete when experiment_description.json is written to project_dir/.evolve/.

Details: Read references/phase_1_clarify.md when you reach this phase.

Phase 2 — Implement

Objective: Generate all project files and verify they work.

Input contract: The ExperimentDescription is the only input to Phase 2. It must contain everything needed to generate all files. If information is missing, Phase 1 was incomplete — go back and fix it.

Gate: Phase 2 is complete when uv run pytest passes in the project directory.

Details: Read references/phase_2_implement.md when you reach this phase.


Critical Rules

  1. Phase 1 is conversation only. Do not create code files during Phase 1. The only file written is experiment_description.json.

  2. Phase 2 requires no user interaction. The ExperimentDescription contains everything needed. If you find yourself wanting to ask a question, Phase 1 was incomplete.

  3. Tests first. In Phase 2, write tests before the code they test.

  4. Never execute user code directly. Syntax-check with uv run python -c "import ast; ast.parse(open('file.py').read())" only. Evaluation happens through uv run pytest which exercises the code in a controlled way.

    CRITICAL: ALWAYS use uv run — NEVER bare python3. Do NOT run python3 evaluator.py, python3 test_program.py, or python3 -c "from evaluator import ...". Always use uv run python or uv run pytest. This ensures the correct virtual environment and dependencies are available and works cross-platform (python3 is not always available on Windows). Using bare python3 can silently use the wrong Python or miss project dependencies.

  5. The evaluator must be CLI-compatible. The evaluator file must be runnable as uv run python evaluator.py --output-file <path> --program-dir <path>. It must export evaluate_program(code, timeout_seconds=30) -> dict for testing (returning {"score": float|None, "insights": [...]}), and include a main() entry point that reads initial_program.py from --program-dir, evaluates it, and writes the result dict to --output-file. Insights capture stdout, stderr, errors, and tracebacks as {"label": str, "text": str} dicts that map to the AlphaEvolveEvaluationInsights API field.

  6. AlphaEvolve always maximizes. If the user wants to minimize a metric, the evaluator must negate the score.

  7. Handle non-finite scores. The evaluator MUST check for NaN and Inf scores (common with neural network training) and return null with an error insight instead. NaN in JSON is invalid and will crash the CLI. Use math.isnan() and math.isinf() before returning.

  8. Validate EVOLVE-BLOCK markers. Before declaring Phase 2 complete, verify that the initial program contains at least one valid # EVOLVE-BLOCK-START / # EVOLVE-BLOCK-END marker pair. The exact syntax matters -- EVOLVE_BLOCK_START (underscores) or other variants will be rejected by the API.

  9. Use uv for all project management. Projects use pyproject.toml, dependencies are installed via uv, tests run via uv run pytest. Never skip pyproject.toml creation or test files. All files listed in Phase 2 Postconditions are mandatory.

  10. Be concise. Do not narrate your internal reasoning. State what you are doing, show results, ask questions when needed.

  11. Never initiate version-control or commit workflows. Experiment files are local working artifacts. Do not stage or commit changes, search for related issues, or draft commit messages. Only create a pull request if the user explicitly requests it.


Show full SKILL.md (266 more words)Show less

Key References

ReferenceWhen to read
references/phase_1_clarify.mdAt the start of Phase 1
references/phase_2_implement.mdAt the start of Phase 2
references/evaluator_patterns.mdWhen designing the evaluator
references/numerical_stability.mdWhen the problem involves
: : neural networks, iterative :
: : optimization, or floating :
: : point arithmetic :
references/evolve_block_guide.mdWhen writing the initial
: : program :
references/multi_file_guide.mdWhen the user points to a
: : directory or multiple files :
resources/experiment_description_schema.pyFor the ExperimentDescription
: : model :
examples/circle_packing/Complete worked example

Constraints

  • Isolation Principle: Never modify original workspace source files directly. Target code MUST be extracted/copied into program files in the experiment directory. For single-file experiments, this means copying into initial_program.py. For multi-file experiments, this means copying into multiple .py files in the experiment directory.
  • No external imports. Program files and evaluator.py must NEVER import from the user's source tree (e.g., from myproject.models import ...). The ae CLI copies files to a temporary directory for evaluation, so imports relative to the original codebase will fail with ModuleNotFoundError. Imports between bundled program files (e.g., import layers where layers.py is another file in the experiment directory) ARE allowed. Only stdlib, pyproject.toml dependencies, and other bundled program files are available at evaluation time. See references/multi_file_guide.md for multi-file import constraints.
  • Never modify the evaluation metric or scoring logic without user consent.
  • Initial programs MUST be standalone runnable Python files.
  • Always use uv run python (or uv run pytest) instead of invoking a bare Python interpreter. This ensures the correct virtual environment and works cross-platform (python3 is not always available on Windows).
  • Reward Hacking: Always verify the semantic validity of evolved code during integration to ensure it represents a genuine discovery rather than an exploit of the scoring function.

© Google-Cloud-AI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 16 other files (references) in skills/alpha_evolve_experiment_design of Google-Cloud-AI/alphaevolve-on-googlecloud.

  • SKILL.md
  • README.md
  • examples/circle_packing/README.md
  • examples/circle_packing/evaluator.py
  • examples/circle_packing/example_evaluation.json
  • examples/circle_packing/initial_program.py
  • examples/circle_packing/problem_description.md
  • examples/circle_packing/pyproject.toml
  • examples/circle_packing/test_evaluator.py
  • examples/circle_packing/test_program.py
  • references/evaluator_patterns.md
  • references/evolve_block_guide.md
  • references/multi_file_guide.md
  • references/numerical_stability.md
  • references/phase_1_clarify.md
  • references/phase_2_implement.md
  • resources/experiment_description_schema.py

Open the folder on GitHubat commit 674dd5e

Compare with similar skills

Alpha Evolve Experiment Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Alpha Evolve Experiment Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Alpha Evolve Experiment Design this skillGoogle-Cloud-AI/alphaevolve-on-googlecloud118—~2.5kAutomated safety check: WarnApache-2.0
Scientific Critical Thinkingweapp-tailwindcss/weapp-tailwindcss1.9k23 repos~5.9kAutomated safety check: NotesMIT
Claim-Driven Experiment PlannerzjYao36/Auto-Research-Refine1287 repos~2.3kAutomated safety check: NotesNone
Benchmark Paper TemplateHKUSTDial/Supervisor-Skills8.5k—~2.8kAutomated safety check: PassCC-BY-4.0
Research Refine PipelinezjYao36/Auto-Research-Refine1286 repos~1.4kAutomated safety check: NotesNone
Metabolic Study Planneraiming-lab/AutoResearchClaw15k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Scientific Critical Thinking

    weapp-tailwindcss/weapp-tailwindcss

    Evaluate research rigor. An agent skill from weapp-tailwindcss/weapp-tailwindcss.

    1.9k GitHub starsUsed in 23 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Claim-Driven Experiment Planner

    zjYao36/Auto-Research-Refine

    Turns a refined research proposal into a claim-to-evidence-to-run-order roadmap instead of a sprawling benchmark wishlist.

    128 GitHub starsUsed in 7 repos~2.3k tokens
    Research & ScienceAuto-check: notes
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.5k GitHub stars~2.8k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Research Refine Pipeline

    zjYao36/Auto-Research-Refine

    Chains research-refine and experiment-plan to turn a vague research direction into a focused proposal and a claim-driven experiment roadmap.

    128 GitHub starsUsed in 6 repos~1.4k tokens
    Research & ScienceAuto-check: notes
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Research & ScienceAuto-check passed
  • Experimental Design

    Oleafly/Oleafly

    Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so results are interpretable.

    206 GitHub starsUsed in 4 repos~3.5k tokens
    Research & ScienceAuto-check: notes

More from Google-Cloud-AI/alphaevolve-on-googlecloud

  • Alpha Evolve Monitor

    Google-Cloud-AI/alphaevolve-on-googlecloud

    Monitor running AlphaEvolve experiments, run the evaluation control loop, and report results using the ae CLI.

    118 GitHub stars~4.4k tokensUpdated 7 days ago
    Auto-check passed
  • Alpha Evolve Orchestrator

    Google-Cloud-AI/alphaevolve-on-googlecloud

    End-to-end AlphaEvolve experiment orchestrator. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud.

    118 GitHub stars~4.1k tokensUpdated 7 days ago
    Auto-check passed
  • Alpha Evolve Post Experiment

    Google-Cloud-AI/alphaevolve-on-googlecloud

    Post-experiment analysis, visualization, and code integration for completed AlphaEvolve experiments.

    118 GitHub stars~9.2k tokensUpdated 7 days ago
    Auto-check passed
  • Alpha Evolve Consultant

    Google-Cloud-AI/alphaevolve-on-googlecloud

    AlphaEvolve expert consultant grounded strictly in the official reference guide.

    118 GitHub stars~3k tokensUpdated 7 days ago
    Auto-check passed
  • Alpha Evolve Runner

    Google-Cloud-AI/alphaevolve-on-googlecloud

    Configure, verify, and launch AlphaEvolve experiments using the ae CLI.

    118 GitHub stars~5.6k tokensUpdated 7 days ago
    Auto-check: warnings

Questions about Alpha Evolve Experiment Design

What does Alpha Evolve Experiment Design do?

Design AlphaEvolve experiments for the Cloud API. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud. Alpha Evolve Experiment Design is an agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud. Design AlphaEvolve experiments for the Cloud API.

When should I use Alpha Evolve Experiment Design?

Alpha Evolve Experiment Design fits situations like: : design an experiment; set up an AlphaEvolve experiment; create an experiment for; I want to evolve.

How do I install Alpha Evolve Experiment Design in Claude Code?

Run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-experiment-design -a claude-code`. Or copy the skill folder (skills/alpha_evolve_experiment_design in Google-Cloud-AI/alphaevolve-on-googlecloud) into .claude/skills/alpha-evolve-experiment-design in your project. Claude Code loads it when a task matches its description.

How do I install Alpha Evolve Experiment Design in Codex?

Run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-experiment-design -a codex`. Or copy the skill folder (skills/alpha_evolve_experiment_design in Google-Cloud-AI/alphaevolve-on-googlecloud) into .agents/skills/alpha-evolve-experiment-design in your project. Codex loads it when a task matches its description.

Can I use Alpha Evolve Experiment Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-experiment-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/alpha-evolve-experiment-design, .gemini/skills/alpha-evolve-experiment-design, .github/skills/alpha-evolve-experiment-design and .opencode/skills/alpha-evolve-experiment-design in your project.

What does Alpha Evolve Experiment Design need to run?

Going by SKILL.md and its folder, Alpha Evolve Experiment Design needs Python for the scripts in its folder and the command-line tools its instructions call (uv and python3). Our summary lists: Python 3.

Does Alpha Evolve Experiment Design access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Alpha Evolve Experiment Design safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): contains instruction-override wording (e.g. “without asking the user”). Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Alpha Evolve Experiment Design use?

Alpha Evolve Experiment Design is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Alpha Evolve Experiment Design use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 18k tokens, read only when the agent opens those files.

What are the alternatives to Alpha Evolve Experiment Design?

Skills that share tags, products or a category with Alpha Evolve Experiment Design: Scientific Critical Thinking (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Claim-Driven Experiment Planner (zjYao36/Auto-Research-Refine, 128 stars), Benchmark Paper Template (HKUSTDial/Supervisor-Skills, 8.5k stars) and Research Refine Pipeline (zjYao36/Auto-Research-Refine, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Alpha Evolve Experiment Design?

Google-Cloud-AI (a GitHub organization) maintains it in Google-Cloud-AI/alphaevolve-on-googlecloud, which has 118 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 1, 2026.

Source: Google-Cloud-AI/alphaevolve-on-googlecloud on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.