Hypothesis Generation
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
Configure, verify, and launch AlphaEvolve experiments using the ae CLI.
The automated check flagged lines worth reading first. See the safety section below.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-runner --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/alpha_evolve_runner .claude/skills/alpha-evolve-runner && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "alpha-evolve-runner" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_runner into .claude/skills/alpha-evolve-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-runner", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_runnerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-runner --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/alpha_evolve_runner .agents/skills/alpha-evolve-runner && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "alpha-evolve-runner" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_runner into .agents/skills/alpha-evolve-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-runner", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-runner --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/alpha_evolve_runner .cursor/skills/alpha-evolve-runner && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "alpha-evolve-runner" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_runner into .cursor/skills/alpha-evolve-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-runner", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git --path skills/alpha_evolve_runner--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-runner --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/alpha_evolve_runner .gemini/skills/alpha-evolve-runner && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "alpha-evolve-runner" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_runner into .gemini/skills/alpha-evolve-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-runner", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-runnerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/alpha_evolve_runner .github/skills/alpha-evolve-runner && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "alpha-evolve-runner" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_runner into .github/skills/alpha-evolve-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-runner", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-runner --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/alpha_evolve_runner .opencode/skills/alpha-evolve-runner && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "alpha-evolve-runner" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_runner into .opencode/skills/alpha-evolve-runner/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-runner", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
alpha-evolve-runnerConfigure, verify, and launch AlphaEvolve experiments using the ae CLI.
Alpha Evolve Runner is an agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud. Configure, verify, and launch AlphaEvolve experiments using the ae CLI. Supports both creating new experiments from Design skill artifacts and launching from user-provided files. Triggers on: "run this experiment", "launch the experiment", "start AlphaEvolve", "create an experiment", "launch experiment", "run AlphaEvolve experiment".
Its SKILL.md is about 5.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `README.md`, `references/cli_reference.md` and `references/debugging.md`).
It sits in Research & Science. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 674dd5e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
curlgcloudFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
discoveryengine.googleapis.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Alpha Evolve Runner loads about 5.6k tokens when it runs, and up to ~9.5k if it reads all its reference files. Until then it costs about 89 tokens; SKILL.md has 2,343 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found patterns that need a careful read before installing.
> `.env` files, or prior experiment directories may be **stale or intended for a> without pausing, **do NOT pause for user confirmation** — proceed directly toAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Google-Cloud-AI/alphaevolve-on-googlecloud at commit 674dd5e, republished under its Apache-2.0 licence (© Google-Cloud-AI). 2,343 words, ~5,645 tokens.
.claude/skills/alpha-evolve-runner/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.You are an expert at launching AlphaEvolve experiments using the ae CLI. Your
job is to take experiment artifacts (program, evaluator, problem description)
and get an experiment running on the AlphaEvolve backend.
ae program evaluate, which handles sandboxing.--json flag when calling ae commands so you can parse
structured output. Present human-readable summaries yourself. Note: --json
is a global flag and must go BEFORE the subcommand, e.g. ae --json config show, NOT ae config show --json.The following examples use Unix shell syntax. Adapt commands for your platform (e.g.,
whereinstead ofwhichon Windows, PowerShell syntax for environment variables).
The ae CLI must be installed and executable. Follow this discovery
sequence — do NOT skip steps or guess paths:
Try the bare command:
ae versionIf that fails, search common install locations:
which ae 2>/dev/null || \
ls ~/.local/bin/ae 2>/dev/null || \
ls ~/.local/share/uv/tools/ae-cli/bin/ae 2>/dev/nullIf found but not on PATH, set and use the full path for all subsequent
commands. For example: AE=/home/user/.local/bin/ae && $AE version
If not found after steps 1-2, tell the user:
The
aeCLI is not installed. This is required before proceeding. Please follow theaeCLI documentation to install it, then try again.
Do NOT spend more than 3 commands searching for the binary. If you cannot
find it after the above steps, ask the user: "Where is your ae binary
installed?"
Before making any claims about network access, verify connectivity:
curl -s -o /dev/null -w "%{http_code}" https://discoveryengine.googleapis.comIf this returns any HTTP status code (200, 404, etc.), you have network access.
NEVER claim you lack internet access without running this check. Only if
curl fails with a connection error should you report network issues.
After installing or updating the CLI or skills, restart the agent session. Most agent runtimes cache skill files for the duration of a session. Changes will not take effect until a new session is started.
Launching an experiment requires 3 separate CLI commands executed in sequence. This is NOT a single command — each step has a distinct purpose:
┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐
│ experiment │ │ experiment │ │ experiment │
│ create │────▶│ start │────▶│ run │
│ │ │ │ │ │
│ Creates the │ │ Uploads the │ │ Runs the local │
│ experiment on │ │ initial program │ │ eval loop: │
│ the backend. │ │ and sets the │ │ acquire → eval │
│ Returns a │ │ baseline score. │ │ → submit score │
│ nickname. │ │ Activates the │ │ → repeat. │
│ │ │ experiment. │ │ │
│ Does NOT start │ │ Does NOT run │ │ This is what │
│ anything yet. │ │ evaluations. │ │ actually drives │
│ │ │ │ │ the experiment. │
└─────────────────┘ └─────────────────┘ └─────────────────┘Why 3 steps? The backend generates candidate programs, but evaluation
happens locally (or in a container). create registers the experiment, start
seeds it with your baseline, and run is the long-running loop that pulls
candidates, evaluates them, and reports scores back.
Do NOT try to combine these into one command. There is no ae experiment create --program-file --evaluator shortcut. Each command has different required
flags (see Stage 4 below for exact syntax).
To launch an experiment, you need three files. These may come from the Design skill or be provided directly by the user:
EVOLVE-BLOCK markers)--output-file and --program-dir)If any of these are missing, ask the user to provide them.
Objective: Set up ae CLI configuration and verify API connectivity.
Run these commands to discover current state:
ae --json config show
ae --json config discoverThe config discover command detects the ambient GCP project from the user's
gcloud configuration without modifying any ae settings. Use its project
field as the default when ae config show has no project set.
WARNING: Stale configuration. Values found in existing
aeprofiles,.envfiles, or prior experiment directories may be stale or intended for a different project. Auto-discovered values are starting points, NOT confirmed truth. You MUST present them for user confirmation in Step 1.2 regardless of whether auto-discovery succeeded (unless operating in a pre-configured environment; see the exception below).
Do NOT skip Step 1.2 unless operating in a pre-configured environment where
ae config show returns a valid profile and ae --json config test succeeds
(see the Pre-configured Environment / Autonomous Execution Exception below).
Otherwise, always present the configuration table for user confirmation before
proceeding.
For any values not already configured, auto-discover them:
Project ID: Use the project field from ae --json config discover.
Engine ID: Use ae --json engine list to list available engines. Pick the
first engine with solutionType: SOLUTION_TYPE_CHAT, or ask the user if
multiple engines exist.
Apply configuration — set all values in a single command:
ae config \
--project=<PROJECT_ID> \
--engine=<ENGINE_ID> \
--location=global \
--models=gemini-3.5-flashIMPORTANT:
ae configsyntax and profile pitfall.All flags use
--flag=valuesyntax. Set them all in one call.Profile pitfall:
ae configwithout--nameupdates thedefaultprofile, which may NOT be the active profile. Always check the active profile first withae --json config show(look at theProfilefield), then pass--name=<active_profile>:
ae config --name=<active_profile> --models=<model_spec>Without
--name, the change silently writes to the wrong profile.To verify changes took effect:
ae --json config show
Default values (use these unless the user specifies otherwise):
--location=global--collection=default_collection (configurable if project uses a custom
collection)--models=gemini-3.5-flash (recommended default; available in all regions)--base-url=https://discoveryengine.googleapis.com (prod default; only
change if directed to a different endpoint)Available models: Read references/models.md for the full list of model IDs
and descriptions. You MUST consult that reference before selecting a model.
Do NOT guess model IDs or read proto files.
<!-- *** MANDATORY USER INTERACTION — NEVER SKIP THIS STEP *** -->
CRITICAL: This confirmation step is MANDATORY for unverified profiles, EXCEPT in pre-configured environments or autonomous execution.
Pre-configured Environment / Autonomous Execution Exception: If
ae config showalready has a pre-configured profile andae --json config testsucceeds, OR if the user explicitly asked to run the experiment end-to-end without pausing, do NOT pause for user confirmation — proceed directly to Step 1.3 and Stage 2.
Present the resolved configuration as a table and ask the user to confirm:
| Setting | Value | Source |
|---|---|---|
| Project | my-project-123 | ae config discover |
| Engine | alpha-evolve-engine | auto-discovered |
| Location | global | default |
| Models | gemini-3.5-flash | default |
| API Endpoint | https://discoveryengine.googleapis.com | default (prod) |
"Does this look correct? (Y/n)"
Accept bare "yes", Enter, "y", or "looks good" as confirmation. If the user
wants to change a value, update with ae config --<flag>=<value> and re-display
the table until confirmed.
Do NOT proceed to Step 1.3 until the user explicitly confirms (unless running under the Pre-configured Environment / Autonomous Execution Exception above).
<!-- *** END MANDATORY USER INTERACTION *** -->
ae --json config testIf this fails, consult references/debugging.md for diagnosis. Common issues:
gcloud auth application-default loginRetry up to 3 times after applying fixes. Do NOT proceed until connectivity works.
Objective: Confirm the evaluator works with the initial program and obtain a baseline score.
Check that the required files exist and are well-formed:
Initial program file:
# EVOLVE-BLOCK-START / # EVOLVE-BLOCK-END
marker pairEvaluator file:
--output-file and
--program-dir flagsevaluate_program(code, timeout_seconds) -> dict for testingProblem description:
.md file for ae experiment createIf any file is missing or invalid, tell the user exactly what is needed.
ae --json program evaluate \
--program-dir <experiment_directory> \
--evaluator <evaluator_file> \
--backend localThe default backend is local. Only suggest podman if:
Parse the JSON output for the score. Display it to the user:
"Baseline evaluation complete. Score: X.XX"
If the score looks problematic:
references/debugging.md. Do NOT proceed until a valid baseline is
obtained.Objective: Present all experiment parameters for user review before creating anything on the backend. This is the user's chance to adjust settings like max_programs, concurrency, or model.
Use these defaults unless the user specified otherwise:
| Parameter | Default | Flag |
|---|---|---|
| Max programs | 100 | --max-programs |
| Concurrency | 4 | --concurrency |
| Title | derived from problem description | --title |
| Models | from config (Stage 1) | --models |
<!-- *** MANDATORY USER INTERACTION *** -->
MANDATORY FOR INTERACTIVE USERS: You MUST present the summary table below (including the verified
Baseline Score) and wait for user confirmation before launching Stage 4.Autonomous / Unattended Benchmark Exception: Only skip this confirmation if the user explicitly instructed you to run without asking for confirmation (e.g., "run autonomously without pausing" or "do not ask for confirmation").
Present a summary table of ALL parameters before creating the experiment:
| Parameter | Value |
|---|---|
| Project | my-project-123 |
| Engine | alpha-evolve-engine |
| Models | gemini-3.5-flash |
| API Endpoint | https://discoveryengine.googleapis.com |
| Max Programs | 100 |
| Concurrency | 4 |
| Program Dir | ./exp_dir/ |
| Evaluator | evaluator.py |
| Baseline Score | 42.5 |
| Eval Backend | local |
"Ready to create and launch? Type 'yes' to proceed, or type any changes you'd like to make (e.g. 'max_programs=50, concurrency=2, models=gemini-3.5-flash')."
Accept bare "yes", Enter, "y", or "looks good" as confirmation. If the user types parameter changes (e.g. "max_programs=200" or "change concurrency to 8"), parse the values, update the table, and re-display it for confirmation.
<!-- *** END MANDATORY USER INTERACTION *** -->
Objective: Execute the 3-step launch sequence (create → start → run).
CRITICAL: Follow these 3 steps exactly. Do NOT guess flags or try to combine steps. Each command has different required arguments.
Creates the experiment resource on the backend. Returns a nickname you will use in all subsequent commands.
Prefer passing --models explicitly. Passing --models keeps the chosen
model unambiguous in the trajectory.
ae --json experiment create \
--max-programs <confirmed_max_programs> \
--concurrency <confirmed_concurrency> \
--problem-file <problem_description_file> \
--title "<experiment_title>" \
--models <confirmed_model>Flags for experiment create:
| Flag | Required | Description |
|---|---|---|
--max-programs | Yes | Max candidate programs |
| : : : to generate (must : | ||
| : : : be > 1) : | ||
--concurrency | No (default 4) | Parallel program |
| : : : generation : | ||
--problem | No* | Inline problem |
| : : : description : | ||
--problem-file | No* | Path to |
: : : problem_description.md : | ||
--title | No | Human-readable |
| : : : experiment title : | ||
--models | Recommended | Model spec: bare name or |
| : : : name=...,weight=... : | ||
| : : : (repeatable) : |
*One of --problem or --problem-file.
Does NOT accept: --program-dir, --evaluator, or --score. These belong
to experiment start and experiment run.
Parse the JSON output for the nickname (e.g., exp-brave-otter).
Uploads the initial program file(s) and baseline score. Transitions the experiment from CREATED to ACTIVE. After this, the backend begins generating candidate programs.
ae --json experiment start <nickname> \
--program-dir <experiment_directory> \
--score <baseline_score>Flags for experiment start:
| Flag | Required | Description |
|---|---|---|
<nickname> | Yes (positional) | From Step 4.1 output |
--program-dir | Yes | Directory with program .py files |
--score | Yes | Baseline score from Stage 2 |
The --program-dir bundles all .py files from the experiment directory
(excluding evaluator.py and test files). The experiment directory should
contain only the cherry-picked program files created by the design skill.
Does NOT accept: --evaluator, --problem-file, --models, or
--max-programs. Those belong to experiment create.
Starts the long-running acquire → evaluate → submit loop that actually drives the experiment forward. Without this step, the experiment has no evaluator and will stall.
ae --json experiment run <nickname> \
--evaluator <evaluator_file> \
--backend local \
--dashboard <nickname>-dashboard.mdUse <nickname>-dashboard.md (e.g., exp-brave-otter-dashboard.md) so that
multiple experiments in the same directory don't overwrite each other's
dashboards.
Flags for experiment run:
| Flag | Required | Description |
|---|---|---|
<nickname> | Yes (positional) | From Step 4.1 output |
--evaluator | Yes | Path to evaluator script |
--backend | No (default local) | local or podman |
--timeout | No (default 60) | Per-evaluation timeout (seconds) |
--dashboard | Yes (always pass) | Path to write live dashboard file |
--max-iterations | No (default 0) | 0 = unlimited |
This is a blocking command that runs until the experiment completes, fails, or is interrupted. The Orchestrator/Monitor skills handle running this in the background.
If you want to confirm the experiment state separately:
ae --json experiment describe <nickname>Verify the status is ACTIVE.
There is no ae experiment pause command. Do NOT try to run it. Experiments
are paused automatically by the backend after ~5 hours of idle time (no
evaluations submitted). To stop evaluations, simply kill the ae experiment run
process. To resume a paused experiment, use ae experiment resume followed by
ae experiment run.
For any ae command failure: 1. Parse the JSON error output 2. Consult
references/debugging.md for known error patterns 3. Suggest a specific fix 4.
Retry after the fix is applied
Common error patterns:
"experiment not found" — Wrong nickname/ID. Run ae experiment list."quota exceeded" — Project quota limit. Request quota increase."already started" — Experiment is active. Use Monitor skill instead."invalid program" — Missing EVOLVE-BLOCK markers. Check marker syntax."evaluation failed" — Evaluator bug or missing deps. Debug locally.CRITICAL: Do NOT retry the same failing step indefinitely.
Follow these escalation rules:
3-strike rule per step. After 3 failed attempts at any single step
(e.g., experiment create, config test, connectivity), STOP retrying
and escalate to the user with a structured diagnosis:
I've tried 3 approaches to [step] and all failed:
- Tried X → Error: Y
- Tried A → Error: B
- Tried C → Error: D
This suggests [root cause hypothesis]. Could you help me with [specific question]?
5-command hard limit. Never execute more than 5 variations of the
same command (e.g., experiment create with different flags) without user
input. If you find yourself trying a 6th variation, you are guessing — stop
and ask.
Track what you've tried. Before each retry, briefly state what you learned from the previous failure and why the next attempt is different. Do NOT blindly retry the same command or try random permutations.
Distinguish fixable vs. unfixable errors. Some errors (auth, wrong project) require user action. Do NOT waste retries on these — escalate immediately after the first occurrence.
See references/cli_reference.md for the full command reference.
# Step 1: Create experiment (returns nickname)
ae --json experiment create \
--max-programs 100 --problem-file problem_description.md \
--title "My Experiment" --models gemini-3.5-flash
# Step 2: Start experiment (uploads program files + baseline)
ae --json experiment start <nickname> \
--program-dir . --score 42.5
# Step 3: Run evaluation loop (blocking, drives the experiment)
ae --json experiment run <nickname> \
--evaluator evaluator.py --backend local --dashboard <nickname>-dashboard.mdae --json config show — Show current configurationae --json config discover — Detect ambient GCP project from gcloudae config --project=X --engine=Y — Set configuration valuesae --json config test — Verify API connectivityae --json program evaluate --program-file P --evaluator E — Test evaluator
locallyae --json experiment describe <exp> — Check experiment statusae --json experiment list — List all experiments© Google-Cloud-AI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (references) in skills/alpha_evolve_runner of Google-Cloud-AI/alphaevolve-on-googlecloud.
Open the folder on GitHubat commit 674dd5e
Alpha Evolve Runner next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Alpha Evolve Runner this skillGoogle-Cloud-AI/alphaevolve-on-googlecloud | 118 | — | ~5.6k | Automated safety check: Warn | Apache-2.0 | |
| Hypothesis Generationspacering-net/codeg | 3.8k | 15 repos | ~3.6k | Automated safety check: Notes | MIT | |
| GitHub Deep Researchbytedance/deer-flow | 83k | 5 repos | ~1.3k | Automated safety check: Pass | MIT | |
| Nature Paper CardYuan1z0825/nature-skills | 46k | 2 repos | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Read arXiv Paperkarpathy/nanochat | 58k | 2 repos | ~494 | Automated safety check: Pass | MIT | |
| Content Research Writerweapp-tailwindcss/weapp-tailwindcss | 1.9k | 25 repos | ~3.5k | Automated safety check: Pass | MIT |
spacering-net/codeg
Structured hypothesis formulation from observations. An agent skill from spacering-net/codeg.
bytedance/deer-flow
Researches a GitHub repository over four rounds using the GitHub API and web search, then writes a structured markdown report with timeline, metrics and Mermaid diagrams.
Yuan1z0825/nature-skills
Builds a structured deep-reading card for one scientific paper, covering methods, how experiments support claims, limitations and research ideas, with a script to prepare the source.
karpathy/nanochat
Fetches the TeX source of an arXiv paper from its URL, reads it and writes a markdown summary tied to the nanochat project.
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
spacering-net/codeg
Structured manuscript/grant review with checklist-based evaluation.
Google-Cloud-AI/alphaevolve-on-googlecloud
Monitor running AlphaEvolve experiments, run the evaluation control loop, and report results using the ae CLI.
Google-Cloud-AI/alphaevolve-on-googlecloud
End-to-end AlphaEvolve experiment orchestrator. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud.
Google-Cloud-AI/alphaevolve-on-googlecloud
Post-experiment analysis, visualization, and code integration for completed AlphaEvolve experiments.
Google-Cloud-AI/alphaevolve-on-googlecloud
AlphaEvolve expert consultant grounded strictly in the official reference guide.
Google-Cloud-AI/alphaevolve-on-googlecloud
Design AlphaEvolve experiments for the Cloud API. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud.
Categories
Configure, verify, and launch AlphaEvolve experiments using the ae CLI. Alpha Evolve Runner is an agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud. Configure, verify, and launch AlphaEvolve experiments using the ae CLI.
Alpha Evolve Runner fits situations like: : run this experiment; launch the experiment; start AlphaEvolve; create an experiment.
Run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a claude-code`. Or copy the skill folder (skills/alpha_evolve_runner in Google-Cloud-AI/alphaevolve-on-googlecloud) into .claude/skills/alpha-evolve-runner in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a codex`. Or copy the skill folder (skills/alpha_evolve_runner in Google-Cloud-AI/alphaevolve-on-googlecloud) into .agents/skills/alpha-evolve-runner in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-runner -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/alpha-evolve-runner, .gemini/skills/alpha-evolve-runner, .github/skills/alpha-evolve-runner and .opencode/skills/alpha-evolve-runner in your project.
Going by SKILL.md and its folder, Alpha Evolve Runner needs the command-line tools its instructions call (curl and gcloud). Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: discoveryengine.googleapis.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.
Alpha Evolve Runner is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.6k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 3.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Alpha Evolve Runner: Hypothesis Generation (spacering-net/codeg, 3.8k stars), GitHub Deep Research (bytedance/deer-flow, 83k stars), Nature Paper Card (Yuan1z0825/nature-skills, 46k stars) and Read arXiv Paper (karpathy/nanochat, 58k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Google-Cloud-AI (a GitHub organization) maintains it in Google-Cloud-AI/alphaevolve-on-googlecloud, which has 118 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 1, 2026.
Source: Google-Cloud-AI/alphaevolve-on-googlecloud on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.