Finishing a Development Branch
obra/superpowers
Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.
Agent skill
by Google-Cloud-AI in Google-Cloud-AI/alphaevolve-on-googlecloud
Post-experiment analysis, visualization, and code integration for completed AlphaEvolve experiments.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-post-experiment --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/alpha_evolve_post_experiment .claude/skills/alpha-evolve-post-experiment && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "alpha-evolve-post-experiment" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_post_experiment into .claude/skills/alpha-evolve-post-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-post-experiment", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_post_experimentType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-post-experiment --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/alpha_evolve_post_experiment .agents/skills/alpha-evolve-post-experiment && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "alpha-evolve-post-experiment" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_post_experiment into .agents/skills/alpha-evolve-post-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-post-experiment", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-post-experiment --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/alpha_evolve_post_experiment .cursor/skills/alpha-evolve-post-experiment && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "alpha-evolve-post-experiment" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_post_experiment into .cursor/skills/alpha-evolve-post-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-post-experiment", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git --path skills/alpha_evolve_post_experiment--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-post-experiment --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/alpha_evolve_post_experiment .gemini/skills/alpha-evolve-post-experiment && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "alpha-evolve-post-experiment" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_post_experiment into .gemini/skills/alpha-evolve-post-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-post-experiment", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-post-experimentInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/alpha_evolve_post_experiment .github/skills/alpha-evolve-post-experiment && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "alpha-evolve-post-experiment" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_post_experiment into .github/skills/alpha-evolve-post-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-post-experiment", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Google-Cloud-AI/alphaevolve-on-googlecloud alpha-evolve-post-experiment --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/alpha_evolve_post_experiment .opencode/skills/alpha-evolve-post-experiment && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "alpha-evolve-post-experiment" agent skill from https://github.com/Google-Cloud-AI/alphaevolve-on-googlecloud/tree/main/skills/alpha_evolve_post_experiment into .opencode/skills/alpha-evolve-post-experiment/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "alpha-evolve-post-experiment", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
alpha-evolve-post-experimentPost-experiment analysis, visualization, and code integration for completed AlphaEvolve experiments.
Alpha Evolve Post Experiment is an agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud. Post-experiment analysis, visualization, and code integration for completed AlphaEvolve experiments. Produces comprehensive results reports with inline visualizations and applies evolved code back to the original codebase. Triggers on: "show experiment results", "analyze experiment", "integrate evolved code", "apply results", "experiment report", "what did the experiment find", "copy results back", "apply the best program", "post-experiment".
Its SKILL.md is about 9.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `README.md`, `references/cli_reference.md` and `references/debugging.md`).
It sits in Development. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 674dd5e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
python3uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Alpha Evolve Post Experiment loads about 9.2k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 119 tokens; SKILL.md has 3,772 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Google-Cloud-AI/alphaevolve-on-googlecloud at commit 674dd5e, republished under its Apache-2.0 licence (© Google-Cloud-AI). 3,772 words, ~9,199 tokens.
.claude/skills/alpha-evolve-post-experiment/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.You are an expert at analyzing completed AlphaEvolve experiments, producing clear visual reports, and integrating evolved code back into the user's codebase. Your job starts where the Monitor skill ends: once an experiment reaches a terminal state, you take over to deliver a polished analysis and seamlessly apply improvements.
--json flag when calling ae commands so you can parse
structured output. Present human-readable summaries yourself. Note: --json
is a global flag and must go BEFORE the subcommand, e.g. ae --json results best <exp>, NOT ae results best --json <exp>.ae program evaluate, which handles sandboxing.brave-otter), a short ID, or a full resource name. Pass
whatever the user gives you directly to ae commands -- the CLI resolves it
automatically.Before doing anything else, verify the ae CLI is installed:
ae versionIf this command fails, tell the user:
The
aeCLI is not installed. This is required before proceeding. Please follow theaeCLI documentation to install it, then try again.
Stop here if ae version fails. Do not proceed.
This skill requires only one input to start:
Always run Stages 1, 2, and 3 immediately and proactively. These stages require no user input beyond the experiment identifier. Do NOT ask the user whether they want analysis or a report — just do it. The flow is: quick summary (Stage 1) → code review (Stage 2) → full report with all three artifacts (Stage 3).
Stage 4 (code integration) is the only part that requires user confirmation. After the report artifacts are generated, ask the user if they want to integrate the evolved code back into their codebase. Only if they confirm, gather the additional inputs needed for integration:
.evolve/experiment_description.json, or from the user
directly.If the user declines integration, skip to Step 4.6 (completion summary). The report artifacts are already generated in Stage 3.
Objective: Give the user a fast, high-level understanding of the experiment outcome. This is a brief inline summary — the detailed report and visualizations are generated in Stage 3.
Run these commands to collect the data:
ae --json experiment describe <EXPERIMENT>
ae --json results best <EXPERIMENT> --top 5If any command fails, consult references/debugging.md for diagnosis.
<!-- *** CRITICAL: AUTHORITATIVE SCORES *** -->
The results best output is sorted by score descending (best first). The
first program in the list is the authoritative best. Record its score
and nickname immediately — these are the authoritative best score and best
nickname for the entire report. Use them verbatim in every section that
references the best score (Result, Results Summary, completion summary, etc.).
Do NOT substitute scores you encounter later during approach analysis.
For the baseline score, read the initialAlphaEvolveProgram field from the
experiment describe output above and fetch it:
ae --json program show <INITIAL_PROGRAM_RESOURCE_NAME>The baseline score is in evaluation.scores.scores[0].score of the response.
<!-- *** END CRITICAL *** -->
Using the authoritative baseline and best scores from Step 1.1, present a brief inline summary:
Experiment
<NICKNAME>—<STATE>Result:
<BASELINE>→<BEST>(+<REL>%) Evaluations:<TOTAL>| Best program:<BEST_NICKNAME>Top 3: 1.
<nickname>—<score>2.<nickname>—<score>3.<nickname>—<score>
This is intentionally brief. Do NOT produce a full formatted report here — the detailed report with charts, approaches table, and failure analysis is generated in Stage 3.
After presenting the summary, proceed immediately to Stage 2.
Objective: Fetch the best program's evolved code, present a clear diff against the initial program, explain the changes, and flag any concerns before integration.
ae --json program show <BEST_NICKNAME> --experiment <EXPERIMENT> --codeIf the code is large (>200 lines), also save it to a file:
ae --json program show <BEST_NICKNAME> --experiment <EXPERIMENT> --code \
--output-file <PROJECT_DIR>/best_evolved_program.pyRun the diff command to show exactly what changed:
ae program diff <BEST_NICKNAME> --experiment <EXPERIMENT>Note: ae program diff does NOT support --json. The output is a
human-readable unified diff. Display it directly.
If program diff fails (known issue with short parent IDs), fall back to a
manual comparison using the agent's built-in file reading capabilities:
diff — it is not cross-platform (e.g., on Windows PowerShell, diff is
aliased to Compare-Object).After showing the diff, provide a concise semantic explanation of what the evolved code does differently. Focus on:
Evaluate whether the evolved code genuinely solves the problem or exploits the scoring function. Flag the result if ANY of these heuristics trigger:
| Heuristic | Red Flag |
|---|---|
| Score suspiciously perfect | Score equals theoretical maximum or a round |
| : : number suggesting hardcoded output : | |
| Hardcoded outputs | Code contains literal values that match |
| : : test case expected outputs : | |
| Evaluator manipulation | Code reads the evaluator file, output file |
| : : path, or test fixtures : | |
| Trivial code | Evolved block is shorter than 3 lines or |
| : : contains only a return statement with a : | |
| : : constant : | |
| Import manipulation | Code imports inspect, sys._getframe, or |
| : : other introspection to detect evaluation : | |
| : : context : |
If any red flag is detected, warn the user:
Reward hacking detected. The evolved code may be exploiting the scoring function rather than genuinely solving the problem. Specifically:
<describe the concern>I recommend reviewing the evolved code carefully before integrating it. Consider re-running with a more robust evaluator.
If no red flags: State "No reward hacking indicators detected" and proceed to Stage 3.
### Code Review: <BEST_NICKNAME> (score: <SCORE>)
**Changes:** <1-2 sentence summary of what changed>
**Mechanism:** <1-2 sentence explanation of why it is better>
**Reward Hacking:** <None detected | WARNING: <concern>>
**Recommendation:** <Integrate | Review carefully before integrating>After the code review summary, present the findings to the user and then proceed immediately to Stage 3 to generate the report artifacts. Do NOT ask the user any questions here — the integration question comes after the report is generated.
Objective: Analyze the experiment results in depth and produce three report artifacts. This stage always runs — it does not require user confirmation.
The three artifacts are produced in sequence, each building on the previous one:
ae results plot <EXPERIMENT> --output <PROJECT_DIR>/score_progression.pngThis must be generated first because the markdown report embeds it.
Examine the top 3-5 programs to understand what strategies were explored and why the winner won.
For each of the top 3-5 programs:
ae --json program show <NICKNAME> --experiment <EXP> --codeBuild an "Approaches Explored" table (this goes into the markdown report):
| Approach | Program | Score | Outcome |
|-----------------------|---------------|--------|-----------------------------------|
| Tournament merge sort | swift-panda | 2.891 | Best -- O(n log n) fits the data |
| Radix sort | calm-fox | 2.756 | Fast but high memory overhead |
| Hybrid quicksort | bold-eagle | 2.512 | Good average case, worse worst |
| Original brute force | (baseline) | 2.107 | O(n^2), baseline |Include the top 3-5 programs plus the baseline. Each row MUST have all four columns filled in.
Write a "Why X won" paragraph. Analyze WHY the winning program outperformed the others. Connect it to the problem's specific characteristics. Consider:
Example:
Why tournament merge sort won: The input data has high variance in element distribution, which causes quicksort's pivot selection to degrade to O(n^2) on ~15% of inputs. Merge sort's guaranteed O(n log n) avoids this worst case. Radix sort was faster on uniform data but consumed 3x memory, exceeding the evaluator's constraints.
Note: If the top programs all use essentially the same approach with minor parameter tweaks (common for numerical optimization), state that in the report rather than trying to differentiate them artificially:
All top programs use the same core approach (gradient descent with momentum). Differences are in hyperparameter tuning: learning rate (0.001-0.01), momentum (0.9-0.99), and batch size (32-128).
ae --json results failed <EXPERIMENT>Group failures by common error patterns (e.g., "SyntaxError", "TypeError", "timeout", "numerical overflow"). This produces the data for the "Failure Analysis" table in the markdown report, formatted as:
| Error Pattern | Count | Example |
|-------------------|-------|--------------------------------------|
| SyntaxError | 5 | Unmatched parenthesis in evolved fn |
| TimeoutError | 3 | Infinite loop in sorting logic |
| ValueError | 2 | NaN in loss computation |
| Total failures | 10/42 | 23.8% failure rate |If no programs failed, state "All evaluations succeeded (0% failure rate)." in the report.
If the metric represents a measurable real-world quantity (latency, throughput, memory usage, accuracy, error rate), translate the raw improvement into practical terms for the "Practical Impact" section.
| Metric type | Include? | Example framing |
|---|---|---|
| Latency/time | Yes | "X ms → Y ms per call, Z% faster" |
| Throughput | Yes | "X ops/sec → Y ops/sec" |
| Memory | Yes | "X MB → Y MB, Z MB freed" |
| Accuracy/F1 | Yes | "X% → Y%, +Z percentage points" |
| Abstract score | No | Pure optimization scores have no real-world unit |
| Mathematical | No | Combinatorial quality metrics don't map to cost/time |
Example format:
| Before | After | Improvement |
|------------------|------------------|--------------------|
| 847ms/call | 187ms/call | 660ms saved (4.5x) |Do NOT fabricate cost estimates. Only include cost/time projections if the user provided context about their usage patterns (e.g., "this runs 10k times/day"). If no such context exists, just show the before/after metric values.
Write the report to <PROJECT_DIR>/experiment_report.md. The report MUST
embed the chart image using the correct relative path to the PNG generated in
Step 3.1. If both files are in the same directory, use . Double-check that the filename matches the
--output path used in the plot command.
Follow this template exactly. Every section below MUST appear in the
generated report in the same order. Replace the <PLACEHOLDER> values with
actual data. Do not omit sections, do not reorder them, and do not add extra
sections before the Reproduction section.
<!-- *** CRITICAL: USE AUTHORITATIVE SCORES *** -->
The <BASELINE>, <BEST>, <REL>, and <BEST_NICKNAME> placeholders in the
template MUST use the authoritative values recorded in Step 1.1 (from results best and program show). Do NOT recompute these values from the approach
analysis in Step 3.2 — when many programs are loaded in context, it is easy to
accidentally swap scores between programs.
<!-- *** END CRITICAL *** -->
<!-- *** CRITICAL: APPROACHES TABLE *** -->
The "Approaches Explored" table is MANDATORY. This is the most commonly skipped section — do NOT skip it. You MUST include a markdown table with at least 3 rows (top programs + baseline) using the data from Step 3.2. Each row MUST have all four columns filled in. If you could not fetch code in Step 3.2, use "Variant N" as the approach label — but the table MUST still appear with scores and nicknames.
<!-- *** END CRITICAL *** -->
# AlphaEvolve Experiment Report: <NICKNAME>
**Date:** <DATE>
**Status:** <STATE>
**Model:** <MODEL>
**Duration:** <DURATION>
## Result
<BASELINE> → <BEST> (+<REL>%)
## Practical Impact (if applicable)
| Before | After | Improvement |
|------------------|------------------|--------------------|
<IMPACT_ROWS>
(Omit this section if the metric is abstract — see Step 3.4.)
## Results Summary
| Metric | Value |
|---------------------|------------------------------------------|
| Baseline Score | <BASELINE> |
| Best Score | <BEST> |
| Improvement | +<ABS> (<REL>%) |
| Total Evaluations | <TOTAL> |
| Success Rate | <RATE>% |
## Score Progression

## Approaches Explored
| Approach | Program | Score | Outcome |
|-----------------------|---------------|--------|----------------------------------|
| <approach_label> | <nickname> | <score>| <1-sentence outcome> |
| <approach_label> | <nickname> | <score>| <1-sentence outcome> |
| ... | ... | ... | ... |
| Original | (baseline) | <score>| Baseline |
**Why <BEST_APPROACH> won:** <EXPLANATION>
## What Changed
\`\`\`python
# Before
<KEY_LINES_FROM_BASELINE — focus on the 5-15 lines that matter>
# After
<KEY_LINES_FROM_EVOLVED — the corresponding changed lines>
\`\`\`
## Best Program: <BEST_NICKNAME>
### Full Evolved Code
\`\`\`python
<evolved code block>
\`\`\`
## Failure Analysis
| Error Pattern | Count | Example |
|-------------------|-------|--------------------------------------|
| <pattern> | <N> | <brief description> |
| ... | ... | ... |
| Total failures | <N>/<TOTAL> | <failure_rate>% |
(Or: "All evaluations succeeded (0% failure rate).")
## All Top Programs
| Rank | Nickname | Score | vs Baseline |
|------|---------------|--------|-------------|
<TOP_PROGRAMS_TABLE>
## Reproduction
To reproduce this experiment:
\`\`\`bash
ae --json experiment describe <NICKNAME>
ae --json results best <NICKNAME> --top 10
ae --json program show <BEST_NICKNAME> --experiment <NICKNAME> --code
\`\`\`ae results report <EXPERIMENT> \
--markdown <PROJECT_DIR>/experiment_report.md \
--output <PROJECT_DIR>/experiment_report.htmlThe HTML report is the final artifact. It combines:
After generating all three artifacts, tell the user where they are:
Report artifacts saved: - Score chart:
<full_absolute_path>/score_progression.png- Markdown report:<full_absolute_path>/experiment_report.md- Interactive report:<full_absolute_path>/experiment_report.html(open in a browser to explore the interactive chart)
Objective: Apply the evolved code back to the user's original source file, validate it works, and leave the codebase in a clean state. This stage only runs if the user confirms.
<!-- *** MANDATORY USER INTERACTION *** -->
After presenting the report artifacts, ask the user whether they want to integrate the evolved code back into their codebase:
Would you like me to integrate the best evolved code back into your codebase? (y/n)
You can also: - Review another program: Specify a rank (e.g., "show me #2") - Compare two programs: "compare #1 and #3"
Accept bare "yes", Enter, "y", or "looks good" as confirmation to proceed with integration. If the user declines, skip to Step 4.6 (completion summary).
<!-- *** END MANDATORY USER INTERACTION *** -->
The source map is the authoritative guide for integration. It records exactly where each code region in the experiment came from and how to apply changes back.
Check for .evolve/source_map.json:
cat <PROJECT_DIR>/.evolve/source_map.jsonIf source_map.json exists, use it to drive integration (Step 4.2a).
If it does NOT exist (older experiments or standalone experiments), fall back to heuristic detection (Step 4.2b).
Parse the source map. Each entry in mappings describes one code region:
| Field | Meaning |
|---|---|
experiment_file | Which file in the experiment directory contains this |
| : : code : | |
symbol | The function/class name (null for full-file mappings) |
is_evolve_block | Whether this region was inside an EVOLVE-BLOCK (i.e., |
| : : evolved) : | |
original_file | Path to the original source file in the user's codebase |
original_lines | [start, end] line range in the original file |
original_symbol | Symbol name in the original file (may differ from |
: : symbol if renamed) : | |
integration_mode | How to apply: function_replacement, |
: : full_file_replacement, evolve_block_replacement, or : | |
: : skip : |
Integration plan: Filter to entries where is_evolve_block is true (these
are the regions that were evolved). Group by original_file -- each unique
original file becomes one integration target. Skip entries with
integration_mode: "skip".
Present the integration plan to the user:
### Integration Plan
| Original File | Symbol | Mode |
|--------------------------------|--------|----------------------|
| src/core/activation.py | relu | function_replacement |
| src/models/layers.py | Dense | skip (not evolved) |If no source map exists, determine integration targets from these sources (in priority order):
original_source_file artifact.<PROJECT_DIR>/.evolve/experiment_description.json and check source_file
or source_files[].path.# ORIGIN: comments in the evolved code to
identify original file paths and symbol names.Also determine the integration mode per target:
| Mode | When | What happens |
|---|---|---|
| **EVOLVE-BLOCK | Original file has | Replace content between |
: replacement** : EVOLVE-BLOCK markers : markers : | ||
| **Function | Specific function was | Replace the function body |
| : replacement** : extracted : : | ||
| **Full file | Entire file was the | Replace the file contents |
| : replacement** : experiment target : : | ||
| Manual (new file) | No original file | Save evolved code as a new |
| : : (standalone experiment) : file : |
Read the best program's code. Extract the evolved regions:
Parse EVOLVE-BLOCK markers: Find all # EVOLVE-BLOCK-START / # EVOLVE-BLOCK-END pairs. Extract the code between each pair.
Strip experiment scaffolding. Before writing any code into the user's source files, remove all experiment artifacts line by line:
# EVOLVE-BLOCK-START and # EVOLVE-BLOCK-END marker lines# ORIGIN: ... comment linesevaluate() function (unless it existed in the original
file)if __name__ == "__main__": block (unless it existed in
the original file)Read references/integration_patterns.md (Scaffolding Cleanup section) for
the full line-by-line stripping procedure, indentation adjustment rules, and
verification steps.
After stripping, verify no scaffolding remains: search for
EVOLVE-BLOCK and # ORIGIN: in the prepared code -- both must return zero
matches.
Match evolved regions to integration targets. Use the source map (Step
4.2a) or ORIGIN comments (Step 4.2b) to match each evolved region to its
original file and location:
ORIGIN comment says src/core/activation.py::relu (lines 12-18),
the evolved code for relu goes back to that file.initial_program.py has multiple EVOLVE-BLOCKs, match them
to their respective ORIGIN comments by position (first block matches
first ORIGIN before a block, etc.).Handle multi-file experiments. If the evolved program has multiple files, compare each against its initial version. Only integrate files that actually changed (files without EVOLVE-BLOCKs should be identical to the initial versions -- skip them).
Prerequisite: Step 4.3 must have already stripped all scaffolding (
EVOLVE-BLOCKmarkers,ORIGINcomments, experiment boilerplate) from the evolved code. The code applied in this step must be clean.
Before applying any changes, consult the provenance metadata to determine exactly where each code region belongs:
.evolve/source_map.json (if it exists) for the authoritative mapping
of experiment symbols to original file paths, line ranges, and integration
modes.# ORIGIN: comments in the evolved code for
file paths and symbol names.references/integration_patterns.md for the full reference on
provenance-driven integration, multi-file handling, scaffolding cleanup, and
rollback procedures.For each integration target from the plan:
Function replacement (function_replacement):
original_symbol from the source
map, which may differ from the experiment's symbol name if it was renamed
for flat imports).EVOLVE-BLOCK replacement (evolve_block_replacement):
# EVOLVE-BLOCK-START / # EVOLVE-BLOCK-END markers.EVOLVE-BLOCK-START and
EVOLVE-BLOCK-END comment lines. They were added during experiment design
and should not persist in the user's codebase.Full file replacement (full_file_replacement):
<original>.bak.ORIGIN and EVOLVE-BLOCK comments from the written file.Manual / new file (no source map, standalone experiment):
<PROJECT_DIR>/evolved_program.py.Important: when multiple original files need updating (e.g., Extract and Isolate inlined functions from 3 different files, and 2 of those functions were in EVOLVE-BLOCKs), apply changes to each file independently. Read each original file, locate the target symbol, replace it, and write. Do NOT try to write all changes to a single file.
After applying changes, validate that the integrated code works:
Step 4.5a: Syntax check
python3 -c "import ast; ast.parse(open('<TARGET_FILE>').read())"On Windows, use
pythoninstead ofpython3.
If syntax check fails, the integration introduced a syntax error. Revert the change and report the issue.
Step 4.5b: Re-run the evaluator
Run the evaluator against the integrated file to confirm the score is preserved:
ae --json program evaluate \
--program-file <TARGET_FILE> \
--evaluator <PROJECT_DIR>/evaluator.py \
--backend localCompare the score against the experiment's best score. If the scores match (within 1% tolerance for floating-point differences), the integration is successful.
If scores diverge significantly:
Score mismatch after integration. The evolved code scored
<X>in the experiment but<Y>after integration. This may indicate: - Import differences between the experiment and the original codebase - Context that was available during experiment evaluation but not in the integrated file - Multi-file dependencies that need to be copied overWould you like me to investigate? (y/n)
Step 4.5c: Run existing tests (if available)
If the original source file has associated tests (e.g., test_<filename>.py in
the same directory, or tests referenced in pyproject.toml):
# For uv-managed projects:
uv run pytest <test_file>
# For general Python projects (use "python" on Windows):
python3 -m pytest <test_file>If tests fail, report the failures and offer to revert:
Existing tests failed after integration.
<N>test(s) failed:<brief summary of failures>Options: 1. Revert the integration (restore original file) 2. Keep the changes and fix the test failures 3. Review the evolved code for compatibility issues
After everything is done, present a completion message that leads with the result and the key insight.
If integration was performed:
Result:
<BASELINE>→<BEST>(+<REL>% improvement)The evolved code from
<BEST_NICKNAME>has been applied to<TARGET_FILE>.What changed: <1-sentence summary of the core algorithmic change>
Why it works: <1-sentence explanation connecting the change to the problem>
<PRACTICAL_IMPACT if applicable, e.g., "This saves X ms per call">
- Validation: <PASSED/FAILED>
- Report:
<REPORT_PATH>- Chart:
<CHART_PATH>- Interactive report:
<HTML_REPORT_PATH>(open in browser, if generated)Next steps: - Review the changes in your editor - Run your full test suite to verify compatibility - Start a new experiment with refined parameters if needed
If integration was skipped:
Result:
<BASELINE>→<BEST>(+<REL>% improvement)What changed: <1-sentence summary>
Why it works: <1-sentence explanation>
- Best program:
<BEST_NICKNAME>- Report:
<REPORT_PATH>- Chart:
<CHART_PATH>- Interactive report:
<HTML_REPORT_PATH>(open in browser, if generated)To integrate later, you can retrieve the best program with:
ae --json program show <BEST_NICKNAME> --experiment <NICKNAME> --code
If the experiment failed (no successful programs):
Experiment did not produce improvements. All
<N>evaluations either failed or scored equal to or below the baseline.Common causes and next steps: - Evaluator too strict: Relax constraints or increase timeout - Search space too narrow: Expand the EVOLVE-BLOCK boundaries - Problem too hard for the model: Try a different model or decompose the problem - Report:
<REPORT_PATH>
For any ae command failure:
references/debugging.md for known error patterns.Common error patterns specific to post-experiment processing:
| Error | Likely Cause | Fix |
|---|---|---|
| "experiment not found" | Wrong nickname/ID | Run `ae --json |
| : : : experiment list` : | ||
| "program not found" | Wrong program nickname | Run `ae --json results |
: : : best <exp>` to find : | ||
| : : : correct name : | ||
| Program code is empty | Backend did not store | Try `ae --json program |
: : code : show <prog> --code : | ||
| : : : --output-file out.py` : | ||
| Diff fails | Parent program ID format | Fall back to manual diff |
| : : issue : (Step 2.2) : | ||
| Eval score mismatch | Environment differences | Check imports and |
| : : : dependencies : | ||
| Integration syntax | Marker mismatch or | Revert and re-extract |
| : error : indentation : carefully : |
See references/cli_reference.md for the full command reference. Key commands
for this skill:
| Command | Purpose |
|---|---|
ae --json experiment describe <exp> | Get experiment metadata |
ae --json results best <exp> --top N | Top N programs by score |
ae --json results history <exp> | All programs by creation time |
ae --json results failed <exp> | Failed programs with errors |
ae --json program show <prog> --code | Get program source code |
`ae --json program show <prog> --code | Save code to file |
: --output-file <f>` : : | |
ae program diff <prog> | Diff vs parent (no --json) |
| `ae --json program evaluate --program-file | Validate integrated code |
: <f> --evaluator <e>` : : |
© Google-Cloud-AI, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (references) in skills/alpha_evolve_post_experiment of Google-Cloud-AI/alphaevolve-on-googlecloud.
Open the folder on GitHubat commit 674dd5e
Alpha Evolve Post Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Alpha Evolve Post Experiment this skillGoogle-Cloud-AI/alphaevolve-on-googlecloud | 118 | — | ~9.2k | Automated safety check: Pass | Apache-2.0 | |
| Finishing a Development Branchobra/superpowers | 296k | 5 repos | ~1.9k | Automated safety check: Pass | MIT | |
| Typescript Advanced Typesrolling-scopes/rsschool-app | 10k | 24 repos | ~4.2k | Automated safety check: Pass | MPL-2.0 | |
| PR Babysitteropeninterpreter/openinterpreter | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Code Review ChecklistshareAI-lab/learn-claude-code | 78k | 5 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Greplooponyx-dot-app/onyx | 32k | 4 repos | ~3.3k | Automated safety check: Pass | MIT |
obra/superpowers
Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.
rolling-scopes/rsschool-app
Master TypeScript's advanced type system including generics, conditional types, mapped types, template literals, and utility types for building type-safe applications.
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
shareAI-lab/learn-claude-code
Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.
onyx-dot-app/onyx
Iteratively improves a PR (GitHub), MR (GitLab), or shelved changelist (Perforce) until Greptile gives it a 5/5 confidence score with zero unresolved comments.
akash-network/node
Behavioral guidelines to reduce common LLM coding mistakes. An agent skill from akash-network/node.
Google-Cloud-AI/alphaevolve-on-googlecloud
Monitor running AlphaEvolve experiments, run the evaluation control loop, and report results using the ae CLI.
Google-Cloud-AI/alphaevolve-on-googlecloud
End-to-end AlphaEvolve experiment orchestrator. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud.
Google-Cloud-AI/alphaevolve-on-googlecloud
AlphaEvolve expert consultant grounded strictly in the official reference guide.
Google-Cloud-AI/alphaevolve-on-googlecloud
Configure, verify, and launch AlphaEvolve experiments using the ae CLI.
Google-Cloud-AI/alphaevolve-on-googlecloud
Design AlphaEvolve experiments for the Cloud API. An agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud.
Categories
Post-experiment analysis, visualization, and code integration for completed AlphaEvolve experiments. Alpha Evolve Post Experiment is an agent skill from Google-Cloud-AI/alphaevolve-on-googlecloud. Post-experiment analysis, visualization, and code integration for completed AlphaEvolve experiments.
Alpha Evolve Post Experiment fits situations like: : show experiment results; analyze experiment; integrate evolved code; experiment report.
Run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a claude-code`. Or copy the skill folder (skills/alpha_evolve_post_experiment in Google-Cloud-AI/alphaevolve-on-googlecloud) into .claude/skills/alpha-evolve-post-experiment in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a codex`. Or copy the skill folder (skills/alpha_evolve_post_experiment in Google-Cloud-AI/alphaevolve-on-googlecloud) into .agents/skills/alpha-evolve-post-experiment in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Google-Cloud-AI/alphaevolve-on-googlecloud --skill alpha-evolve-post-experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/alpha-evolve-post-experiment, .gemini/skills/alpha-evolve-post-experiment, .github/skills/alpha-evolve-post-experiment and .opencode/skills/alpha-evolve-post-experiment in your project.
Going by SKILL.md and its folder, Alpha Evolve Post Experiment needs the command-line tools its instructions call (python3 and uv).
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Alpha Evolve Post Experiment is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 9.2k tokens (SKILL.md is roughly 37k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 8.3k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Alpha Evolve Post Experiment: Finishing a Development Branch (obra/superpowers, 296k stars), Typescript Advanced Types (rolling-scopes/rsschool-app, 10k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars) and Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Google-Cloud-AI (a GitHub organization) maintains it in Google-Cloud-AI/alphaevolve-on-googlecloud, which has 118 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on October 1, 2026.
Source: Google-Cloud-AI/alphaevolve-on-googlecloud on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.