A-Evolve Agent Improvement
aiming-lab/AutoResearchClaw
Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.
A skill your agent uses when the user wants to improve something by repeated search rather than one edit — a program, a prompt, a document, a pipeline, a configuration, an experimental protocol.
$ npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery evolve-design --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evolve-design .claude/skills/evolve-design && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evolve-design" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/evolve-design into .claude/skills/evolve-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evolve-design", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/evolve-designType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery evolve-design --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evolve-design .agents/skills/evolve-design && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evolve-design" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/evolve-design into .agents/skills/evolve-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evolve-design", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery evolve-design --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evolve-design .cursor/skills/evolve-design && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evolve-design" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/evolve-design into .cursor/skills/evolve-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evolve-design", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/openJiuwen-ai/sciencediscovery.git --path skills/evolve-design--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery evolve-design --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evolve-design .gemini/skills/evolve-design && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evolve-design" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/evolve-design into .gemini/skills/evolve-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evolve-design", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install openJiuwen-ai/sciencediscovery evolve-designInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evolve-design .github/skills/evolve-design && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evolve-design" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/evolve-design into .github/skills/evolve-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evolve-design", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install openJiuwen-ai/sciencediscovery evolve-design --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/openJiuwen-ai/sciencediscovery.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evolve-design .opencode/skills/evolve-design && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evolve-design" agent skill from https://github.com/openJiuwen-ai/sciencediscovery/tree/main/skills/evolve-design into .opencode/skills/evolve-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evolve-design", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evolve-designA skill your agent uses when the user wants to improve something by repeated search rather than one edit — a program, a prompt, a document, a pipeline, a configuration, an experimental protocol.
Evolve Design is an agent skill from openJiuwen-ai/sciencediscovery. Use when the user wants to improve something by repeated search rather than one edit — a program, a prompt, a document, a pipeline, a configuration, an experimental protocol. Triggers on /evolve-design, "把这个做得更好", "搜索一个更好的方案", or any request to optimise against a measurable target. Runs a short checkpointed design conversation, verifies the scoring can rank candidates, then calls createevolverun. Not for a single fix, a refactor, or a question about existing code.
Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/custom-script.md`).
The repository describes itself as: ScienceDiscovery is an all‑in‑one agentic workbench built specifically for scientific research. The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ab1403f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evolve Design loads about 3.2k tokens when it runs, and up to ~3.9k if it reads all its reference files. Until then it costs about 122 tokens; SKILL.md has 2,031 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from openJiuwen-ai/sciencediscovery at commit ab1403f, republished under its Apache-2.0 licence (© openJiuwen-ai). 2,031 words, ~3,214 tokens.
.claude/skills/evolve-design/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.A search rewrites a candidate dozens of times and keeps what scores higher. The candidate can be anything text-shaped: a function, a whole program, a prompt, a report section, a config file, a protocol. It succeeds or fails on whether the scoring can tell a good candidate from a bad one — scoring that cannot shows up as a flat run, not as an error.
Algorithm first. The user's /evolve-design command may carry --algorithm puct or
--algorithm openevolve; if unset, ask. PUCT ranks candidates over a tree, OpenEvolve over
MAP-Elites islands. The four steps and the four scoring modes are identical for both, so design
the scorecard, the split and the starting point the same way and only set algorithm in
create_evolve_run.
Four checkpoints. Do the work for a step, show the result, wait for the user, then go on. A checkpoint is a glance, not a form: a few lines they can wave through, with a recommended default for every open question so "looks right" is always a valid answer. Do the work between checkpoints yourself; do not narrate it.
If the user says "you decide" or "just go", state your choices for the remaining steps in one message and run them without stopping. If they correct something, apply it and re-show that step only.
Do not design the whole thing in one thinking pass. Ask, write files, run them, then decide the numbers — a turn spent planning it all produces no candidate and no result.
Read what is free first: this conversation, the workspace, science memory. Settle three things:
The measurable criterion. "More accurate" — which error measure? "Less AI-sounding" — which specific tic? It must separate the complaint from its opposite: "well-structured and readable" fails, because it is true of any document.
What must not change. Inputs unavailable when the thing actually runs, files that define the score, hard limits (budget, runtime, memory, safety).
The starting point. Use what is in the workspace. Otherwise write the simplest thing that already does the job badly — a dozen lines, no tuning, no edge cases. A strong seed spends the search space before the search begins.
But simplest of the right kind: the seed must contain the mechanism the search is meant to improve, in its feeblest form. A seed with no mechanism makes every candidate invent one from nothing, and from-scratch code fails far more often than an edit. Measured: a compression task seeded with identity encoding spent ten expansions on whole compressors written from scratch, seven of which did not run; seed the RLE, not the identity function. A good seed is necessary, not sufficient — candidates may still replace the whole mechanism.
Do not steer by naming the mechanism. "Add a longer match window to the existing RLE" takes the search's own job away and leaves the run tuning your algorithm. The task says what "better" is measured as, never which approach reaches it. A run that finds nothing above the seed is a result. What is yours to reconsider is the scoring — do the cases reward what you care about, is the corpus wide enough to separate approaches — not the approach you would like.
Checkpoint 1. Say back, in a few lines: what will be measured, what is frozen, and what the starting point is. Ask only for what you genuinely could not infer — all of it at once.
Write the starting point and the evaluator, then run the evaluator once, against the starting
point, with run_shell (for example python evaluator.py, selecting a Python-capable
environment_id). You are checking that it executes and emits a number — a typo, a missing
import, a result file never written. That is the whole of the local check.
Do not score a broken copy locally. The server's discrimination probe does exactly that, on
the real shards in the real sandbox, and hands you both numbers when the run starts. Your
run_shell environment and the candidate sandbox are different places with different shard
indices, so the numbers can legitimately differ (one session saw local 0.3853, server 0.0000, and
spent the turn reconciling them). When they disagree, the server's is true — do not
investigate the gap. A probe refusal costs four sandbox evaluations and no model calls, so it is
cheap to be wrong here.
Aim for a starting point in 0.3–0.7: 0 is a floor, solved is a ceiling. The probe reports where it actually landed.
The evaluator must survive bad candidates — including at import. Most candidates are broken,
and the probe deliberately scores a broken one. Guard import candidate itself (a hollowed-out
module can leave a name as None or raise before any per-case try/except can reach it) and
each individual call. On an import failure score every shard worst: 0.0 on a larger-is-better
scale, not 1.0 — an inverted guard made one live run's winner a module that does not load, at a
perfect 1.0000. It is the evaluator that must be robust, never the candidate. If the script
itself crashes, nothing runs and the whole run is refused.
You are checking the ruler, not looking for the answer. Do not go looking for a candidate that beats the starting point — that is the search's entire job, done by hand at the cost of the turn, and succeeding is worse than failing: you then throw the answer away or seed it, and a strong seed spends the search space. A starting point no obvious variation beats is a good one.
Checkpoint 2. Show one number — what the starting point scored — plus whether the evaluator ran, and one sentence on what the scoring rewards. This is where a criterion that measures the wrong thing gets caught, so make it easy to disagree: name what a candidate could do to score higher, and let the user say whether that is what they want.
If the total will not fit, shrink each unit — do not cut the gate, which is the only number the search steers by.
expansions must be at least 4 × workers, or the first sweep forks only the root and the tree
is flat. By search space: a known defect 4–6; swapping approach or restructuring 12–20; writing
something from scratch 20+. Thinking off unless asked — with it on, one whole-candidate
rewrite can exceed the proxy limit and return nothing.
search — leave it out unless this task argues against the defaults. They are upstream's and
usually right; the parameter descriptions on create_evolve_run carry the ranges and what each
does. Set one only for a reason you can give in a line:
cPuct — lower when the gate is large enough to trust and the budget is small; higher when the
score is noisy or coarse or candidates keep tying. A coarse score is a scoring problem first:
widen the gate before reaching for it.priorExponent — when the failure you expect is a whole-mechanism rewrite (five of eighteen
compression candidates replaced a working RLE+Huffman with a from-scratch arithmetic coder, and
each scored 0). It cannot move the reported score; that comes from held-out shards.Checkpoint 3. Show the shape in a few lines: what one unit is, how much is held out, how many expansions with how many workers, and roughly what that costs in time and model calls. Mention
searchonly if you set it, in one line saying why; otherwise leave it out. This is the last point before real money is spent — say so plainly, and default to the smaller option when unsure.
Call create_evolve_run. Set algorithm to what the user chose. howScored is one sentence
for the user: what it measures, how much is held out. risks: at most two, only ones that change
a decision; empty is fine.
Checkpoint 4. Report what came back — the probe's two numbers and what the search will do — and then end the turn. Do not wait for it, do not call
get_evolve_runto check on it, do not loop. The search runs for minutes to hours; its card streams live progress, and sitting on the turn shows the user nothing they cannot already see and burns the run's own budget.get_evolve_runis for later, when the user asks how it went — then report the actual numbers, never an improvement you have not read.
A refusal is design feedback, not an error. "Cannot discriminate" means the scoring needs harder cases or a more mechanical rubric; "no slope" means the starting point is too strong. Fix it and call again — do not hand the user the server's refusal text as a work order.
Take the first that fits. custom_script is the fallback, not the default: you write and
maintain the measuring apparatus yourself, so every mistake in it is yours. Judge by what the task
is, not by what is in the workspace right now — "there is no table yet" is not a reason to skip
dataset_metric when the task is to predict a column; write the table, then use it.
dataset_metric. Deterministic and
cheapest, and the framework owns the split so you cannot get it wrong. Metrics: accuracy / mae /
r2 / rmse / seconds.test_gate. Take it whenever the user says "write
tests", "make these cases pass", or describes behaviour case by case. The failure text is the
learning signal, and frozenGlobs must include the test paths, or the shortest way to a
higher score is to weaken the tests. Shards ≤ half the case count.custom_script, which you write: simulations, optimisation heuristics,
anything whose quality is a computation with no natural table and no test suite. Read
references/custom-script.md before writing the evaluator — it holds the contract, how it
runs, and the error rule the probe enforces (a crash must say where, not just what).llm_judge. For prose and explanations, anything whose
quality is a reading. The most gameable and the only non-deterministic mode; ask once whether it
could be a custom_script instead.Score on a gradient, not a cliff. For hard limits prefer "stop and score what you have" over "violation scores zero": zeroing lands every failure on the same 0 and leaves nothing to climb.
Never reward a property the candidate can fake. Rewarding "gives specific numbers" produces invented numbers; reward agreement with the given source instead.
Candidates get numpy / pandas / scipy / sklearn and the standard library. Anything else goes in
packages — bare names only (optionally ==version), no paths, URLs, or pip options. Your
run_shell environment and the candidate sandbox are not the same.
© openJiuwen-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file (references) in skills/evolve-design of openJiuwen-ai/sciencediscovery.
Open the folder on GitHubat commit ab1403f
Evolve Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evolve Design this skillopenJiuwen-ai/sciencediscovery | 159 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | |
| A-Evolve Agent Improvementaiming-lab/AutoResearchClaw | 15k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Harness Evolveruvnet/ruflo | 74k | — | ~1.6k | Automated safety check: Notes | MIT | |
| A-Evolve Agent EvolutionOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.6k | Automated safety check: Pass | MIT | |
| Skill Improversickn33/agentic-awesome-skills | 47k | 2 repos | ~1.5k | Automated safety check: Pass | MIT | |
| Improve Oaselastic/kibana | 21k | — | ~4.8k | Automated safety check: Pass | Custom licence |
aiming-lab/AutoResearchClaw
Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.
ruvnet/ruflo
Run @metaharness/darwin evolve {repo} to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and…
Orchestra-Research/AI-Research-SKILLs
Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.
sickn33/agentic-awesome-skills
Iteratively improve a Claude Code skill using the skill-reviewer agent until it meets quality standards.
elastic/kibana
Add or improve OpenAPI descriptions, examples, and code samples for a Kibana API area.
Yeachan-Heo/oh-my-claudecode
Runs an autonomous improvement loop on a repository: agents propose and execute plans, a tournament picks the winner by benchmark, and each round is recorded and plotted.
openJiuwen-ai/sciencediscovery
A skill your agent uses when you need to write and execute Python/R code to process, transform, and analyze data, delivering reproducible computational results with complete code-level methodology…
openJiuwen-ai/sciencediscovery
Operate GitCode issues, PRs, wikis, code/MR refs, and cached org templates.
openJiuwen-ai/sciencediscovery
Inspect a local PDB structure, summarize chains and residue composition, and identify protein atoms near a user-specified ligand or pocket center.
openJiuwen-ai/sciencediscovery
Prepare, launch, monitor, and summarize the real RFdiffusion to ProteinMPNN to Protenix antibody pipeline on a local or remote ScienceDiscovery Runner with sandboxed Ascend NPUs.
openJiuwen-ai/sciencediscovery
A skill your agent uses to orchestrate a multi-domain research team for literature/evidence research and data analysis.
openJiuwen-ai/sciencediscovery
A skill your agent uses when a research workflow needs verified academic source retrieval through literature-search MCP interfaces available in the current session before evidence extraction.
A skill your agent uses when the user wants to improve something by repeated search rather than one edit — a program, a prompt, a document, a pipeline, a configuration, an experimental protocol. Evolve Design is an agent skill from openJiuwen-ai/sciencediscovery. Use when the user wants to improve something by repeated search rather than one edit — a program, a prompt, a document, a pipeline, a configuration, an experimental protocol.
Evolve Design fits situations like: the user wants to improve something by repeated search rather than one edit — a program; A configuration; an experimental protocol; any request to optimise against a measurable target.
Run `npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a claude-code`. Or copy the skill folder (skills/evolve-design in openJiuwen-ai/sciencediscovery) into .claude/skills/evolve-design in your project. Claude Code loads it when a task matches its description.
Run `npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a codex`. Or copy the skill folder (skills/evolve-design in openJiuwen-ai/sciencediscovery) into .agents/skills/evolve-design in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add openJiuwen-ai/sciencediscovery --skill evolve-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evolve-design, .gemini/skills/evolve-design, .github/skills/evolve-design and .opencode/skills/evolve-design in your project.
Going by SKILL.md and its folder, Evolve Design needs the command-line tools its instructions call (python). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evolve Design is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 659 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Evolve Design: A-Evolve Agent Improvement (aiming-lab/AutoResearchClaw, 15k stars), Harness Evolve (ruvnet/ruflo, 74k stars), A-Evolve Agent Evolution (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Skill Improver (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
openJiuwen-ai (a GitHub organization) maintains it in openJiuwen-ai/sciencediscovery, which has 159 GitHub stars. The repository holds 23 skills in this directory. The repository was last updated on October 10, 2026.
Source: openJiuwen-ai/sciencediscovery on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.