Agent skill

A-Evolve Agent Improvement

by aiming-lab in aiming-lab/AutoResearchClaw

Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.

MITAuto-check passedAI & LLM Engineering

Install A-Evolve Agent Improvement

skills CLI
$ npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aiming-lab/AutoResearchClaw a-evolve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/a-evolve .claude/skills/a-evolve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
a-evolve
GitHub stars
15k
Token cost
~1.8k tokens
SKILL.md length
703 words
Files
1
Skills in repo
34
Repo updated
First seen
Licence
MIT

At a glance

Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.

  • Works in 5 steps: Solve (Collect Evidence) → Observe (Diagnose) → Evolve (Propose Mutations) → …
  • Working out why an agent keeps failing the same kind of task
  • SKILL.md covers Core Loop, Usage with AutoResearchClaw, Anti-Patterns and Relationship to MetaClaw
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill applies a five-step loop named Solve, Observe, Evolve, Gate, Reload. It is prompt-based, so it needs no external dependencies and no harness changes. The agent first gathers evidence such as run logs, error traces, pass or fail results per task and metric values, and inside an AutoResearchClaw pipeline it also looks at artifacts/rc-* outputs, evolve.log and reviews.md.

In the Observe step each failed or weak task gets an error category, a root cause, a frequency and a severity, written as a structured list of observations. In the Evolve step the agent proposes changes: a new SKILL.md for a recurring pattern seen at least 3 times, which targets one failure category and stays under 100 lines, or a short system prompt patch for ambiguity or missing guidance. The result is durable artifacts the agent can load in later runs.

When your agent uses it

  • Working out why an agent keeps failing the same kind of task
  • Turning recurring error patterns into a new targeted skill
  • Patching a system prompt that leaves the agent's behavior ambiguous
  • Reviewing logs from an AutoResearchClaw run to see what went wrong

Example prompts

  • “Diagnose the failures in these logs and write a skill for the API pagination errors.”
  • “Evolve my agent's system prompt based on the pass or fail results in results/run.json.”
  • “What went wrong in the last pipeline run, and how do we fix it?”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Solve (Collect Evidence)
  2. Observe (Diagnose)
  3. Evolve (Propose Mutations)
  4. Gate (Validate)
  5. Reload (Apply and Record)

What it can do on your machine

Read from SKILL.md and the folder at commit be4ba47. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

A-Evolve Agent Improvement loads about 1.8k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 703 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aiming-lab/AutoResearchClaw at commit be4ba47, republished under its MIT licence (© aiming-lab). 703 words, ~1,810 tokens.

Download SKILL.mdSave it as .claude/skills/a-evolve/SKILL.md (or your agent's skills folder).
name
a-evolve
description
Apply A-Evolve's agentic evolution methodology to improve AI agent performance across runs. Use when the user wants to diagnose agent failures, generate targeted skills from error patterns, evolve system prompts, or accumulate episodic knowledge. Works standalone or inside AutoResearchClaw pipelines. Triggers on: "evolve", "self-improve", "diagnose failures", "generate skills from errors", "what went wrong and how to fix it", or any mention of A-Evolve.

A-Evolve: Agentic Evolution Skill

Apply the Solve → Observe → Evolve → Gate → Reload methodology from A-Evolve to iteratively improve agent performance. This skill is prompt-based — no external dependencies, no harness changes. You analyze failures, propose workspace mutations, and generate durable artifacts (skills, prompt patches, knowledge entries) that the agent can load in future runs.

Core Loop

When asked to evolve or improve agent performance, follow this 5-step loop:

1. Solve (Collect Evidence)

Gather the agent's execution artifacts. Ask the user for or locate:

  • Run logs, error traces, or experiment outputs
  • Pass/fail results per task
  • Metric values (accuracy, reward, success rate)
  • Any existing session files from previous runs

If inside AutoResearchClaw, look at:

  • artifacts/rc-*/ — experiment outputs, charts, reviews
  • evolve.log or stage-specific logs
  • reviews.md — peer review feedback
  • Sentinel watchdog reports
2. Observe (Diagnose)

Analyze the collected evidence to produce structured observations:

For each failed or underperforming task, identify:

  • Error category: code bug, timeout, wrong approach, missing knowledge, API misuse, hallucinated reference, prompt ambiguity, etc.
  • Root cause: What specifically went wrong and why
  • Frequency: Is this a one-off or a recurring pattern across tasks?
  • Severity: blocking (pipeline crash) / degrading (wrong result) / cosmetic (formatting issue)

Write observations as a structured list:

## Observations (Batch N)

### OBS-1: [Category] Short description
- Tasks affected: task_001, task_005, task_012
- Root cause: ...
- Frequency: 3/50 tasks (6%)
- Severity: degrading

### OBS-2: ...
3. Evolve (Propose Mutations)

Based on observations, propose one or more of these mutation types:

A. Generate a Skill (for recurring patterns, frequency ≥ 3)

Write a new SKILL.md file that teaches the agent how to handle this pattern. A good evolved skill:

  • Targets a specific failure category, not generic advice
  • Contains concrete steps the agent should follow
  • Includes a "when to apply" trigger condition
  • Is short (under 100 lines) and self-contained

Example — if the agent keeps failing at API pagination:

markdown
---
name: api-pagination-handler
description: >
  Handle paginated API responses correctly. Use when making API calls
  that may return partial results, or when results seem truncated.
---

When calling any API that supports pagination:

1. Check response for pagination indicators: `next_page`, `offset`,
   `has_more`, `cursor`, or truncated result counts.
2. If paginated, loop until all pages are collected.
3. Concatenate results before processing.
4. Set a max-page safety limit (default: 20) to prevent infinite loops.
5. Log total items collected vs expected count if available.

B. Patch the System Prompt (for prompt ambiguity or missing guidance)

Write a short addendum to the system prompt that addresses the gap. Keep patches minimal — one paragraph per issue. Format:

## Prompt Patch: [Issue]
Append to system prompt:
> When [specific situation], always [specific action] because [reason].

C. Add a Knowledge Entry (for factual gaps or learned heuristics)

Record a reusable insight as a knowledge entry:

json
{
  "id": "know-001",
  "category": "experiment_design",
  "insight": "Synthetic benchmarks with <100 samples produce high-variance results. Always use ≥500 samples or report confidence intervals.",
  "source": "observation OBS-3 from batch 2",
  "confidence": 0.85
}

D. Do Nothing (if observation is a one-off, severity is cosmetic, or the fix would be too broad / risky)

4. Gate (Validate)

Before accepting any mutation, check:

  • Specificity: Does it target the observed failure without being so broad it could cause regressions elsewhere?
  • Testability: Could you verify this mutation helps by re-running the failed tasks?
  • Blast radius: How much of the agent's behavior does this change? Prefer small, targeted mutations over large rewrites.
  • Consistency: Does it contradict existing skills or prompt guidance?

If a mutation fails the gate, either refine it or discard it. Explain your reasoning to the user.

Show full SKILL.md (278 more words)Show less
5. Reload (Apply and Record)

Present the accepted mutations to the user. For each:

  • State what changed and why
  • Show the artifact (skill file, prompt patch, knowledge entry)
  • Suggest where to place it in the project

For AutoResearchClaw projects, recommended locations:

ArtifactLocation
Evolved skill.claude/skills/evolved/<skill-name>/SKILL.md
Prompt patchAppend to prompts.default.yaml or custom prompts file
Knowledge entrydocs/kb/evolved_knowledge/<id>.json
Observation logevolution/observations/<batch>.md

Keep a running version log so the user can track what evolved and when:

## Evolution Log
- evo-1 (2026-03-30): Generated `api-pagination-handler` skill from OBS-1
- evo-2 (2026-03-30): Prompt patch for citation format from OBS-4

Usage with AutoResearchClaw

This skill maps to ARC's pipeline stages:

ARC StageEvolution Role
12 EXPERIMENT_RUNSource of Solve artifacts
13 ITERATIVE_REFINEMain Observe + Evolve trigger point
15 RESEARCH_DECISIONNatural Gate — PROCEED = accept, REFINE = retry
18 PEER_REVIEWAdditional Observe signal for writing quality

When the user says "evolve my research pipeline" or similar:

  1. Ask which run to analyze (or find the latest artifacts/rc-*/)
  2. Run the Observe step on experiment outputs + review feedback
  3. Propose mutations targeting the weakest pipeline stages
  4. Generate skill files that ARC can load via .claude/skills/

Anti-Patterns

Do NOT:

  • Generate vague, generic skills ("always be careful", "check your work")
  • Propose mutations for one-off errors that won't recur
  • Rewrite the entire system prompt — patch it surgically
  • Generate more than 3 skills per evolution cycle (quality over quantity)
  • Mutate tool code unless the user explicitly asks for it

Relationship to MetaClaw

If AutoResearchClaw has MetaClaw enabled (metaclaw_bridge.enabled: true), evolved skills from this process can be placed in ~/.metaclaw/skills/arc-*/ so MetaClaw injects them into future runs automatically. The two systems are complementary:

  • A-Evolve skill: Deep, targeted mutation from structured observation
  • MetaClaw lesson: Broad pattern captured from pipeline warnings/errors

Both can coexist. Skills generated here are higher-precision; MetaClaw lessons are higher-recall.

© aiming-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/a-evolve of aiming-lab/AutoResearchClaw.

Open the folder on GitHubat commit be4ba47

Compare with similar skills

A-Evolve Agent Improvement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

A-Evolve Agent Improvement compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
A-Evolve Agent Improvement this skillaiming-lab/AutoResearchClaw15k—~1.8kAutomated safety check: PassMIT
Copilot Session Failure Analysisdotnet/maui23k—~3.4kAutomated safety check: PassMIT
Workflow Schema Tuningbreaking-brake/cc-wf-studio5.4k—~1.3kAutomated safety check: PassCustom licence
Microsoft Foundrymicrosoft/GitHub-Copilot-for-Azure2551 repos~6.7kAutomated safety check: PassMIT
Skill With Prompt EngineeringLeoYeAI/openclaw-master-skills2.2k—~4.1kAutomated safety check: PassMIT
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Workflow Schema Tuning

    breaking-brake/cc-wf-studio

    Guides edits to cc-wf-studio's workflow schema so AI agents generate better workflows, treating schema text as prompt engineering rather than validation.

    5.4k GitHub stars~1.3k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Microsoft Foundry

    microsoft/GitHub-Copilot-for-Azure

    Official

    Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end.

    255 GitHub starsUsed in 1 repo~6.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Skill With Prompt Engineering

    LeoYeAI/openclaw-master-skills

    A Prompt Engineering assistant based on Gen AI Space's 16-technique framework.

    2.2k GitHub stars~4.1k tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed

More from aiming-lab/AutoResearchClaw

All 34 skills in this repo
  • Metabolic Study Planner

    aiming-lab/AutoResearchClaw

    Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.

    15k GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • MFA Pipeline Orchestrator

    aiming-lab/AutoResearchClaw

    Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.

    15k GitHub stars~923 tokensUpdated 1 mo ago
    Auto-check passed
  • Qiskit 2.x Quantum ML Reference

    aiming-lab/AutoResearchClaw

    Reference patterns for writing qiskit 2.x code for variational quantum machine learning: feature maps, VQC training, VQE for chemistry, MPS circuits and noise models.

    15k GitHub stars~4.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Genome-Scale Metabolic Model Builder

    aiming-lab/AutoResearchClaw

    Builds or loads a genome-scale metabolic model in COBRApy, sets its growth medium and objective, and exports it as a validated JSON file for flux analysis.

    15k GitHub stars~1.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Biopython Bioinformatics

    aiming-lab/AutoResearchClaw

    Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.

    15k GitHub stars~810 tokensUpdated 1 mo ago
    Auto-check passed
  • RDKit Cheminformatics Practices

    aiming-lab/AutoResearchClaw

    Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.

    15k GitHub stars~708 tokensUpdated 1 mo ago
    Auto-check passed

Questions about A-Evolve Agent Improvement

What does A-Evolve Agent Improvement do?

Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop. The skill applies a five-step loop named Solve, Observe, Evolve, Gate, Reload. It is prompt-based, so it needs no external dependencies and no harness changes.

When should I use A-Evolve Agent Improvement?

A-Evolve Agent Improvement fits situations like: working out why an agent keeps failing the same kind of task; turning recurring error patterns into a new targeted skill; patching a system prompt that leaves the agent's behavior ambiguous; reviewing logs from an AutoResearchClaw run to see what went wrong.

How do I install A-Evolve Agent Improvement in Claude Code?

Run `npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a claude-code`. Or copy the skill folder (.claude/skills/a-evolve in aiming-lab/AutoResearchClaw) into .claude/skills/a-evolve in your project. Claude Code loads it when a task matches its description.

How do I install A-Evolve Agent Improvement in Codex?

Run `npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a codex`. Or copy the skill folder (.claude/skills/a-evolve in aiming-lab/AutoResearchClaw) into .agents/skills/a-evolve in your project. Codex loads it when a task matches its description.

Can I use A-Evolve Agent Improvement in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/a-evolve, .gemini/skills/a-evolve, .github/skills/a-evolve and .opencode/skills/a-evolve in your project.

What does A-Evolve Agent Improvement need to run?

SKILL.md names no scripts, command-line tools or credentials: A-Evolve Agent Improvement is instructions for the agent only.

Does A-Evolve Agent Improvement access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is A-Evolve Agent Improvement safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does A-Evolve Agent Improvement use?

A-Evolve Agent Improvement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does A-Evolve Agent Improvement use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to A-Evolve Agent Improvement?

Skills that share tags, products or a category with A-Evolve Agent Improvement: Copilot Session Failure Analysis (dotnet/maui, 23k stars), Workflow Schema Tuning (breaking-brake/cc-wf-studio, 5.4k stars), Microsoft Foundry (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Skill With Prompt Engineering (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains A-Evolve Agent Improvement?

aiming-lab (a GitHub organization) maintains it in aiming-lab/AutoResearchClaw, which has 14,587 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on August 19, 2026.

Source: aiming-lab/AutoResearchClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.