Agent skill

A-Evolve Agent Evolution

by Orchestra-Research in Orchestra-Research/AI-Research-SKILLs

Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.

MITAuto-check passedAI & LLM Engineering

Install A-Evolve Agent Evolution

skills CLI
$ npx skills add Orchestra-Research/AI-Research-SKILLs --skill evolving-ai-agents -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Orchestra-Research/AI-Research-SKILLs evolving-ai-agents --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Orchestra-Research/AI-Research-SKILLs.git skills-src && mkdir -p .claude/skills && cp -r skills-src/14-agents/a-evolve .claude/skills/evolving-ai-agents && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evolving-ai-agents
GitHub stars
13k
Token cost
~3.6k tokens
SKILL.md length
937 words
Files
9 (incl. references)
Skills in repo
96
Repo updated
First seen
Licence
MIT

At a glance

Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.

  • Works in 5 steps: Solve — Agent processes a batch of tasks… → Observe — Benchmark evaluates… → Evolve — Evolution engine mutates… → …
  • Optimizing an agent's prompts and skills against a measurable benchmark
  • SKILL.md covers Overview, When to Use A-Evolve, Quick Start and Core Concepts, plus 8 more sections
  • Calls git and pip; needs ANTHROPIC_API_KEY

What it does

A-Evolve keeps everything that can evolve about an agent as files in a workspace, with a `manifest.yaml` plus folders for prompts, skills, memory and tools. Each cycle has the agent solve a batch of benchmark tasks, the benchmark returns feedback on its trajectories, and an LLM-driven evolution engine mutates the workspace files. Changes are gated and can be rolled back, and every change is kept in git history.

A three-line example creates an `Evolver` for the built-in SWE agent and the SWE-bench Verified benchmark and runs 10 cycles. Installation is `pip install a-evolve`, with an anthropic extra or an all-providers extra. The skill presents A-Evolve as an optimizer that sits on top of an existing agent framework, and sends multi-agent orchestration to CrewAI or LangGraph, one-shot tasks to LangChain or LlamaIndex, and prompt-only tuning to DSPy. Reference files cover the API, architecture, design patterns, examples, issues, releases and tutorials.

When your agent uses it

  • Optimizing an agent's prompts and skills against a measurable benchmark
  • Building a self-improving agent with automated gating and rollback
  • Keeping a git-versioned history of every change made to an agent

Example prompts

  • “Set up A-Evolve to improve my coding agent's prompts against SWE-bench Verified.”
  • “Create an agent workspace with a manifest, prompts and skills folders for evolution.”
  • “Run 10 evolution cycles and show me what changed in the workspace.”

Requirements

  • Python with `a-evolve` installed through pip
  • A benchmark that scores the agent's runs

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Solve — Agent processes a batch of tasks from the benchmark
  2. Observe — Benchmark evaluates trajectories, producing (task, trajectory, feedback) triples
  3. Evolve — Evolution engine mutates workspace files based on observations
  4. Gate — Validate mutations (git snapshot before/after for rollback)
  5. Reload — Agent reinitializes from evolved filesystem state

What it can do on your machine

Read from SKILL.md and the folder at commit 773a529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • ANTHROPIC_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

A-Evolve Agent Evolution loads about 3.6k tokens when it runs, and up to ~36k if it reads all its reference files. Until then it costs about 73 tokens; SKILL.md has 937 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~3.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~36k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Orchestra-Research/AI-Research-SKILLs at commit 773a529, republished under its MIT licence (© Orchestra-Research). 937 words, ~3,616 tokens.

Download SKILL.mdSave it as .claude/skills/evolving-ai-agents/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
evolving-ai-agents
description
Provides guidance for automatically evolving and optimizing AI agents across any domain using LLM-driven evolution algorithms. Use when building self-improving agents, optimizing agent prompts and skills against benchmarks, or implementing automated agent evaluation loops.
version
1.0.0
author
A-EVO Lab
license
MIT
tags
Agent Evolution, Self-Improving Agents, Prompt Optimization, LLM, Benchmark Evaluation, Skill Discovery, Agentic AI
dependencies
a-evolve>=0.1.0, pyyaml>=6.0

Evolving AI Agents with A-Evolve

Overview

A-Evolve is universal infrastructure for evolving any AI agent across any domain using any evolution algorithm with zero manual engineering. It represents all evolvable agent state as files (prompts, skills, memory, tools), runs iterative solve-observe-evolve cycles against benchmarks, and uses LLM-driven mutation to improve agent performance automatically.

Benchmark results (Claude Opus 4.6):

  • MCP-Atlas: 79.4% (#1)
  • SWE-bench Verified: 76.8% (~#5)
  • Terminal-Bench 2.0: 76.5% (~#7)
  • SkillsBench: 34.9% (#2)

When to Use A-Evolve

Use A-Evolve when:

  • Optimizing agent prompts, skills, or memory against a measurable benchmark
  • Building self-improving agents with automated gating and rollback
  • Evolving domain-specific tool usage and procedures through LLM-driven mutation
  • Running iterative solve-observe-evolve loops to maximize agent performance
  • Needing reproducible, git-versioned evolution history for every change

Key differentiator: Other frameworks build agents; A-Evolve optimizes them. It sits on top of any agent framework and makes it better through automated evolution.

Do NOT use A-Evolve for:

  • Building multi-agent orchestration from scratch (use CrewAI, LangGraph)
  • One-shot agent tasks with no iteration needed (use LangChain, LlamaIndex)
  • RAG pipeline optimization (use LlamaIndex, Chroma)
  • Prompt-only optimization without skill/memory evolution (use DSPy)

Quick Start

Installation
bash
pip install a-evolve                    # Core
pip install a-evolve[anthropic]         # With Claude support
pip install a-evolve[all]               # All providers
Three-Line Evolution
python
import agent_evolve as ae

evolver = ae.Evolver(agent="swe", benchmark="swe-verified")
results = evolver.run(cycles=10)
print(f"Final score: {results.final_score}")

This copies the built-in SWE seed workspace, runs 10 evolution cycles against SWE-bench Verified, and returns the optimized agent.

Core Concepts

The Agent Workspace

All evolvable state lives as files in a workspace directory:

my-agent/
├── manifest.yaml          # Metadata + entrypoint
├── prompts/
│   ├── system.md          # Main system prompt (evolved)
│   └── fragments/         # Modular prompt pieces
├── skills/
│   └── skill-name/
│       └── SKILL.md       # Reusable procedure with frontmatter
├── memory/
│   ├── episodic.jsonl     # Lessons from failures
│   └── semantic.jsonl     # General knowledge
├── tools/
│   ├── registry.yaml      # Tool manifest
│   └── tool_name.py       # Tool implementations
└── evolution/             # Managed by engine (metrics, history)
The Evolution Loop

Each cycle follows five phases:

  1. Solve — Agent processes a batch of tasks from the benchmark
  2. Observe — Benchmark evaluates trajectories, producing (task, trajectory, feedback) triples
  3. Evolve — Evolution engine mutates workspace files based on observations
  4. Gate — Validate mutations (git snapshot before/after for rollback)
  5. Reload — Agent reinitializes from evolved filesystem state
Three Pluggable Interfaces
python
# 1. Agent — implements solve()
class MyAgent(ae.BaseAgent):
    def solve(self, task: ae.Task) -> ae.Trajectory:
        # Domain-specific solving logic
        return ae.Trajectory(task_id=task.id, output=result, steps=steps)

# 2. Benchmark — implements get_tasks() and evaluate()
class MyBenchmark(ae.BenchmarkAdapter):
    def get_tasks(self, split="train", limit=None) -> list[ae.Task]:
        return [ae.Task(id="1", input="...")]

    def evaluate(self, task: ae.Task, trajectory: ae.Trajectory) -> ae.Feedback:
        return ae.Feedback(success=True, score=0.95, detail="Passed")

# 3. Engine — implements step()
class MyEngine(ae.EvolutionEngine):
    def step(self, workspace, observations, history, trial):
        # Mutate workspace based on observations
        return ae.StepResult(mutated=True, summary="Updated prompts")

Workflow 1: Evolve an Existing Agent

Use when: You have a working agent and want to optimize it against a benchmark.

Critical Requirements:

  • Agent implements BaseAgent.solve() returning Trajectory
  • Benchmark implements BenchmarkAdapter with get_tasks() and evaluate()
  • Seed workspace has manifest.yaml with entrypoint and evolvable layers
  • System prompt exists at prompts/system.md
  • Workspace is a git repo (run git init && git add -A && git commit -m "init")
Steps
python
import agent_evolve as ae

# Configure evolution parameters
config = ae.EvolveConfig(
    batch_size=10,           # Tasks per solve round
    max_cycles=20,           # Maximum evolution iterations
    evolve_prompts=True,     # Mutate system prompt
    evolve_skills=True,      # Discover and refine skills
    evolve_memory=True,      # Build episodic memory
    evolver_model="us.anthropic.claude-opus-4-6-v1",
)

# Point to your agent workspace and benchmark
evolver = ae.Evolver(
    agent="./my-agent-workspace",
    benchmark="swe-verified",     # Or custom BenchmarkAdapter instance
    config=config,
)

# Run evolution
results = evolver.run(cycles=10)

# Inspect results
print(f"Cycles completed: {results.cycles_completed}")
print(f"Final score: {results.final_score}")
print(f"Converged: {results.converged}")
for cycle_num, score in enumerate(results.score_history):
    print(f"  Cycle {cycle_num + 1}: {score:.3f}")
Post-Evolution

The workspace is now optimized. Inspect what changed:

bash
cd my-agent-workspace
git log --oneline              # See evo-1, evo-2, ... tags
git diff evo-1 evo-10          # Compare first and last evolution
cat prompts/system.md          # Read evolved prompt
ls skills/                     # See discovered skills

Workflow 2: Add a Custom Benchmark

Use when: You want to evolve agents on your own domain-specific tasks.

Critical Requirements:

  • Define task format (inputs, expected outputs)
  • Implement scoring logic (0.0–1.0 scale)
  • Prepare task dataset (train + holdout split)
Steps
python
import agent_evolve as ae

class CodeReviewBenchmark(ae.BenchmarkAdapter):
    """Evaluate agents on code review quality."""

    def get_tasks(self, split="train", limit=None):
        tasks = load_review_dataset(split)
        if limit:
            tasks = tasks[:limit]
        return [
            ae.Task(id=t["id"], input=t["diff"], metadata={"expected": t["comments"]})
            for t in tasks
        ]

    def evaluate(self, task, trajectory):
        expected = task.metadata["expected"]
        actual = trajectory.output
        precision, recall = compute_review_metrics(expected, actual)
        f1 = 2 * precision * recall / (precision + recall + 1e-9)
        return ae.Feedback(
            success=f1 > 0.7,
            score=f1,
            detail=f"P={precision:.2f} R={recall:.2f} F1={f1:.2f}",
        )

# Use with any agent
evolver = ae.Evolver(agent="./my-agent", benchmark=CodeReviewBenchmark())
results = evolver.run(cycles=5)

Workflow 3: Create a Custom Evolution Engine

Use when: The default LLM-driven mutation doesn't suit your domain.

Steps
python
import agent_evolve as ae

class RuleBasedEngine(ae.EvolutionEngine):
    def step(self, workspace, observations, history, trial):
        failures = [o for o in observations if not o.feedback.success]
        if not failures:
            return ae.StepResult(mutated=False, summary="No failures to address")

        # Analyze failure patterns
        error_types = categorize_errors(failures)
        prompt = workspace.read_prompt()

        # Append learned rules to prompt
        new_rules = generate_rules(error_types)
        workspace.write_prompt(prompt + "\n" + new_rules)

        return ae.StepResult(
            mutated=True,
            summary=f"Added {len(new_rules)} rules from {len(failures)} failures",
        )

evolver = ae.Evolver(
    agent="./my-agent",
    benchmark="my-benchmark",
    engine=RuleBasedEngine(),
)

Built-in Components

Seed Agents
AgentDomainModelKey Feature
sweSWE-benchClaude Opus 4.6Verify-fix loop, skill proposals
terminalTerminal-BenchClaude Sonnet 4Concurrent timeout, env discovery
mcpMCP-AtlasClaude Opus 4.6MCP server integration
Benchmarks
NameDomainMetric
swe-verifiedCode patchingPass rate
mcp-atlasTool callingAccuracy
terminal2Shell tasksPass rate
skill-benchMulti-step proceduresAccuracy
arc-agi-3Interactive gamesRHAE score
Evolution Algorithms
AlgorithmStrategyBest For
A-Evolve/SkillForgeLLM-driven workspace mutationGeneral-purpose
Guided SynthesisMemory-first, curated skillsSkill discovery
Adaptive EvolutionReward tracking, filtered observationsFine-grained control
Adaptive SkillSkill-centric refinementSkill-heavy domains

Configuration Reference

python
ae.EvolveConfig(
    batch_size=10,              # Tasks per solve round
    max_cycles=20,              # Max evolution iterations
    holdout_ratio=0.2,          # Test set split for gating
    evolve_prompts=True,        # Mutate system prompts
    evolve_skills=True,         # Discover/refine skills
    evolve_memory=True,         # Build episodic memory
    evolve_tools=False,         # Mutate tool implementations
    trajectory_only=False,      # Hide scores from evolver
    evolver_model="us.anthropic.claude-opus-4-6-v1",
    evolver_max_tokens=16384,
    egl_threshold=0.05,         # Convergence epsilon
    egl_window=3,               # Cycles for plateau detection
)

Convergence: Evolution stops early when score improvement is less than egl_threshold over the last egl_window cycles.

Skill Format

Skills are reusable procedures discovered and refined during evolution:

markdown
---
name: verify-edge-cases
description: "TRIGGER when: checking boundary conditions. DO NOT TRIGGER: for happy-path tests."
---

## Pattern
Test all falsy-but-valid values: 0, False, "", [], {}

## Process
1. List all input boundaries
2. Run each against the implementation
3. Check both output AND side effects

Skills accumulate in the workspace skills/ directory. The evolver curates them: ACCEPT new skills, MERGE overlapping ones, SKIP redundant proposals. Target: 5–10 broad skills, not 30 narrow ones.

Common Issues

Show full SKILL.md (374 more words)Show less
Evolution score plateaus early

Cause: Batch size too small or evolver doesn't see enough failure diversity. Fix: Increase batch_size (try 15–20) and ensure benchmark tasks cover diverse failure modes. Set trajectory_only=False so the evolver sees scores.

Agent workspace grows too large

Cause: Skill library bloat from accepting every proposal. Fix: The default SkillForge engine curates skills automatically. If using a custom engine, implement merging logic to consolidate overlapping skills.

Git conflicts during evolution

Cause: Multiple evolution runs on the same workspace. Fix: Each evolver.run() should operate on its own workspace copy. Use Evolver(agent="seed-name") to auto-copy the seed each time.

LLM provider errors during evolution

Cause: Rate limits or authentication issues with the evolver model. Fix: Check evolver_model config. For Bedrock, ensure AWS credentials are configured. For Anthropic, set ANTHROPIC_API_KEY.

Custom agent not picking up evolved state

Cause: Agent doesn't implement reload_from_fs(). Fix: Override reload_from_fs() in your BaseAgent subclass to re-read prompts, skills, and memory from the workspace after each evolution cycle.

Usage Instructions for Agents

When this skill is loaded:

  1. Read this entire file before implementing any evolution workflow
  2. Start with the Quick Start — get a minimal evolution running before customizing
  3. Use built-in seeds when possible — "swe", "terminal", "mcp" have battle-tested configurations
  4. Always initialize git in custom workspaces before running evolution
  5. Check convergence settings — default egl_threshold=0.05 with egl_window=3 may be too aggressive for your domain
  6. Inspect evolved state after each run — read prompts/system.md and skills/ to understand what the evolver learned

Pro Tips:

  • Set trajectory_only=False (default) so the evolver sees scores — this accelerates learning
  • Start with batch_size=10 and adjust based on task diversity
  • Use holdout_ratio=0.2 to prevent overfitting to training tasks
  • After evolution, git diff evo-1 evo-N shows the cumulative effect of all mutations
  • If the evolver isn't finding skills, enrich feedback.detail strings with specific failure reasons

Warning Signs:

  • Score oscillating between cycles → benchmark evaluation may be non-deterministic
  • Skills directory growing past 15+ skills → engine isn't merging/curating properly
  • Prompt growing past 10K chars → evolution is appending without refactoring
  • converged=True after 2-3 cycles → increase egl_window and decrease egl_threshold

References

© Orchestra-Research, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in 14-agents/a-evolve of Orchestra-Research/AI-Research-SKILLs.

  • SKILL.md
  • references/README.md
  • references/api.md
  • references/architecture.md
  • references/design-patterns.md
  • references/examples.md
  • references/issues.md
  • references/releases.md
  • references/tutorials.md

Open the folder on GitHubat commit 773a529

Compare with similar skills

A-Evolve Agent Evolution next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

A-Evolve Agent Evolution compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
A-Evolve Agent Evolution this skillOrchestra-Research/AI-Research-SKILLs13k—~3.6kAutomated safety check: PassMIT
Swarms Multi-Agent Frameworkkyegomez/swarms7.2k—~5.5kAutomated safety check: PassApache-2.0
Uipath FunctionsUiPath/skills168—~3.6kAutomated safety check: NotesMIT
Strandsstrands-agents/harness-sdk8.7k—~1kAutomated safety check: PassApache-2.0
Analyzing Claude Code Sessionsamd/gaia1.6k—~2.3kAutomated safety check: PassMIT
Clawpathy AutoresearchClawBio/ClawBio1.2k—~1.4kAutomated safety check: PassMIT

Similar skills

  • Teaches the Swarms Python framework: the Agent class, tools, loops, memory and multi-agent structures such as sequential, concurrent and graph workflows.

    7.2k GitHub stars~5.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Uipath Functions

    UiPath/skills

    UiPath Coded Functions — deterministic Python or TypeScript/JavaScript units built with the uip function CLI (new -l py|ts|js, init, serve, run, pack, publish); the functions map in uipath.json…

    168 GitHub stars~3.6k tokensUpdated today
    AI & LLM EngineeringAuto-check: notes
  • Strands

    strands-agents/harness-sdk

    Build, extend, evaluate, or migrate applications with Strands Agents in Python or TypeScript.

    8.7k GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Mines local Claude Code session transcripts with a deterministic Python pipeline to show what the agent is actually used for, how often it fails and what it costs.

    1.6k GitHub stars~2.3k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Clawpathy Autoresearch

    ClawBio/ClawBio

    Eval-driven skill tuning. An agent skill from ClawBio/ClawBio.

    1.2k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Pydantic AI Harness

    pydantic/pydantic-ai

    Official

    Adds optional capabilities to Pydantic AI agents from pydantic-ai-harness, led by Code Mode, which runs many tool calls as one sandboxed Python script.

    20k GitHub stars~4.9k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from Orchestra-Research/AI-Research-SKILLs

All 96 skills in this repo
  • AudioCraft Audio Generation

    Orchestra-Research/AI-Research-SKILLs

    Generates music from text descriptions with MusicGen and sound effects with AudioGen, using Meta's AudioCraft PyTorch library with melody and style conditioning.

    13k GitHub starsUsed in 8 repos~3.9k tokens
    Auto-check passed
  • LLM Benchmarking with lm-evaluation-harness

    Orchestra-Research/AI-Research-SKILLs

    Runs lm-evaluation-harness to benchmark language models on academic suites such as MMLU, GSM8K and HumanEval, compare models and track training checkpoints.

    13k GitHub starsUsed in 8 repos~3k tokens
    Auto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 8 repos~3.3k tokens
    Auto-check passed
  • Chroma Vector Database

    Orchestra-Research/AI-Research-SKILLs

    Shows how to store documents and embeddings in Chroma, query them by similarity with metadata filters, and persist them to disk for RAG and semantic search projects.

    13k GitHub starsUsed in 7 repos~2.3k tokens
    Auto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 7 repos~1.7k tokens
    Auto-check passed
  • Whisper Speech Recognition

    Orchestra-Research/AI-Research-SKILLs

    Transcribes audio with OpenAI's Whisper: 99 languages, translation to English, language detection, six model sizes and word-level timestamps, from Python or the CLI.

    13k GitHub starsUsed in 7 repos~1.9k tokens
    Auto-check: notes

Works with

Questions about A-Evolve Agent Evolution

What does A-Evolve Agent Evolution do?

Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles. yaml` plus folders for prompts, skills, memory and tools. Each cycle has the agent solve a batch of benchmark tasks, the benchmark returns feedback on its trajectories, and an LLM-driven evolution engine mutates the workspace files.

When should I use A-Evolve Agent Evolution?

A-Evolve Agent Evolution fits situations like: optimizing an agent's prompts and skills against a measurable benchmark; building a self-improving agent with automated gating and rollback; keeping a git-versioned history of every change made to an agent.

How do I install A-Evolve Agent Evolution in Claude Code?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evolving-ai-agents -a claude-code`. Or copy the skill folder (14-agents/a-evolve in Orchestra-Research/AI-Research-SKILLs) into .claude/skills/evolving-ai-agents in your project. Claude Code loads it when a task matches its description.

How do I install A-Evolve Agent Evolution in Codex?

Run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evolving-ai-agents -a codex`. Or copy the skill folder (14-agents/a-evolve in Orchestra-Research/AI-Research-SKILLs) into .agents/skills/evolving-ai-agents in your project. Codex loads it when a task matches its description.

Can I use A-Evolve Agent Evolution in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Orchestra-Research/AI-Research-SKILLs --skill evolving-ai-agents -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evolving-ai-agents, .gemini/skills/evolving-ai-agents, .github/skills/evolving-ai-agents and .opencode/skills/evolving-ai-agents in your project.

What does A-Evolve Agent Evolution need to run?

Going by SKILL.md and its folder, A-Evolve Agent Evolution needs the command-line tools its instructions call (git and pip) and credentials named ANTHROPIC_API_KEY. Our summary lists: Python with `a-evolve` installed through pip; A benchmark that scores the agent's runs.

Does A-Evolve Agent Evolution access the network?

SKILL.md contains no URLs. Its commands use git and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is A-Evolve Agent Evolution safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does A-Evolve Agent Evolution use?

A-Evolve Agent Evolution is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does A-Evolve Agent Evolution use?

About 3.6k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 32k tokens, read only when the agent opens those files.

What are the alternatives to A-Evolve Agent Evolution?

Skills that share tags, products or a category with A-Evolve Agent Evolution: Swarms Multi-Agent Framework (kyegomez/swarms, 7.2k stars), Uipath Functions (UiPath/skills, 168 stars), Strands (strands-agents/harness-sdk, 8.7k stars) and Analyzing Claude Code Sessions (amd/gaia, 1.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains A-Evolve Agent Evolution?

Orchestra-Research (a GitHub organization) maintains it in Orchestra-Research/AI-Research-SKILLs, which has 13,374 GitHub stars. The repository holds 96 skills in this directory. The repository was last updated on June 16, 2026.

Source: Orchestra-Research/AI-Research-SKILLs on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.