Copilot Session Failure Analysis
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.
$ npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install aiming-lab/AutoResearchClaw a-evolve --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/a-evolve .claude/skills/a-evolve && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "a-evolve" agent skill from https://github.com/aiming-lab/AutoResearchClaw/tree/main/.claude/skills/a-evolve into .claude/skills/a-evolve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "a-evolve", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/aiming-lab/AutoResearchClaw/tree/main/.claude/skills/a-evolveType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install aiming-lab/AutoResearchClaw a-evolve --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/a-evolve .agents/skills/a-evolve && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "a-evolve" agent skill from https://github.com/aiming-lab/AutoResearchClaw/tree/main/.claude/skills/a-evolve into .agents/skills/a-evolve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "a-evolve", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install aiming-lab/AutoResearchClaw a-evolve --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/a-evolve .cursor/skills/a-evolve && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "a-evolve" agent skill from https://github.com/aiming-lab/AutoResearchClaw/tree/main/.claude/skills/a-evolve into .cursor/skills/a-evolve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "a-evolve", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/aiming-lab/AutoResearchClaw.git --path .claude/skills/a-evolve--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install aiming-lab/AutoResearchClaw a-evolve --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/a-evolve .gemini/skills/a-evolve && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "a-evolve" agent skill from https://github.com/aiming-lab/AutoResearchClaw/tree/main/.claude/skills/a-evolve into .gemini/skills/a-evolve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "a-evolve", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install aiming-lab/AutoResearchClaw a-evolveInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/a-evolve .github/skills/a-evolve && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "a-evolve" agent skill from https://github.com/aiming-lab/AutoResearchClaw/tree/main/.claude/skills/a-evolve into .github/skills/a-evolve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "a-evolve", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install aiming-lab/AutoResearchClaw a-evolve --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/aiming-lab/AutoResearchClaw.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/a-evolve .opencode/skills/a-evolve && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "a-evolve" agent skill from https://github.com/aiming-lab/AutoResearchClaw/tree/main/.claude/skills/a-evolve into .opencode/skills/a-evolve/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "a-evolve", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
a-evolveDiagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop.
The skill applies a five-step loop named Solve, Observe, Evolve, Gate, Reload. It is prompt-based, so it needs no external dependencies and no harness changes. The agent first gathers evidence such as run logs, error traces, pass or fail results per task and metric values, and inside an AutoResearchClaw pipeline it also looks at artifacts/rc-* outputs, evolve.log and reviews.md.
In the Observe step each failed or weak task gets an error category, a root cause, a frequency and a severity, written as a structured list of observations. In the Evolve step the agent proposes changes: a new SKILL.md for a recurring pattern seen at least 3 times, which targets one failure category and stays under 100 lines, or a short system prompt patch for ambiguity or missing guidance. The result is durable artifacts the agent can load in later runs.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit be4ba47. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and json).
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
A-Evolve Agent Improvement loads about 1.8k tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 703 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from aiming-lab/AutoResearchClaw at commit be4ba47, republished under its MIT licence (© aiming-lab). 703 words, ~1,810 tokens.
.claude/skills/a-evolve/SKILL.md (or your agent's skills folder).Apply the Solve → Observe → Evolve → Gate → Reload methodology from A-Evolve to iteratively improve agent performance. This skill is prompt-based — no external dependencies, no harness changes. You analyze failures, propose workspace mutations, and generate durable artifacts (skills, prompt patches, knowledge entries) that the agent can load in future runs.
When asked to evolve or improve agent performance, follow this 5-step loop:
Gather the agent's execution artifacts. Ask the user for or locate:
If inside AutoResearchClaw, look at:
artifacts/rc-*/ — experiment outputs, charts, reviewsevolve.log or stage-specific logsreviews.md — peer review feedbackAnalyze the collected evidence to produce structured observations:
For each failed or underperforming task, identify:
Write observations as a structured list:
## Observations (Batch N)
### OBS-1: [Category] Short description
- Tasks affected: task_001, task_005, task_012
- Root cause: ...
- Frequency: 3/50 tasks (6%)
- Severity: degrading
### OBS-2: ...Based on observations, propose one or more of these mutation types:
A. Generate a Skill (for recurring patterns, frequency ≥ 3)
Write a new SKILL.md file that teaches the agent how to handle this
pattern. A good evolved skill:
Example — if the agent keeps failing at API pagination:
---
name: api-pagination-handler
description: >
Handle paginated API responses correctly. Use when making API calls
that may return partial results, or when results seem truncated.
---
When calling any API that supports pagination:
1. Check response for pagination indicators: `next_page`, `offset`,
`has_more`, `cursor`, or truncated result counts.
2. If paginated, loop until all pages are collected.
3. Concatenate results before processing.
4. Set a max-page safety limit (default: 20) to prevent infinite loops.
5. Log total items collected vs expected count if available.B. Patch the System Prompt (for prompt ambiguity or missing guidance)
Write a short addendum to the system prompt that addresses the gap. Keep patches minimal — one paragraph per issue. Format:
## Prompt Patch: [Issue]
Append to system prompt:
> When [specific situation], always [specific action] because [reason].C. Add a Knowledge Entry (for factual gaps or learned heuristics)
Record a reusable insight as a knowledge entry:
{
"id": "know-001",
"category": "experiment_design",
"insight": "Synthetic benchmarks with <100 samples produce high-variance results. Always use ≥500 samples or report confidence intervals.",
"source": "observation OBS-3 from batch 2",
"confidence": 0.85
}D. Do Nothing (if observation is a one-off, severity is cosmetic, or the fix would be too broad / risky)
Before accepting any mutation, check:
If a mutation fails the gate, either refine it or discard it. Explain your reasoning to the user.
Present the accepted mutations to the user. For each:
For AutoResearchClaw projects, recommended locations:
| Artifact | Location |
|---|---|
| Evolved skill | .claude/skills/evolved/<skill-name>/SKILL.md |
| Prompt patch | Append to prompts.default.yaml or custom prompts file |
| Knowledge entry | docs/kb/evolved_knowledge/<id>.json |
| Observation log | evolution/observations/<batch>.md |
Keep a running version log so the user can track what evolved and when:
## Evolution Log
- evo-1 (2026-03-30): Generated `api-pagination-handler` skill from OBS-1
- evo-2 (2026-03-30): Prompt patch for citation format from OBS-4This skill maps to ARC's pipeline stages:
| ARC Stage | Evolution Role |
|---|---|
| 12 EXPERIMENT_RUN | Source of Solve artifacts |
| 13 ITERATIVE_REFINE | Main Observe + Evolve trigger point |
| 15 RESEARCH_DECISION | Natural Gate — PROCEED = accept, REFINE = retry |
| 18 PEER_REVIEW | Additional Observe signal for writing quality |
When the user says "evolve my research pipeline" or similar:
artifacts/rc-*/).claude/skills/Do NOT:
If AutoResearchClaw has MetaClaw enabled (metaclaw_bridge.enabled: true),
evolved skills from this process can be placed in ~/.metaclaw/skills/arc-*/
so MetaClaw injects them into future runs automatically. The two systems
are complementary:
Both can coexist. Skills generated here are higher-precision; MetaClaw lessons are higher-recall.
© aiming-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/a-evolve of aiming-lab/AutoResearchClaw.
Open the folder on GitHubat commit be4ba47
A-Evolve Agent Improvement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| A-Evolve Agent Improvement this skillaiming-lab/AutoResearchClaw | 15k | — | ~1.8k | Automated safety check: Pass | MIT | |
| Copilot Session Failure Analysisdotnet/maui | 23k | — | ~3.4k | Automated safety check: Pass | MIT | |
| Workflow Schema Tuningbreaking-brake/cc-wf-studio | 5.4k | — | ~1.3k | Automated safety check: Pass | Custom licence | |
| Microsoft Foundrymicrosoft/GitHub-Copilot-for-Azure | 255 | 1 repos | ~6.7k | Automated safety check: Pass | MIT | |
| Skill With Prompt EngineeringLeoYeAI/openclaw-master-skills | 2.2k | — | ~4.1k | Automated safety check: Pass | MIT | |
| Diagnosing Superpowers Sessionsobra/superpowers | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT |
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
breaking-brake/cc-wf-studio
Guides edits to cc-wf-studio's workflow schema so AI agents generate better workflows, treating schema text as prompt engineering rather than validation.
microsoft/GitHub-Copilot-for-Azure
Build, deploy, evaluate, optimize, fine-tune, and manage Microsoft Foundry agents, models, and resources end to end.
LeoYeAI/openclaw-master-skills
A Prompt Engineering assistant based on Gen AI Space's 16-technique framework.
obra/superpowers
Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
aiming-lab/AutoResearchClaw
Turns a broad metabolic modelling topic into a concrete, paper-shaped plan with organism, model, perturbations, metrics and figures before any FBA code is written.
aiming-lab/AutoResearchClaw
Runs a metabolic flux analysis from model loading to phenotype prediction and figures by handing work to four sub-agents in sequence.
aiming-lab/AutoResearchClaw
Reference patterns for writing qiskit 2.x code for variational quantum machine learning: feature maps, VQC training, VQE for chemistry, MPS circuits and noise models.
aiming-lab/AutoResearchClaw
Builds or loads a genome-scale metabolic model in COBRApy, sets its growth medium and objective, and exports it as a validated JSON file for flux analysis.
aiming-lab/AutoResearchClaw
Quick reference for Biopython work: sequence operations, SeqIO file parsing, BLAST searches, Entrez queries, phylogenetic trees and PDB structure analysis.
aiming-lab/AutoResearchClaw
Reference guide for working with molecules in RDKit: reading SMILES and SDF files, computing descriptors and fingerprints, and searching substructures.
Categories
Diagnoses where an agent failed across runs and turns the findings into new skills, system prompt patches and knowledge entries, using the A-Evolve loop. The skill applies a five-step loop named Solve, Observe, Evolve, Gate, Reload. It is prompt-based, so it needs no external dependencies and no harness changes.
A-Evolve Agent Improvement fits situations like: working out why an agent keeps failing the same kind of task; turning recurring error patterns into a new targeted skill; patching a system prompt that leaves the agent's behavior ambiguous; reviewing logs from an AutoResearchClaw run to see what went wrong.
Run `npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a claude-code`. Or copy the skill folder (.claude/skills/a-evolve in aiming-lab/AutoResearchClaw) into .claude/skills/a-evolve in your project. Claude Code loads it when a task matches its description.
Run `npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a codex`. Or copy the skill folder (.claude/skills/a-evolve in aiming-lab/AutoResearchClaw) into .agents/skills/a-evolve in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aiming-lab/AutoResearchClaw --skill a-evolve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/a-evolve, .gemini/skills/a-evolve, .github/skills/a-evolve and .opencode/skills/a-evolve in your project.
SKILL.md names no scripts, command-line tools or credentials: A-Evolve Agent Improvement is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
A-Evolve Agent Improvement is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with A-Evolve Agent Improvement: Copilot Session Failure Analysis (dotnet/maui, 23k stars), Workflow Schema Tuning (breaking-brake/cc-wf-studio, 5.4k stars), Microsoft Foundry (microsoft/GitHub-Copilot-for-Azure, 255 stars) and Skill With Prompt Engineering (LeoYeAI/openclaw-master-skills, 2.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
aiming-lab (a GitHub organization) maintains it in aiming-lab/AutoResearchClaw, which has 14,587 GitHub stars. The repository holds 34 skills in this directory. The repository was last updated on August 19, 2026.
Source: aiming-lab/AutoResearchClaw on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.