Agent skill

Skill Evolution

by vibeeval in vibeeval/vibecosystem

Self-evolving skill system. An agent skill from vibeeval/vibecosystem.

MITAuto-check passed

Install Skill Evolution

skills CLI
$ npx skills add vibeeval/vibecosystem --skill skill-evolution -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install vibeeval/vibecosystem skill-evolution --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-evolution .claude/skills/skill-evolution && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-evolution
GitHub stars
531
Token cost
~1.8k tokens
SKILL.md length
562 words
Files
1
Skills in repo
144
Repo updated
First seen
Licence
MIT

At a glance

Self-evolving skill system. An agent skill from vibeeval/vibecosystem.

  • Works in 5 steps: Verify scores in… → Add locked: true to the skill's… → Apply git tag → …
  • SKILL.md covers The 5 Scoring Dimensions, Skill Lifecycle, Score Storage Format and Crystallization Protocol, plus 4 more sections
  • Calls python3, git and node

What it does

Skill Evolution is an agent skill from vibeeval/vibecosystem. Self-evolving skill system. Skills are scored after execution (0-100) on 5 dimensions. Score 90+ over 5 runs = crystallized (locked). Score below 30 = auto-repair attempted. Skills improve themselves through usage feedback.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution. The licence is MIT.

Example prompts

  • “/skill-evolution”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Verify scores in ~/.claude/skill-scores.jsonl -- confirm no outliers inflating the average
  2. Add locked: true to the skill's frontmatter
  3. Apply git tag
  4. Log the crystallization in thoughts/SKILL-EVOLUTION.md
  5. Notify via canavar cross-training so all agents know this skill is stable

What it can do on your machine

Read from SKILL.md and the folder at commit 3b763b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • git
    • node

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Evolution loads about 1.8k tokens when it runs. Until then it costs about 60 tokens; SKILL.md has 562 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~60
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from vibeeval/vibecosystem at commit 3b763b1, republished under its MIT licence (© vibeeval). 562 words, ~1,814 tokens.

Download SKILL.mdSave it as .claude/skills/skill-evolution/SKILL.md (or your agent's skills folder).
name
skill-evolution
description
Self-evolving skill system. Skills are scored after execution (0-100) on 5 dimensions. Score 90+ over 5 runs = crystallized (locked). Score below 30 = auto-repair attempted. Skills improve themselves through usage feedback.

Skill Evolution

Darwinian selection for skills. Skills that produce good outcomes are crystallized and protected. Skills that produce poor outcomes are repaired or archived. Every execution generates a score that drives the next generation of the skill.

The 5 Scoring Dimensions

Each skill execution is scored 0-100 on five dimensions:

DimensionWeightWhat It Measures
Accuracy25%Did the skill produce the correct result for the task?
Relevance20%Was the skill content applicable to the actual use case?
Token Efficiency20%Did the skill guide the agent without bloat or repetition?
User Satisfaction20%Did the outcome meet or exceed user expectations?
Reusability15%Could another agent use this skill in a similar situation?

Composite score = weighted average of all five dimensions (0-100).

Scoring Rubric
90-100: Excellent -- candidate for crystallization
70-89:  Good -- active skill, no action needed
50-69:  Adequate -- flag for review after 3 more runs
30-49:  Poor -- schedule auto-repair attempt
0-29:   Critical -- immediate auto-repair or archive

Skill Lifecycle

DRAFT          ACTIVE         CRYSTALLIZED      ARCHIVED
  |               |                |                |
New skill   In regular use   Proven stable    Deprecated/replaced
  |               |                |                |
  +-- first run ->+-- score >90   ++-- score <30    |
                  |   for 5+ runs  |   (3 attempts)  |
                  +-- score <30 -->+ auto-repair      |
                  |   auto-repair  |   fails 3x -->--+
                  +-- score >90 -->+
Draft

New skills enter as Draft. They receive no special protection and are evaluated critically on first use. A Draft skill that scores below 30 on its very first run is discarded rather than repaired.

Active

Skills in regular use. Scores are tracked in ~/.claude/skill-scores.jsonl. No action unless scores trend below 30 or above 90 over a rolling window of 5 runs.

Crystallized

A skill that maintains an average composite score above 90 over 5 or more consecutive runs is crystallized:

  • Git tag applied: skill/<name>/crystallized-v<N>
  • Read-only flag added to frontmatter: locked: true
  • Skill is excluded from auto-repair
  • Changes require explicit human unlock + PR
Archived

A skill that fails auto-repair 3 times is archived:

  • Moved to skills/_archived/<name>/
  • Git tag applied: skill/<name>/archived
  • Replacement skill drafted by catalyst agent if the capability is still needed

Score Storage Format

Append one record per execution to ~/.claude/skill-scores.jsonl:

jsonl
{"skill":"experiment-loop","ts":"2026-04-07T10:00:00Z","session":"abc123","scores":{"accuracy":88,"relevance":92,"token_efficiency":75,"user_satisfaction":90,"reusability":85},"composite":86.5,"feedback":"Loop ran 4 iterations successfully, target nearly met"}
{"skill":"experiment-loop","ts":"2026-04-07T14:30:00Z","session":"def456","scores":{"accuracy":95,"relevance":90,"token_efficiency":82,"user_satisfaction":95,"reusability":88},"composite":90.4,"feedback":"Bundle size reduced 28%, target exceeded"}
Score CLI (quick check)
bash
# Average scores for a skill (last 10 runs)
cat ~/.claude/skill-scores.jsonl | python3 -c "
import sys, json, statistics
skill = '$1'
runs = [json.loads(l) for l in sys.stdin if json.loads(l).get('skill') == skill][-10:]
if runs:
    avg = statistics.mean(r['composite'] for r in runs)
    print(f'{skill}: {avg:.1f} avg over {len(runs)} runs')
"

Crystallization Protocol

When a skill reaches 90+ composite score over 5+ consecutive runs:

  1. Verify scores in ~/.claude/skill-scores.jsonl -- confirm no outliers inflating the average
  2. Add locked: true to the skill's frontmatter
  3. Apply git tag:
    bash
    git tag skill/<name>/crystallized-v1 -m "Crystallized: avg score 92.3 over 7 runs"
    git push origin skill/<name>/crystallized-v1
  4. Log the crystallization in thoughts/SKILL-EVOLUTION.md
  5. Notify via canavar cross-training so all agents know this skill is stable
Show full SKILL.md (233 more words)Show less

Auto-Repair Protocol

When a skill's composite score drops below 30:

Diagnosis
  1. Identify the lowest-scoring dimension (the primary failure mode)
  2. Read the last 3 session feedback notes from ~/.claude/skill-scores.jsonl
  3. Summarize what went wrong (specific, not vague)
Repair

The catalyst agent rewrites the failing section(s) of the skill:

  • Only the sections relevant to the low-scoring dimension
  • Preserve all high-scoring sections unchanged
  • Add a concrete example for the repaired section
Validation

After repair, the skill is re-scored on a synthetic test case by the verifier agent:

  • Synthetic score must be 50+ to proceed to Active state
  • If synthetic score < 50, attempt 2 of 3 repairs begins
Escalation

After 3 failed auto-repairs:

  • Archive the skill
  • Alert via thoughts/SKILL-EVOLUTION.md
  • Spawn catalyst to draft a replacement from scratch

Evolution Log Format

Append events to thoughts/SKILL-EVOLUTION.md:

markdown
## 2026-04-07

### skill: experiment-loop
- Status change: Active -> Crystallized
- Trigger: avg composite 91.2 over 6 consecutive runs
- Git tag: skill/experiment-loop/crystallized-v1
- Notable strength: Token Efficiency dimension consistently 85+

### skill: legacy-deploy-helper
- Status change: Active -> Auto-Repair (attempt 1/3)
- Trigger: composite 24 on last run
- Lowest dimension: Relevance (12) -- skill referenced outdated Heroku patterns
- Repair: catalyst rewrote "Deployment Targets" section with Vercel/Railway focus
- Post-repair synthetic score: 71 -- promoted back to Active

Integration with Canavar Cross-Training

Skill evolution data feeds into canavar's cross-training pipeline:

  • A crystallized skill is injected into canavar's skill-matrix.json with trust: locked
  • An archived skill is marked trust: deprecated -- agents stop referencing it
  • Auto-repair failures are logged to error-ledger.jsonl with source: skill-evolution
  • The canavar leaderboard tracks which agents most frequently produce high-scoring skill executions
bash
# View crystallized skills
node ~/.claude/hooks/dist/canavar-cli.mjs leaderboard --filter crystallized

# View skills needing repair
cat ~/.claude/skill-scores.jsonl | python3 -c "
import sys, json, collections
runs = [json.loads(l) for l in sys.stdin]
low = {r['skill'] for r in runs if r['composite'] < 30}
print('Skills needing repair:', low)
"

Activation

This skill activates automatically when:

  • A skill completes an execution (PostToolUse hook)
  • A skill is referenced in a session that ends with user dissatisfaction
  • The verifier agent reports a skill-guided task as failed

Agents involved: catalyst (repair), verifier (validation), self-learner (feedback extraction), canavar (cross-training propagation).

© vibeeval, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/skill-evolution of vibeeval/vibecosystem.

Open the folder on GitHubat commit 3b763b1

Compare with similar skills

Skill Evolution next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Evolution compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Evolution this skillvibeeval/vibecosystem531—~1.8kAutomated safety check: PassMIT
A-Evolve Agent EvolutionOrchestra-Research/AI-Research-SKILLs13k1 repos~3.6kAutomated safety check: PassMIT
Harness Evolveruvnet/ruflo74k—~1.6kAutomated safety check: NotesMIT
Evolver Agent Self-EvolutionEvoMap/evolver9.1k1 repos~2.2kAutomated safety check: PassGPL-3.0
Evolutionsickn33/agentic-awesome-skills47k2 repos~3.1kAutomated safety check: PassMIT
Darwin Mode Harness Evolutionruvnet/RuView97k—~693Automated safety check: PassMIT

Similar skills

  • A-Evolve Agent Evolution

    Orchestra-Research/AI-Research-SKILLs

    Guidance for using A-Evolve to improve an AI agent automatically, evolving its prompts, skills and memory against a benchmark through solve, observe and evolve cycles.

    13k GitHub starsUsed in 1 repo~3.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Harness Evolve

    ruvnet/ruflo

    Run @metaharness/darwin evolve <repo to mutate a harness's seven policy surfaces (planner/contextBuilder/reviewer/retryPolicy/toolPolicy/memoryPolicy/scorePolicy), sandbox-score each variant, and…

    74k GitHub stars~1.6k tokensUpdated today
    DevelopmentAuto-check: notes
  • Self-evolution engine for agents: studies runtime history for failures and inefficiencies, writes improvements, and syncs with the EvoMap Hub through a local proxy mailbox.

    9.1k GitHub starsUsed in 1 repo~2.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Evolution

    sickn33/agentic-awesome-skills

    This skill enables makepad-skills to self-improve continuously during development.

    47k GitHub starsUsed in 2 repos~3.1k tokens
    Auto-check passed
  • Runs Darwin Mode on an agent harness: mutates one policy file per generation in a sandbox, scores each variant against your tests, and archives only variants that measurably improve.

    97k GitHub stars~693 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Harness Score

    ruvnet/ruflo

    5-dimension harness readiness scorecard from metaharness score <path.

    74k GitHub stars~605 tokensUpdated today
    DevelopmentAuto-check: notes

More from vibeeval/vibecosystem

All 144 skills in this repo
  • Agent Benchmark

    vibeeval/vibecosystem

    Framework for measuring and tracking agent response quality over time.

    531 GitHub stars~2.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Differential Review

    vibeeval/vibecosystem

    Security-focused differential code review with blast radius analysis, risk-adaptive depth (DEEP/FOCUSED/SURGICAL), git history correlation, and structured finding format.

    531 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • Factcheck Guard

    vibeeval/vibecosystem

    A skill your agent uses when making any factual claim about the codebase — existence, absence, or behavior.

    531 GitHub stars~2.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Fp Check

    vibeeval/vibecosystem

    Systematic false positive verification for security findings.

    531 GitHub stars~1.6k tokensUpdated 2 mo ago
    Auto-check passed
  • N8n Workflows

    vibeeval/vibecosystem

    n8n otomasyon workflow'lari. An agent skill from vibeeval/vibecosystem.

    531 GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Notepad System

    vibeeval/vibecosystem

    A skill your agent uses when context compression is imminent, when resuming a session, or when preserving critical decisions across long tasks.

    531 GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Skill Evolution

What does Skill Evolution do?

Self-evolving skill system. An agent skill from vibeeval/vibecosystem. Skill Evolution is an agent skill from vibeeval/vibecosystem. Self-evolving skill system.

How do I install Skill Evolution in Claude Code?

Run `npx skills add vibeeval/vibecosystem --skill skill-evolution -a claude-code`. Or copy the skill folder (skills/skill-evolution in vibeeval/vibecosystem) into .claude/skills/skill-evolution in your project. Claude Code loads it when a task matches its description.

How do I install Skill Evolution in Codex?

Run `npx skills add vibeeval/vibecosystem --skill skill-evolution -a codex`. Or copy the skill folder (skills/skill-evolution in vibeeval/vibecosystem) into .agents/skills/skill-evolution in your project. Codex loads it when a task matches its description.

Can I use Skill Evolution in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vibeeval/vibecosystem --skill skill-evolution -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-evolution, .gemini/skills/skill-evolution, .github/skills/skill-evolution and .opencode/skills/skill-evolution in your project.

What does Skill Evolution need to run?

Going by SKILL.md and its folder, Skill Evolution needs the command-line tools its instructions call (python3, git and node). Our summary lists: Python 3.

Does Skill Evolution access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Skill Evolution safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Evolution use?

Skill Evolution is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Evolution use?

About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Evolution?

Skills that share tags, products or a category with Skill Evolution: A-Evolve Agent Evolution (Orchestra-Research/AI-Research-SKILLs, 13k stars), Harness Evolve (ruvnet/ruflo, 74k stars), Evolver Agent Self-Evolution (EvoMap/evolver, 9.1k stars) and Evolution (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Evolution?

vibeeval (a GitHub user) maintains it in vibeeval/vibecosystem, which has 531 GitHub stars. The repository holds 144 skills in this directory. The repository was last updated on August 8, 2026.

Source: vibeeval/vibecosystem on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.