Agent skill

Self-Improvement Tournament Loop

by zereight in zereight/gitlab-mcp

Runs an autonomous evolutionary loop that improves a codebase against a measurable benchmark, using agent roles, tournament selection and recorded history until a stop condition.

MITAuto-check: warningsDevelopment

Install Self-Improvement Tournament Loop

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add zereight/gitlab-mcp --skill self-improve -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install zereight/gitlab-mcp self-improve --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/zereight/gitlab-mcp.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.github/skills/self-improve .claude/skills/self-improve && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
self-improve
GitHub stars
2k
Used in
1 other repo
Token cost
~1.8k tokens
SKILL.md length
703 words
Files
1
Skills in repo
24
Repo updated
First seen
Licence
MIT

At a glance

Runs an autonomous evolutionary loop that improves a codebase against a measurable benchmark, using agent roles, tournament selection and recorded history until a stop condition.

  • Works in 12 steps: Stale Worktree Cleanup → Refresh State → Check Stop Request → …
  • Iteratively optimizing code toward a measurable benchmark goal
  • SKILL.md covers When to Use, When NOT to Use, Autonomous Execution Policy and State Tracking, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Meant for goals with a measurable benchmark, such as speed, bundle size, test coverage or accuracy, the skill manages the full cycle of setup, research, planning, execution, tournament selection, history recording and stop-condition checks. Setup confirms the target repo, requires an explicit trust confirmation before benchmarks run, where declining aborts, runs a short interview on objective, metric, target and scope that is saved to `goal.md`, creates or wraps a benchmark and validates it three times to record a baseline, and confirms default harness rules.

Once the gate passes, the loop runs without pausing to ask you: a failed agent is retried once and then skipped, and rejected plans are logged. State lives under `.omc/self-improve/`. Roles map to agents: explore and architect for hypotheses, planner for plans, architect again for an advisory review, critic for approval against harness rules, executor to implement and benchmark, and git-master for merges, tags and PRs. One-shot fixes and interactive coding are sent to `/omg-autopilot` and `/ralph` instead.

When your agent uses it

  • Iteratively optimizing code toward a measurable benchmark goal
  • Improving performance, bundle size, test coverage or accuracy with metrics
  • Running several competing improvement plans and keeping the winner through tournament selection

Example prompts

  • “Self-improve this parser until the benchmark runtime drops.”
  • “Run a benchmark loop to shrink the bundle size, using tournament selection.”
  • “Evolve the code to raise test coverage and stop when the target is reached.”

Requirements

  • A target repository with a benchmark, or one the agent can create
  • Agent roles for explore, architect, planner, critic, executor and git-master

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Stale Worktree Cleanup
  2. Refresh State
  3. Check Stop Request
  4. Check User Ideas
  5. Research
  6. Plan
  7. Review
  8. Execute
  9. Tournament Selection
  10. Record & Visualize
  11. Cleanup
  12. Stop Condition Check

What it can do on your machine

Read from SKILL.md and the folder at commit 0109168. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Self-Improvement Tournament Loop loads about 1.8k tokens when it runs. Until then it costs about 54 tokens; SKILL.md has 703 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:28
    - Do not ask for confirmation between iterations

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from zereight/gitlab-mcp at commit 0109168, republished under its MIT licence (© zereight). 703 words, ~1,757 tokens.

Download SKILL.mdSave it as .claude/skills/self-improve/SKILL.md (or your agent's skills folder).
name
self-improve
description
Autonomous evolutionary code improvement engine with tournament selection. Activate when user says: self-improve, self improve, evolve code, improve iteratively, tournament, benchmark loop, optimize code.
argument-hint
[repo path] [--resume]

Self-Improvement Orchestrator

Autonomous loop controller for evolutionary code improvement. Manages the full lifecycle: setup, research, planning, execution, tournament selection, history recording, and stop-condition evaluation.

When to Use

  • You want to iteratively improve a codebase toward a measurable benchmark goal
  • Optimization tasks: performance, bundle size, test coverage, accuracy
  • Code quality improvement with measurable metrics

When NOT to Use

  • No measurable benchmark available
  • One-shot fix or feature request → use /omg-autopilot
  • Manual, interactive coding → use /ralph

Autonomous Execution Policy

NEVER stop or pause to ask the user during the improvement loop. Once the gate check passes and the loop begins, run fully autonomously until a stop condition is met.

  • Do not ask for confirmation between iterations
  • On agent failure: retry once, then skip and continue
  • On all plans rejected: log it, continue to next iteration
  • The only things that stop the loop are the stop conditions in Step 11

State Tracking

All state lives under .omc/self-improve/:

.omc/self-improve/
├── config/
│   ├── settings.json          # agents, benchmark, thresholds, sealed_files
│   ├── goal.md                # Improvement objective + target metric
│   ├── harness.md             # Guardrail rules (H001/H002/H003)
│   └── idea.md                # User experiment ideas
├── state/
│   ├── agent-settings.json    # iterations, best_score, status, counters
│   ├── iteration_state.json   # Within-iteration progress (resumability)
│   ├── research_briefs/       # Research output per round
│   ├── iteration_history/     # Full history per round
│   ├── merge_reports/         # Tournament results
│   └── plan_archive/          # Archived plans (permanent)
├── plans/                     # Active plans (current round)
└── tracking/
    ├── raw_data.json          # All candidate scores
    ├── baseline.json          # Initial benchmark score
    └── events.json            # Config changes

Agent Mapping

StepRoleAgentPurpose
ResearchCodebase analysis@explore + @architectHypothesis generation
PlanningHypothesis → plan@plannerStructured plan per agent
Architecture Review6-point review@architectAdvisory review
Critic ReviewHarness enforcement@criticApprove/reject plans
ExecutionImplement + benchmark@executorImplement plan faithfully
Git OperationsMerge/tag/PR@git-masterAtomic merge operations

Setup Phase

  1. Check if target repo path exists. If not configured, ask user.
  2. Create .omc/self-improve/ directory structure.
  3. Read agent-settings.json. Check setup flags.
  4. Trust confirmation (mandatory):
    • Display target repo path, ask user to confirm benchmark execution.
    • If declined: abort.
    • Record consent: trust_confirmed: true
  5. If goal not set → Run Socratic interview (Objective, Metric, Target, Scope). Write to goal.md.
  6. If benchmark not set → Survey repo, create/wrap benchmark, validate 3x, record baseline.
  7. If harness not set → Confirm default harness rules (H001/H002/H003).
  8. Gate: All settings + trust must be true.
  9. Create improvement branch: improve/{goal_slug} from target branch.
  10. Write initial state via omg_write_state.

Improvement Loop

Gate: All settings must be true. Execute continuously without stopping.

Step 0 — Stale Worktree Cleanup

Remove orphaned worktrees from prior iterations.

Step 1 — Refresh State

Update state to reset TTL.

Step 2 — Check Stop Request

If state is cleared or status is user_stopped: exit gracefully.

Step 3 — Check User Ideas

Read idea.md. If non-empty, pass to planners.

Step 4 — Research

Spawn @explore + @architect to analyze codebase and generate hypotheses based on goal, history, and prior briefs.

Step 5 — Plan

Spawn N @planner agents in parallel (N = number_of_agents). Each produces a plan with one testable hypothesis, approach_family tag, and history_reference.

Show full SKILL.md (294 more words)Show less
Step 6 — Review

For each plan:

  • 6a. Architecture Review: @architect with 6-point checklist (testability, novelty, scope, target files, implementation clarity, expected outcome). Advisory only.
  • 6b. Critic Review: @critic with harness rules (H001: one hypothesis, H002: no approach_family streak ≥3, H003: intra-round diversity). Sets critic_approved: true/false.
Step 7 — Execute

For each approved plan, spawn @executor in parallel. Each executor works in a git worktree, implements the plan, runs validation, and benchmarks.

Step 8 — Tournament Selection
  1. Collect results, filter to status: "success"
  2. Rank by benchmark_score (respecting direction)
  3. For each candidate (best first):
    • No-regression check vs best_score
    • Merge via @git-master with --no-ff
    • Re-benchmark on merged state
    • If confirmed: accept winner, break
    • If regression: revert merge, try next
    • If conflict: abort merge, try next
  4. Archive non-winner branches
Step 9 — Record & Visualize

Write iteration history, update agent-settings (scores, plateau count, circuit breaker), append tracking data.

Step 10 — Cleanup

Remove worktrees, update iteration state to completed.

Step 11 — Stop Condition Check
ConditionCheck
User stopstatus == "user_stopped"
Target reachedbest_score meets/exceeds target_value
Plateauplateau_consecutive_count >= plateau_window
Max iterationsiterations >= max_iterations
Circuit breakercircuit_breaker_count >= circuit_breaker_threshold

If NO stop condition: immediately go back to Step 1.

Resumability

On invocation:

  1. Always run Step 0 (stale worktree cleanup)
  2. Check agent-settings.json:
    • user_stopped: ask to resume
    • running: crashed — resume automatically
    • idle: fresh start
  3. Check iteration_state.json: resume from last step if in-progress

Completion

  1. Update final status
  2. Print summary (status, iterations, best score, baseline, improvement %)
  3. Run /cancel for clean state cleanup

Approach Family Taxonomy

Every plan must be tagged with exactly one:

TagDescription
architectureModel/component structure changes
training_configOptimizer, LR, scheduler, batch size
dataData loading, augmentation, preprocessing
infrastructureMixed precision, distributed training
optimizationAlgorithmic/numerical optimizations
testingEvaluation methodology changes
documentationDocumentation-only changes
otherDoes not fit above

© zereight, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .github/skills/self-improve of zereight/gitlab-mcp.

Open the folder on GitHubat commit 0109168

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in zereight/gitlab-mcp, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Self-Improvement Tournament Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Self-Improvement Tournament Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Self-Improvement Tournament Loop this skillzereight/gitlab-mcp2k1 repos~1.8kAutomated safety check: WarnMIT
Effective Harnessesliangdabiao/exa-research-mcp-skill110—~1.5kAutomated safety check: PassNone
Branch Standup Facilitatorthedotmack/claude-mem98k—~1.7kAutomated safety check: NotesApache-2.0
Measured Optimization LoopEveryInc/compound-engineering-plugin25k—~2kAutomated safety check: PassMIT
Multi Agent Orchestrationcat-xierluo/legal-skills713—~6.6kAutomated safety check: PassMIT
Polyphonyalinaqi/maggy707—~980Automated safety check: PassMIT

Similar skills

  • Effective Harnesses

    liangdabiao/exa-research-mcp-skill

    Long-running agent project harness for Codex, OpenClaw, Claude Code, and other coding agents.

    110 GitHub stars~1.5k tokensUpdated 2 mo ago
    DevelopmentAuto-check passed
  • Branch Standup Facilitator

    thedotmack/claude-mem

    Facilitates a read-only standup between git worktrees, branches or PRs, where each acts as an agent in a shared markdown chat to agree one consolidation plan.

    98k GitHub stars~1.7k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Measured Optimization Loop

    EveryInc/compound-engineering-plugin

    Optimizes a named target with a measured loop, attributing a workload's cost or scoring variants and keeping winners, on a dedicated branch with a disk log.

    25k GitHub stars~2k tokensUpdated today
    DevelopmentAuto-check passed
  • Multi Agent Orchestration

    cat-xierluo/legal-skills

    编排两个以上边界独立的本地 worker,使用 Orca Run/Task/Dispatch、独立 worktree/session 或 tmux 回退,由 PM 负责拆解、派发、巡检、429 停滞恢复、独立验收、PR 收口与临时资源清理;也用于用户明确要求“并行推进”“多个 worker”“PM 总控”“Wave Autopilot”或防止 PM…

    713 GitHub stars~6.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Polyphony

    alinaqi/maggy

    Multi-agent orchestration with container-isolated workspaces — each agent session runs in its own Docker container with independent git branches

    707 GitHub stars~980 tokensUpdated 14 days ago
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 5 repos~1.1k tokens
    DevelopmentAuto-check passed

More from zereight/gitlab-mcp

All 24 skills in this repo
  • OMG Mode Canceller

    zereight/gitlab-mcp

    Detects which autonomous OMG mode is currently active - Autopilot, Ralph, Ultrawork, UltraQA, Team or Self-Improve - and shuts it down cleanly.

    2k GitHub starsUsed in 1 repo~690 tokens
    Auto-check passed
  • CCG Tri-Model Orchestration

    zereight/gitlab-mcp

    Runs a task through Codex and Gemini CLIs in parallel alongside Claude, then synthesizes the three outputs into one answer with agreements and conflicts called out.

    2k GitHub starsUsed in 1 repo~657 tokens
    Auto-check passed
  • Shared reference for naming, function size, complexity and error handling rules that reviewer agents apply across TypeScript, Python, Go, Rust, Java, C# and Swift.

    2k GitHub starsUsed in 1 repo~1.4k tokens
    Auto-check passed
  • Remember Project Knowledge

    zereight/gitlab-mcp

    Sorts what you learned in a session into the right memory surface, filtering out ephemeral notes and duplicates before anything is stored.

    2k GitHub starsUsed in 1 repo~846 tokens
    Auto-check passed
  • Pre-Commit Security Scan

    zereight/gitlab-mcp

    Runs a fast security sweep of recent code changes before a commit or PR, checking for leaked secrets, vulnerable dependencies, unsafe input handling and auth gaps.

    2k GitHub starsUsed in 1 repo~859 tokens
    Auto-check: notes
  • Skill Inventory Stocktake

    zereight/gitlab-mcp

    Audits a project's skill directories for broken frontmatter, missing template sync and quality gaps, then produces a health report of what needs fixing.

    2k GitHub starsUsed in 1 repo~1k tokens
    Auto-check passed

Works with

Questions about Self-Improvement Tournament Loop

What does Self-Improvement Tournament Loop do?

Runs an autonomous evolutionary loop that improves a codebase against a measurable benchmark, using agent roles, tournament selection and recorded history until a stop condition. Meant for goals with a measurable benchmark, such as speed, bundle size, test coverage or accuracy, the skill manages the full cycle of setup, research, planning, execution, tournament selection, history recording and stop-condition checks.md`, creates or wraps a benchmark and validates it three times to record a baseline, and confirms default harness rules.

When should I use Self-Improvement Tournament Loop?

Self-Improvement Tournament Loop fits situations like: iteratively optimizing code toward a measurable benchmark goal; improving performance, bundle size, test coverage or accuracy with metrics; running several competing improvement plans and keeping the winner through tournament selection.

How do I install Self-Improvement Tournament Loop in Claude Code?

Run `npx skills add zereight/gitlab-mcp --skill self-improve -a claude-code`. Or copy the skill folder (.github/skills/self-improve in zereight/gitlab-mcp) into .claude/skills/self-improve in your project. Claude Code loads it when a task matches its description.

How do I install Self-Improvement Tournament Loop in Codex?

Run `npx skills add zereight/gitlab-mcp --skill self-improve -a codex`. Or copy the skill folder (.github/skills/self-improve in zereight/gitlab-mcp) into .agents/skills/self-improve in your project. Codex loads it when a task matches its description.

Can I use Self-Improvement Tournament Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add zereight/gitlab-mcp --skill self-improve -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/self-improve, .gemini/skills/self-improve, .github/skills/self-improve and .opencode/skills/self-improve in your project.

What does Self-Improvement Tournament Loop need to run?

SKILL.md names no scripts, command-line tools or credentials: Self-Improvement Tournament Loop is instructions for the agent only. Our summary lists: A target repository with a benchmark, or one the agent can create; Agent roles for explore, architect, planner, critic, executor and git-master.

Does Self-Improvement Tournament Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Self-Improvement Tournament Loop safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Self-Improvement Tournament Loop use?

Self-Improvement Tournament Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Self-Improvement Tournament Loop use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Self-Improvement Tournament Loop?

Skills that share tags, products or a category with Self-Improvement Tournament Loop: Effective Harnesses (liangdabiao/exa-research-mcp-skill, 110 stars), Branch Standup Facilitator (thedotmack/claude-mem, 98k stars), Measured Optimization Loop (EveryInc/compound-engineering-plugin, 25k stars) and Multi Agent Orchestration (cat-xierluo/legal-skills, 713 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Self-Improvement Tournament Loop?

zereight (a GitHub user) maintains it in zereight/gitlab-mcp, which has 2,029 GitHub stars. The repository holds 24 skills in this directory. The repository was last updated on October 6, 2026.

Source: zereight/gitlab-mcp on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.