Agent skill

Experiment

by SethGammon in SethGammon/Citadel

Automated optimization loop with scalar fitness function. An agent skill from SethGammon/Citadel.

MITAuto-check passedDevelopment

Install Experiment

skills CLI
$ npx skills add SethGammon/Citadel --skill experiment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install SethGammon/Citadel experiment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/SethGammon/Citadel.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment .claude/skills/experiment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment
GitHub stars
922
Token cost
~1.5k tokens
SKILL.md length
624 words
Files
3
Skills in repo
48
Repo updated
First seen
Licence
MIT

At a glance

Automated optimization loop with scalar fitness function. An agent skill from SethGammon/Citadel.

  • Works in 4 steps: BASELINE → ITERATE → CONVERGENCE CHECK → …
  • Tasks that involve Git worktrees
  • SKILL.md covers Inputs, Protocol, Common Metrics and When to Use, plus 5 more sections
  • Calls npm, node and npx

What it does

Experiment is an agent skill from SethGammon/Citadel. Automated optimization loop with scalar fitness function. Proposes changes in isolated worktrees, measures with a metric command, keeps improvements, discards failures. Supports convergence detection and diminishing returns.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `__benchmarks__/no-metric.md` and `__benchmarks__/with-metric.md`).

It sits in Development, covering Git worktrees. It works with npm. The repository describes itself as: The operating layer for Claude Code + OpenAI Codex: persistent project memory, intent routing, safety hooks, cost telemetry, and parallel agent fleets. The licence is MIT.

When your agent uses it

  • Tasks that involve Git worktrees

Example prompts

  • “/experiment”

Requirements

  • Node.js

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. BASELINE
  2. ITERATE
  3. CONVERGENCE CHECK
  4. REPORT

What it can do on your machine

Read from SKILL.md and the folder at commit e41ff1d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • node
    • npx
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, npx and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment loads about 1.5k tokens when it runs. Until then it costs about 59 tokens; SKILL.md has 624 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~59
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from SethGammon/Citadel at commit e41ff1d, republished under its MIT licence (© SethGammon). 624 words, ~1,527 tokens.

Download SKILL.mdSave it as .claude/skills/experiment/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
experiment
description
Automated optimization loop with scalar fitness function. Proposes changes in isolated worktrees, measures with a metric command, keeps improvements, discards failures. Supports convergence detection and diminishing returns.
license
MIT
user-invocable
true
auto-trigger
false
trigger_keywords
experiment, optimize, try, A/B, measure
last-updated
2026-03-21

/experiment — Metric-Driven Optimization Loop

Inputs

The user provides three things:

  1. scope: Files to modify (glob pattern, e.g., "src/api/**/*.ts")
  2. metric: Shell command that outputs a single number (e.g., npm run build 2>&1 | tail -1 | grep -oP '\d+')
  3. budget: Iteration cap (default: 5) or time cap (e.g., "10 minutes")

If any input is missing, ask for it. The metric MUST output a single number to stdout.

Protocol

Step 1: BASELINE
  1. Stash any uncommitted changes (restore on exit)
  2. Run the metric command. Record the baseline value.
  3. Determine direction: does lower = better (bundle size, error count) or higher = better (FPS, test count)? Ask the user if ambiguous.
  4. Log: Baseline: {value} ({metric command})
Step 2: ITERATE

For each iteration (up to budget):

  1. Create isolation: Spawn a sub-agent in a worktree (isolation: "worktree")
  2. Propose change: The agent modifies files within scope to improve the metric. Provide context: baseline value, metric direction, scope, what previous iterations tried.
  3. Measure: Run the metric command in the worktree (via node scripts/run-with-timeout.js 300)
  4. Gate: Run typecheck (also via timeout wrapper). If it fails, discard immediately.
  5. Evaluate:
    • Improved? → KEEP. Merge the worktree branch. New baseline = new value.
    • Same or worse? → DISCARD. Delete the worktree.
  6. Log iteration:
    Iteration {N}: {value} ({delta from baseline}) → {KEEP|DISCARD}
    Change: {one-line description of what was tried}
Step 3: CONVERGENCE CHECK

After each iteration, check:

  • Local optimum: Last 3 iterations all discarded → stop ("no more improvements found")
  • Diminishing returns: Last kept improvement was < 0.5% → stop ("diminishing returns")
  • Budget exhausted: Iteration count or time exceeded → stop
Step 4: REPORT

Write results to .planning/research/experiment-{slug}.md:

# Experiment: {Description}

> Metric: `{command}`
> Direction: {lower|higher} is better
> Scope: {glob pattern}
> Budget: {N iterations}
> Date: {ISO date}

## Results

| Iteration | Value | Delta | Verdict | Change |
|-----------|-------|-------|---------|--------|
| baseline  | {N}   | —     | —       | —      |
| 1         | {N}   | {+/-} | KEEP    | {desc} |
| 2         | {N}   | {+/-} | DISCARD | {desc} |

## Outcome
- **Start**: {baseline}
- **End**: {final value}
- **Improvement**: {percentage}
- **Iterations**: {kept}/{total}
- **Stop reason**: {convergence|diminishing|budget}

## Kept Changes
{List of changes that were kept, with commit hashes}

Also log to .planning/telemetry/agent-runs.jsonl:

json
{"event":"experiment-complete","slug":"{slug}","baseline":0,"final":0,"improvement":"0%","kept":0,"total":0,"timestamp":"ISO"}

Common Metrics

GoalMetric Command
Reduce bundle sizenpm run build 2>&1 | grep -oP 'Total size: \K\d+'
Reduce type errorsnpx tsc --noEmit 2>&1 | grep -c 'error TS'
Increase test pass ratenpm test 2>&1 | grep -oP '\d+ passing'
Reduce file countfind src -name '*.ts' | wc -l
Reduce line countwc -l src/**/*.ts | tail -1 | awk '{print $1}'

When to Use

  • When you want to optimize a measurable metric (bundle size, error count, test coverage, FPS)
  • When you have a clear hypothesis but aren't sure which of several approaches wins
  • When manual A/B testing would be too slow or error-prone
  • NOT when the goal is subjective ("make it feel better") — the metric must be a number
Show full SKILL.md (251 more words)Show less

Safety Rules

  • NEVER modify files outside scope
  • ALWAYS use worktree isolation for changes
  • ALWAYS run typecheck before keeping a change
  • Restore stashed changes on exit (even on error)
  • If the metric command fails, treat as DISCARD (not crash)

Contextual Gates

Disclosure: "Running experiment loop on [target] with fitness: [function]. Each iteration commits. Budget: [N iterations]." Reversibility: amber — modifies source files across iterations; each iteration is committed; undo with git revert on kept commits. Trust gates:

  • Familiar (5+ sessions): iterates and commits autonomously; novices should use /improve with manual review between steps.

Quality Gates

  • Baseline was measured before any iterations ran
  • Every kept iteration improved the metric AND passed typecheck
  • Every discarded iteration has a logged reason
  • The stop reason is one of: convergence, diminishing returns, or budget exhausted
  • The experiment report exists at .planning/research/experiment-{slug}.md with all iteration rows filled

Fringe Cases

Metric command outputs nothing or non-numeric text: Treat as a metric failure. Ask the user to provide a command that outputs a single number to stdout before starting iterations.

No worktree support (e.g., shallow clone): Fall back to branch isolation. Create a branch, run changes there, measure, then delete or merge the branch. Never modify the working tree directly.

If .planning/research/ does not exist: Create it before writing the experiment report. If .planning/ itself doesn't exist, create the full path or output the report inline.

Budget exhausted with zero kept iterations: Report outcome as "no improvement found". This is a valid result — do not continue past the budget.

Exit Protocol

---HANDOFF---
- Experiment: {description}
- Result: {baseline} → {final} ({improvement}%)
- Kept: {N}/{total} iterations
- Stop reason: {reason}
- Report: .planning/research/experiment-{slug}.md
- Reversibility: amber — undo kept iterations with `git revert` on each kept commit
---

© SethGammon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in skills/experiment of SethGammon/Citadel.

  • SKILL.md
  • __benchmarks__/no-metric.md
  • __benchmarks__/with-metric.md

Open the folder on GitHubat commit e41ff1d

Compare with similar skills

Experiment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment this skillSethGammon/Citadel922—~1.5kAutomated safety check: PassMIT
Dev Instanceyc-software/qm15k—~3.7kAutomated safety check: NotesMIT
Harness ContributingFairladyZ625/harness-anything225—~4.2kAutomated safety check: PassAGPL-3.0
Squad Git Branching Workflowmicrosoft/waza1.4k4 repos~1.5kAutomated safety check: PassMIT
Cabloy Worktree Environmentcabloy/cabloy982—~2.9kAutomated safety check: NotesMIT
Move To Worktreefoyzulkarim/claude-lens250—~663Automated safety check: NotesMIT

Similar skills

  • Dev Instance

    yc-software/qm

    Run the current worktree as a production-shaped local dev instance with web, Slack, or both, on a real LLM + Postgres.

    15k GitHub stars~3.7k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • Harness Contributing

    FairladyZ625/harness-anything

    Contribute a public change to Harness Anything from a GitHub issue through an isolated worktree, scoped tests, manifest-selected gates, a complete bilingual PR body, review triage, and maintainer…

    225 GitHub stars~4.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Official

    Dev-first branching model for the Squad project: feature work branches from dev, issue branches follow a naming rule and parallel issues use git worktrees.

    1.4k GitHub starsUsed in 4 repos~1.5k tokens
    DevelopmentAuto-check passed
  • This skill must be used only when the user explicitly invokes /cabloy-worktree-environment or explicitly asks to perform the named Cabloy worktree-environment setup.

    982 GitHub stars~2.9k tokensUpdated today
    DevelopmentAuto-check: notes
  • Move To Worktree

    foyzulkarim/claude-lens

    After /start-task: park the current clean, pushed feature branch in its own issue-numbered nested worktree (.worktrees/<issue) and return the primary checkout to current main, so the next parallel…

    250 GitHub stars~663 tokensUpdated 2 mo ago
    DevelopmentAuto-check: notes
  • Devcontainer Dev

    stacklok/toolhive-studio

    Spin up and interact with ToolHive Studio's containerized dev environment (Xvfb + noVNC + DinD).

    170 GitHub stars~3.8k tokensUpdated today
    Agent WorkflowsAuto-check: notes

More from SethGammon/Citadel

All 48 skills in this repo
  • Create Skill

    SethGammon/Citadel

    Creates new skills from the user's repeating patterns. An agent skill from SethGammon/Citadel.

    922 GitHub stars~1.9k tokensUpdated 6 days ago
    Auto-check passed
  • Houseclean

    SethGammon/Citadel

    Cross-drive storage audit and cleanup. An agent skill from SethGammon/Citadel.

    922 GitHub stars~2.2k tokensUpdated 6 days ago
    Auto-check passed
  • Loop

    SethGammon/Citadel

    Bounded foreground repetition for the current session. An agent skill from SethGammon/Citadel.

    922 GitHub stars~1.4k tokensUpdated 6 days ago
    Auto-check passed
  • Triage

    SethGammon/Citadel

    GitHub issue and PR investigator. An agent skill from SethGammon/Citadel.

    922 GitHub stars~2.7k tokensUpdated 6 days ago
    Auto-check passed
  • Watch

    SethGammon/Citadel

    File sentinel that monitors the working directory for changes and marker comments, then auto-triggers appropriate skills.

    922 GitHub stars~2.9k tokensUpdated 6 days ago
    Auto-check passed
  • Archon

    SethGammon/Citadel

    Autonomous multi-session campaign agent. An agent skill from SethGammon/Citadel.

    922 GitHub stars~5.4k tokensUpdated 6 days ago
    Auto-check passed

Works with

Questions about Experiment

What does Experiment do?

Automated optimization loop with scalar fitness function. An agent skill from SethGammon/Citadel. Experiment is an agent skill from SethGammon/Citadel. Automated optimization loop with scalar fitness function.

When should I use Experiment?

Experiment fits situations like: tasks that involve Git worktrees.

How do I install Experiment in Claude Code?

Run `npx skills add SethGammon/Citadel --skill experiment -a claude-code`. Or copy the skill folder (skills/experiment in SethGammon/Citadel) into .claude/skills/experiment in your project. Claude Code loads it when a task matches its description.

How do I install Experiment in Codex?

Run `npx skills add SethGammon/Citadel --skill experiment -a codex`. Or copy the skill folder (skills/experiment in SethGammon/Citadel) into .agents/skills/experiment in your project. Codex loads it when a task matches its description.

Can I use Experiment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add SethGammon/Citadel --skill experiment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment, .gemini/skills/experiment, .github/skills/experiment and .opencode/skills/experiment in your project.

What does Experiment need to run?

Going by SKILL.md and its folder, Experiment needs the command-line tools its instructions call (npm, node, npx and git). Our summary lists: Node.js.

Does Experiment access the network?

SKILL.md contains no URLs. Its commands use npm, npx and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Experiment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment use?

Experiment is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment use?

About 1.5k tokens (SKILL.md is roughly 6.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Experiment?

Skills that share tags, products or a category with Experiment: Dev Instance (yc-software/qm, 15k stars), Harness Contributing (FairladyZ625/harness-anything, 225 stars), Squad Git Branching Workflow (microsoft/waza, 1.4k stars) and Cabloy Worktree Environment (cabloy/cabloy, 982 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment?

SethGammon (a GitHub user) maintains it in SethGammon/Citadel, which has 922 GitHub stars. The repository holds 48 skills in this directory. The repository was last updated on October 1, 2026.

Source: SethGammon/Citadel on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.