Profile
verl-project/verl-omni
Route a verl-omni performance investigation to the right tool and capture a usable trace.
Autonomous experiment loop: hypothesize modify test evaluate keep/discard repeat.
$ npx skills add vibeeval/vibecosystem --skill experiment-loop -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install vibeeval/vibecosystem experiment-loop --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment-loop .claude/skills/experiment-loop && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "experiment-loop" agent skill from https://github.com/vibeeval/vibecosystem/tree/main/skills/experiment-loop into .claude/skills/experiment-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "experiment-loop", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/vibeeval/vibecosystem/tree/main/skills/experiment-loopType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add vibeeval/vibecosystem --skill experiment-loop -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install vibeeval/vibecosystem experiment-loop --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/experiment-loop .agents/skills/experiment-loop && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "experiment-loop" agent skill from https://github.com/vibeeval/vibecosystem/tree/main/skills/experiment-loop into .agents/skills/experiment-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "experiment-loop", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vibeeval/vibecosystem --skill experiment-loop -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install vibeeval/vibecosystem experiment-loop --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/experiment-loop .cursor/skills/experiment-loop && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "experiment-loop" agent skill from https://github.com/vibeeval/vibecosystem/tree/main/skills/experiment-loop into .cursor/skills/experiment-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "experiment-loop", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/vibeeval/vibecosystem.git --path skills/experiment-loop--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add vibeeval/vibecosystem --skill experiment-loop -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install vibeeval/vibecosystem experiment-loop --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/experiment-loop .gemini/skills/experiment-loop && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "experiment-loop" agent skill from https://github.com/vibeeval/vibecosystem/tree/main/skills/experiment-loop into .gemini/skills/experiment-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "experiment-loop", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install vibeeval/vibecosystem experiment-loopInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add vibeeval/vibecosystem --skill experiment-loop -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/experiment-loop .github/skills/experiment-loop && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "experiment-loop" agent skill from https://github.com/vibeeval/vibecosystem/tree/main/skills/experiment-loop into .github/skills/experiment-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "experiment-loop", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add vibeeval/vibecosystem --skill experiment-loop -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install vibeeval/vibecosystem experiment-loop --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/vibeeval/vibecosystem.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/experiment-loop .opencode/skills/experiment-loop && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "experiment-loop" agent skill from https://github.com/vibeeval/vibecosystem/tree/main/skills/experiment-loop into .opencode/skills/experiment-loop/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "experiment-loop", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
experiment-loopAutonomous experiment loop: hypothesize modify test evaluate keep/discard repeat.
Experiment Loop is an agent skill from vibeeval/vibecosystem. Autonomous experiment loop: hypothesize modify test evaluate keep/discard repeat. Run N experiments automatically with measurable metrics. Works for performance optimization, A/B testing, prompt engineering, and any measurable improvement task.
Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering Performance optimization and A/B testing. The repository describes itself as: AI software team for Claude Code - 138 agents, 295 skills, 73 hooks. Self-learning, multi-agent swarm, autonomous skill evolution. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 3b763b1. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Experiment Loop loads about 1.8k tokens when it runs. Until then it costs about 68 tokens; SKILL.md has 475 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from vibeeval/vibecosystem at commit 3b763b1, republished under its MIT licence (© vibeeval). 475 words, ~1,822 tokens.
.claude/skills/experiment-loop/SKILL.md (or your agent's skills folder).Autonomous, iterative improvement inspired by Karpathy's autoresearch methodology. Define a metric, set a target, and let the loop run until the target is met or the iteration limit is reached.
1. HYPOTHESIZE -> Form a specific, falsifiable improvement hypothesis
2. MODIFY -> Apply the minimal code/config/prompt change
3. TEST -> Run the measurement suite (benchmarks, tests, evals)
4. EVALUATE -> Compare result against baseline and previous best
5. DECIDE -> KEEP if better, DISCARD (git stash pop --index) if worse
|
Repeat until target met OR max_iterations reachedEach iteration is atomic: one hypothesis, one change, one measurement, one decision.
Define an experiment in your task or in thoughts/EXPERIMENTS.md:
experiment:
name: "reduce-api-latency"
metric: "p95 response time (ms)"
baseline: 340
target: 200
direction: minimize # minimize | maximize
max_iterations: 10 # hard cap, never exceed
measurement_cmd: "npm run bench:api"
measurement_key: "p95" # JSON key from bench output
scope: "src/api/" # files the loop is allowed to touch| Field | Description |
|---|---|
metric | Human-readable name of what you are measuring |
baseline | Measured value before any changes (run this first) |
target | Success condition -- loop exits when this is met |
direction | minimize for latency/size, maximize for coverage/score |
max_iterations | Safety cap, default 10, absolute maximum 10 |
measurement_cmd | Shell command that produces JSON with the metric value |
scope | Directories/files the loop is allowed to modify |
Before every experiment iteration:
# Save current state
git stash push -u -m "experiment-loop: iteration N baseline"
# Run experiment
# ... apply hypothesis change ...
# ... run measurement ...
# Decision
if result is better:
git stash drop # keep changes, discard stash
else:
git stash pop --index # restore exactly: staged + unstagedNever skip the stash. Never accumulate multiple iterations without a decision checkpoint. If the measurement command fails or times out, treat it as DISCARD.
The experiment loop coordinates three vibecosystem agents:
| Phase | Agent | Role |
|---|---|---|
| Hypothesize | profiler | Identify bottlenecks, suggest what to change |
| Modify | spark | Apply the focused code change |
| Test + Evaluate | verifier / tdd-guide | Run benchmarks, tests, evals and parse results |
Spawn profiler once at the start to get the initial hypothesis queue. Then run spark + verifier in tight loops per iteration.
experiment:
name: "optimize-bundle-size"
metric: "gzipped bundle size (KB)"
baseline: 420
target: 300
direction: minimize
max_iterations: 10
measurement_cmd: "npm run build && node scripts/measure-bundle.js"
measurement_key: "gzipped_kb"
scope: "src/"Hypothesis queue to try in order:
moment with date-fns (smaller footprint)import() at route boundariesusedExports: true in webpack/rollup configaxios with native fetch wrapperexperiment:
name: "reduce-api-latency"
metric: "p95 response time (ms)"
baseline: 340
target: 200
direction: minimize
max_iterations: 8
measurement_cmd: "npm run bench:api"
measurement_key: "p95"
scope: "src/api/"Hypothesis queue:
max: 20)Promise.all)experiment:
name: "improve-test-coverage"
metric: "line coverage (%)"
baseline: 64
target: 80
direction: maximize
max_iterations: 10
measurement_cmd: "npm test -- --coverage --json > coverage.json"
measurement_key: "coverageMap.total.lines.pct"
scope: "src/"experiment:
name: "improve-extraction-accuracy"
metric: "extraction F1 score"
baseline: 0.71
target: 0.85
direction: maximize
max_iterations: 10
measurement_cmd: "python eval/run_evals.py --output eval/results.json"
measurement_key: "f1"
scope: "prompts/"Append each iteration result to thoughts/EXPERIMENTS.md:
## Experiment: reduce-api-latency
Started: 2026-04-07T10:00:00Z
Baseline: 340ms | Target: 200ms | Direction: minimize
### Iteration 1
- Hypothesis: Add Redis cache for repeated DB reads
- Change: `src/api/users.ts` lines 45-67 -- wrap DB call with cache layer
- Result: 280ms (improvement: -60ms, -17.6%)
- Decision: KEEP
- Cumulative best: 280ms
### Iteration 2
- Hypothesis: Replace N+1 queries with JOIN
- Change: `src/api/users.ts` lines 89-102 -- rewrite fetchWithPosts()
- Result: 210ms (improvement: -70ms, -25%)
- Decision: KEEP
- Cumulative best: 210ms
### Iteration 3
- Hypothesis: Add connection pool sizing max:20
- Change: `src/db/pool.ts` line 12 -- max: 10 -> 20
- Result: 215ms (regression: +5ms)
- Decision: DISCARD (restored via git stash pop)
- Cumulative best: 210ms
### Final Result
- Target: 200ms | Achieved: 210ms | Status: NEAR_MISS (within 5%)
- Iterations: 3 of 10 used
- Total improvement: -38% from baseline| Condition | Action |
|---|---|
| Target met | EXIT -- log SUCCESS, keep all accumulated changes |
| max_iterations reached | EXIT -- log PARTIAL, keep best achieved state |
| 3 consecutive DISCARDs | PAUSE -- re-run profiler for new hypothesis queue |
| Measurement command fails | DISCARD current iteration, continue loop |
| Git stash fails | STOP -- do not continue, report error |
Invoke this skill by describing the experiment:
Use experiment-loop to reduce the API p95 latency from 340ms to under 200ms.
Baseline measurement: npm run bench:api
Max iterations: 8
Scope: src/api/The loop will:
thoughts/EXPERIMENTS.md for prior runs on the same metricprofiler for an ordered hypothesis queue© vibeeval, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/experiment-loop of vibeeval/vibecosystem.
Open the folder on GitHubat commit 3b763b1
Experiment Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Experiment Loop this skillvibeeval/vibecosystem | 531 | — | ~1.8k | Automated safety check: Pass | MIT | |
| Profileverl-project/verl-omni | 1.2k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| Performance Optimizationalbumentations-team/AlbumentationsX | 566 | — | ~1.7k | Automated safety check: Pass | AGPL-3.0 | |
| Spec Optimizeleo-kuang-ai/spec-first | 107 | — | ~13k | Automated safety check: Pass | MIT | |
| Distribution Profilerai-analyst-lab/ai-analyst | 304 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Perfupraullenchai/Rapid-MLX | 3.9k | — | ~1.6k | Automated safety check: Notes | Custom licence |
verl-project/verl-omni
Route a verl-omni performance investigation to the right tool and capture a usable trace.
albumentations-team/AlbumentationsX
Systematic performance audit for AlbumentationsX runtime code.
leo-kuang-ai/spec-first
Run metric-driven iterative optimization loops. An agent skill from leo-kuang-ai/spec-first.
ai-analyst-lab/ai-analyst
Single-column distribution deep-dive. An agent skill from ai-analyst-lab/ai-analyst.
raullenchai/Rapid-MLX
Autonomous performance optimization: research, PoC, benchmark, implement, review, PR
stas00/the-art-of-debugging
Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.
vibeeval/vibecosystem
Framework for measuring and tracking agent response quality over time.
vibeeval/vibecosystem
Security-focused differential code review with blast radius analysis, risk-adaptive depth (DEEP/FOCUSED/SURGICAL), git history correlation, and structured finding format.
vibeeval/vibecosystem
A skill your agent uses when making any factual claim about the codebase — existence, absence, or behavior.
vibeeval/vibecosystem
Systematic false positive verification for security findings.
vibeeval/vibecosystem
n8n otomasyon workflow'lari. An agent skill from vibeeval/vibecosystem.
vibeeval/vibecosystem
A skill your agent uses when context compression is imminent, when resuming a session, or when preserving critical decisions across long tasks.
Autonomous experiment loop: hypothesize modify test evaluate keep/discard repeat. Experiment Loop is an agent skill from vibeeval/vibecosystem. Autonomous experiment loop: hypothesize modify test evaluate keep/discard repeat.
Experiment Loop fits situations like: tasks that involve Performance optimization; tasks that involve A/B testing.
Run `npx skills add vibeeval/vibecosystem --skill experiment-loop -a claude-code`. Or copy the skill folder (skills/experiment-loop in vibeeval/vibecosystem) into .claude/skills/experiment-loop in your project. Claude Code loads it when a task matches its description.
Run `npx skills add vibeeval/vibecosystem --skill experiment-loop -a codex`. Or copy the skill folder (skills/experiment-loop in vibeeval/vibecosystem) into .agents/skills/experiment-loop in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add vibeeval/vibecosystem --skill experiment-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-loop, .gemini/skills/experiment-loop, .github/skills/experiment-loop and .opencode/skills/experiment-loop in your project.
Going by SKILL.md and its folder, Experiment Loop needs the command-line tools its instructions call (git). Our summary lists: Python 3.
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Experiment Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.8k tokens (SKILL.md is roughly 7.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Experiment Loop: Profile (verl-project/verl-omni, 1.2k stars), Performance Optimization (albumentations-team/AlbumentationsX, 566 stars), Spec Optimize (leo-kuang-ai/spec-first, 107 stars) and Distribution Profiler (ai-analyst-lab/ai-analyst, 304 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
vibeeval (a GitHub user) maintains it in vibeeval/vibecosystem, which has 531 GitHub stars. The repository holds 144 skills in this directory. The repository was last updated on August 8, 2026.
Source: vibeeval/vibecosystem on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.