Agent skill

Optimize Loop

by gaasher in gaasher/Agent-Loop-Skills

A skill your agent uses when the user wants to iteratively improve an artifact under a hard correctness bound while minimizing a measured cost — refactoring a code module to cut complexity while its…

MITAuto-check: warningsDatabases

Install Optimize Loop

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add gaasher/Agent-Loop-Skills --skill optimize-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gaasher/Agent-Loop-Skills optimize-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/loops/optimize-loop .claude/skills/optimize-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
optimize-loop
GitHub stars
174
Token cost
~2.4k tokens
SKILL.md length
1,211 words
Files
5
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user wants to iteratively improve an artifact under a hard correctness bound while minimizing a measured cost — refactoring a code module to cut complexity while its…

  • Speeding up a SQL query while it returns the same rows
  • SKILL.md covers When to use, Setup, The loop (until plateau or ) and Ledger, plus 1 more section
  • Runs Python scripts from its folder; calls python3
  • Tasks that involve SQL

What it does

Optimize Loop is an agent skill from gaasher/Agent-Loop-Skills. Use when the user wants to iteratively improve an artifact under a hard correctness bound while minimizing a measured cost — refactoring a code module to cut complexity while its test suite stays green, OR speeding up a SQL query while it returns the same rows. Each iteration applies one focused change, checks a correctness gate that must pass, measures a metric that must drop, and keeps the change only if both hold, else reverts; loops to a plateau or budget. Not for adding features, fixing bugs, or any change…

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files (for example `examples/refactor.run.yaml`, `examples/sql.run.yaml` and `tools/bench.py`). Compatibility notes: Requires Python 3.9+

It sits in Databases, covering SQL, Test generation and Refactoring. It works with SQL. The repository describes itself as: Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills… The licence is MIT.

When your agent uses it

  • Speeding up a SQL query while it returns the same rows
  • Tasks that involve SQL
  • Tasks that involve Test generation

Example prompts

  • “/optimize-loop”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.9+

What it can do on your machine

Read from SKILL.md and the folder at commit f1169e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.9+

    From compatibility in the SKILL.md frontmatter.

Context cost

Optimize Loop loads about 2.4k tokens when it runs. Until then it costs about 144 tokens; SKILL.md has 1,211 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~144
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:21
    budget runs out. Once the loop starts, do not pause for permission.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gaasher/Agent-Loop-Skills at commit f1169e6, republished under its MIT licence (© gaasher). 1,211 words, ~2,427 tokens.

Download SKILL.mdSave it as .claude/skills/optimize-loop/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
optimize-loop
description
Use when the user wants to iteratively improve an artifact under a hard correctness bound while minimizing a measured cost — refactoring a code module to cut complexity while its test suite stays green, OR speeding up a SQL query while it returns the same rows. Each iteration applies one focused change, checks a correctness gate that must pass, measures a metric that must drop, and keeps the change only if both hold, else reverts; loops to a plateau or budget. Not for adding features, fixing bugs, or any change that is allowed to alter behaviour or results.
compatibility
Requires Python 3.9+
metadata.version
0.1.0

Optimize Loop

An evaluator-optimizer loop with a pluggable correctness gate + minimized metric. The artifact is some editable thing (a code module or a SQL query); the feedback signal is two-part: a bound gate that must pass (behaviour/results unchanged) and a bound metric that must drop (the cost you minimize). You apply one change, check the gate, measure the metric, and keep the change only if the gate passes AND the metric improves — otherwise you revert. Repeat until the metric stops improving or the budget runs out. Once the loop starts, do not pause for permission.

Two ready bindings ship in tools/ (both vendored, stdlib-only):

  • code mode — gate: <gate_cmd> (the test suite) exits 0; metric: tools/metrics.py prints complexity (primary), max_nesting, loc (lexicographic tie-breakers). Lower is better.
  • sql mode — gate: the result-set hash from tools/bench.py matches the baseline; metric: the same tool's median_ms. Lower is better.

The gate is non-negotiable in both modes: a change that fails it is a regression, not an improvement. Never edit the ground truth (the tests / tools/metrics.py in code mode, the database / tools/bench.py in sql mode) — editing what measures you to move the number defeats the loop.

When to use

Use when there is a clear correctness bound to hold and a number to minimize: refactoring code that has a passing test suite (cut complexity), or tuning a SQL query that has a fixed result-set (cut latency). The default is the matching shipped tool; the escape hatch is to bind any <gate_cmd> that exits 0 on pass and any <metric_cmd> that prints a single number to minimize (e.g. a linter's issue count, or a non-SQLite engine's timing + result fingerprint). Not for adding features or fixing bugs — those intend to change behaviour, which this loop is built to forbid.

Setup

Resolve bindings interactively. If loop.run.yaml exists in the working dir, load it, confirm the values back in one line, and skip to the loop. Otherwise pick <mode> first (it selects the gate + metric), then on Claude Code (the AskUserQuestion tool is available) infer a likely value for each binding and present it as the recommended option; on other hosts ask each as a quoted plain-text prompt. Then write loop.run.yaml and confirm the values before creating any other files.

<gate_cmd> and <metric_cmd> are the pluggable core: bind them per <mode> from the table. In sql mode one bench command supplies both — its hash is the gate, its median_ms is the metric.

bindingmeaningdefaulthow to infer
<mode>code (refactor under test) or sql (query, fixed results)—the artifact's kind
<editable_files>the file(s) the loop may change—code: source files (not tests/configs/the tool); sql: the query file (+ optional indexes file)
<gate_cmd>the bound gate that must PASS, else revert—code: the test command (exits 0 on pass); sql: implicit — candidate hash must equal the baseline hash from the bench command
<metric_cmd>the bound metric printing a number to minimize—code: python3 <skill_dir>/tools/metrics.py <editable_files> (→ complexity, then max_nesting, loc); sql: python3 <skill_dir>/tools/bench.py --db <db> --query <query_file> --setup <indexes_file> --repeat 5 (→ median_ms, hash)
<sandbox_root>where snapshots + the ledger live./sandbox—
<budget>max iterations (hard cap)8—
<patience>stop after N consecutive no-improvement iterations3—

<skill_dir> is this skill's installed folder; substitute the real path when writing loop.run.yaml. For non-default engines/languages, bind any <gate_cmd>/<metric_cmd> meeting the contract above. Two worked configs: examples/refactor.run.yaml (code) and examples/sql.run.yaml (sql).

Show full SKILL.md (658 more words)Show less

The loop (until plateau or <budget>)

Copy this checklist and tick items off:

  • Iteration 0 — baseline: run <gate_cmd> (code) — if not green, stop (the loop needs a passing gate to protect behaviour). Run <metric_cmd>; record the metric as the current best, and in sql mode record the baseline hash as the correctness reference. Log the baseline row.
  • Snapshot every file in <editable_files> to <sandbox_root>/iter<N>/ so the iteration reverts.
  • Apply one focused change (see the per-mode ideas below) — one idea per iteration so each metric delta is attributable.
  • Check the gate: code — run <gate_cmd>; sql — read the candidate's hash from <metric_cmd>.
  • If the gate fails (tests non-zero / hash ≠ baseline / the tool errored), discard: restore from the snapshot, log the reason, continue.
  • Measure the metric and compare to the best: code — the triple (complexity, max_nesting, loc) lexicographically (complexity first; only on a tie consult max_nesting, then loc); sql — median_ms, keeping only on a margin clear of timing noise (default ≥ 3% relative).
  • Keep if the metric strictly improves the best (update the best, leave the files in place), else discard (restore from the snapshot).
  • Append a ledger row; stop on plateau (<patience>) or <budget>.

Change ideas — code mode: flatten nested if/else into guard clauses, replace a hand-rolled loop with a stdlib call (sum, min, max, statistics.*), collapse duplicated branches, remove dead code. Preserve public behaviour — names, signatures, return shapes, raised exceptions; the test suite is the contract.

Change ideas — sql mode: add an index to <indexes_file> covering filtered/joined/grouped columns; rewrite the query (correlated subquery → JOIN + GROUP BY, hoist a repeated computation, replace SELECT * with needed columns, push a filter earlier, drop a redundant DISTINCT/sort). The hash is over the multiset of rows, so it does not catch a changed row order — if ORDER BY is part of the contract, eyeball that the rewrite preserves it.

Lexicographic keep (code mode), current best (18, 3, 64): (15, 3, 45) keep (lower complexity); (18, 2, 70) keep (tie complexity, lower nesting); (18, 3, 61) keep (tie, fewer lines); (18, 3, 64) discard (no progress); (19, 1, 20) discard (higher complexity outweighs simpler nesting/loc).

Plateau counting: increment the no-improvement counter on every iteration that does not set a new best — discarded for a failed gate, a broken change, or an insufficient metric gain — and reset it to 0 on each keep. <patience> fruitless iterations in a row ends the run; <budget> is the hard cap. On stop, restore the working files to the best iteration (if the latest was a discard) and report: baseline vs best metric (and, in sql mode, the speedup factor), the trajectory, and the winning change set. If you run low on ideas before the budget, look harder rather than stopping early.

Ledger

<sandbox_root>/ledger.tsv, tab-separated, never commas in the description. status ∈ {keep, discard, baseline}. Use the columns for the active <mode>.

Code mode header iter complexity max_nesting loc status description:

iter	complexity	max_nesting	loc	status	description
0	23	7	86	baseline	unmodified module
1	19	5	78	keep	flatten summarize guard clauses
2	19	5	80	discard	extract helper (no complexity gain)
3	13	3	40	keep	use statistics + min/max/median

SQL mode header iter median_ms rows hash_ok status description (hash_ok ∈ {yes,no,-}):

iter	median_ms	rows	hash_ok	status	description
0	1121.06	10	-	baseline	correlated subquery no index
1	6.82	10	yes	keep	rewrite correlated subquery as JOIN + GROUP BY
2	1.18	10	yes	keep	add index orders(customer_id, amount)
4	0.40	10	no	discard	drop ORDER BY — changed result set

Report the best iteration, not necessarily the last.

Constraints

  • Only edit files in <editable_files> — the gate's ground truth is read-only: the tests and tools/metrics.py (code), the database and tools/bench.py (sql). Editing what measures you to move the number invalidates the run.
  • The correctness gate is non-negotiable. A change that fails it (tests red, or hash ≠ baseline) is a regression, not an optimization — revert it regardless of the metric. A green gate after a behaviour change means the gate is too weak, not that the change is safe; prefer holding behaviour identical over trusting a thin gate.
  • One change per iteration, so each metric delta is attributable.
  • Measure every candidate the same way (code: the same <gate_cmd>; sql: the same --repeat); compare the metric, not a single noisy run.
  • Keep changes within the existing dependency set; do not add imports the project lacks (stdlib is fine).
  • The sandbox is self-contained — no ../ escapes beyond the bound <sandbox_root>.
  • Do not pause the loop to ask whether to continue; run until plateau or budget.

© gaasher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in loops/optimize-loop of gaasher/Agent-Loop-Skills.

  • SKILL.md
  • examples/refactor.run.yaml
  • examples/sql.run.yaml
  • tools/bench.py
  • tools/metrics.py

Open the folder on GitHubat commit f1169e6

Compare with similar skills

Optimize Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Optimize Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Optimize Loop this skillgaasher/Agent-Loop-Skills174—~2.4kAutomated safety check: WarnMIT
Rift Backend EffectCompound-inc/rift124—~1.8kAutomated safety check: PassCustom licence
Protheus Data Dictionary Lookuptotvs/engpro-advpl-tlpp-skills143—~1.5kAutomated safety check: PassMIT
Coding Agentmastra-ai/mastra29k—~2.3kAutomated safety check: PassCustom licence
Design Itsmallnest/goal-workflow289—~1.1kAutomated safety check: PassMIT
Relational Query ProcessorFoundationDB/fdb-record-layer675—~1kAutomated safety check: PassApache-2.0

Similar skills

  • Rift Backend Effect

    Compound-inc/rift

    A skill your agent uses when adding, reviewing, or refactoring backend code in Rift's TanStack Start app that should follow apps/start/BACKENDEFFECTPLAYBOOK.md.

    124 GitHub stars~1.8k tokensUpdated 10 days ago
    DevOps & CloudAuto-check passed
  • Protheus Data Dictionary Lookup

    totvs/engpro-advpl-tlpp-skills

    Queries the TOTVS Protheus ERP data dictionary for tables, fields, indexes, parameters, triggers and lookups, including impact checks during refactoring.

    143 GitHub stars~1.5k tokensUpdated 3 days ago
    DevelopmentAuto-check passed
  • Coding Agent

    mastra-ai/mastra

    Authoring playbook for building agents that write, edit, review, or refactor code.

    29k GitHub stars~2.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Design It

    smallnest/goal-workflow

    A skill your agent uses when turning a requirement, spec, or feature brief into a single self-contained HTML design document in a fixed house style — one styled HTML page with a table-of-contents…

    289 GitHub stars~1.1k tokensUpdated 25 days ago
    Frontend & DesignAuto-check passed
  • Relational Query Processor

    FoundationDB/fdb-record-layer

    Specialized skill for working in the fdb-relational-core SQL processing layer — parser, plan generator, and Cascades planner.

    675 GitHub stars~1k tokensUpdated today
    DatabasesAuto-check passed
  • Diesel Guard

    ayarotsky/diesel-guard

    Lints Diesel and SQLx Postgres migrations for unsafe schema changes that lock tables or cause downtime, and authors custom Rhai checks.

    121 GitHub stars~3.1k tokensUpdated 10 days ago
    DatabasesAuto-check passed

More from gaasher/Agent-Loop-Skills

All 21 skills in this repo
  • Alpha Evolve

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel…

    174 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Karpathy

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

    174 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Tournament Autoresearch

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants an autonomous ML research loop that pressure-tests competing ideas before spending compute — several research subagents each propose one architecture…

    174 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Dueling Autoresearch

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.

    174 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check: warnings
  • Anomaly Investigation

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

    174 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Auto-check passed

Works with

Questions about Optimize Loop

What does Optimize Loop do?

A skill your agent uses when the user wants to iteratively improve an artifact under a hard correctness bound while minimizing a measured cost — refactoring a code module to cut complexity while its…. Optimize Loop is an agent skill from gaasher/Agent-Loop-Skills. Use when the user wants to iteratively improve an artifact under a hard correctness bound while minimizing a measured cost — refactoring a code module to cut complexity while its test suite stays green, OR speeding up a SQL query while it returns the same rows.

When should I use Optimize Loop?

Optimize Loop fits situations like: speeding up a SQL query while it returns the same rows; tasks that involve SQL; tasks that involve Test generation.

How do I install Optimize Loop in Claude Code?

Run `npx skills add gaasher/Agent-Loop-Skills --skill optimize-loop -a claude-code`. Or copy the skill folder (loops/optimize-loop in gaasher/Agent-Loop-Skills) into .claude/skills/optimize-loop in your project. Claude Code loads it when a task matches its description.

How do I install Optimize Loop in Codex?

Run `npx skills add gaasher/Agent-Loop-Skills --skill optimize-loop -a codex`. Or copy the skill folder (loops/optimize-loop in gaasher/Agent-Loop-Skills) into .agents/skills/optimize-loop in your project. Codex loads it when a task matches its description.

Can I use Optimize Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaasher/Agent-Loop-Skills --skill optimize-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimize-loop, .gemini/skills/optimize-loop, .github/skills/optimize-loop and .opencode/skills/optimize-loop in your project.

What does Optimize Loop need to run?

Going by SKILL.md and its folder, Optimize Loop needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.9+.

Does Optimize Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Optimize Loop safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Optimize Loop use?

Optimize Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Optimize Loop use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Optimize Loop?

Skills that share tags, products or a category with Optimize Loop: Rift Backend Effect (Compound-inc/rift, 124 stars), Protheus Data Dictionary Lookup (totvs/engpro-advpl-tlpp-skills, 143 stars), Coding Agent (mastra-ai/mastra, 29k stars) and Design It (smallnest/goal-workflow, 289 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Optimize Loop?

gaasher (a GitHub user) maintains it in gaasher/Agent-Loop-Skills, which has 174 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on June 30, 2026.

Source: gaasher/Agent-Loop-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.