Agent skill

Iterative Refinement

by sammcj in sammcj/agentic-coding

Disciplined, measurable iteration for a substantial refinement or investigation: loop against verifiable pass/fail conditions, fan work out to subagents, and keep the main context lean.

Apache-2.0Auto-check passedAgent Workflows

Install Iterative Refinement

skills CLI
$ npx skills add sammcj/agentic-coding --skill iterative-refinement -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sammcj/agentic-coding iterative-refinement --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sammcj/agentic-coding.git skills-src && mkdir -p .claude/skills && cp -r skills-src/Skills_disabled/iterative-refinement .claude/skills/iterative-refinement && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
iterative-refinement
GitHub stars
162
Token cost
~3.2k tokens
SKILL.md length
1,880 words
Files
3 (incl. references)
Skills in repo
62
Repo updated
First seen
Licence
Apache-2.0

At a glance

Disciplined, measurable iteration for a substantial refinement or investigation: loop against verifiable pass/fail conditions, fan work out to subagents, and keep the main context lean.

  • Works in 6 steps: Turn the goal into a pass/fail rubric… → Bake the rubric into the artifact as a… → Change one layer, re-run the check.… → …
  • Improving something measurable over repeated cycles (tuning a metric
  • SKILL.md covers The loop, Writing good conditions, Context-window economics and Delegation and subagents, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Iterative Refinement is an agent skill from sammcj/agentic-coding. Disciplined, measurable iteration for a substantial refinement or investigation: loop against verifiable pass/fail conditions, fan work out to subagents, and keep the main context lean. Use when improving something measurable over repeated cycles (tuning a metric or detector, refactoring against a regression bar), chasing a surprising or suspicious number, or driving a long multi-step task where delegation and context discipline matter. Not for one-shot edits or quick lookups that don't warrant a loop.

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `evals/trigger_evals.json` and `references/worked_example.md`).

It sits in Agent Workflows, covering Subagents and Refactoring. The repository describes itself as: Agentic Coding Rules, Templates etc... The licence is Apache-2.0.

When your agent uses it

  • Improving something measurable over repeated cycles (tuning a metric
  • Refactoring against a regression bar)
  • Chasing a surprising
  • Suspicious number

Example prompts

  • “/iterative-refinement”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Turn the goal into a pass/fail rubric with explicit thresholds. "Useful" isn't checkable; "contamination < 2%, both naive and effective…
  2. Bake the rubric into the artifact as a self-check it prints. The artifact computes its own metrics and prints [PASS]/[FAIL] per condition…
  3. Change one layer, re-run the check. Catch each regression at the edit that caused it, not three edits later.
  4. Loop on a fast slice, not the full dataset. Size the slice so a run takes seconds, not minutes. Run the full corpus only at checkpoints…
  5. Keep a frozen held-out slice the loop never touches, and confirm against it before sign-off. Looping hard on one slice optimises the…
  6. When all conditions pass but value remains, write a stricter rubric and loop again. Stop when a cycle yields nothing material.

What it can do on your machine

Read from SKILL.md and the folder at commit 62ba5a2. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Iterative Refinement loads about 3.2k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 132 tokens; SKILL.md has 1,880 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~132
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sammcj/agentic-coding at commit 62ba5a2, republished under its Apache-2.0 licence (© sammcj). 1,880 words, ~3,247 tokens.

Download SKILL.mdSave it as .claude/skills/iterative-refinement/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
iterative-refinement
description
Disciplined, measurable iteration for a substantial refinement or investigation: loop against verifiable pass/fail conditions, fan work out to subagents, and keep the main context lean. Use when improving something measurable over repeated cycles (tuning a metric or detector, refactoring against a regression bar), chasing a surprising or suspicious number, or driving a long multi-step task where delegation and context discipline matter. Not for one-shot edits or quick lookups that don't warrant a loop.

Iterative refinement with verifiable conditions

Improve any system (script, pipeline, prompt, doc, config, dataset) by looping against measurable pass/fail conditions while keeping the main thread's context lean. Task-agnostic. The aim is to make "is it good yet?" a single repeatable command, catch each regression at the edit that caused it, and keep load-bearing reasoning in cheap, auditable steps.

Use this as a toolkit, not a script. Each method below earns its place by what it prevents, and that reasoning is stated inline so you can judge when it applies. Reach for the methods the situation calls for, scale them to the stakes, and adapt or skip what doesn't fit; they compose well, but no fixed subset is mandatory and this skill can't anticipate every task you'll point it at. When a method clearly fits, lean into it fully rather than half-applying it. The judgement of which to use, and how hard, stays yours.

The loop

  1. Turn the goal into a pass/fail rubric with explicit thresholds. "Useful" isn't checkable; "contamination < 2%, both naive and effective ratios reported, every input row classified, prints Overall: PASS" is.
  2. Bake the rubric into the artifact as a self-check it prints. The artifact computes its own metrics and prints [PASS]/[FAIL] per condition, so the check can't drift out of sync with the code the way an external checklist does. Now any party (you, a subagent, the user) re-verifies with one command.
  3. Change one layer, re-run the check. Catch each regression at the edit that caused it, not three edits later.
  4. Loop on a fast slice, not the full dataset. Size the slice so a run takes seconds, not minutes. Run the full corpus only at checkpoints, when a layer is structurally complete and at sign-off, since that's where rare signals and period splits actually appear and the cost is justified.
  5. Keep a frozen held-out slice the loop never touches, and confirm against it before sign-off. Looping hard on one slice optimises the rubric for that slice (Goodhart on your own metric); the held-out slice is what proves the gain generalised rather than memorised.
  6. When all conditions pass but value remains, write a stricter rubric and loop again. Stop when a cycle yields nothing material.

Writing good conditions

  • Cheap to evaluate: exit code, grep, a printed PASS line, a line count. If checking costs as much as the work, you won't loop.
  • Falsifiable with a number in it. "< 2%", not "low".
  • Tied to the goal, not a proxy. A metric can pass while the output is useless, so pair aggregates with a meaningfulness audit (below) that samples real outputs.
  • Reconciliation invariants are programmatic gates between steps that catch silent bugs: classified counts must sum to the total; naive and effective measures must bracket reality.
  • Track cost alongside correctness. A rubric that watches only accuracy hides regressions in the dimensions that also matter: tool-call count, token consumption, wall-clock, and error rate. A change that "passes" but doubles tokens is a finding, not a success.

Context-window economics

Every byte a tool returns stays in context and taxes every later turn. This is context engineering: treat the window as the scarcest resource, because reasoning quality degrades with depth (context rot) well before any hard limit. Three mitigations, highest leverage first:

  1. Think-in-code. To process output (filter/count/parse/aggregate), run code over it and print only the answer; raw bytes never enter the conversation. Reading bytes is justified only when you'll edit them. The context-mode ctx_execute tools do this in a sandbox when available.
  2. Return conclusions, not transcripts. A subagent's tool history stays in its context; only its final message returns. A task worth 50 file-reads costs you one paragraph. Forbid the subagent from pasting file bodies or full reports.
  3. Locate before reading. Print the line numbers or symbols you need with one targeted query, then read only that span.

Reasoning degrades in bands: roughly 0-100k tokens is peak, 100-150k still strong, 150-200k noticeably softer, and auto-compaction looms around a third of the window. Budget against the bands, not the ceiling. Bring the main thread's irreversible state (task list, strategy notes, decisions) up to date before you approach compaction; a summary written ahead of time survives, working memory you were relying on may not. If a single task can't fit the smart zone, that's a signal to decompose and delegate, not to push through.

Delegation and subagents

Push work off the main thread whenever it produces more bytes than its conclusion is worth.

  • Delegate bulk reads/scans (raw bytes stay in the subagent) and well-scoped edits.
  • For a delegated edit, give a tight spec: the objective, the exact variables/contracts it may touch, hard "don't touch X" boundaries, the output format, and the SAME machine-checkable acceptance loop you'd run. It inherits your verification standard and returns a compact summary.
  • State invariants up front so framing can't drift (for example: assess on effective metrics, never relabel the workflow "degraded", keep output lean). Cheaper to state than to repair.
  • Named subagent (fresh context) for research, inspection, or anything needing an unbiased read. Fork (inherits your context) only when accumulated nuance helps and a fresh view wouldn't. Never fork a review, audit, or premise-check; a fork inherits your blind spots.
  • Disjoint ownership when running several at once; overlapping write scope means they erase each other.
  • Verify what returns: a summary is a claim. Spot-check load-bearing results with a cheap command, or have a second independent agent check.

Fan-out topology

Match it to the write-pattern:

  • Sequential single subagent when edits hit the same region of one file; parallelism only causes merge collisions.
  • Parallel read-only subagents when subtasks mutate nothing shared. Best fit for a meaningfulness audit: one samples classifier false-negatives, one judges detector precision, one sanity-checks counts, same data, no writes, can't conflict. Send them in one batch so they truly run in parallel.
  • Git-worktree agents only when agents must write concurrently, or to race competing variants for a winner. They cost setup and disk; don't use when read-only or sequential suffices.
  • Scale the fleet to the task's complexity, and say so when you cap coverage. Rough budget: a simple lookup is one agent and a few tool calls; a broad audit is many agents in parallel. Fan-out multiplies token spend by agent count (a parallel multi-agent sweep can run on the order of 10-15x a single-threaded pass), so "ran 5 of 20" must never read as "covered everything".
Show full SKILL.md (812 more words)Show less

Accuracy and verifiability

  • Replicate before trusting. Reproduce a surprising number independently before acting on it; matching the original exactly proves you're measuring the same thing, then you can decompose it. A surprising aggregate is a hypothesis, not a finding.
  • Fix the measurement before the output. Suspect the instrument first. A measurement artifact is a construct-validity threat: the metric isn't measuring what its name claims. The fix is usually upstream (classification, provenance, sampling), not a better keyword list downstream.
  • Question imported benchmarks. Borrowed thresholds often assume a workflow you're not running (a Read:Edit ratio looks "degraded" only if you ignore that reads were routed through other tools).
  • Report naive and effective side by side. Keep the flawed metric visible next to the corrected one so the gap is auditable rather than hidden.
  • Meaningfulness audit. Aggregates hide false positives. Sample flagged items and judge them; estimate precision/recall on real data, not in theory.
  • Calibrate any judge before you trust it. An LLM-as-judge has its own failure modes (rewards verbosity, favours its own phrasing, drifts across a run). Check a sample of its verdicts against your own judgement first; only then let its aggregate stand in for yours.
  • Evaluate end state, not just per-step output. For a loop that mutates state (files, a database, a running system), assert on the final result and the invariants it must satisfy, not only on each intermediate artifact. Output audits miss bugs that surface only in the accumulated state.
  • Make outputs auditable. Emit example matches alongside counts so a human can spot-check the classifier without rerunning it.

Regression recovery

When a fix won't hold after a few tries, stop pushing on the same line:

  • Re-read the contract. Is the metric measuring what you think?
  • Prefer fixing input/measurement over patching output.
  • Escalate a genuinely stuck subtask to a fresh subagent running a systematic-debugging / Fagan-inspection pass: give it the symptom, the failed attempts, and the violated invariant, and let it root-cause from scratch. A fresh agent with an explicit protocol breaks loops repeated patching won't, and keeps the thrashing out of the main context.

Staying on task

A live task list is the backbone of completeness, not bureaucracy. Stand one up before the first loop iteration and keep it current as you work, marking items done and adding new ones the moment they surface. It's the one piece of state that reliably survives compaction, so treat it as the source of truth for what's done and what's left, not an afterthought you reconstruct at the end.

  • Decompose generously; err toward too many tasks, not too few. A task you can complete in one step is the right size; "fix everything" is not.
  • Exactly one item in-progress at a time. Mark it before starting, complete it before moving on.
  • Add a follow-up the instant you discover it, even mid-edit. Discoveries that live only in working memory die at compaction.
  • Encode delegated work as tasks too: what the agent is doing and its acceptance condition. An interrupted session then resumes from the list without re-deriving state.
  • Put load-bearing findings in the task description, not just the chat; task text survives compaction, conversation may not.

Prior art (why this works)

This playbook assembles established techniques. Reach for the named version when you want to go deeper or justify the approach.

Practice hereEstablished nameOrigin
The rubric-and-self-check loopevaluator-optimizer pattern; eval-driven developmentAnthropic, Building Effective Agents; Hamel Husain, Your AI Product Needs Evals
Critique-then-revise iterationReflexion; Self-RefineShinn et al. 2023; Madaan et al. 2023
Reconciliation invariants between stepsprogrammatic gatesAnthropic, Building Effective Agents
Fix the measurement before the outputconstruct validity; Goodhart's lawCronbach & Meehl 1955
Meaningfulness audit and judge calibrationLLM-as-judge failure modesZheng et al. 2023 (MT-Bench)
Context economics, conclusions-not-transcriptscontext engineering; context rot; compactionAnthropic, Effective Context Engineering for AI Agents
Read-only parallel auditors that voteself-consistency; orchestrator-workersWang et al. 2022; Anthropic, Multi-Agent Research System

For a worked example of the whole chain (a "frustration spike" that turned out to be a measurement artifact), see references/worked_example.md.

Checklist

A prompt to confirm what applies to the task in front of you, not a gate every task must clear. Skip the lines that don't fit.

  • Goal is a pass/fail rubric with numeric thresholds
  • Rubric baked into the artifact as a printed self-check
  • Cost dimensions tracked alongside accuracy (tokens, tool calls, errors)
  • Fast slice for the loop; full run plus a frozen held-out slice at sign-off
  • Bulk reads/edits delegated; only conclusions return
  • Invariants stated so delegates can't undo deliberate choices
  • Independent checks fanned out read-only; worktrees only for parallel writes or competing variants
  • Surprising numbers replicated before action
  • Naive and corrected metrics reported side by side
  • Meaningfulness audit on sampled outputs; judge calibrated before its aggregate is trusted
  • End state asserted, not just per-step output
  • On regression: question the measurement, fix upstream, escalate to structured debugging
  • Plateau -> stricter rubric, or stop

© sammcj, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in Skills_disabled/iterative-refinement of sammcj/agentic-coding.

  • SKILL.md
  • evals/trigger_evals.json
  • references/worked_example.md

Open the folder on GitHubat commit 62ba5a2

Compare with similar skills

Iterative Refinement next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Iterative Refinement compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Iterative Refinement this skillsammcj/agentic-coding162—~3.2kAutomated safety check: PassApache-2.0
Antigravity Agentsmarkfulton/claude-antigravity-agents130—~2.1kAutomated safety check: PassMIT
Subagent Coordinatorflyxl/datazen114—~908Automated safety check: PassGPL-3.0
Orchestrate Batch RefactorDimillian/Skills4k—~889Automated safety check: PassMIT
Batch Orchestrationrohitg00/pro-workflow2.9k—~1.2kAutomated safety check: PassNone
Codex CLIkortix-ai/suna20k—~1.6kAutomated safety check: PassCustom licence

Similar skills

  • Antigravity Agents

    markfulton/claude-antigravity-agents

    Delegate coding, code review, analysis, and research jobs to Google Antigravity CLI (agy) sub-agents that run alongside your own work.

    130 GitHub stars~2.1k tokensUpdated 2 mo ago
    Agent WorkflowsAuto-check passed
  • Orchestrate multi-track parallel feature development with subagents and git worktrees.

    114 GitHub stars~908 tokensUpdated 9 days ago
    Agent WorkflowsAuto-check passed
  • Plan and execute large refactor or rewrite efforts efficiently with parallel multi-agent analysis and implementation.

    4k GitHub stars~889 tokensUpdated 6 mo ago
    Agent WorkflowsAuto-check passed
  • Batch Orchestration

    rohitg00/pro-workflow

    Decompose large-scale changes into independent units and spawn parallel agents in isolated worktrees.

    2.9k GitHub stars~1.2k tokensUpdated 9 days ago
    Agent WorkflowsAuto-check passed
  • Codex CLI

    kortix-ai/suna

    Drive OpenAI's Codex CLI (codex exec) as a non-interactive coding sub-agent from inside Claude Code.

    20k GitHub stars~1.6k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Batch Refactor With Sub Agents

    r3bl-org/r3bl-open-core

    Use a sub-agent (like generalist) to perform repetitive code transformations across multiple files in a single turn.

    485 GitHub stars~816 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from sammcj/agentic-coding

All 62 skills in this repo
  • Yue2 Music

    sammcj/agentic-coding

    A skill your agent uses when generating songs with YuE2, covering a recording via SheetSage2 audio-to-ABC, editing a score or lyrics with melody preservation, or building a reproducible listening…

    162 GitHub stars~2.3k tokensUpdated yesterday
    Auto-check passed
  • Bento Slides

    sammcj/agentic-coding

    A skill your agent uses when creating or editing Bento (.bento.html) slide decks, including any request for a single-file HTML slide deck.

    162 GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Idrive Backup

    sammcj/agentic-coding

    A skill your agent uses whenever the user wants you to manage, discuss or diagnose iDrive Backup configuration on macOS

    162 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check: notes
  • Piper Tts Training

    sammcj/agentic-coding

    Train custom TTS voices for Piper (ONNX format) using fine-tuning or from-scratch approaches.

    162 GitHub stars~1.4k tokensUpdated yesterday
    Auto-check passed
  • PPTX To Md

    sammcj/agentic-coding

    Convert a PPTX slide deck into per-slide markdown that preserves both the verbatim text and the meaning of embedded screenshots, diagrams and charts in their original layout positions.

    162 GitHub stars~1.8k tokensUpdated yesterday
    Auto-check passed
  • Skill Creator Primer

    sammcj/agentic-coding

    You MUST load this skill before the skill-creator skill AND before making ANY change to, or conducting a review of ANY Agent Skill.

    162 GitHub stars~9.8k tokensUpdated yesterday
    Auto-check passed

Questions about Iterative Refinement

What does Iterative Refinement do?

Disciplined, measurable iteration for a substantial refinement or investigation: loop against verifiable pass/fail conditions, fan work out to subagents, and keep the main context lean. Iterative Refinement is an agent skill from sammcj/agentic-coding. Disciplined, measurable iteration for a substantial refinement or investigation: loop against verifiable pass/fail conditions, fan work out to subagents, and keep the main context lean.

When should I use Iterative Refinement?

Iterative Refinement fits situations like: improving something measurable over repeated cycles (tuning a metric; refactoring against a regression bar); chasing a surprising; suspicious number.

How do I install Iterative Refinement in Claude Code?

Run `npx skills add sammcj/agentic-coding --skill iterative-refinement -a claude-code`. Or copy the skill folder (Skills_disabled/iterative-refinement in sammcj/agentic-coding) into .claude/skills/iterative-refinement in your project. Claude Code loads it when a task matches its description.

How do I install Iterative Refinement in Codex?

Run `npx skills add sammcj/agentic-coding --skill iterative-refinement -a codex`. Or copy the skill folder (Skills_disabled/iterative-refinement in sammcj/agentic-coding) into .agents/skills/iterative-refinement in your project. Codex loads it when a task matches its description.

Can I use Iterative Refinement in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sammcj/agentic-coding --skill iterative-refinement -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/iterative-refinement, .gemini/skills/iterative-refinement, .github/skills/iterative-refinement and .opencode/skills/iterative-refinement in your project.

What does Iterative Refinement need to run?

SKILL.md names no scripts, command-line tools or credentials: Iterative Refinement is instructions for the agent only.

Does Iterative Refinement access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Iterative Refinement safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Iterative Refinement use?

Iterative Refinement is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Iterative Refinement use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 447 tokens, read only when the agent opens those files.

What are the alternatives to Iterative Refinement?

Skills that share tags, products or a category with Iterative Refinement: Antigravity Agents (markfulton/claude-antigravity-agents, 130 stars), Subagent Coordinator (flyxl/datazen, 114 stars), Orchestrate Batch Refactor (Dimillian/Skills, 4k stars) and Batch Orchestration (rohitg00/pro-workflow, 2.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Iterative Refinement?

sammcj (a GitHub user) maintains it in sammcj/agentic-coding, which has 162 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on October 7, 2026.

Source: sammcj/agentic-coding on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.