Agent skill

Experiment Iterative Coder

by EvoScientist in EvoScientist/EvoSkills

Iterative code refinement through plan → code → evaluate → refine cycles.

Apache-2.0Auto-check passedDevelopment

Install Experiment Iterative Coder

skills CLI
$ npx skills add EvoScientist/EvoSkills --skill experiment-iterative-coder -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install EvoScientist/EvoSkills experiment-iterative-coder --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/EvoScientist/EvoSkills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/experiment-iterative-coder .claude/skills/experiment-iterative-coder && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
experiment-iterative-coder
GitHub stars
474
Used in
3 other repos
Token cost
~2.5k tokens
SKILL.md length
1,080 words
Files
3 (incl. references, assets)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Iterative code refinement through plan → code → evaluate → refine cycles.

  • Works in 6 steps: Plan → Code → Evaluate → …
  • : the main agent delegates a code task with MODE: MOREEFFORT
  • SKILL.md covers When to Use This Skill, The Iteration Mindset, Before Starting: Load Context and Phase Decomposition, plus 5 more sections
  • Calls ruff, python and pip

What it does

Experiment Iterative Coder is an agent skill from EvoScientist/EvoSkills. Iterative code refinement through plan → code → evaluate → refine cycles. Runs lint checks (ruff), tests (pytest), and structured self-evaluation each cycle, then diagnoses failures and refines. Decomposes complex tasks into sequential phases, iterates up to 3 times per phase (10 total). Use when: the main agent delegates a code task with 'MODE: MOREEFFORT', the user selects 'More Effort' code generation mode, or the task explicitly requests iterative refinement for higher code quality. Do NOT use for single-pass…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files and assets (for example `assets/iteration-log-template.md` and `references/evaluation-protocol.md`).

It sits in Development, covering Linting and formatting, Code quality and Unit testing. It works with Ruff and pytest. The repository describes itself as: 🧬 Extend EvoScientist with Installable Skill & Knowledge Packs. The licence is Apache-2.0.

When your agent uses it

  • : the main agent delegates a code task with MODE: MOREEFFORT
  • The user selects More Effort code generation mode
  • The task explicitly requests iterative refinement for higher code quality
  • Single-pass code generation (Lite mode)

Example prompts

  • “MODE: MOREEFFORT”
  • “More Effort”
  • “/experiment-iterative-coder”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): write_file, edit_file, read_file, think_tool, execute

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Plan
  2. Code
  3. Evaluate
  4. Score
  5. Decide
  6. Log

What it can do on your machine

Read from SKILL.md and the folder at commit 9a9f8cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • write_file
    • edit_file
    • read_file
    • think_tool
    • execute

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • ruff
    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Experiment Iterative Coder loads about 2.5k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 176 tokens; SKILL.md has 1,080 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~176
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from EvoScientist/EvoSkills at commit 9a9f8cf, republished under its Apache-2.0 licence (© EvoScientist). 1,080 words, ~2,452 tokens.

Download SKILL.mdSave it as .claude/skills/experiment-iterative-coder/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
experiment-iterative-coder
description
Iterative code refinement through plan → code → evaluate → refine cycles. Runs lint checks (ruff), tests (pytest), and structured self-evaluation each cycle, then diagnoses failures and refines. Decomposes complex tasks into sequential phases, iterates up to 3 times per phase (10 total). Use when: the main agent delegates a code task with 'MODE: MORE_EFFORT', the user selects 'More Effort' code generation mode, or the task explicitly requests iterative refinement for higher code quality. Do NOT use for single-pass code generation (Lite mode), experiment pipeline orchestration (use experiment-pipeline), or diagnosing a specific experiment failure (use experiment-craft).
allowed-tools
write_file, edit_file, read_file, think_tool, execute
metadata.author
EvoScientist
metadata.version
1.0.0
metadata.tags
core, code-generation, iteration, refinement

Iterative Coder

Iterative code refinement through structured plan → code → evaluate → refine cycles. Each cycle runs objective checks (lint, tests) and self-evaluation, then diagnoses failures and plans targeted improvements. Reaches production quality in 3-8 iterations.

When to Use This Skill

  • Main agent delegates a code task prefixed with "MODE: MORE_EFFORT"
  • User selected "More Effort" mode for code generation
  • Task requires high code quality with verified correctness
  • Task involves complex implementation (5+ files, multiple modules)
  • You want to iterate on code quality rather than submit first-pass code
  • You mention "iterative refinement", "code quality loop", "plan-code-evaluate"

The Iteration Mindset

Code quality comes from fast feedback loops, not careful first attempts. A fast plan → code → evaluate → fix cycle beats spending 30 minutes on a "perfect" first implementation. The evaluate step reveals problems you cannot predict by thinking alone — lint errors, import failures, test regressions, and missing edge cases all surface immediately when you actually run the code.

Before Starting: Load Context

  1. Read /memory/experiment-memory.md for proven strategies from past cycles (skip if it doesn't exist)
  2. Identify existing tests, linting config (pyproject.toml, ruff.toml), or CI setup in the workspace
  3. Check available tools:
    bash
    ruff --version 2>&1; echo "---"; python -m pytest --version 2>&1
    If either is missing, you will skip that check during evaluation (do not fail the iteration).

Phase Decomposition

Before iterating, analyze the task and break it into sequential phases:

Task ComplexityRecommended Phases
Single file, well-defined function1 phase
2-4 files, clear interfaces2 phases
5+ files, multiple interacting modules3-5 phases

For each phase, define:

  • Name: concise label (e.g., "Data loading pipeline")
  • Goal: what "done" looks like for this phase
  • Verification signal: how to confirm the phase is complete (specific test, lint clean, output matches)

Order phases by dependency — later phases may build on earlier ones.

The Iteration Loop

For each phase, iterate up to 3 times. Global maximum: 10 iterations across all phases.

Step 1: Plan

Read current code and previous evaluation feedback (if any). Write a concise improvement plan.

First iteration of a phase: Write an initial implementation plan based on the phase goal.

Subsequent iterations: Analyze the last evaluation's feedback and diagnose the root cause of failures before planning changes. Do not repeat the same approach that already failed.

Adapt your plan based on the failure mode from the last evaluation:

Last FailurePlanned Response
TimeoutAdd --quick/--smoke mode, reduce data size, add early stopping
Syntax ErrorSimplify logic, run python -c "import ast; ast.parse(open('file.py').read())" to validate before running
Import ErrorCheck pip list, use only installed packages, add missing deps to requirements
Test FailureFocus on the specific failing test, make minimal targeted changes
Lint FailureRun ruff check --fix . && ruff format . before any logic changes
Low self-assessmentRe-read the original task requirements, check for missing functionality
Step 2: Code

Implement the plan. Keep changes focused on what the plan specifies.

  • Do not rewrite working files unless the plan explicitly requires it
  • After writing code, do a quick sanity read of the changed files
Step 3: Evaluate

CRITICAL: You MUST run these commands every iteration. Do not skip evaluation.

bash
# 1. Lint check
ruff check . 2>&1 | tail -20
echo "LINT_EXIT: $?"

# 2. Format check
ruff format --check . 2>&1 | tail -10
echo "FORMAT_EXIT: $?"

# 3. Run tests (only if test files exist in workspace)
python -m pytest -x -q --tb=short 2>&1 | tail -30
echo "TEST_EXIT: $?"

If ruff is not installed, skip checks 1-2. If pytest is not installed or no test files exist, skip check 3. Record which checks were skipped.

Step 4: Score

Compute a composite score from objective signals and self-assessment.

Objective signals (from Step 3 exit codes):

  • LINT_EXIT=0 → lint_score = 1.0, else lint_score = 0.0
  • FORMAT_EXIT=0 → format_score = 1.0, else format_score = 0.0
  • TEST_EXIT=0 → test_score = 1.0, else parse pass ratio from pytest output (e.g., "3 passed, 1 failed" → 0.75)

Self-assessment (rate 0.0 – 1.0): Evaluate on: correctness (does the code do what was asked?), completeness (all requirements addressed?), error handling (reasonable edge cases covered?), readability (clear names, structure).

Composite score — dynamic weighting based on available signals:

  • Lint + tests available: 0.2 × lint + 0.1 × format + 0.3 × test + 0.4 × self
  • Lint only (no tests): 0.3 × lint + 0.1 × format + 0.6 × self
  • Tests only (no ruff): 0.4 × test + 0.6 × self
  • Neither available: 1.0 × self

Self-assessment hard caps — prevent score inflation from self-assessment:

  • If lint check FAILED → composite capped at 0.4, regardless of self-assessment
  • If any test FAILED → composite capped at 0.6
  • Only claim composite ≥ 0.85 if BOTH lint and tests pass AND implementation is complete
  • Deductions: missing error handling for obvious cases (−0.1), hardcoded absolute paths (−0.05)

See references/evaluation-protocol.md for detailed scoring edge cases.

Show full SKILL.md (375 more words)Show less
Step 5: Decide
  • Composite score ≥ 0.85 → advance to next phase (or finish if last phase)
  • Composite score < 0.85 → return to Step 1 with evaluation feedback
  • Phase iteration limit reached (3 per phase) → advance to next phase anyway, note remaining issues
  • Global iteration limit reached (10 total) → stop, output current best result
Step 6: Log

CRITICAL: Append to /artifacts/iteration_log.md after every iteration.

Use the template at assets/iteration-log-template.md:

markdown
## Iteration {N} (Phase {M}/{T})
- **Score**: {composite} (lint={X} format={X} test={X} self={X})
- **Lint**: passed/failed ({N} issues)
- **Tests**: passed/failed ({passed}/{total})
- **Changes**: [{files changed}]
- **Feedback**: [{key evaluation findings}]
- **Next**: continue / next_phase / done

Completion

After all phases complete or global iteration limit is reached:

  1. Report to the caller:
    • Total iterations used
    • Final composite score
    • Key improvements per phase (1-2 sentences each)
  2. List all output file paths (code, configs, tests)
  3. Note remaining issues: lint warnings, missing tests, known limitations, TODOs

Counterintuitive Iteration Rules

  1. Fix lint before logic: Lint errors compound — one import error masks all test failures downstream. Always run ruff check --fix . before investigating logic bugs.

  2. 3 iterations is enough per phase: If you cannot fix it in 3 targeted iterations, the problem is architectural (wrong decomposition), not incremental. Advance to the next phase or re-plan rather than iterating further.

  3. Tests reveal more than reading: Running tests for 10 seconds teaches you more about correctness than reading code for 5 minutes. Always run tests, even when you are confident the code is correct.

  4. Score drops are information: If your composite score drops after a change, that is a signal about what matters. Analyze why it dropped before undoing the change.

  5. Don't gold-plate: 0.85 is the target, not 1.0. Diminishing returns kick in hard above 0.9. Ship and iterate in the next conversation if needed.

Skill Integration

Before Starting (load memory)

Refer to evo-memory → Read /memory/experiment-memory.md for prior strategies

On Failure (stuck after max iterations)

Refer to experiment-craft → 5-step diagnostic flow to understand the root cause before retrying

On Success (all phases complete, score ≥ 0.85)

Report to the main agent → main agent continues pipeline (data-analysis, writing, etc.)

Handoff Artifacts
ArtifactLocationUsed By
Iteration log/artifacts/iteration_log.mdMain agent summary, evo-memory ESE
Final codeWorkspace rootNext pipeline step
Test resultsIteration log entriesdata-analysis-agent

Reference Navigation

TopicReference FileWhen to Use
Scoring rules and edge casesevaluation-protocol.mdWhen scoring edge cases arise (partial tests, missing tools)
Iteration log templateiteration-log-template.mdEvery iteration (Step 6)

© EvoScientist, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references, assets) in skills/experiment-iterative-coder of EvoScientist/EvoSkills.

  • SKILL.md
  • assets/iteration-log-template.md
  • references/evaluation-protocol.md

Open the folder on GitHubat commit 9a9f8cf

Used in 3 other repositories

We found 3 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 3 other GitHub owners. This page covers the copy in EvoScientist/EvoSkills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Experiment Iterative Coder next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Experiment Iterative Coder compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Experiment Iterative Coder this skillEvoScientist/EvoSkills4743 repos~2.5kAutomated safety check: PassApache-2.0
Kedro Babysitkedro-org/kedro11k—~4kAutomated safety check: PassCustom licence
Cb Code QualityBlkLeg/CircuitBreaker201—~1.9kAutomated safety check: PassMIT
Python ProJeffallan/claude-skills12k—~1.6kAutomated safety check: PassMIT
Modern Pythonantoinebou12/uml-mcp105—~1kAutomated safety check: PassMIT
Specx Project Toolingmaksimzayats/specx202—~965Automated safety check: PassMIT

Similar skills

  • Kedro Babysit

    kedro-org/kedro

    Run Kedro's local lint / format / type-check / tests on changed files (uses the project's pre-commit hooks, ruff, mypy, pytest, lint-imports, detect-secrets, Make targets — in the right venv), or…

    11k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Cb Code Quality

    BlkLeg/CircuitBreaker

    Circuit Breaker code conventions and the quality gates that actually block a push — ruff, mypy, eslint, the pytest coverage ratchet, and the make verify tiers.

    201 GitHub stars~1.9k tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Python Pro

    Jeffallan/claude-skills

    Writes type-annotated Python 3.11+ with async patterns, dataclasses and pytest suites, validated with mypy in strict mode, black and ruff.

    12k GitHub stars~1.6k tokensUpdated 4 days ago
    DevelopmentAuto-check passed
  • Modern Python

    antoinebou12/uml-mcp

    Modern Python tooling and best practices using uv, ruff, ty, and pytest.

    105 GitHub stars~1k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Specx Project Tooling

    maksimzayats/specx

    Add strict Python project tooling for a specx service. An agent skill from maksimzayats/specx.

    202 GitHub stars~965 tokensUpdated 2 mo ago
    DevelopmentAuto-check passed
  • Python

    alinaqi/maggy

    Python development with ruff, mypy, pytest - TDD and type safety

    707 GitHub stars~1.1k tokensUpdated 13 days ago
    DevelopmentAuto-check passed

More from EvoScientist/EvoSkills

All 16 skills in this repo
  • Evomath Tao

    EvoScientist/EvoSkills

    A skill your agent uses whenever the user submits a non-trivial mathematical claim that needs a rigorous proof or audit.

    474 GitHub starsUsed in 2 repos~3.8k tokens
    Auto-check passed
  • Paper Figures

    EvoScientist/EvoSkills

    A skill your agent uses to produce standalone, publication-ready PNG graphics and reproducible matplotlib scripts from tabular data (CSVs or DataFrames).

    474 GitHub starsUsed in 1 repo~4.4k tokens
    Auto-check passed
  • Paper Planning

    EvoScientist/EvoSkills

    Guides pre-writing planning for academic papers with 4 structured steps: story design (task-challenge-insight-contribution-advantage), experiment planning (comparisons + ablations), figure design…

    474 GitHub starsUsed in 3 repos~2.4k tokens
    Auto-check passed
  • Research Survey

    EvoScientist/EvoSkills

    Generates structured literature survey reports from collected papers using a multi-stage pipeline: outline generation (query-type adaptive) → draft survey → section-by-section expansion → summary…

    474 GitHub starsUsed in 3 repos~2.5k tokens
    Auto-check passed
  • Paper Navigator

    EvoScientist/EvoSkills

    Find and read academic papers (S2 + arXiv). An agent skill from EvoScientist/EvoSkills.

    474 GitHub stars~6.3k tokensUpdated 7 days ago
    Auto-check: notes
  • Academic Slides

    EvoScientist/EvoSkills

    A skill your agent uses for creating or refining an academic slide deck and the talk built around it: structuring a conference talk, thesis defense, lab meeting, or paper-to-slides deck; deciding…

    474 GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed

Works with

Categories

Questions about Experiment Iterative Coder

What does Experiment Iterative Coder do?

Iterative code refinement through plan → code → evaluate → refine cycles. Experiment Iterative Coder is an agent skill from EvoScientist/EvoSkills. Iterative code refinement through plan → code → evaluate → refine cycles.

When should I use Experiment Iterative Coder?

Experiment Iterative Coder fits situations like: : the main agent delegates a code task with MODE: MOREEFFORT; the user selects More Effort code generation mode; the task explicitly requests iterative refinement for higher code quality; single-pass code generation (Lite mode).

How do I install Experiment Iterative Coder in Claude Code?

Run `npx skills add EvoScientist/EvoSkills --skill experiment-iterative-coder -a claude-code`. Or copy the skill folder (skills/experiment-iterative-coder in EvoScientist/EvoSkills) into .claude/skills/experiment-iterative-coder in your project. Claude Code loads it when a task matches its description.

How do I install Experiment Iterative Coder in Codex?

Run `npx skills add EvoScientist/EvoSkills --skill experiment-iterative-coder -a codex`. Or copy the skill folder (skills/experiment-iterative-coder in EvoScientist/EvoSkills) into .agents/skills/experiment-iterative-coder in your project. Codex loads it when a task matches its description.

Can I use Experiment Iterative Coder in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add EvoScientist/EvoSkills --skill experiment-iterative-coder -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/experiment-iterative-coder, .gemini/skills/experiment-iterative-coder, .github/skills/experiment-iterative-coder and .opencode/skills/experiment-iterative-coder in your project.

What does Experiment Iterative Coder need to run?

Going by SKILL.md and its folder, Experiment Iterative Coder needs the command-line tools its instructions call (ruff, python and pip). Our summary lists: Python 3. Its frontmatter pre-approves these tools: write_file, edit_file, read_file, think_tool, execute.

Does Experiment Iterative Coder access the network?

SKILL.md contains no URLs. Its commands use pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Experiment Iterative Coder safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Experiment Iterative Coder use?

Experiment Iterative Coder is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Experiment Iterative Coder use?

About 2.5k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 940 tokens, read only when the agent opens those files.

What are the alternatives to Experiment Iterative Coder?

Skills that share tags, products or a category with Experiment Iterative Coder: Kedro Babysit (kedro-org/kedro, 11k stars), Cb Code Quality (BlkLeg/CircuitBreaker, 201 stars), Python Pro (Jeffallan/claude-skills, 12k stars) and Modern Python (antoinebou12/uml-mcp, 105 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Experiment Iterative Coder?

EvoScientist (a GitHub organization) maintains it in EvoScientist/EvoSkills, which has 474 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on September 30, 2026.

Source: EvoScientist/EvoSkills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.