Agent skill

Test Iterate Loop

by flonat in flonat/flonat-research

Autonomously diagnose a codebase, apply minimal fixes, and rerun tests until they pass or a real blocker is reached.

MITAuto-check passedDevelopment

Install Test Iterate Loop

skills CLI
$ npx skills add flonat/flonat-research --skill test-iterate-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install flonat/flonat-research test-iterate-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/flonat/flonat-research.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/test-iterate-loop .claude/skills/test-iterate-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-iterate-loop
GitHub stars
145
Token cost
~2.2k tokens
SKILL.md length
871 words
Files
1
Skills in repo
83
Repo updated
First seen
Licence
MIT

At a glance

Autonomously diagnose a codebase, apply minimal fixes, and rerun tests until they pass or a real blocker is reached.

  • Works in 4 steps: Detect test runner → Initial run → Iteration loop → …
  • The user explicitly requests an iterative fix-until-green loop across Python
  • SKILL.md covers Hard Rules, When to Use, When NOT to Use and Modes, plus 9 more sections
  • Calls git, uv and npm

What it does

Test Iterate Loop is an agent skill from flonat/flonat-research. Autonomously diagnose a codebase, apply minimal fixes, and rerun tests until they pass or a real blocker is reached. Use when the user explicitly requests an iterative fix-until-green loop across Python, R, Julia, or HPC workflows.

Its SKILL.md is about 2.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development. It works with Python. The repository describes itself as: Shareable Claude Code + Codex infrastructure for PhD researchers — skills, agents, hooks, and rules for academic workflows. The licence is MIT.

When your agent uses it

  • The user explicitly requests an iterative fix-until-green loop across Python

Example prompts

  • “/test-iterate-loop”

Requirements

  • Python 3
  • Docker
  • Pre-approved tools (allowed-tools): Read, Write, Edit, Glob, Grep, Bash(uv*), Bash(pytest*), Bash(Rscript*), Bash(julia*), Bash(docker*), Bash(make*), Bash(git*), TaskCreate, TaskUpdate, AskUserQuestion

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Detect test runner
  2. Initial run
  3. Iteration loop
  4. Termination + report

What it can do on your machine

Read from SKILL.md and the folder at commit da27600. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit
    • Glob
    • Grep
    • Bash(uv*)
    • Bash(pytest*)
    • Bash(Rscript*)
    • Bash(julia*)
    • Bash(docker*)

    …and 5 more on the same allowed-tools line.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • uv
    • npm
    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, uv and npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Iterate Loop loads about 2.2k tokens when it runs. Until then it costs about 62 tokens; SKILL.md has 871 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~62
When it runs · the whole SKILL.md, loaded when a task matches
~2.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from flonat/flonat-research at commit da27600, republished under its MIT licence (© flonat). 871 words, ~2,230 tokens.

Download SKILL.mdSave it as .claude/skills/test-iterate-loop/SKILL.md (or your agent's skills folder).
name
test-iterate-loop
description
Autonomously diagnose a codebase, apply minimal fixes, and rerun tests until they pass or a real blocker is reached. Use when the user explicitly requests an iterative fix-until-green loop across Python, R, Julia, or HPC workflows.
allowed-tools
Read, Write, Edit, Glob, Grep, Bash(uv*), Bash(pytest*), Bash(Rscript*), Bash(julia*), Bash(docker*), Bash(make*), Bash(git*), TaskCreate, TaskUpdate, AskUserQuestion
argument-hint
[project-path] [--max-iter N] [--mock-hpc] [--container <image>]
skill-dependencies
computational-experiments

Test-Iterate Loop — Autonomous Bug Fix Cycle

Autonomous loop: run tests → root-cause failures → apply minimal fix → retry. Bounded by iteration cap and same-error-repeat detector. Never commits — leaves clean working tree + markdown report. Generic across Python (pytest), R (testthat), Julia (Pkg.test), and HPC pipelines (mock-HPC Docker).

Hard Rules

Existential — block proceed
  1. Never git commit, git push, or modify .git/ state. All changes go into the working tree only. The user reviews the final clean tree + report and decides what to commit. Enforced via the standard forbid-list per subagent-write-guard.md if dispatching sub-agents.
  2. Bounded iteration: max 10 iterations OR 3-same-error-repeat — whichever first. Hardcoded ceiling. Past that, stop and summarise. Don't ask "continue?" — exhaustion is the signal.
  3. Each iteration must change ≥1 file. A no-op iteration (Claude couldn't propose a fix) counts as a same-error-repeat.
  4. Memory bug detector: flag any iteration that adds a model load, env var change, or version pin without a corresponding test — these are common failure-cause patterns from past HPC sessions.
Format — catch in review
  1. Every iteration logged to log/test-iterate/<project>-YYYY-MM-DD-HHMM.md with: hypothesis, fix applied, result.
  2. Final report at the same path summarises iterations, terminal state, and remaining failures (if any).
  3. Use TaskCreate/TaskUpdate for live progress tracking — the user can see what iteration is running and why.

When to Use

  • A pipeline / library has failing tests and you want them fixed autonomously
  • Pre-flight before submitting an HPC job (catch torch/transformers/CUDA mismatches in mock-HPC Docker before the real job queues)
  • Refactor cycles: change → run tests → fix breaks → repeat
  • Reproducing a failure on a fresh checkout

When NOT to Use

  • The test suite itself is broken (fix the tests first, or you'll loop trying to fix code to match wrong tests)
  • The failure root cause is external (network, third-party API, hardware) — the active agent cannot fix those
  • You want to write tests, not fix them — use computational-experiments or direct work
  • The fix requires research / design decisions, not just code adjustment

Modes

InvocationBehaviour
test-iterate-loopAuto-detects test runner; iterates up to 10
test-iterate-loop --max-iter 5Lower iteration cap
test-iterate-loop --mock-hpcRun tests inside the mock-HPC Docker container (catches HPC-specific torch/CUDA bugs pre-submission)
test-iterate-loop --container <image>Custom Docker image
test-iterate-loop --no-fixRun tests once, root-cause failures, stop without applying fixes (diagnostic mode)

Architecture

Phase 1 (detect)    → which test runner? pytest, testthat, julia, custom?
Phase 2 (initial)   → run tests, capture full failure log
Phase 3 (loop)      → for each iteration: hypothesis → fix → re-run → log
Phase 4 (terminate) → all-pass | iter-cap | same-error-3x → final report

Phase 1: Detect test runner

Heuristics in priority order:

SignalRunner
pyproject.toml with [tool.pytest] or tests/ dir + *.pyuv run pytest --maxfail=1 -x
package.json with "scripts": {"test": ...}npm test
DESCRIPTION (R package) + tests/testthat/Rscript -e 'devtools::test()'
Project.toml (Julia)julia --project -e 'using Pkg; Pkg.test()'
Makefile with test: targetmake test
noxfile.py or tox.ininox / tox
Custom runner specified by user (--runner '<cmd>')use that

If multiple match, ask the user which.

Phase 2: Initial run

Run the test command. Capture:

  • Exit code
  • Full stdout + stderr to /tmp/test-iterate-<run-id>.log
  • Failure count, error types, first few failing test names

If exit 0: print "All green ✓" and exit. No iteration needed.

Show full SKILL.md (297 more words)Show less

Phase 3: Iteration loop

Per iteration:

  1. Hypothesis — read the failure log; identify the root cause. Use the Read tool on the test file + the implementation file to confirm. Write the hypothesis to the iteration log:

    Iteration 3 — 2026-05-10 14:23
    Hypothesis: pytest fails on test_inventory_split because the function
    returns a list when it should return a dict (matches old API).
  2. Fix — make the minimal edit to address the hypothesis. Document the file + line range in the iteration log.

  3. Re-run — same test command.

  4. Log result — pass / new failure / same failure:

    • All pass → exit loop, terminal state PASS.
    • New failure → fresh hypothesis next iteration.
    • Same failure (3rd time consecutively) → exit loop, terminal state STUCK.
  5. Memory-bug check — if this iteration's fix added a from <model> import, a .cuda() call, an os.environ[] set, or a version pin (torch==X.Y), flag in the log with [MEMORY-BUG-RISK]. These are the patterns that bit past [HPC cluster] runs.

Phase 4: Termination + report

Final report at log/test-iterate/<project>-YYYY-MM-DD-HHMM.md:

markdown
# Test-Iterate Loop — <project>

**Started:** YYYY-MM-DD HH:MM
**Ended:** YYYY-MM-DD HH:MM
**Terminal state:** PASS | STUCK (same error 3x) | EXHAUSTED (hit iter cap)

## Iterations

| # | Hypothesis | Fix | Result |
|---|---|---|---|
| 1 | ... | ... | new failure |
| 2 | ... | ... | new failure |
| 3 | ... | ... | PASS |

## Final test output
<paste of last test run>
```

Files modified

  • src/foo.py (lines 42-58) — fix iteration 1
  • tests/conftest.py (lines 12-15) — fix iteration 2
  • ...

Working tree

Clean (no commits made). User decides whether to commit, what to amend, or what to revert.

Memory-bug flags

  • Iteration 2: [MEMORY-BUG-RISK] — added torch.cuda.empty_cache() without a paired test

## HPC mock mode (`--mock-hpc`)

Use when iterating before submitting to [HPC cluster]. Runs inside a Docker container that mirrors [HPC cluster]'s environment:

- Default image: `user/hpc-mock:latest` (built from `scripts/hpc/Dockerfile.hpc-mock`)
- Container has same CUDA, torch, transformers, slurm-mock as [HPC cluster]
- Test command runs inside the container; failures bubble out to the iteration log

**Hard rule:** never run `--mock-hpc` against a project whose data lives outside the project directory — the Docker mount won't see it. Verify data paths first.

If `user/hpc-mock:latest` doesn't exist on the current machine, print the build instructions and exit.

## Standard forbid-list (when dispatching sub-agents)

If for very large iteration loops (>5 files modified concurrently) the orchestrator decides to dispatch sub-agents, each gets:

This sub-agent has a narrow scope. It does NOT inherit the orchestrator's authorisation for any other action.

  • Do NOT run git add, git commit, git push, or any other git write command.
  • Do NOT modify files outside the scope I've named below.
  • Do NOT run latexmk, pip install, uv add, or any package-management command.
  • Do NOT modify CI configs or pyproject.toml dependencies.
  • Read-only access except for the explicit edit scope.

Scope: <file paths> Task: <hypothesis + fix>

Report your changes as a diff. The orchestrator runs the tests.


## Cross-References

| Skill / Rule | Relationship |
|---|---|
| `subagent-write-guard.md` | Sub-agent dispatch (when needed) follows this rule |
| the `code-review` agent | Run AFTER test-iterate-loop terminates PASS — quality scorecard |
| `computational-experiments` | The skill that *writes* the tests this loop iterates on |
| `code-paper-auditor` agent | Code-paper consistency check — orthogonal concern |

## Anti-Patterns

- **Don't** loop without a same-error-repeat detector — an agent can spin on the same issue forever.
- **Don't** auto-commit even on PASS — leave the clean tree for the user to review and stage as they want.
- **Don't** treat warnings as failures — only test exit code 0 means PASS. Warnings get logged but don't trigger iterations.
- **Don't** apply fixes to test files — that's a smell that you're fitting tests to broken code. Fix the code.
- **Don't** skip the memory-bug flag — past HPC sessions had `torch.cuda` and `TRANSFORMERS_OFFLINE` bugs that would have been caught with this signal.

© flonat, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/test-iterate-loop of flonat/flonat-research.

Open the folder on GitHubat commit da27600

Compare with similar skills

Test Iterate Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Iterate Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Iterate Loop this skillflonat/flonat-research145—~2.2kAutomated safety check: PassMIT
Merge Dependabot PRsonyx-dot-app/onyx32k1 repos~2.2kAutomated safety check: PassMIT
Kedro Babysitkedro-org/kedro11k—~4kAutomated safety check: PassCustom licence
Adk Sample Creatorgoogle/adk-python22k—~1.3kAutomated safety check: PassApache-2.0
Mirage VFS Adapter Authoringstrukto-ai/mirage3.7k—~2.4kAutomated safety check: PassApache-2.0
Create Vibe Featuremistralai/mistral-vibe5.1k—~1.2kAutomated safety check: PassApache-2.0

Similar skills

  • Merge Dependabot PRs

    onyx-dot-app/onyx

    Triages and lands a batch of open Dependabot PRs in the Onyx repo, where main is gated exclusively by GitHub's merge queue: approves and enqueues green PRs, closes superseded duplicates, fixes…

    32k GitHub starsUsed in 1 repo~2.2k tokens
    DevelopmentAuto-check passed
  • Kedro Babysit

    kedro-org/kedro

    Run Kedro's local lint / format / type-check / tests on changed files (uses the project's pre-commit hooks, ruff, mypy, pytest, lint-imports, detect-secrets, Make targets — in the right venv), or…

    11k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Adk Sample Creator

    google/adk-python

    Official

    Creates a new sample agent in the ADK Python repository — the sample directory, its agent.py, and its README.md — following the conventions the existing samples already use.

    22k GitHub stars~1.3k tokensUpdated today
    DevelopmentAuto-check passed
  • Builds or extends a custom Mirage virtual filesystem adapter for an API, database, object store or app data, with a working mount configuration and filesystem tests.

    3.7k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • Create Vibe Feature

    mistralai/mistral-vibe

    Official

    Guides feature work in the Mistral Vibe Python CLI so each change lands in the right module and matches the project's architecture decision records.

    5.1k GitHub stars~1.2k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Adk Setup

    google/adk-python

    Official

    Sets up a local ADK Python development environment in a git clone of the open-source adk-python repository: a uv virtual environment, all dependency extras, pre-commit hooks, and a first unit-test…

    22k GitHub stars~993 tokensUpdated today
    DevelopmentAuto-check: notes

More from flonat/flonat-research

All 83 skills in this repo
  • Latex Posters

    flonat/flonat-research

    Create a large-format academic poster in LaTeX using beamerposter, tikzposter, or baposter.

    145 GitHub stars~1.5k tokensUpdated 8 days ago
    Auto-check: notes
  • Skill Creator

    flonat/flonat-research

    Create, revise, and evaluate reusable AI workflow skills, including trigger-quality tests.

    145 GitHub stars~4.4k tokensUpdated 8 days ago
    Auto-check passed
  • DOCX

    flonat/flonat-research

    Create, read, edit, or convert Microsoft Word documents while preserving professional document structure.

    145 GitHub stars~1.2k tokensUpdated 8 days ago
    Auto-check passed
  • PDF

    flonat/flonat-research

    Read, create, combine, split, rotate, OCR, watermark, secure, or extract content from PDF files.

    145 GitHub stars~488 tokensUpdated 8 days ago
    Auto-check passed
  • Init Project Orchestration

    flonat/flonat-research

    Create or migrate project-level agents, repeatable project workflows, and planning state from one client-neutral contract, then render repository-scoped adapters for both Claude Code and Codex.

    145 GitHub stars~1.6k tokensUpdated 8 days ago
    Auto-check passed
  • Pre Commit Audit

    flonat/flonat-research

    Deliver a fast pre-commit safety scan: file size, anonymity (author / affiliation strings in tex/bib), hardcoded secrets, and invisible-Unicode carriers.

    145 GitHub stars~2.8k tokensUpdated 8 days ago
    Auto-check: notes

Works with

Questions about Test Iterate Loop

What does Test Iterate Loop do?

Autonomously diagnose a codebase, apply minimal fixes, and rerun tests until they pass or a real blocker is reached. Test Iterate Loop is an agent skill from flonat/flonat-research. Autonomously diagnose a codebase, apply minimal fixes, and rerun tests until they pass or a real blocker is reached.

When should I use Test Iterate Loop?

Test Iterate Loop fits situations like: the user explicitly requests an iterative fix-until-green loop across Python.

How do I install Test Iterate Loop in Claude Code?

Run `npx skills add flonat/flonat-research --skill test-iterate-loop -a claude-code`. Or copy the skill folder (skills/test-iterate-loop in flonat/flonat-research) into .claude/skills/test-iterate-loop in your project. Claude Code loads it when a task matches its description.

How do I install Test Iterate Loop in Codex?

Run `npx skills add flonat/flonat-research --skill test-iterate-loop -a codex`. Or copy the skill folder (skills/test-iterate-loop in flonat/flonat-research) into .agents/skills/test-iterate-loop in your project. Codex loads it when a task matches its description.

Can I use Test Iterate Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add flonat/flonat-research --skill test-iterate-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-iterate-loop, .gemini/skills/test-iterate-loop, .github/skills/test-iterate-loop and .opencode/skills/test-iterate-loop in your project.

What does Test Iterate Loop need to run?

Going by SKILL.md and its folder, Test Iterate Loop needs the command-line tools its instructions call (git, uv, npm and make). Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Read, Write, Edit, Glob, Grep, Bash(uv*), Bash(pytest*), Bash(Rscript*), Bash(julia*), Bash(docker*), Bash(make*), Bash(git*), TaskCreate, TaskUpdate, AskUserQuestion.

Does Test Iterate Loop access the network?

SKILL.md contains no URLs. Its commands use git, uv and npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Test Iterate Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Iterate Loop use?

Test Iterate Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Iterate Loop use?

About 2.2k tokens (SKILL.md is roughly 8.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Iterate Loop?

Skills that share tags, products or a category with Test Iterate Loop: Merge Dependabot PRs (onyx-dot-app/onyx, 32k stars), Kedro Babysit (kedro-org/kedro, 11k stars), Adk Sample Creator (google/adk-python, 22k stars) and Mirage VFS Adapter Authoring (strukto-ai/mirage, 3.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Iterate Loop?

flonat (a GitHub user) maintains it in flonat/flonat-research, which has 145 GitHub stars. The repository holds 83 skills in this directory. The repository was last updated on September 29, 2026.

Source: flonat/flonat-research on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.