Agent skill

Loop Design Check

by affaan-m in affaan-m/ECC

Design a goal-oriented agent loop or review one for failure modes: spinning, Goodhart-gaming the verifier, or running a wrong answer to completion.

MITAuto-check passedAgent Workflows

Install Loop Design Check

skills CLI
$ npx skills add affaan-m/ECC --skill loop-design-check -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC loop-design-check --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/loop-design-check .claude/skills/loop-design-check && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
loop-design-check
GitHub stars
276k
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,710 words
Files
1
Skills in repo
673
Repo updated
First seen
Licence
MIT

At a glance

Design a goal-oriented agent loop or review one for failure modes: spinning, Goodhart-gaming the verifier, or running a wrong answer to completion.

  • Works in 6 steps: Subtract first: should you even build… → Define a machine-decidable goal (the… → Pick the loop type → …
  • Checking an agent loop
  • SKILL.md covers When to use / not, Red-line premise: two levels…, Action 1 — Write a loop (5… and Action 2 — Review a loop…, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Loop Design Check is an agent skill from affaan-m/ECC. Design a goal-oriented agent loop or review one for failure modes: spinning, Goodhart-gaming the verifier, or running a wrong answer to completion. Covers machine-decidable goals, loop types, plan/build/judge skeletons, and runaway prevention; mechanism wiring lives in autonomous-loops. Use when designing, writing, or checking an agent loop. 中文触发:写 loop、设计 loop、做一个 loop、检查 loop 对不对、loop 体检、loop 会不会跑飞、可判定目标、五个崩法、plan build judge。

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Autonomous loops. The repository describes itself as: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. The licence is MIT.

When your agent uses it

  • Checking an agent loop
  • Tasks that involve Autonomous loops

Example prompts

  • “/loop-design-check”

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Subtract first: should you even build it? (4-condition gate, any miss = veto)
  2. Define a machine-decidable goal (the hard part — the loop lives or dies here)
  3. Pick the loop type
  4. Pick a skeleton
  5. Add damping (against oscillation / runaway)
  6. Land in three stages (don't go fully automatic on day one)

What it can do on your machine

Read from SKILL.md and the folder at commit ef648e0. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Loop Design Check loads about 2.9k tokens when it runs. Until then it costs about 113 tokens; SKILL.md has 1,710 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~113
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit ef648e0, republished under its MIT licence (© affaan-m). 1,710 words, ~2,928 tokens.

Download SKILL.mdSave it as .claude/skills/loop-design-check/SKILL.md (or your agent's skills folder).
name
loop-design-check
description
Design a goal-oriented agent loop or review one for failure modes: spinning, Goodhart-gaming the verifier, or running a wrong answer to completion. Covers machine-decidable goals, loop types, plan/build/judge skeletons, and runaway prevention; mechanism wiring lives in autonomous-loops. Use when designing, writing, or checking an agent loop. 中文触发:写 loop、设计 loop、做一个 loop、检查 loop 对不对、loop 体检、loop 会不会跑飞、可判定目标、五个崩法、plan build judge。
metadata.origin
ECC

Loop Design + Review

Premise. An LLM is a feed-forward system: prompt in → tokens out, with no built-in "steer toward the goal" across turns. To make it behave like a goal-oriented system, you wrap a feedback loop around it. This skill helps you write that loop correctly and review it so it won't run away.

When to use / not

Use it when:

  • You want to hand a repeating task to an agent that runs over and over (write→test, test→fix, fix→verify…).
  • You already have a loop and worry it spins, cheats, or runs a wrong answer to completion.

Don't use it for:

  • A one-off task → just do it; don't wrap a loop around it.
  • A plain timer / poll → use /loop; no design needed.
  • How to wire the loop architecture (pipelines → DAGs, long-run recovery) → that's the mechanism layer; see autonomous-loops / continuous-agent-loop. This skill only covers "is the goal right, and will it run away" — it does not re-explain mechanism.

Red-line premise: two levels of feedback

LevelWho owns itWhat it does
Execution (low)machine / agentMeasures "how far from the literal goal" and grinds it to zero. The machine is strong here.
Judgment (high)humanDecides "is this goal itself right, should it change, should it stop." The machine can't step outside its own loop to question the goal.

A thermostat can feed back "how far from 26°C," but when you have a fever and want 28°C it can't judge whether 26 is the right target — it just grinds toward 26. "What to set today" is always the human's call. Handing judgment / sign-off / the last switch to the machine = removing the high-level feedback = it sprints, fast and hard, toward a goal no one questioned → wrong output.


Action 1 — Write a loop (5 steps)

Step 0 · Subtract first: should you even build it? (4-condition gate, any miss = veto)

① the task repeats weekly or more ② verification can be automated ③ the token budget can take it ④ the agent has tools that actually run and see the result

Miss any one → don't build a loop; do it by hand or another way.

What stops most people isn't "can I write a loop," it's "does my repo deserve one." A repo that deserves a loop has a reconciliation baseline (golden sample / upstream total) + tests + a lint guard. A repo that doesn't deserve a loop will only have its errors amplified by one.

Step 1 · Define a machine-decidable goal (the hard part — the loop lives or dies here)

The whole loop rides on the comparator's "is it done yet?" The comparator can only work if your exit condition can be judged yes/no by a machine.

  • Bad: Vague ("make it good," "write it sharper") → the comparator can't judge → either it never passes (stuck retrying) or it guesses (passes/blocks at random).
  • Good: Decidable ("all 96 unit tests green AND a change-list is produced," "module-02 fields filled, pytest passes, business logic untouched") → one check settles it; the loop converges cleanly.

Five-point goal framework:

  1. Done-criterion is machine-verifiable.
  2. Boundary conditions defined alongside the done-criterion ("what it must NOT do") — anti-Goodhart; missing boundaries = a license to cheat.
  3. Has a failure fallback — retry cap N + escalate to a human when exceeded.
  4. Goal is layered.
  5. Prefer reconciliation over assertion for the done-criterion — anchor to external fact (golden sample / upstream total / financial tie-out / platform back-office numbers) before your own assertions. "All tests pass" can be gamed (loosen asserts, fake mocks, swallow exceptions); "diff vs the reference < 0.01" can't.

Self-check: read the goal to someone who doesn't know the domain — can they run one command and tell whether it's done? If not, it isn't decidable enough. Go back.

Step 2 · Pick the loop type
Your taskLoop type (cybernetic)How it stops
Has a clear "done" test (write to done / a batch of images processed)servo (/goal-style closed-loop)stops on reaching the goal
No endpoint, must keep maintaining a state (inventory alert / scheduled health check)regulator (/loop-style thermostat)never stops; acts only on change (dead-band suppresses noise)
Periodic sampling, stop on a condition (watch a PR until CI is green)regulator with an exitstops when the exit condition holds
Must "ensure something happens on time"wrap the above in /schedulecron fires it

Rule of thumb: clear "done" test → servo; must keep maintaining, no endpoint → regulator; must "happen on time" → wrap a regulator in schedule.

Step 3 · Pick a skeleton

Maintenance type (tend something that exists) → document-driven dispatch. The loop isn't "run a fixed check on a timer," it's "read a doc on a timer, and dispatch only when the doc changed." The doc is the task queue + state machine + human interface. Three disciplines: ① the problem column is human-write-only, the result column is loop-write-only, state advances one-way and never rolls back; ② the exit code is final (if the script says exit 1, the script wins); ③ state advances only as far as "awaiting verification" — the "done" cell is flipped by a human only. The loop is the worker, not the acceptance officer.

Greenfield type (build from scratch) → plan / build / judge, three roles.

RoleDoesKey
Planbreak the goal into a spec + decidable acceptance conditionsacceptance must be script-judgeable
Buildwrite to the specmust not change the acceptance conditions
Judgerun acceptance independently; pass → stop, fail → return with the failure reason to Buildindependent + deterministic

Three iron rules (all bet on the judge): ① the judge must be independent — not the same agent as Build (grading your own homework always inflates); ② deterministic rules — pytest / reconciliation diff / type check / diff, never "looks right"; ③ Build may not edit the acceptance conditions to pass. Three failed retries → escalate to a human.

Step 4 · Add damping (against oscillation / runaway)

Retry cap, hard stop, human flips the last switch = damping. Negative feedback with no damping oscillates (the Ralph-Wiggum loop: spinning in place, burning tokens).

Step 5 · Land in three stages (don't go fully automatic on day one)

① Run it once by hand (forces you to state exactly "how the judge decides") → ② harden into a skill / Claude Code sub-agents (a main Claude loops, dispatching plan/build/judge) → ③ hang it on cron for full automation.


Show full SKILL.md (684 more words)Show less

Action 2 — Review a loop (checklist = five failure modes)

Run the loop past each row. Hitting any one = this loop will misfire; send it back. These five are negative experience (gotchas) — worth more than positive rules.

#Failure mode (how it breaks)Review question (a hit = red)Antibody
1Goal is a correct platitude → spins, burns moneyCan the exit condition be machine-judged yes/no? Or is it "manage it well / make it good"?Replace with a decidable result condition (Action 1·Step 1)
2"Verification" written as "check if it looks ok" → agent confidently says fine and stopsIs the judge the defendant itself? Does verification rest on "looks right" or deterministic rules?Reconcile + exit code rules + independent judge
3(worst) Only gates on "all tests pass" → agent deletes the testsIs there a boundary ("what it must NOT do")? Or only a done-criterion?Done-criterion + boundary together (the Goodhart antibody)
4Counts on the agent asking mid-run → it won't; it runs the wrong answer to the endIs there any "clarify only at runtime" point?Front-load every clarification; settle it once before launch
5Bloated CLAUDE.md + stale memory → the faster it loops, the more it errsAre the docs/memory it depends on fresh? Who maintains them?Layered memory + periodic lint

Plus three red lines (violate any = not allowed to go automatic):

  • Keep judgment with the human. Acceptance / the "done" cell is flipped by a human; the loop is not the acceptance officer.
  • Responsibility doesn't transfer. Anything whose failure you can't afford (merge the wrong PR / publish the wrong thing / misallocate money) → don't hand over the authority automatically.
  • Counter-intuitive warning. The more "self-improving / rewrites-its-own-rules" a loop is, the stricter the human review it needs (to see what it rewrote the rules into) — not looser. The machine is too fast to intercept after the fact, so the human's judgment must sit before the action (a hard gate), not as a post-hoc patch.

Worked example — reviewing a "nightly green-keeper" loop

You want a loop that runs every night and fixes whatever tests are failing.

  • Naive goal: "make all tests pass." → Step-1 self-check fails: this is the bait for failure mode #3.
  • Decidable goal (fixed): "all tests green AND no test file deleted or weakened AND coverage not lowered AND a change-list produced." Boundary now defined alongside the done-criterion.
  • Type: servo with a retry cap of 3 (Step 2 + Step 4).
  • Skeleton: plan/build/judge — the judge is CI run independently, never the fixing agent (Step 3).

Now run the review checklist, and it catches what the naive version would have missed:

  • #3 hit → the naive "all tests pass" lets the agent delete a failing test to "win." Fixed by the boundary "no test file deleted/weakened."
  • #2 hit → if the fixing agent also judged its own fix, it would pass itself. Fixed by "judge = independent CI, deterministic."
  • #4 hit → if a fix is ambiguous, the agent won't stop to ask at 2 a.m.; it'll commit a guess. Fixed by front-loading: ambiguous fixes are left for the human, not guessed.
  • Red line → the loop opens a PR but does not auto-merge; the human flips the last switch (responsibility doesn't transfer).

The naive loop and the reviewed loop differ by four lines of constraint — and that's the difference between "wakes you to a deleted test suite" and "wakes you to a clean PR."


One-line close

The hard part of writing a loop isn't "can I write a loop," it's defining a goal a machine can reconcile — decidable, bounded, reconciliation-based. The controller must be deterministic and external; keep judgment and the standard with the human; the system tends toward entropy, so maintain it. A loop only rewards someone who has already thought it through. Count on it to think for you, and it will happily think wrong, with you, at scale.


Lineage: Wiener's two-level feedback (The Human Use of Human Beings, 1950) for the judgment/execution split and red lines; the plan/build/judge pattern from Anatoli's Loops explained and Addy's Loop Engineering. Mechanism layer (how to wire the loop architecture): see autonomous-loops / continuous-agent-loop. This skill does not re-implement mechanism; it covers goal definition and runaway prevention only.

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/loop-design-check of affaan-m/ECC.

Open the folder on GitHubat commit ef648e0

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Loop Design Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Loop Design Check compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Loop Design Check this skillaffaan-m/ECC276k1 repos~2.9kAutomated safety check: PassMIT
Show Me Your Work Decision Logcursor/plugins10k8 repos~1.6kAutomated safety check: PassNone
Autoresearch Iteration Loopuditgoenka/autoresearch6.5k1 repos~2kAutomated safety check: PassMIT
Install Loop Engineeringcobusgreyling/loop-engineering11k1 repos~648Automated safety check: PassMIT
LoopyForward-Future/loopy3.2k—~3.9kAutomated safety check: PassMIT
AI Performance Improvement Plantanweai/pua20k2 repos~6.9kAutomated safety check: PassMIT

Similar skills

  • Official

    Keeps a TSV decision log for long or unattended agent runs, one row per decision with what, why, evidence and result, so a reviewer can check the work later.

    10k GitHub starsUsed in 8 repos~1.6k tokens
    Agent WorkflowsAuto-check passed
  • Autoresearch Iteration Loop

    uditgoenka/autoresearch

    Runs an autonomous modify, verify, keep-or-discard loop against any metric, with subcommands for planning, debugging, fixing, security audits, shipping and more.

    6.5k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Install Loop Engineering

    cobusgreyling/loop-engineering

    Installs Loop Engineering into a project through the single @cobusgreyling/loop CLI, scaffolding a report-only loop and a readiness score.

    11k GitHub starsUsed in 1 repo~648 tokens
    Agent WorkflowsAuto-check passed
  • Loopy

    Forward-Future/loopy

    Discover, find, compare, audit, repair, adapt, craft, run, debrief, save, and prepare repeatable AI-agent loops for publication.

    3.2k GitHub stars~3.9k tokensUpdated 29 days ago
    Agent WorkflowsAuto-check passed
  • Pushes an agent to exhaust every option, investigate before asking and take initiative beyond the literal request, instead of giving up or waiting passively.

    20k GitHub starsUsed in 2 repos~6.9k tokens
    Agent WorkflowsAuto-check passed
  • LoopX Self Repair

    loopx-project/loopx

    Diagnoses surprising LoopX behavior, such as stale recommendations or tiny progress, assigns it to the responsible layer and repairs it at the lowest durable level.

    6.2k GitHub stars~2.2k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from affaan-m/ECC

All 673 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    276k GitHub starsUsed in 5 repos~1.9k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    276k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    276k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    276k GitHub stars~2.9k tokensUpdated 4 days ago
    Auto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Auto-check passed
  • Instinct-based learning system that observes sessions via hooks, creates atomic instincts with confidence scoring, and evolves them into skills/commands/agents.

    276k GitHub stars~3.5k tokensUpdated 4 days ago
    Auto-check passed

Categories

Questions about Loop Design Check

What does Loop Design Check do?

Design a goal-oriented agent loop or review one for failure modes: spinning, Goodhart-gaming the verifier, or running a wrong answer to completion. Loop Design Check is an agent skill from affaan-m/ECC. Design a goal-oriented agent loop or review one for failure modes: spinning, Goodhart-gaming the verifier, or running a wrong answer to completion.

When should I use Loop Design Check?

Loop Design Check fits situations like: checking an agent loop; tasks that involve Autonomous loops.

How do I install Loop Design Check in Claude Code?

Run `npx skills add affaan-m/ECC --skill loop-design-check -a claude-code`. Or copy the skill folder (skills/loop-design-check in affaan-m/ECC) into .claude/skills/loop-design-check in your project. Claude Code loads it when a task matches its description.

How do I install Loop Design Check in Codex?

Run `npx skills add affaan-m/ECC --skill loop-design-check -a codex`. Or copy the skill folder (skills/loop-design-check in affaan-m/ECC) into .agents/skills/loop-design-check in your project. Codex loads it when a task matches its description.

Can I use Loop Design Check in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill loop-design-check -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/loop-design-check, .gemini/skills/loop-design-check, .github/skills/loop-design-check and .opencode/skills/loop-design-check in your project.

What does Loop Design Check need to run?

SKILL.md names no scripts, command-line tools or credentials: Loop Design Check is instructions for the agent only.

Does Loop Design Check access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Loop Design Check safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Loop Design Check use?

Loop Design Check is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Loop Design Check use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Loop Design Check?

Skills that share tags, products or a category with Loop Design Check: Show Me Your Work Decision Log (cursor/plugins, 10k stars), Autoresearch Iteration Loop (uditgoenka/autoresearch, 6.5k stars), Install Loop Engineering (cobusgreyling/loop-engineering, 11k stars) and Loopy (Forward-Future/loopy, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Loop Design Check?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 275,546 GitHub stars. The repository holds 673 skills in this directory. The repository was last updated on October 5, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.