Agent skill

Plan Loop

by gaasher in gaasher/Agent-Loop-Skills

A skill your agent uses when the user has a coding or engineering prompt and wants it refined into a detailed, executable plan before any code is written — the planning stage of a prompt → plan →…

MITAuto-check passedDevelopment

Install Plan Loop

skills CLI
$ npx skills add gaasher/Agent-Loop-Skills --skill plan-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gaasher/Agent-Loop-Skills plan-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gaasher/Agent-Loop-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/loops/plan-loop .claude/skills/plan-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
plan-loop
GitHub stars
174
Token cost
~2.6k tokens
SKILL.md length
1,085 words
Files
6
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when the user has a coding or engineering prompt and wants it refined into a detailed, executable plan before any code is written — the planning stage of a prompt → plan →…

  • The user has a coding
  • SKILL.md covers When to use, Setup, The loop and Plan artifacts, plus 3 more sections
  • Runs Python scripts from its folder; calls python3
  • Engineering prompt and wants it refined into a detailed

What it does

Plan Loop is an agent skill from gaasher/Agent-Loop-Skills. Use when the user has a coding or engineering prompt and wants it refined into a detailed, executable plan before any code is written — the planning stage of a prompt → plan → execute → debug pipeline. It decomposes the prompt from first principles (objective, end state, environment, building blocks, tools, packages), breaks the work into PR-sized tasks each tied to a component with its files, tests, and dependencies, orders them topologically, splits each into atomic subtasks, then a separate principal-engineer…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files (for example `examples/run.example.yaml`, `roles/principal-engineer.md` and `schemas/critique.schema.json`). Compatibility notes: Requires Python 3.9+.

It sits in Development, covering Proposals and quotes, Task breakdown and Project scaffolding. The repository describes itself as: Loop until it's better — drop-in agentic loops (autoresearch, scientific writing, data analysis, code/SQL/prompt optimization, red-teaming) as open-standard Agent Skills… The licence is MIT.

When your agent uses it

  • The user has a coding
  • Engineering prompt and wants it refined into a detailed
  • Executable plan before any code is written — the planning stage of a prompt → plan → execute → debug pipeline

Example prompts

  • “/plan-loop”

Requirements

  • Python 3
  • Compatibility (from SKILL.md): Requires Python 3.9+.

What it can do on your machine

Read from SKILL.md and the folder at commit f1169e6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Requires Python 3.9+.

    From compatibility in the SKILL.md frontmatter.

Context cost

Plan Loop loads about 2.6k tokens when it runs. Until then it costs about 224 tokens; SKILL.md has 1,085 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~224
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from gaasher/Agent-Loop-Skills at commit f1169e6, republished under its MIT licence (© gaasher). 1,085 words, ~2,614 tokens.

Download SKILL.mdSave it as .claude/skills/plan-loop/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
plan-loop
description
Use when the user has a coding or engineering prompt and wants it refined into a detailed, executable plan before any code is written — the planning stage of a prompt → plan → execute → debug pipeline. It decomposes the prompt from first principles (objective, end state, environment, building blocks, tools, packages), breaks the work into PR-sized tasks each tied to a component with its files, tests, and dependencies, orders them topologically, splits each into atomic subtasks, then a separate principal-engineer agent critiques the plan for alignment, coverage, sizing, and executability; it revises until the critique passes, emitting a plan.md layout and a structured tasks.json that a junior engineer or a smaller model can execute correctly. Not for executing, scaffolding, or debugging the plan (those are downstream loops), and not for research proposals or experiment plans.
compatibility
Requires Python 3.9+.
metadata.version
0.1.0

Plan Loop

A planning loop: it turns a prompt into a plan detailed and correct enough to hand to a lower-tier model. The artifact is the plan — plan.md (the layout) + tasks.json (PR-sized tasks, each with files, tests, dependencies, and atomic subtasks). The feedback signal is two-part, like the repo's other evaluator loops: an objective gate (tools/validate_plan.py — schema shape, an acyclic dependency graph, a valid topological order, full component coverage) and a qualitative gate (a separate principal engineer agent that critiques alignment, decomposition, testability, and whether a junior could execute each task without guessing). You build the plan from first principles, validate it, critique it, and revise until the critique passes. This is the plan stage of a larger prompt → plan → execute → debug pipeline; it stops once the plan is ready to delegate.

When to use

Use to convert a feature/bug/refactor prompt into an executable plan grounded in a real repository — when the goal is a hand-off artifact a downstream executor (or a smaller model) can implement task-by-task. The plan is only as good as its weakest task for a literal-minded implementer, so the loop optimizes for executability, not prose.

Default: ground the plan in the <repo> you are given and let the principal-engineer critique drive the revisions. Escape hatch: if a key decision can't be resolved from the prompt or the repo, record it as an open_question for the human rather than guessing. Not for writing the code (a downstream execute loop), and not for research/experiment proposals (use research-proposal).

Setup

Resolve bindings interactively. If loop.run.yaml exists, load it, confirm the values in one line, and skip to the loop. Otherwise: on Claude Code (the AskUserQuestion tool is available) infer a likely value per binding and recommend it; on other hosts ask each as a quoted prompt. Then write loop.run.yaml (format: examples/run.example.yaml) and confirm before creating any other files.

bindingmeaningdefaulthow to infer
<prompt>the task to plan — a file path or inline text—the user's request
<repo>the project the plan targets; read for ground-truth env + conventions (never edited).the repo being worked on
<plan_file>the human-readable plan layout (markdown)<sandbox_root>/plan.md—
<tasks_file>the structured tasks, validated against schemas/plan.schema.json<sandbox_root>/tasks.json—
<pr_loc>target lines of code per task (PR-sized; a task may exceed it)100-400—
<sandbox_root>where plan, tasks, and the ledger live./sandbox—
<budget>max refine cycles5—

<skill_dir> is this skill's installed folder; substitute the real path when writing loop.run.yaml. The objective gate runs each cycle:

python3 <skill_dir>/tools/validate_plan.py --tasks <tasks_file>

It prints one JSON object {ok, errors, warnings, stats} — ok must be true (no errors) before a plan is considered ready.

The loop

Copy this checklist and tick items off.

Build the plan (iteration 0 — first principles):

  • Ground. Read <prompt> and inspect <repo> for ground truth: language, package manager, runtime/OS, the test command, and existing modules/conventions to reuse. Check what you can; never assume what you can read.
  • Decompose. Lay out the components from first principles — objective; end state / definition of done; non-goals; environment; building blocks; interfaces/contracts between blocks; tools; packages; data/assets; reuse (existing code + installed skills); open questions. Recursively split any component too big to reason about in one piece.
  • Tasks (PR-sized). Break the work into tasks, each ~<pr_loc> and equivalent to one PR. For each record: which components it serves, a description, building_blocks/tools/packages, the exact files it creates/modifies, tests that prove it, acceptance_criteria, and estimated_loc.
  • Order. Fill each task's depends_on, then compute a topological order (every task after its dependencies).
  • Subtasks. Split each task into atomic subtasks a junior can do with no further decisions — "create file X", "define function f(args) -> T", "wire f into Y".
  • Write + validate. Write <plan_file> and <tasks_file>, then run tools/validate_plan.py; fix every error and weigh every warning before the first critique. Log the baseline ledger row.

Refine (repeat until the plan passes or <budget>):

  • Critique. Spawn the principal engineer (spawn-or-degrade, roles/principal-engineer.md) with the prompt, the repo, plan.md, tasks.json, the latest validate_plan.py output, and the list of installed skills on this host. It returns a structured critique (schemas/critique.schema.json): verdict, score, issues (blocking/major/minor, each with a fix), coverage gaps, and suggested skills.
  • Revise. Apply the fix for every blocking and major issue — re-decompose, re-size, re-order, add tests, pin packages, add concrete signatures/paths/data shapes so a junior can't go wrong. Re-run tools/validate_plan.py. Append a ledger row.
  • Stop when the critique returns pass (no blocking or major issues) and the validator is clean — the plan is ready to delegate. Else loop, up to <budget>; if the score plateaus with only minor issues, stop and record them as open notes.

On stop, the deliverable is <plan_file> + <tasks_file> (plus any open_questions for the human), built to be executed task-by-task in order by a downstream execute loop or a lower-tier model.

Show full SKILL.md (306 more words)Show less

Plan artifacts

<tasks_file> (tasks.json) — the machine-executable plan; full contract in schemas/plan.schema.json. Compact shape:

json
{
  "objective": "...", "end_state": "...", "non_goals": ["..."],
  "environment": {"language": "python", "package_manager": "uv", "test_command": "pytest -q"},
  "components": [{"id": "c1", "name": "config", "description": "load + validate config"}],
  "open_questions": ["which auth provider?"],
  "tasks": [
    {"id": "T1", "title": "config loader", "serves": ["c1"], "description": "...",
     "building_blocks": ["dataclass Config"], "tools": ["pytest"], "packages": ["pyyaml"],
     "files": [{"path": "src/config.py", "action": "create", "what": "Config + load()"}],
     "subtasks": [{"id": "T1.1", "description": "define load(path) -> Config"}],
     "tests": [{"description": "load() parses a valid file", "kind": "unit"}],
     "acceptance_criteria": ["invalid config raises ConfigError"],
     "depends_on": [], "estimated_loc": 180, "suggested_skills": []}
  ],
  "order": ["T1"]
}

<plan_file> (plan.md) — the human-readable layout: objective, end state, non-goals, environment, the components, a task table (id · title · serves · depends_on · est. LOC) in order, risks/open questions, and a one-line "how to execute" pointer to tasks.json. It mirrors tasks.json; tasks.json is the source of truth the executor consumes.

Ledger

<sandbox_root>/ledger.tsv, tab-separated, never commas in free text. Header iter phase verdict score blocking major change:

iter	phase	verdict	score	blocking	major	change
0	build	-	-	-	-	first-principles decomposition: 6 components, 8 tasks, validator ok
1	critique	revise	72	1	2	PE: T3 bundles 2 PRs; T5 has no real test; nothing covers config loading
2	revise	-	-	-	-	split T3 -> T3a/T3b; added retry test to T5; added T9 config loader; reordered
3	critique	pass	90	0	0	PE: solid; one minor naming nit recorded as an open note

Report the final plan at the cycle the critique passed (or the best score reached at <budget>).

Constraints

  • Plan only — do not implement. Write no production code, scaffolding, or files in <repo>; the output is plan.md + tasks.json. Execution and debugging are downstream loops.
  • Ground every claim in the real repo/environment. Do not invent packages, files, APIs, or commands; verify against <repo>. A genuine unknown is an open_question for the human, not a guess — a plan that confidently states something false is worse than one that flags the gap.
  • Keep tasks PR-sized and self-contained. Each task links to ≥1 component and has tests and acceptance criteria; each subtask is atomic. Optimize for a literal implementer who will not fill gaps.
  • Keep the critic independent. The principal engineer reviews the plan it did not write (spawn-or-degrade gives real isolation on Claude Code); never let the author pass its own plan.
  • Edit only inside <sandbox_root> (plan, tasks, ledger); <repo> is read-only context. Run the loop to a passing critique or <budget> without pausing to ask whether to continue.

Roles

roles/principal-engineer.md — the adversarial plan critic. Spawn-or-degrade: a real isolated subagent on Claude Code (the Agent/Task tool), else adopt the role inline. It is read-only, judges against a fixed rubric, and returns JSON validated against schemas/critique.schema.json. Pass it the host's installed-skill list so it can recommend reuse; if that list is unavailable, it simply skips suggested_skills.

© gaasher, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files in loops/plan-loop of gaasher/Agent-Loop-Skills.

  • SKILL.md
  • examples/run.example.yaml
  • roles/principal-engineer.md
  • schemas/critique.schema.json
  • schemas/plan.schema.json
  • tools/validate_plan.py

Open the folder on GitHubat commit f1169e6

Compare with similar skills

Plan Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Plan Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Plan Loop this skillgaasher/Agent-Loop-Skills174—~2.6kAutomated safety check: PassMIT
Plugin BuilderAIDotNet/NextCoWork638—~2.4kAutomated safety check: PassApache-2.0
Python Pep Authorpproenca/dot-skills215—~2.1kAutomated safety check: PassMIT
Codex Agentmajiayu000/spellbook286—~1.9kAutomated safety check: PassMIT
Task Createjiangzhe/doradb121—~581Automated safety check: PassApache-2.0
Corvus Standalone Evaluatorcorvus-dotnet/Corvus.JsonSchema199—~988Automated safety check: PassApache-2.0

Similar skills

  • Plugin Builder

    AIDotNet/NextCoWork

    Author, package, and debug NextCoWork plugins — the single entry point.

    638 GitHub stars~2.4k tokensUpdated 5 days ago
    DevelopmentAuto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    215 GitHub stars~2.1k tokensUpdated 1 mo ago
    DevelopmentAuto-check passed
  • Codex Agent

    majiayu000/spellbook

    A skill your agent uses when you want a second-opinion review via Codex CLI, cross-verification after another agent implements changes, debugging help, or alternative implementation proposals.

    286 GitHub stars~1.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Task Create

    jiangzhe/doradb

    Design and create implementation-ready Doradb task documents through deep repository research, strict RFC complexity gating, two proposal and review rounds, explicit user approval, and isolated task…

    121 GitHub stars~581 tokensUpdated 2 days ago
    DevelopmentAuto-check passed
  • Corvus Standalone Evaluator

    corvus-dotnet/Corvus.JsonSchema

    Generate and use standalone schema evaluators (validation and annotation collection without full type generation), and understand the schema evaluation program that every generated type validates…

    199 GitHub stars~988 tokensUpdated today
    DevelopmentAuto-check passed
  • Contribution Architect

    majiayu000/spellbook

    A skill your agent uses when a contributor wants to move beyond simple bug fixes into architectural improvements, technical debt discovery, design proposals, or module ownership opportunities.

    286 GitHub stars~1.8k tokensUpdated today
    DevelopmentAuto-check: notes

More from gaasher/Agent-Loop-Skills

All 21 skills in this repo
  • Alpha Evolve

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants to evolve an ML model/program through population-based search rather than a single sequential refine loop — a generational evolution where parallel…

    174 GitHub starsUsed in 1 repo~3.4k tokens
    Auto-check passed
  • Karpathy

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants the LLM to do its own ML research: a fully-autonomous loop that hacks the training code, runs it, and keeps changes that lower a single scalar metric (e.g.

    174 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed
  • Tournament Autoresearch

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants an autonomous ML research loop that pressure-tests competing ideas before spending compute — several research subagents each propose one architecture…

    174 GitHub starsUsed in 1 repo~3k tokens
    Auto-check passed
  • Dueling Autoresearch

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user wants two approaches raced head-to-head on a single shared metric — e.g.

    174 GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check: warnings
  • Anomaly Investigation

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has a known, already-observed anomaly in their data — a metric spike or drop, an outlier, an unexpected number — and wants its root cause diagnosed, not guessed.

    174 GitHub stars~2.1k tokensUpdated 3 mo ago
    Auto-check passed
  • Blue Team

    gaasher/Agent-Loop-Skills

    A skill your agent uses when the user has concrete failing cases in code or a guardrail/classifier/filter/prompt/API they own — a red-team failure catalogue OR a CI/CD test-failure report (failing…

    174 GitHub stars~3.6k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Plan Loop

What does Plan Loop do?

A skill your agent uses when the user has a coding or engineering prompt and wants it refined into a detailed, executable plan before any code is written — the planning stage of a prompt → plan →…. Plan Loop is an agent skill from gaasher/Agent-Loop-Skills. Use when the user has a coding or engineering prompt and wants it refined into a detailed, executable plan before any code is written — the planning stage of a prompt → plan → execute → debug pipeline.

When should I use Plan Loop?

Plan Loop fits situations like: the user has a coding; engineering prompt and wants it refined into a detailed; executable plan before any code is written — the planning stage of a prompt → plan → execute → debug pipeline.

How do I install Plan Loop in Claude Code?

Run `npx skills add gaasher/Agent-Loop-Skills --skill plan-loop -a claude-code`. Or copy the skill folder (loops/plan-loop in gaasher/Agent-Loop-Skills) into .claude/skills/plan-loop in your project. Claude Code loads it when a task matches its description.

How do I install Plan Loop in Codex?

Run `npx skills add gaasher/Agent-Loop-Skills --skill plan-loop -a codex`. Or copy the skill folder (loops/plan-loop in gaasher/Agent-Loop-Skills) into .agents/skills/plan-loop in your project. Codex loads it when a task matches its description.

Can I use Plan Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gaasher/Agent-Loop-Skills --skill plan-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/plan-loop, .gemini/skills/plan-loop, .github/skills/plan-loop and .opencode/skills/plan-loop in your project.

What does Plan Loop need to run?

Going by SKILL.md and its folder, Plan Loop needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3. Compatibility (from SKILL.md): Requires Python 3.9+..

Does Plan Loop access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Plan Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Plan Loop use?

Plan Loop is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Plan Loop use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Plan Loop?

Skills that share tags, products or a category with Plan Loop: Plugin Builder (AIDotNet/NextCoWork, 638 stars), Python Pep Author (pproenca/dot-skills, 215 stars), Codex Agent (majiayu000/spellbook, 286 stars) and Task Create (jiangzhe/doradb, 121 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Plan Loop?

gaasher (a GitHub user) maintains it in gaasher/Agent-Loop-Skills, which has 174 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on June 30, 2026.

Source: gaasher/Agent-Loop-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.