Agent skill

Loop Plan Evaluator

by Ibrahim-3d in Ibrahim-3d/orchestrator-supaconductor

Evaluate-Loop Step 2: EVALUATE PLAN. An agent skill from Ibrahim-3d/orchestrator-supaconductor.

AGPL-3.0Auto-check passed

Install Loop Plan Evaluator

skills CLI
$ npx skills add Ibrahim-3d/orchestrator-supaconductor --skill loop-plan-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Ibrahim-3d/orchestrator-supaconductor loop-plan-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Ibrahim-3d/orchestrator-supaconductor.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/loop-plan-evaluator .claude/skills/loop-plan-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
loop-plan-evaluator
GitHub stars
380
Token cost
~2.6k tokens
SKILL.md length
474 words
Files
1
Skills in repo
27
Repo updated
First seen
Licence
AGPL-3.0

At a glance

Evaluate-Loop Step 2: EVALUATE PLAN. An agent skill from Ibrahim-3d/orchestrator-supaconductor.

  • Works in 5 steps: Track's plan.md — the plan to evaluate… → Track's spec.md — requirements to check… → conductor/tracks.md — completed tracks… → …
  • SKILL.md covers Inputs Required, Evaluation Passes, Verdict and Metadata Checkpoint Updates, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Loop Plan Evaluator is an agent skill from Ibrahim-3d/orchestrator-supaconductor. Evaluate-Loop Step 2: EVALUATE PLAN. Use this agent to verify an execution plan before any code is written. Checks scope alignment, overlap with completed work, DAG validity, dependency correctness, task clarity, and invokes Board of Directors for major tracks. Outputs PASS/FAIL verdict. Triggered by: 'evaluate plan', 'review plan', 'check plan before executing'. Always runs after loop-planner and before loop-executor.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Multi-agent orchestration system for Claude Code with parallel execution, automated quality gates, Board of Directors, and bundled Superpowers skills. The licence is AGPL-3.0.

Example prompts

  • “evaluate plan”
  • “review plan”
  • “check plan before executing”
  • “/loop-plan-evaluator”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Track's plan.md — the plan to evaluate (including DAG)
  2. Track's spec.md — requirements to check against
  3. conductor/tracks.md — completed tracks (overlap check)
  4. Track's metadata.json — track type and priority
  5. Codebase state — what files/components already exist

What it can do on your machine

Read from SKILL.md and the folder at commit 76c9b10. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown, json, python and typescript).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Loop Plan Evaluator loads about 2.6k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 474 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Ibrahim-3d/orchestrator-supaconductor at commit 76c9b10, republished under its AGPL-3.0 licence (© Ibrahim-3d). 474 words, ~2,630 tokens.

Download SKILL.mdSave it as .claude/skills/loop-plan-evaluator/SKILL.md (or your agent's skills folder).
name
loop-plan-evaluator
description
Evaluate-Loop Step 2: EVALUATE PLAN. Use this agent to verify an execution plan before any code is written. Checks scope alignment, overlap with completed work, DAG validity, dependency correctness, task clarity, and invokes Board of Directors for major tracks. Outputs PASS/FAIL verdict. Triggered by: 'evaluate plan', 'review plan', 'check plan before executing'. Always runs after loop-planner and before loop-executor.

Loop Plan Evaluator Agent — Step 2: EVALUATE PLAN

Pre-execution quality gate. Verifies the plan is correct and scoped before any implementation begins. This prevents the exact problem that caused the PLAN-005 design system rebuild — an agent executing work that was already done.

For major tracks (architecture, features with 5+ tasks, integrations, infrastructure), this step also invokes the Board of Directors for multi-perspective expert review.

Inputs Required

  1. Track's plan.md — the plan to evaluate (including DAG)
  2. Track's spec.md — requirements to check against
  3. conductor/tracks.md — completed tracks (overlap check)
  4. Track's metadata.json — track type and priority
  5. Codebase state — what files/components already exist

Evaluation Passes

Pass 1: Scope Alignment

Check every task against spec.md:

For Each TaskCheck
Is it in spec?Task must trace to a specific spec requirement
Is it needed?Would removing this task leave a spec requirement unmet?
Is it scoped?Does the task do only what spec asks, not more?

Output:

markdown
### Scope Alignment: PASS ✅ / FAIL ❌
- Tasks in spec: [X]/[Y]
- Tasks NOT in spec (scope creep): [list]
- Spec requirements NOT covered: [list]
Pass 2: Overlap Detection

Cross-reference with tracks.md and the codebase:

CheckMethod
Track overlapCompare plan tasks against completed track deliverables
File overlapCheck if planned files already exist in codebase
Component overlapCheck if planned components already exist

Output:

markdown
### Overlap Detection: PASS ✅ / FAIL ❌
- Overlapping tasks: [list with which track already did them]
- Files that already exist: [list]
- Recommendation: [SKIP/MODIFY/PROCEED for each overlap]
Pass 3: Dependency Check

Verify task ordering and prerequisites:

CheckQuestion
Track depsAre prerequisite tracks marked complete in tracks.md?
Task orderingDo later tasks depend on earlier tasks being done first?
External depsAre required packages/APIs available?

Output:

markdown
### Dependencies: PASS ✅ / FAIL ❌
- Missing track dependencies: [list]
- Misordered tasks: [list]
- Missing external dependencies: [list]
Pass 4: Task Quality

Evaluate each task for clarity and completeness:

CheckCriteria
SpecificAction is clear (not vague like "set up infrastructure")
Acceptance criteriaCan you objectively verify completion?
File targetsExpected file paths are listed?
Session-sizedCan be completed in one sitting?

Output:

markdown
### Task Quality: PASS ✅ / FAIL ❌
- Vague tasks: [list with suggestions to clarify]
- Missing acceptance criteria: [list]
- Oversized tasks (should split): [list]
Show full SKILL.md (195 more words)Show less
Pass 5: DAG Validation

Verify the dependency graph is valid for parallel execution:

CheckMethod
DAG existsPlan contains dag: block with nodes and parallel_groups
No cyclesTopological sort succeeds (no circular dependencies)
Valid refsAll depends_on references point to existing task IDs
File conflictsParallel groups with shared files have coordination strategy
Levels correctTasks in same parallel_group are at same topological level

Cycle Detection Algorithm:

python
def detect_cycles(dag):
    """Returns True if cycle exists, False otherwise."""
    visited = set()
    rec_stack = set()

    def dfs(node_id):
        visited.add(node_id)
        rec_stack.add(node_id)

        node = next((n for n in dag['nodes'] if n['id'] == node_id), None)
        for dep in node.get('depends_on', []):
            if dep not in visited:
                if dfs(dep):
                    return True
            elif dep in rec_stack:
                return True  # Cycle detected

        rec_stack.remove(node_id)
        return False

    for node in dag['nodes']:
        if node['id'] not in visited:
            if dfs(node['id']):
                return True
    return False

Output:

markdown
### DAG Validation: PASS ✅ / FAIL ❌
- DAG present: yes/no
- Nodes: [count]
- Parallel groups: [count]
- Cycle detected: yes/no (list cycle path if yes)
- Invalid references: [list of broken depends_on]
- Conflict issues: [list parallel groups with unhandled file conflicts]
Pass 6: Board of Directors Review (Major Tracks Only)

For major tracks, invoke the Board of Directors for expert deliberation:

When to invoke Board:

  • Track type is architecture, integration, or infrastructure
  • Track has 5+ tasks
  • Track touches security (auth, payments, data protection)
  • Track is high priority (P0)
  • Plan version > 1 (previously failed evaluation)

Board Invocation:

typescript
// If track qualifies for board review
if (isMajorTrack(metadata)) {
  // Initialize board session via message bus
  const boardResult = await invokeBoardMeeting(
    proposal: plan.md content,
    context: { spec, metadata, dag }
  );

  // Store board session in metadata
  metadata.loop_state.board_sessions.push({
    session_id: boardResult.session_id,
    checkpoint: "EVALUATE_PLAN",
    verdict: boardResult.verdict,
    vote_summary: boardResult.votes,
    conditions: boardResult.conditions,
    timestamp: new Date().toISOString()
  });

  // Board verdict affects overall evaluation
  if (boardResult.verdict === "REJECTED") {
    return FAIL with board conditions;
  }
}

Output:

markdown
### Board Review: PASS ✅ / FAIL ❌ / SKIPPED ⏭️
- Board invoked: yes/no (reason if no)
- Directors voted: [CA, CPO, CSO, COO, CXO]
- Verdict: APPROVED / APPROVED_WITH_REVIEW / REJECTED
- Vote breakdown: [X] APPROVE / [Y] REJECT
- Conditions from board:
  1. [Condition 1] (from [Director])
  2. [Condition 2] (from [Director])

Verdict

markdown
## Plan Evaluation Report

**Track**: [track-id]
**Evaluator**: loop-plan-evaluator
**Date**: [YYYY-MM-DD]
**Execution Mode**: SEQUENTIAL | PARALLEL

### Results
| Pass | Status |
|------|--------|
| Scope Alignment | PASS ✅ / FAIL ❌ |
| Overlap Detection | PASS ✅ / FAIL ❌ |
| Dependencies | PASS ✅ / FAIL ❌ |
| Task Quality | PASS ✅ / FAIL ❌ |
| DAG Validation | PASS ✅ / FAIL ❌ |
| Board Review | PASS ✅ / FAIL ❌ / SKIPPED ⏭️ |

### Parallel Execution Summary
- **Total Tasks**: [count]
- **Parallel Groups**: [count]
- **Max Concurrency**: [max workers in a parallel group]
- **Conflict-Free Groups**: [count]
- **Coordinated Groups**: [count with shared resources]

### Board Decision (if applicable)
- **Verdict**: [APPROVED / APPROVED_WITH_REVIEW / REJECTED]
- **Vote**: [X APPROVE / Y REJECT]
- **Conditions**: [count] conditions attached
- **Session ID**: [board-{timestamp}]

### Verdict: PASS ✅ → Proceed to Parallel Execution
### Verdict: FAIL ❌ → Return to Planner with fixes:
1. [Fix 1]
2. [Fix 2]

### Board Conditions (carry forward):
1. [Condition from board that must be verified in EVALUATE_EXECUTION]

Metadata Checkpoint Updates

The plan evaluator MUST update the track's metadata.json at key points:

On Start
json
{
  "loop_state": {
    "current_step": "EVALUATE_PLAN",
    "step_status": "IN_PROGRESS",
    "step_started_at": "[ISO timestamp]",
    "checkpoints": {
      "EVALUATE_PLAN": {
        "status": "IN_PROGRESS",
        "started_at": "[ISO timestamp]",
        "agent": "loop-plan-evaluator"
      }
    }
  }
}
On PASS
json
{
  "loop_state": {
    "current_step": "PARALLEL_EXECUTE",
    "step_status": "NOT_STARTED",
    "execution_mode": "PARALLEL",
    "checkpoints": {
      "EVALUATE_PLAN": {
        "status": "PASSED",
        "completed_at": "[ISO timestamp]",
        "verdict": "PASS",
        "checks": {
          "scope_alignment": true,
          "overlap_detection": true,
          "dependencies": true,
          "task_quality": true,
          "dag_validation": true,
          "board_review": true
        },
        "cto_review": {
          "status": "PASSED",
          "reviewed_at": "[timestamp if run]"
        },
        "dag_summary": {
          "total_tasks": 8,
          "parallel_groups": 3,
          "max_concurrency": 4,
          "conflict_free_groups": 2,
          "coordinated_groups": 1
        }
      },
      "PARALLEL_EXECUTE": {
        "status": "NOT_STARTED"
      }
    },
    "board_sessions": [
      {
        "session_id": "board-20260201-123456",
        "checkpoint": "EVALUATE_PLAN",
        "verdict": "APPROVED",
        "vote_summary": {
          "CA": "APPROVE",
          "CPO": "APPROVE",
          "CSO": "APPROVE",
          "COO": "APPROVE",
          "CXO": "APPROVE"
        },
        "conditions": [
          "Add caching layer (CA)",
          "Security audit before launch (CSO)"
        ],
        "timestamp": "[ISO timestamp]"
      }
    ]
  }
}
On FAIL
json
{
  "loop_state": {
    "current_step": "PLAN",
    "step_status": "NOT_STARTED",
    "checkpoints": {
      "EVALUATE_PLAN": {
        "status": "FAILED",
        "completed_at": "[ISO timestamp]",
        "verdict": "FAIL",
        "checks": {
          "scope_alignment": true,
          "overlap_detection": false,
          "dependencies": true,
          "task_quality": false
        },
        "failure_reasons": [
          "Overlap with existing track: component already built",
          "Task 3 is too vague"
        ]
      },
      "PLAN": {
        "status": "NOT_STARTED",
        "plan_version": 2
      }
    }
  }
}
Update Protocol
  1. read_file current metadata.json
  2. Update loop_state.checkpoints.EVALUATE_PLAN with verdict and checks
  3. If PASS: Advance current_step to EXECUTE
  4. If FAIL: Reset current_step to PLAN, increment plan_version
  5. write_file back to metadata.json

Handoff

  • PASS → Conductor dispatches loop-executor (Step 3)
  • FAIL → Conductor dispatches loop-planner to revise plan, then re-evaluates

© Ibrahim-3d, AGPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/loop-plan-evaluator of Ibrahim-3d/orchestrator-supaconductor.

Open the folder on GitHubat commit 76c9b10

Compare with similar skills

Loop Plan Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Loop Plan Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Loop Plan Evaluator this skillIbrahim-3d/orchestrator-supaconductor380—~2.6kAutomated safety check: PassAGPL-3.0
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence
Agent Evaluation Reportingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

    40k GitHub stars~6.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from Ibrahim-3d/orchestrator-supaconductor

All 27 skills in this repo
  • Cto Advisor

    Ibrahim-3d/orchestrator-supaconductor

    Technical leadership guidance for engineering teams, architecture decisions, and technology strategy.

    380 GitHub starsUsed in 4 repos~2.4k tokens
    Auto-check passed
  • Context Driven Development

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when working with Conductor's context-driven development methodology, managing project context artifacts, or understanding the relationship between product.md, tech-stack.md…

    380 GitHub starsUsed in 9 repos~2.9k tokens
    Auto-check passed
  • Agent Factory

    Ibrahim-3d/orchestrator-supaconductor

    Creates specialized worker agents dynamically from templates.

    380 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Board Of Directors

    Ibrahim-3d/orchestrator-supaconductor

    Simulate a 5-member expert board deliberation for major decisions.

    380 GitHub stars~1.9k tokensUpdated 10 days ago
    Auto-check passed
  • Business Docs Sync

    Ibrahim-3d/orchestrator-supaconductor

    A skill your agent uses when completing a track that changes pricing, AI models, product features, or asset pipelines — syncs business context documents across all tiers.

    380 GitHub stars~2.1k tokensUpdated 10 days ago
    Auto-check passed
  • Context Loader

    Ibrahim-3d/orchestrator-supaconductor

    Load project context efficiently for Conductor workflows. An agent skill from Ibrahim-3d/orchestrator-supaconductor.

    380 GitHub stars~830 tokensUpdated 10 days ago
    Auto-check passed

Questions about Loop Plan Evaluator

What does Loop Plan Evaluator do?

Evaluate-Loop Step 2: EVALUATE PLAN. An agent skill from Ibrahim-3d/orchestrator-supaconductor. Loop Plan Evaluator is an agent skill from Ibrahim-3d/orchestrator-supaconductor. Evaluate-Loop Step 2: EVALUATE PLAN.

How do I install Loop Plan Evaluator in Claude Code?

Run `npx skills add Ibrahim-3d/orchestrator-supaconductor --skill loop-plan-evaluator -a claude-code`. Or copy the skill folder (skills/loop-plan-evaluator in Ibrahim-3d/orchestrator-supaconductor) into .claude/skills/loop-plan-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Loop Plan Evaluator in Codex?

Run `npx skills add Ibrahim-3d/orchestrator-supaconductor --skill loop-plan-evaluator -a codex`. Or copy the skill folder (skills/loop-plan-evaluator in Ibrahim-3d/orchestrator-supaconductor) into .agents/skills/loop-plan-evaluator in your project. Codex loads it when a task matches its description.

Can I use Loop Plan Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Ibrahim-3d/orchestrator-supaconductor --skill loop-plan-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/loop-plan-evaluator, .gemini/skills/loop-plan-evaluator, .github/skills/loop-plan-evaluator and .opencode/skills/loop-plan-evaluator in your project.

What does Loop Plan Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Loop Plan Evaluator is instructions for the agent only. Our summary lists: Python 3.

Does Loop Plan Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Loop Plan Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Loop Plan Evaluator use?

Loop Plan Evaluator is published under the AGPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Loop Plan Evaluator use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Loop Plan Evaluator?

Skills that share tags, products or a category with Loop Plan Evaluator: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Evaluators (Arize-ai/phoenix, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Loop Plan Evaluator?

Ibrahim-3d (a GitHub user) maintains it in Ibrahim-3d/orchestrator-supaconductor, which has 380 GitHub stars. The repository holds 27 skills in this directory. The repository was last updated on September 27, 2026.

Source: Ibrahim-3d/orchestrator-supaconductor on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.