Alignment
Ovid/paad
A skill your agent uses when verifying that requirements/specs/PRDs and their implementation plans match — before starting work, after a spec or plan update, or when suspecting coverage gaps, scope…
Orchestrate parallel judge agent execution, aggregate CaseScore results, write plan-judges.json, code-judges.json, prd-judges.json, or feature-judges.json, and validate output.
$ npx skills add closedloop-ai/claude-plugins --skill run-judges -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install closedloop-ai/claude-plugins run-judges --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/judges/skills/run-judges .claude/skills/run-judges && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run-judges" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judges into .claude/skills/run-judges/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-judges", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judgesType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add closedloop-ai/claude-plugins --skill run-judges -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install closedloop-ai/claude-plugins run-judges --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/judges/skills/run-judges .agents/skills/run-judges && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run-judges" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judges into .agents/skills/run-judges/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-judges", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add closedloop-ai/claude-plugins --skill run-judges -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install closedloop-ai/claude-plugins run-judges --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/judges/skills/run-judges .cursor/skills/run-judges && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run-judges" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judges into .cursor/skills/run-judges/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-judges", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/closedloop-ai/claude-plugins.git --path plugins/judges/skills/run-judges--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add closedloop-ai/claude-plugins --skill run-judges -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install closedloop-ai/claude-plugins run-judges --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/judges/skills/run-judges .gemini/skills/run-judges && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run-judges" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judges into .gemini/skills/run-judges/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-judges", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install closedloop-ai/claude-plugins run-judgesInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add closedloop-ai/claude-plugins --skill run-judges -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/judges/skills/run-judges .github/skills/run-judges && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run-judges" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judges into .github/skills/run-judges/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-judges", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add closedloop-ai/claude-plugins --skill run-judges -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install closedloop-ai/claude-plugins run-judges --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/closedloop-ai/claude-plugins.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/judges/skills/run-judges .opencode/skills/run-judges && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run-judges" agent skill from https://github.com/closedloop-ai/claude-plugins/tree/main/plugins/judges/skills/run-judges into .opencode/skills/run-judges/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run-judges", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
run-judgesOrchestrate parallel judge agent execution, aggregate CaseScore results, write plan-judges.json, code-judges.json, prd-judges.json, or feature-judges.json, and validate output.
Run Judges is an agent skill from closedloop-ai/claude-plugins. Orchestrate parallel judge agent execution, aggregate CaseScore results, write plan-judges.json, code-judges.json, prd-judges.json, or feature-judges.json, and validate output. Supports evaluating implementation plans (16 judges, 4 batches), code artifacts (11 judges, 3 batches), PRD artifacts (5 judges, 2 batches), or Feature artifacts (3 judges, 1 batch) via --artifact-type parameter.
Its SKILL.md is about 15k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `references/judge-input-contract.md`, `scripts/ensure_agents_snapshot.sh` and `scripts/judge_input_mapping.py`).
It sits in Product & Project Management, covering PRD writing and Planning. The repository describes itself as: Open-source Claude Code plugins for multi-agent software delivery. Plan-first SDLC workflow, code review, LLM quality judges, and self-learning — grounded in your codebase… The licence is Apache-2.0.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4600742. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
uvjqbashcurlshpython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
astral.shFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Run Judges loads about 15k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 100 tokens; SKILL.md has 5,307 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
curl -LsSf https://astral.sh/uv/install.sh | shAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from closedloop-ai/claude-plugins at commit 4600742, republished under its Apache-2.0 licence (© closedloop-ai). 5,307 words, ~15,018 tokens.
.claude/skills/run-judges/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.Execute specialized judge agents in parallel to evaluate implementation plan quality (16 judges, 4 batches), code quality (11 judges, 3 batches), PRD quality (5 judges, 2 batches), or Feature quality (3 judges, 1 batch). All batches respect the Task tool's 4-concurrent-agent limit. Aggregates results into $CLOSEDLOOP_WORKDIR/plan-judges.json (plan), $CLOSEDLOOP_WORKDIR/code-judges.json (code), $CLOSEDLOOP_WORKDIR/prd-judges.json (prd), or $CLOSEDLOOP_WORKDIR/feature-judges.json (feature) with validated output format.
--workdir: Path to the working directory containing judge artifacts (optional)
--workdir argument → $CLOSEDLOOP_WORKDIR environment variable → .closedloop-ai/judges (default, relative to current working directory)plan-judges.json, code-judges.json, prd-judges.json, judge-input.json, perf.jsonl, etc.) are written to this resolved directory--artifact-type: Artifact category to evaluate (plan | code | prd | feature), default: plan
judge-input.json)The judge input contract is maintained in:
skills/run-judges/references/judge-input-contract.md (resolve to an absolute path at runtime via Glob)
This keeps orchestration flow readable while preserving a single source of truth for contract fields and semantics.
run-judges is the producer chokepoint for judge-input.json. After mode-specific context preparation and before launching any judge agent, invoke the deterministic mapper:
uv run "${CLAUDE_PLUGIN_ROOT}/skills/run-judges/scripts/judge_input_mapping.py" \
--workdir "$CLOSEDLOOP_WORKDIR" \
--artifact-type "$ARTIFACT_TYPE" \
--schema "${CLAUDE_PLUGIN_ROOT}/schemas/judge-input.schema.json"The mapper builds from the runtime workdir contract: primary artifacts under <runDir>, supporting context under <runDir>/.closedloop-ai/context, and attachments under <runDir>/.closedloop-ai/work/attachments. It validates the generated envelope against schemas/judge-input.schema.json before judge launch. If mapping fails, emit a clear warning and use the documented one-run legacy fallback paths (prd.md, plan.md, or existing compatibility artifacts) only for that run.
You are orchestrating quality evaluation for a ClosedLoop artifact (implementation plan, code, or PRD). Your responsibilities:
For plan artifacts (default):
judge-input.json with plan task/context mapping$CLOSEDLOOP_WORKDIR/plan-judges.jsonFor code artifacts (--artifact-type code):
judge-input.json with code task/context mapping$CLOSEDLOOP_WORKDIR/code-judges.jsonFor PRD artifacts (--artifact-type prd):
$CLOSEDLOOP_WORKDIR/prd.md exists (graceful exit if missing)judge-input.json by invoking scripts/judge_input_mapping.py$CLOSEDLOOP_WORKDIR/prd-judges.jsonFor Feature artifacts (--artifact-type feature):
$CLOSEDLOOP_WORKDIR/feature.md exists, or $CLOSEDLOOP_WORKDIR/prd.md exists for legacy Feature inputs (graceful exit code 0 if both are missing)judge-input.json by invoking scripts/judge_input_mapping.py$CLOSEDLOOP_WORKDIR/feature-judges.jsonFeature mode judge selection rationale:
prd-auditor is excluded because it assumes US-###/AC-#.# numbering and multi-story traceability, which Feature artifacts do not followprd-scope-judge is excluded because it assumes In/Out-of-Scope sections that are not present in Feature artifactsFeature mode preamble: Feature mode uses the dedicated feature_preamble.md so judges receive a Feature-shaped contract (evaluation_type=feature, lightweight structure, no PRD-only sections). Do NOT substitute prd_preamble.md — it would frame the input as a full PRD and contradict the envelope's evaluation_type.
Success criteria:
The run-judges skill supports per-artifact-type threshold customization via JSON configuration files. This allows you to adjust evaluation strictness for different artifact types (e.g., applying a lower threshold for test-judge when evaluating code vs plan).
Threshold overrides are defined in a JSON file with the following structure:
{
"overrides": {
"artifact_type:judge_name": <threshold_float>
}
}Where:
"artifact_type:judge_name" (e.g., "code:test-judge", "plan:technical-accuracy-judge")[0.0, 1.0]Example configuration:
{
"overrides": {
"code:test-judge": 0.75,
"plan:technical-accuracy-judge": 0.85
}
}The skill checks the following locations in order, using the first valid configuration found:
Run-specific overrides (highest precedence):
$CLOSEDLOOP_WORKDIR/.closedloop-ai/settings/threshold-overrides.jsonRepo-level defaults (fallback):
<project-root>/.closedloop-ai/settings/threshold-overrides.jsonHardcoded defaults (graceful degradation):
The following default overrides apply when evaluating code artifacts:
| Judge | Code Threshold | Plan Threshold | Rationale |
|---|---|---|---|
test-judge | 0.75 | 0.8 | Code may have tests written separately from implementation, lower threshold accounts for incremental test development |
All other judges use the same threshold (typically 0.8) across artifact types.
When loading threshold overrides, the skill applies the following validation rules:
Schema Validation:
"overrides" keyartifact_type:judge_name[0.0, 1.0]plan, code, prd) and judge namesError Behavior:
Warning: Invalid threshold-overrides.json, skipping overrides: {error}Error recovery ensures the skill always completes judge execution, even if threshold configuration is incorrect.
When executing judges:
final_status (pass/fail) based on metric scoresWhen artifact type is code:
You MUST emit a pipeline_step event to $CLOSEDLOOP_WORKDIR/perf.jsonl at the end of each phase below. This keeps perf telemetry in the canonical schema and adds nested metadata for judge/sub-agent work.
Context: CLOSEDLOOP_WORKDIR, CLOSEDLOOP_RUN_ID, and CLOSEDLOOP_ITERATION are set by the run-loop. CLOSEDLOOP_PARENT_STEP and CLOSEDLOOP_PARENT_STEP_NAME are set as env vars on the claude invocation by run-loop; they are inherited by all Bash tool calls — no sourcing needed.
Use sub_step as numeric phase order and optional sub_step_name to capture the judge/sub-agent name when applicable (for batch-level phases where many judges run, use the batch label).
Sub-step numbering:
| Artifact | sub_step | sub_step_name |
|---|---|---|
| plan | 0 | context_manager |
| plan | 1–4 | batch_1 … batch_4 |
| plan | 5 | aggregate |
| plan | 6 | validate |
| code | 0 | context_manager |
| code | 1–3 | batch_1 … batch_3 |
| code | 4 | aggregate |
| code | 5 | validate |
| prd | 0 | context_prep (skipped — prd mode does not use context-manager-for-judges) |
| prd | 1–2 | batch_1, batch_2 |
| prd | 3 | aggregate |
| prd | 4 | validate |
| feature | 0 | context_prep (skipped — feature mode does not use context-manager-for-judges) |
| feature | 1 | batch_1 |
| feature | 2 | aggregate |
| feature | 3 | validate |
Start of phase (run Bash once at the beginning of each phase): Set the two sub-step variables at the top for the current phase, then run the block. It writes start time to a temp file so the end-of-phase Bash can compute duration. CLOSEDLOOP_PARENT_STEP and CLOSEDLOOP_PARENT_STEP_NAME are already in the environment (set by run-loop on the claude invocation).
# Set these two values for the current phase:
SUB_STEP_NUM=0
SUB_STEP_LABEL="context_manager" # context_manager | batch_1 … | aggregate | validate
mkdir -p "$CLOSEDLOOP_WORKDIR/.closedloop-ai"
{
echo "SUB_STEP=${SUB_STEP_NUM}"
echo "SUB_STEP_NAME=${SUB_STEP_LABEL}"
echo "PARENT_STEP=${CLOSEDLOOP_PARENT_STEP:-0}"
echo "PARENT_STEP_NAME=${CLOSEDLOOP_PARENT_STEP_NAME:-unknown}"
echo "STARTED_AT=$(date -u +%Y-%m-%dT%H:%M:%SZ)"
echo "START_EPOCH=$(date +%s)"
} > "$CLOSEDLOOP_WORKDIR/.closedloop-ai/perf-substep-start.env"End of phase (run Bash once at the end of each phase, after the phase work is done): Read start time, compute duration, append one line to perf.jsonl, then remove the temp file.
source "$CLOSEDLOOP_WORKDIR/.closedloop-ai/perf-substep-start.env"
END_EPOCH=$(date +%s)
ENDED_AT=$(date -u +%Y-%m-%dT%H:%M:%SZ)
DURATION=$((END_EPOCH - START_EPOCH))
jq -n -c \
--arg event "pipeline_step" \
--arg run_id "${CLOSEDLOOP_RUN_ID:-unknown}" \
--argjson iteration "${CLOSEDLOOP_ITERATION:-0}" \
--argjson step "$PARENT_STEP" \
--arg step_name "$PARENT_STEP_NAME" \
--argjson sub_step "$SUB_STEP" \
--arg sub_step_name "$SUB_STEP_NAME" \
--arg started_at "$STARTED_AT" \
--arg ended_at "$ENDED_AT" \
--argjson duration_s "$DURATION" \
--argjson exit_code 0 \
--argjson skipped false \
'{event:$event,run_id:$run_id,iteration:$iteration,step:$step,step_name:$step_name,sub_step:$sub_step,sub_step_name:$sub_step_name,started_at:$started_at,ended_at:$ended_at,duration_s:$duration_s,exit_code:$exit_code,skipped:$skipped}' >> "$CLOSEDLOOP_WORKDIR/perf.jsonl"
rm -f "$CLOSEDLOOP_WORKDIR/.closedloop-ai/perf-substep-start.env"Order of operations per phase: Run the "start of phase" Bash first (set SUB_STEP_NUM and SUB_STEP_LABEL at the top, then run the block), then perform the phase work, then run the "end of phase" Bash.
Before any other step, resolve the working directory and export it as CLOSEDLOOP_WORKDIR:
# Resolve working directory (precedence: --workdir arg > env var > default)
if [ -n "$ARG_WORKDIR" ]; then
WORKDIR="$ARG_WORKDIR"
elif [ -n "$CLOSEDLOOP_WORKDIR" ]; then
WORKDIR="$CLOSEDLOOP_WORKDIR"
else
WORKDIR="$(pwd)/.closedloop-ai/judges"
fi
mkdir -p "$WORKDIR"
export CLOSEDLOOP_WORKDIR="$WORKDIR"Where $ARG_WORKDIR is the value passed via --workdir in the invocation prompt. All subsequent references to $CLOSEDLOOP_WORKDIR use this resolved value.
Before any judge execution, ensure a snapshot of judge agent definitions exists in $CLOSEDLOOP_WORKDIR/agents-snapshot/. This preserves the exact agent versions used for each evaluation run.
Action: Run the snapshot script via Bash:
bash "${CLAUDE_PLUGIN_ROOT}/skills/run-judges/scripts/ensure_agents_snapshot.sh" "$CLOSEDLOOP_WORKDIR"The script is idempotent — it skips if manifest.json already exists.
Error handling: If the script fails or is not found, log a warning and continue — snapshot failure must not block judge execution.
Before any judge execution, validate the agent registry to ensure all judge agents required for the current artifact type are resolvable. This prevents launching batches only to discover agents are missing mid-run.
Action: Run validate_agent_registry.py via Bash:
uv run "${CLAUDE_PLUGIN_ROOT}/tools/python/validate_agent_registry.py" \
--artifact-type "$ARTIFACT_TYPE" \
--workdir "$CLOSEDLOOP_WORKDIR"Exit behavior:
0 — all required agents are registered; proceed with judge executionOn failure:
Before any prerequisite checks or judge launches:
Glob with:**/skills/run-judges/references/judge-input-contract.mdjudge-input-contract.md file in full.$CLOSEDLOOP_WORKDIR/judge-input.json.Performance: At the start of this phase run the "start of phase" Bash with SUB_STEP_NUM=0 and SUB_STEP_LABEL=context_manager for both plan and code modes. For prd and feature modes, emit sub_step=0 with SUB_STEP_LABEL=context_prep and skipped=true immediately (no context manager runs). At the end of the phase run the "end of phase" Bash.
Before starting, verify required inputs exist:
For plan artifacts (default):
# Validate input files exist
if [ ! -f "$CLOSEDLOOP_WORKDIR/prd.md" ]; then
echo "WARNING: $CLOSEDLOOP_WORKDIR/prd.md not found. Skipping judges."
exit 0 # Graceful skip - do not fail workflow
fi
if [ ! -f "$CLOSEDLOOP_WORKDIR/plan.json" ]; then
echo "WARNING: $CLOSEDLOOP_WORKDIR/plan.json not found. Skipping judges."
exit 0
fiInvestigation log resolution (plan mode):
After validating prd.md and plan.json, resolve supporting context for plan judges:
Use existing file first
$CLOSEDLOOP_WORKDIR/investigation-log.md exists, use it as-is.Check @code:pre-explorer availability before invoking
@code:pre-explorer in the active Claude/plugin environment.Task() call targeting @code:pre-explorer.If available, invoke pre-explorer
@code:pre-explorer with WORKDIR=$CLOSEDLOOP_WORKDIR to generate missing pre-exploration artifacts.$CLOSEDLOOP_WORKDIR/investigation-log.md after completion.If unavailable or invocation failed, run internal fallback
investigation-log.md with a lightweight local-only investigation.prd.md and extract top entities/actions as search seeds.Glob/Grep against the local repository for likely implementation files.Files Discovered / Key Findings.Requirements Mapping.## Search Strategy## Files Discovered## Key Findings## Requirements Mapping## UncertaintiesNever block plan context preparation on investigation context
Prepare plan-context.json via context-manager-for-judges
@judges:context-manager-for-judges with artifact_type=plan.$CLOSEDLOOP_WORKDIR/plan-context.json exists.plan.json + prd.md.Plan-mode source-of-truth policy
plan-context.json is primary and required.plan.json + prd.md may be used for this run only.Build plan-mode judge-input.json
scripts/judge_input_mapping.py with --artifact-type plan.evaluation_type, task, primary_artifact, supporting_artifacts, source_of_truth, fallback_mode, and metadata from the runtime workdir contract.prd.md as supporting evidence when available.prd.md plus plan.md or the existing compatibility artifact is readable.For code artifacts (--artifact-type code):
# Resolve investigation context for code judges (best effort)
if [ ! -f "$CLOSEDLOOP_WORKDIR/investigation-log.md" ]; then
echo "INFO: investigation-log.md missing. Attempting best-effort generation via @code:pre-explorer..."
# Launch @code:pre-explorer with WORKDIR=$CLOSEDLOOP_WORKDIR
# If unavailable/fails, continue with warning (non-blocking for code judges)
fi
# Launch context-manager-for-judges agent to prepare compressed context
# This agent reads code artifacts (git diff, changed-files.json, etc.)
# and produces .closedloop-ai/context/code-context.json with token-budgeted compression
# investigation-log.md is optional secondary context for code judging
if [ ! -f "$CLOSEDLOOP_WORKDIR/investigation-log.md" ]; then
echo "WARNING: investigation-log.md unavailable. Continuing code judges with canonical code context only."
fi
# Verify canonical code context exists after context manager completes. The root
# code-context.json path is fallback-only for old runs.
if [ ! -f "$CLOSEDLOOP_WORKDIR/.closedloop-ai/context/code-context.json" ] && [ ! -f "$CLOSEDLOOP_WORKDIR/code-context.json" ]; then
echo "ERROR: Context preparation failed - .closedloop-ai/context/code-context.json not found"
# Abort with error CaseScore for all judges
# Generate error report with final_status=3, justification="Context preparation failed"
exit 1
fi
# Build and validate code-mode judge-input.json with scripts/judge_input_mapping.py.
# The mapper prefers .closedloop-ai/context/code-context.json as primary and
# preserves root code-context.json as a one-run legacy fallback when needed.For PRD artifacts (--artifact-type prd):
PRD mode does NOT use context-manager-for-judges. Context preparation is lightweight: verify the PRD document exists, then build judge-input.json directly from it.
# PRD mode context prep: check prd.md exists
if [ ! -f "$CLOSEDLOOP_WORKDIR/prd.md" ]; then
echo "WARNING: $CLOSEDLOOP_WORKDIR/prd.md not found. Skipping PRD judges."
exit 0 # Graceful exit — do not fail parent workflow
fi
# Build and validate prd-mode judge-input.json with scripts/judge_input_mapping.py.
# The mapper sets primary_artifact to primary_prd and includes mapped context,
# prompt, repo metadata, prior summaries, and attachments in source_of_truth order.PRD context prep notes:
prd.md results in a WARNING and graceful exit (code 0), not an errorjudge-input.json is built by scripts/judge_input_mapping.py and validated against schemas/judge-input.schema.jsonFor Feature artifacts (--artifact-type feature):
Feature mode does NOT use context-manager-for-judges. Context preparation is lightweight: verify feature.md exists, or prd.md exists for legacy Feature inputs, then build judge-input.json from the mapper.
# Feature mode context prep: check feature.md or legacy prd.md exists
if [ ! -f "$CLOSEDLOOP_WORKDIR/feature.md" ] && [ ! -f "$CLOSEDLOOP_WORKDIR/prd.md" ]; then
echo "WARNING: neither $CLOSEDLOOP_WORKDIR/feature.md nor legacy $CLOSEDLOOP_WORKDIR/prd.md found. Skipping Feature judges."
exit 0 # Graceful exit — do not fail parent workflow
fi
# Build and validate feature-mode judge-input.json with scripts/judge_input_mapping.py.
# The mapper prefers feature.md and marks fallback_mode.active=true when it must
# use the legacy prd.md Feature path.Feature context prep notes:
feature.md and legacy prd.md results in a WARNING and graceful exit (code 0), not an errorjudge-input.json is built by scripts/judge_input_mapping.py with evaluation_type="feature"feature_preamble.md for all 3 feature judgesIf required files are missing:
The run-judges skill supports three artifact types with different judge configurations:
plan-judges.json{RUN_ID}-plan-judges--category plan (16 judges expected)code-judges.json{RUN_ID}-code-judges--category code (11 judges expected)Code Judge Batches:
Batch 1: Core Principles (4 judges)
judges:dry-judgejudges:ssot-judgejudges:kiss-judgejudges:code-organization-judgeBatch 2: Best Practices + SOLID Principles (4 judges)
judges:custom-best-practices-judgejudges:readability-judgejudges:solid-isp-dip-judgejudges:solid-liskov-substitution-judgeBatch 3: Technical Quality + Testing (3 judges)
judges:solid-open-closed-judgejudges:technical-accuracy-judgejudges:test-judgeprd-judges.json{RUN_ID}-prd-judges--category prd (5 judges expected)$CLOSEDLOOP_WORKDIR/judge-input.json produced by scripts/judge_input_mapping.py, with primary_prd normally pointing to $CLOSEDLOOP_WORKDIR/prd.mdfeature-judges.json{RUN_ID}-feature-judges--category feature (3 judges expected)$CLOSEDLOOP_WORKDIR/judge-input.json produced by scripts/judge_input_mapping.py, with primary_feature normally pointing to feature.md and legacy fallback to prd.mdfeature_preamble.md (Feature-shaped contract; do NOT substitute prd_preamble.md)Feature Mode Execution:
Batch 1: Feature Quality (sub_step=1)
judges:feature-completeness-judge — evaluates Feature request completeness and clarityjudges:prd-testability-judge — evaluates requirement testabilityjudges:prd-dependency-judge — evaluates dependency clarity and completenessPRD Mode Execution:
Batch 1: Structure & Completeness (sub_step=1)
judges:feature-completeness-judge — evaluates Feature request completeness and clarityjudges:prd-auditor — structural completeness audit of the PRDjudges:prd-scope-judge — evaluates scope definition and boundary clarityBatch 2: Quality Gates (sub_step=2)
judges:prd-dependency-judge — evaluates dependency clarity and completenessjudges:prd-testability-judge — evaluates requirement testabilityPerformance: For each batch/phase, run "start of phase" Bash before launching the batch and "end of phase" Bash after the batch completes. Plan: batch_1=sub_step 1, batch_2=sub_step 2, batch_3=sub_step 3, batch_4=sub_step 4. Code: batch_1=sub_step 1, batch_2=sub_step 2, batch_3=sub_step 3. PRD: batch_1=sub_step 1, batch_2=sub_step 2. Feature: batch_1=sub_step 1.
Constraint: The Task tool supports maximum 4 concurrent agents per batch.
Action: Launch judges in sequential batches based on artifact type.
<judge_batches>
Batch 1: Core Principles (DRY/SSOT/KISS + Organization)
| Agent Type | Evaluates |
|---|---|
judges:dry-judge | Don't Repeat Yourself violations |
judges:ssot-judge | Single Source of Truth violations |
judges:kiss-judge | Keep It Simple violations |
judges:code-organization-judge | File and folder structure organization |
Batch 2: Best Practices + Response Quality
| Agent Type | Evaluates |
|---|---|
judges:custom-best-practices-judge | Adherence to custom best practices documents |
judges:goal-alignment-judge | Alignment with stated health goals |
judges:readability-judge | Plan readability, clarity, structure, template adherence |
judges:verbosity-judge | Verbosity calibration to problem complexity |
Batch 3: SOLID Principles
| Agent Type | Evaluates |
|---|---|
judges:solid-isp-dip-judge | Interface Segregation & Dependency Inversion Principles |
judges:solid-liskov-substitution-judge | Liskov Substitution Principle adherence |
judges:solid-open-closed-judge | Open/Closed Principle adherence |
judges:technical-accuracy-judge | Technical accuracy (API usage, algorithms) |
Batch 4: Plan Grounding + Testing
| Agent Type | Evaluates |
|---|---|
judges:test-judge | Test coverage, assertions, structure, best practices |
judges:brownfield-accuracy-judge | Reuse vs reimplementation, integration-point accuracy, scope accuracy against investigation findings |
judges:codebase-grounding-judge | File-path/module-reference accuracy and existing-code awareness grounded in investigation findings |
judges:convention-adherence-judge | Alignment with established naming, structural, and tooling conventions in the codebase |
Batch 1: Structure & Completeness (sub_step=1)
| Agent Type | Evaluates |
|---|---|
judges:feature-completeness-judge | Feature request completeness and clarity |
judges:prd-auditor | Structural completeness, section coverage, clarity |
judges:prd-scope-judge | Scope definition and boundary clarity |
Batch 2: Quality Gates (sub_step=2)
| Agent Type | Evaluates |
|---|---|
judges:prd-dependency-judge | Dependency clarity and completeness |
judges:prd-testability-judge | Requirement testability and measurability |
Batch 1: Feature Quality (sub_step=1)
| Agent Type | Evaluates |
|---|---|
judges:feature-completeness-judge | Feature request completeness and clarity |
judges:prd-testability-judge | Requirement testability and measurability |
judges:prd-dependency-judge | Dependency clarity and completeness |
Excluded judges (feature mode):
judges:prd-auditor — excluded because it assumes US-###/AC-#.# numbering and multi-story traceability that Feature artifacts do not followjudges:prd-scope-judge — excluded because it assumes In/Out-of-Scope sections that are not present in Feature artifacts</judge_batches>
<prompt_template>
Before invoking each judge, prepend the common and artifact-specific preambles:
Locate preamble files:
skills/artifact-type-tailored-context/preambles/common_input_preamble.mdskills/artifact-type-tailored-context/preambles/{artifact_type}_preamble.md**/artifact-type-tailored-context/preambles/*.mdRead preamble content:
common_input_preamble.md{artifact_type}_preamble.mdConcatenate:
common_input_preamble + "\n\n---\n\n" + artifact_preamble + "\n\n---\n\n" + judge_promptcommon_input_preamble.md is the only runtime source of judge input-loading contract text; judge-specific agent files should not duplicate that contract.Pass to judge: Use concatenated prompt as judge's full prompt
If either preamble file is missing:
final_status=3, justification="Preamble file not found: {path}"NOTE — Feature Mode: When
--artifact-type featureis used, resolve{artifact_type}_preamble.mdasfeature_preamble.md(notprd_preamble.md). The Feature preamble frames the input as a Feature artifact (evaluation_type=feature, lightweight structure, no PRD-only sections such as US-###/AC-#.# numbering or In/Out-of-Scope) and aligns with the envelope built by feature mode. Substitutingprd_preamble.mdwould inject contradictory contract instructions and may cause judges to error or evaluate against PRD-only expectations.
For plan artifacts:
WORKDIR=$CLOSEDLOOP_WORKDIR. Read $CLOSEDLOOP_WORKDIR/judge-input.json first.
Evaluate according to `task` and `source_of_truth` ordering.
Treat the envelope's `primary_artifact` as authoritative.
If `fallback_mode.active=true`, use fallback artifacts specified in the envelope.For code artifacts:
WORKDIR=$CLOSEDLOOP_WORKDIR. Read $CLOSEDLOOP_WORKDIR/judge-input.json first.
Evaluate according to `task` and `source_of_truth` ordering.
Treat the envelope's `primary_artifact` as authoritative.
Apply your {judge_name} criteria to assess code quality.For PRD artifacts:
WORKDIR=$CLOSEDLOOP_WORKDIR. Read $CLOSEDLOOP_WORKDIR/judge-input.json first.
Evaluate according to `task` and `source_of_truth` ordering.
Treat the envelope's `primary_artifact` as the authoritative PRD document and load supporting descriptors as source-of-truth evidence.
Apply your {judge_name} criteria to assess PRD quality.For Feature artifacts:
WORKDIR=$CLOSEDLOOP_WORKDIR. Read $CLOSEDLOOP_WORKDIR/judge-input.json first.
Evaluate according to `task` and `source_of_truth` ordering.
Treat the envelope's `primary_artifact` as the authoritative Feature document and load supporting descriptors as source-of-truth evidence.
Apply your {judge_name} criteria to assess Feature quality.</prompt_template>
<expected_output> Each judge returns a CaseScore JSON object:
{
"type": "case_score",
"case_id": "dry-judge",
"final_status": 1,
"metrics": [
{
"metric_name": "dry_score",
"threshold": 0.8,
"score": 0.85,
"justification": "Plan follows DRY principles..."
}
]
}Status Code Semantics:
| Code | Meaning | When to Use |
|---|---|---|
1 | Pass | Score meets or exceeds threshold |
2 | Fail | Score below threshold |
3 | Error | Judge execution failed |
</expected_output>
<error_handling>
CRITICAL REQUIREMENT: If a judge Task call fails, you MUST construct an error CaseScore.
Error CaseScore Template:
{
"type": "case_score",
"case_id": "{judge-name}",
"final_status": 3,
"error_reason": "Brief human-readable description of what failed",
"metrics": [
{
"metric_name": "{metric}_score",
"threshold": 0.8,
"score": 0.0,
"justification": "Judge execution failed: {error message}"
}
]
}error_reason field guidance:
When to set it: Set error_reason whenever final_status=3. Common cases include:
{artifact_type}_preamble.md missing)What to put in it: A brief, human-readable string describing the specific failure. Examples:
"Task tool error: agent not found""Parse error: response was not valid JSON""Timeout: judge did not complete within 5 minutes""Preamble file not found: plan_preamble.md"Effect on aggregation: CaseScores with final_status=3 are excluded by compute_average_excluding_errors, which then averages MetricStatistics.score across every metric of every remaining (non-errored) CaseScore. error_reason is informational and does not control exclusion (see field docstring at validate_judge_report.py:46). Errored judges do not drag down the aggregate score for judges that did execute successfully.
Aggregation rules when errors are present:
final_status=3, compute_average_excluding_errors returns the average of MetricStatistics.score across only the non-errored judges (return type Optional[float]). Callers rendering this for humans should annotate the value as "avg of N/M judges" by separately computing N (non-errored CaseScore count) and M (total CaseScore count) from the input list — the function itself does not return the annotation.final_status=3, or no non-errored judge contributes any metric, compute_average_excluding_errors returns None — no meaningful average can be computed.Continue-on-failure semantics:
</error_handling>
When displaying the evaluation results summary (e.g., in the final output or any human-readable report), follow these conventions for errored scores:
Errored score display:
ERR marker in place of a numeric score for any judge whose CaseScore has final_status=3. error_reason, when present, can be displayed in a hover/tooltip or separate column but does not control whether ERR is shown.Example summary table:
| Judge | Score | Status |
|---|---|---|
| dry-judge | 0.92 | PASS |
| ssot-judge | ERR | ERROR |
| kiss-judge | 0.75 | FAIL |
| readability-judge | ERR | ERROR |
Average annotation:
"avg of N/M judges", where N is the number of non-errored judges and M is the total number of judges.avg of 14/16 judgesFooter line:
When one or more judges are excluded, add a footer line to the summary:
X of Y judges excluded due to errorswhere X is the count of errored judges and Y is the total expected judge count.
Example: 2 of 16 judges excluded due to errors
When ALL judges errored:
ERR for every judge rowN/A (not a number) for the aggregate average — do not attempt to compute or display an averageY of Y judges excluded due to errorsPerformance: Run "start of phase" with sub_step 5 (plan), 4 (code), 3 (prd), or 2 (feature), sub_step_name=aggregate. Emit 'end of phase' after the aggregation step regardless of file write outcome.
Task: Collect all CaseScore outputs and structure them into an EvaluationReport.
<output_structure>
Output file logic:
if artifact_type == 'code':
report_filename = 'code-judges.json'
report_id = f'{RUN_ID}-code-judges'
elif artifact_type == 'prd':
report_filename = 'prd-judges.json'
report_id = f'{RUN_ID}-prd-judges'
elif artifact_type == 'feature':
report_filename = 'feature-judges.json'
report_id = f'{RUN_ID}-feature-judges'
else:
report_filename = 'plan-judges.json'
report_id = f'{RUN_ID}-plan-judges'
output_path = $CLOSEDLOOP_WORKDIR / report_filenamePlan artifact report structure (plan-judges.json):
{
"report_id": "{RUN_ID}-plan-judges",
"timestamp": "2024-02-03T15:45:30Z",
"stats": [
{ /* CaseScore from dry-judge */ },
{ /* CaseScore from ssot-judge */ },
{ /* CaseScore from kiss-judge */ },
{ /* CaseScore from code-organization-judge */ },
{ /* CaseScore from custom-best-practices-judge */ },
{ /* CaseScore from goal-alignment-judge */ },
{ /* CaseScore from readability-judge */ },
{ /* CaseScore from verbosity-judge */ },
{ /* CaseScore from solid-isp-dip-judge */ },
{ /* CaseScore from solid-liskov-substitution-judge */ },
{ /* CaseScore from solid-open-closed-judge */ },
{ /* CaseScore from technical-accuracy-judge */ },
{ /* CaseScore from test-judge */ },
{ /* CaseScore from brownfield-accuracy-judge */ },
{ /* CaseScore from codebase-grounding-judge */ },
{ /* CaseScore from convention-adherence-judge */ }
]
}Code artifact report structure (code-judges.json):
{
"report_id": "{RUN_ID}-code-judges",
"timestamp": "2024-02-03T15:45:30Z",
"stats": [
{ /* CaseScore from dry-judge */ },
{ /* CaseScore from ssot-judge */ },
{ /* CaseScore from kiss-judge */ },
{ /* CaseScore from code-organization-judge */ },
{ /* CaseScore from custom-best-practices-judge */ },
{ /* CaseScore from readability-judge */ },
{ /* CaseScore from solid-isp-dip-judge */ },
{ /* CaseScore from solid-liskov-substitution-judge */ },
{ /* CaseScore from solid-open-closed-judge */ },
{ /* CaseScore from technical-accuracy-judge */ },
{ /* CaseScore from test-judge */ }
]
}PRD artifact report structure (prd-judges.json):
{
"report_id": "{RUN_ID}-prd-judges",
"timestamp": "2024-02-03T15:45:30Z",
"stats": [
{ /* CaseScore from feature-completeness-judge */ },
{ /* CaseScore from prd-auditor */ },
{ /* CaseScore from prd-dependency-judge */ },
{ /* CaseScore from prd-testability-judge */ },
{ /* CaseScore from prd-scope-judge */ }
]
}Feature artifact report structure (feature-judges.json):
{
"report_id": "{RUN_ID}-feature-judges",
"timestamp": "2024-02-03T15:45:30Z",
"stats": [
{ /* CaseScore from feature-completeness-judge */ },
{ /* CaseScore from prd-testability-judge */ },
{ /* CaseScore from prd-dependency-judge */ }
]
}Field requirements:
| Field | Format | How to Derive |
|---|---|---|
report_id | {RUN_ID}-plan-judges, {RUN_ID}-code-judges, {RUN_ID}-prd-judges, or {RUN_ID}-feature-judges | Extract RUN_ID from $CLOSEDLOOP_WORKDIR directory name, append suffix based on artifact type |
timestamp | ISO 8601 | Generate with date -u +%Y-%m-%dT%H:%M:%SZ |
stats | Array[CaseScore] | 16 CaseScore objects for plan, 11 for code, 5 for prd, 3 for feature (one per judge) |
</output_structure>
Performance: Run "start of phase" with sub_step 6 (plan), 5 (code), 4 (prd), or 3 (feature), sub_step_name=validate. Emit 'end of phase' after each validation attempt regardless of exit code, then apply failure recovery logic.
CRITICAL: You MUST run the validation script after writing the judge report. Do not consider the task complete until validation passes.
<validation_workflow>
Step 3.1: Locate the Validation Script
The script is in this skill's scripts/ directory:
SCRIPT_PATH="scripts/validate_judge_report.py"Step 3.2: Ensure uv is Installed
if ! command -v uv &> /dev/null; then
# Install uv — alternatives: brew install uv, pip install uv
curl -LsSf https://astral.sh/uv/install.sh | sh
fiStep 3.3: Run Validation
# CRITICAL: Run from script's directory so uv can find inline dependencies
cd "$(dirname "$SCRIPT_PATH")"
# Determine category based on artifact type
CATEGORY="plan" # default
if [ "$ARTIFACT_TYPE" = "code" ]; then
CATEGORY="code"
elif [ "$ARTIFACT_TYPE" = "prd" ]; then
CATEGORY="prd"
elif [ "$ARTIFACT_TYPE" = "feature" ]; then
CATEGORY="feature"
fi
# Run validation with appropriate category
uv run "$SCRIPT_PATH" --workdir "$CLOSEDLOOP_WORKDIR" --category "$CATEGORY"Argument requirements:
--workdir must be the absolute path to $CLOSEDLOOP_WORKDIR--category must be plan (16 judges), code (11 judges), prd (5 judges), or feature (3 judges)plan-judges.json, code-judges.json, prd-judges.json, or feature-judges.json is located</validation_workflow>
<validation_checks>
The script validates using strict Pydantic models:
| Check | Requirement |
|---|---|
| JSON syntax | Valid JSON format |
| Required fields | report_id, timestamp, stats array |
| Judge coverage | All expected judges present (16 for plan, 11 for code, 5 for prd, 3 for feature) |
| Status values | final_status ∈ {1, 2, 3} |
| Metric completeness | Each judge has ≥1 metric |
| Report ID format | Ends with '-judges' (plan), '-code-judges' (code), '-prd-judges' (prd), or '-feature-judges' (feature) |
Expected judge case_ids for plan artifacts (16 total):
brownfield-accuracy-judge
code-organization-judge
codebase-grounding-judge
convention-adherence-judge
custom-best-practices-judge
dry-judge
goal-alignment-judge
kiss-judge
readability-judge
solid-isp-dip-judge
solid-liskov-substitution-judge
solid-open-closed-judge
ssot-judge
technical-accuracy-judge
test-judge
verbosity-judgeExpected judge case_ids for code artifacts (11 total):
code-organization-judge
custom-best-practices-judge
dry-judge
kiss-judge
readability-judge
solid-isp-dip-judge
solid-liskov-substitution-judge
solid-open-closed-judge
ssot-judge
technical-accuracy-judge
test-judgeNote: Code artifacts exclude: goal-alignment-judge, verbosity-judge
Expected judge case_ids for PRD artifacts (5 total):
feature-completeness-judge
prd-auditor
prd-dependency-judge
prd-testability-judge
prd-scope-judgeNote: PRD judges run in 2 sequential batches (3 + 2) to respect the Task tool's 4-concurrent-agent limit.
Expected judge case_ids for Feature artifacts (3 total):
feature-completeness-judge
prd-dependency-judge
prd-testability-judgeNote: Feature judges run in 1 batch. prd-auditor and prd-scope-judge are excluded — see Feature mode judge selection rationale in Task Context section.
</validation_checks>
| Code | Meaning | Action |
|---|---|---|
0 | Valid | Task complete ✓ |
1 | Invalid | Read error, fix report JSON, re-validate |
<failure_recovery>
Follow this sequence:
</failure_recovery>
<pydantic_schema>
The validation script uses these strict Pydantic models:
class MetricStatistics(BaseModel):
"""A single metric evaluation result."""
metric_name: str
threshold: Optional[float] = None
score: float
justification: str
class CaseScore(BaseModel):
"""Score for a single judge evaluation."""
type: Optional[str] = "case_score"
case_id: str
final_status: int # 1=pass, 2=fail, 3=error
metrics: List[MetricStatistics]
error_reason: Optional[str] = None # set when final_status=3; excluded from aggregation averages
class EvaluationReport(BaseModel):
"""Top-level report containing all judge evaluations."""
report_id: str
timestamp: str
stats: List[CaseScore]Model constraints:
ConfigDict(strict=True) enforces exact type matchingfinal_status validator rejects values outside {1, 2, 3}</pydantic_schema>
<completion_criteria>
Before marking this task complete, verify:
For all artifact types:
agents-snapshot/manifest.json exists in $CLOSEDLOOP_WORKDIR (created if missing, skipped if present)For plan artifacts (default):
artifact_type=planplan-context.json exists, or compatibility mode explicitly activatedjudge-input.json exists with required fieldsinvestigation-log.md reused, generated via pre-explorer, or best-effort generated internallyplan-judges.json written to $CLOSEDLOOP_WORKDIR--category planFor code artifacts (--artifact-type code):
.closedloop-ai/context/code-context.json exists at $CLOSEDLOOP_WORKDIR, or root code-context.json fallback is explicitly used for an old runjudge-input.json exists with required fieldsinvestigation-log.md reused or generated best-effort; missing file does not block code judgingcode-judges.json written to $CLOSEDLOOP_WORKDIR--category codeFor PRD artifacts (--artifact-type prd):
$CLOSEDLOOP_WORKDIR/prd.md found, or graceful exit with WARNING (code 0)scripts/judge_input_mapping.py wrote schema-valid judge-input.json with evaluation_type="prd" and primary_artifact.id="primary_prd"prd-judges.json written to $CLOSEDLOOP_WORKDIR--category prd (sub_step=4)For Feature artifacts (--artifact-type feature):
$CLOSEDLOOP_WORKDIR/feature.md or legacy $CLOSEDLOOP_WORKDIR/prd.md found, or emit sub_step=0 (skipped=true) perf event, emit WARNING, and graceful exit with WARNING (code 0)scripts/judge_input_mapping.py wrote schema-valid judge-input.json with evaluation_type="feature" and primary_artifact.id="primary_feature"feature-judges.json written to $CLOSEDLOOP_WORKDIR--category feature (sub_step=3)</completion_criteria>
<troubleshooting>
| Error Message | Root Cause | Solution |
|---|---|---|
| "Report file does not exist" | File not written to correct location | Verify $CLOSEDLOOP_WORKDIR is set; check write path matches artifact type (plan-judges.json, code-judges.json, prd-judges.json, or feature-judges.json) |
| "Invalid JSON" | Syntax error in output file | Run python3 -m json.tool "$CLOSEDLOOP_WORKDIR/{plan,code,prd,feature}-judges.json" to identify syntax error |
| "Missing expected judges" | Incomplete batch execution | Verify all batches launched (4 for plan, 3 for code, 2 for prd, 1 for feature); check error CaseScores for failures; plan expects 16 judges, code expects 11, prd expects 5, feature expects 3 |
| "final_status must be 1, 2, or 3" | Invalid status code | Use only: 1 (pass), 2 (fail), 3 (error) |
| "report_id should end with '-plan-judges'" | Incorrect ID format for plan | Use pattern: {RUN_ID}-plan-judges for plan artifacts |
| "report_id should end with '-code-judges'" | Incorrect ID format for code | Use pattern: {RUN_ID}-code-judges for code artifacts |
| "Judge {name} has no metrics" | Empty metrics array | Each CaseScore must have ≥1 MetricStatistics entry |
| "Context preparation failed" | context-manager-for-judges failed | Check context-manager agent output; verify artifact files exist |
| "judge-input.json missing" | Orchestrator did not generate envelope | Run scripts/judge_input_mapping.py before launching judges |
| "judge-input schema invalid" | Missing required envelope fields | Re-run scripts/judge_input_mapping.py; it validates required fields: evaluation_type, task, primary_artifact, supporting_artifacts, source_of_truth, fallback_mode, metadata |
| "plan-context.json not found" | plan context manager did not produce output | Run @judges:context-manager-for-judges with artifact_type=plan; if still missing, activate one-run compatibility fallback to plan.json + prd.md |
| "Preamble file not found" | Missing common or artifact preamble .md file | Verify both skills/artifact-type-tailored-context/preambles/common_input_preamble.md and skills/artifact-type-tailored-context/preambles/{artifact_type}_preamble.md exist |
| "pre-explorer unavailable" | @code:pre-explorer not installed/resolvable | Log warning and use internal fallback investigation to create investigation-log.md |
| "investigation-log.md missing after fallback" | Both pre-explorer and internal fallback failed | Log warning and continue; do not block context preparation |
| "investigation-log.md missing in code mode" | pre-explorer unavailable or generation failed during code preflight | Log warning and continue with .closedloop-ai/context/code-context.json only (non-blocking), using root code-context.json only as legacy fallback |
| "Invalid --artifact-type value" | Unsupported artifact type | Use only 'plan', 'code', 'prd', or 'feature' |
| "prd.md not found" | PRD document missing from workdir | Emit WARNING and exit gracefully (code 0); do not fail the parent workflow |
| "report_id should end with '-prd-judges'" | Incorrect ID format for prd | Use pattern: {RUN_ID}-prd-judges for PRD artifacts |
| "report_id should end with '-feature-judges'" | Incorrect ID format for feature | Use pattern: {RUN_ID}-feature-judges for Feature artifacts |
| "feature_preamble.md not found" | feature_preamble.md missing from preambles directory | Verify skills/artifact-type-tailored-context/preambles/feature_preamble.md exists; do NOT fall back to prd_preamble.md (it injects contradictory contract instructions for feature mode) |
| "Missing expected judges (feature)" | Incomplete batch execution for feature mode | Verify batch_1 launched all 3 judges: feature-completeness-judge, prd-testability-judge, prd-dependency-judge |
</troubleshooting>
If --artifact-type value is not 'plan', 'code', 'prd', or 'feature':
If context-manager-for-judges agent exceeds 5 minutes:
final_status=3, error_reason="Timeout: context preparation exceeded 5 minutes", justification="Context preparation timeout" (see error_reason guidance above)If context-manager-for-judges agent exceeds 5 minutes in plan mode:
plan.json + prd.mdIf a single judge Task call fails during execution:
final_status=3 and a populated error_reason describing the specific failure (e.g. "Task tool error: agent not found", "Parse error: response was not valid JSON") per the error_reason guidance aboveWhen --artifact-type is not specified or equals 'plan':
plan-judges.json (not code-judges.json)plan-context.json as primary input; use one-run compatibility fallback only if context preparation failsjudge-input.json envelope to judges--category planThis is the standard plan mode flow; orchestrators must support context-manager launch, judge-input.json construction, and preamble injection. The compatibility fallback (raw plan.json + prd.md) activates only when context preparation fails (e.g., context-manager timeout), not for orchestrators that have not been updated.
When --artifact-type prd is specified:
$CLOSEDLOOP_WORKDIR/prd.md exists; emit WARNING and exit gracefully (code 0) if missingjudge-input.json with scripts/judge_input_mapping.py --artifact-type prdprd-judges.json--category prd (sub_step=4)When --artifact-type feature is specified:
$CLOSEDLOOP_WORKDIR/feature.md exists, or legacy $CLOSEDLOOP_WORKDIR/prd.md exists; emit sub_step=0 (context_prep, skipped=true) perf event, emit WARNING, and exit gracefully (code 0) if both are missingjudge-input.json with scripts/judge_input_mapping.py --artifact-type featurefeature_preamble.md for all 3 feature judges (Feature-shaped contract; do NOT substitute prd_preamble.md)feature-judges.json--category feature (sub_step=3)© closedloop-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (scripts, references) in plugins/judges/skills/run-judges of closedloop-ai/claude-plugins.
Open the folder on GitHubat commit 4600742
Run Judges next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Run Judges this skillclosedloop-ai/claude-plugins | 122 | — | ~15k | Automated safety check: Notes | Apache-2.0 | |
| AlignmentOvid/paad | 131 | — | ~4.9k | Automated safety check: Pass | MIT | |
| Clarification AnalystGulajavaMinistudio/Mayukai-Theme | 139 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Feature SpecPackmindHub/packmind | 317 | — | ~1.8k | Automated safety check: Pass | Apache-2.0 | |
| Mcaf Feature Specmanagedcode/Storage | 138 | — | ~959 | Automated safety check: Pass | MIT | |
| Dex Improvedavekilleen/Dex | 493 | — | ~2.8k | Automated safety check: Pass | MIT |
Ovid/paad
A skill your agent uses when verifying that requirements/specs/PRDs and their implementation plans match — before starting work, after a spec or plan update, or when suspecting coverage gaps, scope…
GulajavaMinistudio/Mayukai-Theme
Helps interrogate Product Requirements (PRD), Technical Specifications, and Implementation Plans to find ambiguities, missing edge cases, and hidden assumptions.
PackmindHub/packmind
Generate a Packmind feature specification from a GitHub issue, file, URL, or direct description.
managedcode/Storage
Create or update a feature spec under docs/Features/ with business rules, user flows, system behaviour, verification, and Definition of Done.
davekilleen/Dex
Workshop one improvement idea into an implementation plan. An agent skill from davekilleen/Dex.
Dwlad90/stylex-swc-plugin
Turn a PRD into a multi-phase implementation plan using tracer-bullet vertical slices, saved as a local Markdown file in ./plans/.
closedloop-ai/claude-plugins
Run Codex to review a plan file and return structured feedback with a verdict.
closedloop-ai/claude-plugins
Check if critic reviews are still valid before re-running Phase 2.5 critics.
closedloop-ai/claude-plugins
Check if cross-repo coordinator results can be reused, avoiding redundant Sonnet agent launches.
closedloop-ai/claude-plugins
Check for a cached plan-evaluation.json result before launching the plan-evaluator agent.
closedloop-ai/claude-plugins
This skill should be used when needing to locate files within the Claude Code plugins cache directory (~/.claude/plugins/cache).
closedloop-ai/claude-plugins
Finish a vibe session in symphony-alpha and hand it to whoever picks it up next, usually design and then engineering.
Orchestrate parallel judge agent execution, aggregate CaseScore results, write plan-judges.json, code-judges.json, prd-judges.json, or feature-judges.json, and validate output. Run Judges is an agent skill from closedloop-ai/claude-plugins.json, and validate output.
Run Judges fits situations like: tasks that involve PRD writing; tasks that involve Planning.
Run `npx skills add closedloop-ai/claude-plugins --skill run-judges -a claude-code`. Or copy the skill folder (plugins/judges/skills/run-judges in closedloop-ai/claude-plugins) into .claude/skills/run-judges in your project. Claude Code loads it when a task matches its description.
Run `npx skills add closedloop-ai/claude-plugins --skill run-judges -a codex`. Or copy the skill folder (plugins/judges/skills/run-judges in closedloop-ai/claude-plugins) into .agents/skills/run-judges in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add closedloop-ai/claude-plugins --skill run-judges -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-judges, .gemini/skills/run-judges, .github/skills/run-judges and .opencode/skills/run-judges in your project.
Going by SKILL.md and its folder, Run Judges needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (uv, jq, bash, curl, sh and python3). Our summary lists: Python 3; A Bash shell.
SKILL.md names 1 domain. In commands or code: astral.sh; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pipes a well-known installer script into a shell), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Run Judges is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 15k tokens (SKILL.md is roughly 60k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 697 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Run Judges: Alignment (Ovid/paad, 131 stars), Clarification Analyst (GulajavaMinistudio/Mayukai-Theme, 139 stars), Feature Spec (PackmindHub/packmind, 317 stars) and Mcaf Feature Spec (managedcode/Storage, 138 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
closedloop-ai (a GitHub organization) maintains it in closedloop-ai/claude-plugins, which has 122 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on October 8, 2026.
Source: closedloop-ai/claude-plugins on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.