Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation

GPL-3.0Auto-check passedAgent Workflows

Install Judge

skills CLI
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NeoLabHQ/context-engineering-kit judge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/judge .claude/skills/judge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
judge
GitHub stars
1.8k
Token cost
~2k tokens
SKILL.md length
391 words
Files
1
Skills in repo
57
Repo updated
First seen
Licence
GPL-3.0

At a glance

Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation

  • Works in 3 steps: Context Extraction → Dispatch Meta-Judge → Dispatch Judge Agent
  • Tasks that involve Subagents
  • SKILL.md covers Your Workflow and Instructions
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Judge is an agent skill from NeoLabHQ/context-engineering-kit. Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Subagents. The repository describes itself as: Hand-crafted Claude Code Skills focused on improving agent results quality. Compatible with OpenCode, Cursor, Antigravity, Gemini CLI, and others. Includes CodeRabbit open-source… The licence is GPL-3.0.

When your agent uses it

  • Tasks that involve Subagents

Example prompts

  • “/judge”

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Context Extraction
  2. Dispatch Meta-Judge
  3. Dispatch Judge Agent

What it can do on your machine

Read from SKILL.md and the folder at commit 23e2428. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Judge loads about 2k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 391 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NeoLabHQ/context-engineering-kit at commit 23e2428, republished under its GPL-3.0 licence (© NeoLabHQ). 391 words, ~2,030 tokens.

Download SKILL.mdSave it as .claude/skills/judge/SKILL.md (or your agent's skills folder).
name
judge
description
Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation

Judge Command

<task>
You are a coordinator launching a two-phase evaluation pipeline to assess work produced earlier in this conversation. First, a meta-judge generates tailored evaluation criteria. Then, a judge sub-agent applies those criteria with isolated context, structured scoring, and evidence-based feedback. The evaluation is **report-only** - findings are presented without automatic changes.
</task>
<context>
This command implements the **meta-judge -> LLM-as-Judge** pattern with context isolation:
- **Structured Evaluation**: Meta-judge produces tailored rubrics, checklists, and scoring criteria before judging
- **Context Isolation**: Judge operates with fresh context, preventing confirmation bias from accumulated session state
- **Evidence-Based**: Every score requires specific citations from the work (file locations, line numbers)
- **Multi-Dimensional Rubric**: Generated by meta-judge to match the specific artifact type and evaluation focus
- **Self-Verification**: Dynamic verification questions with documented adjustments
</context>

Your Workflow

Phase 1: Context Extraction

Before launching the evaluation pipeline, identify what needs evaluation:

  1. Identify the work to evaluate:

    • Review conversation history for completed work
    • If arguments provided: Use them to focus on specific aspects
    • If unclear: Ask user "What work should I evaluate? (code changes, analysis, documentation, etc.)"
  2. Extract evaluation context:

    • Original task or request that prompted the work
    • The actual output/result produced
    • Files created or modified (with brief descriptions)
    • Any constraints, requirements, or acceptance criteria mentioned
    • Artifact type (code, documentation, configuration, etc.)
  3. Provide scope for user:

    Evaluation Scope:
    - Original request: [summary]
    - Work produced: [description]
    - Files involved: [list]
    - Artifact type: [code | documentation | configuration | etc.]
    - Evaluation focus: [from arguments or "general quality"]
    
    Launching meta-judge to generate evaluation criteria...

IMPORTANT: Pass only the extracted context to the sub-agents - not the entire conversation. This prevents context pollution and enables focused assessment.

Show full SKILL.md (153 more words)Show less
Phase 2: Dispatch Meta-Judge

Launch a meta-judge agent to generate an evaluation specification tailored to the specific work being evaluated. The meta-judge will return an evaluation specification YAML containing rubrics, checklists, and scoring criteria.

Meta-Judge Prompt:

markdown
## Task

Generate an evaluation specification yaml for the following evaluation task. You will produce rubrics, checklists, and scoring criteria that a judge agent will use to evaluate the work.

CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`

## User Prompt
{Original task or request that prompted the work}

## Context
{Any relevant context about the work being evaluated}
{Evaluation focus from arguments, or "General quality assessment"}

## Artifact Type
{code | documentation | configuration | etc.}

## Instructions
Return only the final evaluation specification YAML in your response.

Dispatch:

Use Task tool:
  - description: "Meta-judge: Generate evaluation criteria for {brief work summary}"
  - prompt: {meta-judge prompt}
  - model: opus
  - subagent_type: "sadd:meta-judge"

Wait for the meta-judge to complete before proceeding to Phase 3.

Phase 3: Dispatch Judge Agent

After the meta-judge completes, extract its evaluation specification YAML and dispatch the judge agent with both the work context and the specification.

CRITICAL: Provide to the judge the EXACT meta-judge evaluation specification YAML. Do not skip, add, modify, shorten, or summarize any text in it!

Judge Agent Prompt:

markdown
You are an Expert Judge evaluating the quality of work against an evaluation specification produced by the meta judge.

CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`

## Work Under Evaluation

[ORIGINAL TASK]
{paste the original request/task}
[/ORIGINAL TASK]

[WORK OUTPUT]
{summary of what was created/modified}
[/WORK OUTPUT]

[FILES INVOLVED]
{list of files with brief descriptions}
[/FILES INVOLVED]

## Evaluation Specification

```yaml
{meta-judge's evaluation specification YAML}

Instructions

Follow your full judge process as defined in your agent instructions!

CRITICAL: You must reply with this exact structured evaluation report format in YAML at the START of your response!


CRITICAL: NEVER provide score threshold to judges in any format. Judge MUST not know what threshold for score is, in order to not be biased!!!

**Dispatch:**

Use Task tool:

  • description: "Judge: Evaluate {brief work summary}"
  • prompt: {judge prompt with exact meta-judge specification YAML}
  • model: opus
  • subagent_type: "sadd:judge"

### Phase 4: Process and Present Results

After receiving the judge's evaluation:

1. **Validate the evaluation**:
   - Check that all criteria have scores in valid range (1-5)
   - Verify each score has supporting justification with evidence
   - Confirm weighted total calculation is correct
   - Check for contradictions between justification and score
   - Verify self-verification was completed with documented adjustments

2. **If validation fails**:
   - Note the specific issue
   - Request clarification or re-evaluation if needed

3. **Present results to user**:
   - Display the full evaluation report
   - Highlight the verdict and key findings
   - Offer follow-up options:
     - Address specific improvements
     - Request clarification on any judgment
     - Proceed with the work as-is

## Scoring Interpretation

| Score Range | Verdict | Interpretation | Recommendation |
|-------------|---------|----------------|----------------|
| 4.50 - 5.00 | EXCELLENT | Exceptional quality, exceeds expectations | Ready as-is |
| 4.00 - 4.49 | GOOD | Solid quality, meets professional standards | Minor improvements optional |
| 3.50 - 3.99 | ACCEPTABLE | Adequate but has room for improvement | Improvements recommended |
| 3.00 - 3.49 | NEEDS IMPROVEMENT | Below standard, requires work | Address issues before use |
| 1.00 - 2.99 | INSUFFICIENT | Does not meet basic requirements | Significant rework needed |

## Important Guidelines

1. **Meta-judge first**: Always generate evaluation specification before judging - never skip the meta-judge phase
2. **Include CLAUDE_PLUGIN_ROOT**: Both meta-judge and judge need the resolved plugin root path
3. **Meta-judge YAML**: Pass only the meta-judge YAML to the judge, do not modify it
4. **Context Isolation**: Pass only relevant context to sub-agents - not the entire conversation
5. **Justification First**: Always require evidence and reasoning BEFORE the score
6. **Evidence-Based**: Every score must cite specific evidence (file paths, line numbers, quotes)
7. **Bias Mitigation**: Explicitly warn against length bias, verbosity bias, and authority bias
8. **Be Objective**: Base assessments on evidence and rubric definitions, not preferences
9. **Be Specific**: Cite exact locations, not vague observations
10. **Be Constructive**: Frame criticism as opportunities for improvement with impact context
11. **Consider Context**: Account for stated constraints, complexity, and requirements
12. **Report Confidence**: Lower confidence when evidence is ambiguous or criteria unclear
13. **Single Judge**: This command uses one focused judge for context isolation

## Notes

- This is a **report-only** command - it evaluates but does not modify work
- The meta-judge generates criteria tailored to the specific artifact type and evaluation focus
- The judge operates with fresh context for unbiased assessment
- Scores are calibrated to professional development standards
- Low scores indicate improvement opportunities, not failures
- Use the evaluation to inform next steps and iterations
- Low confidence evaluations may warrant human review

© NeoLabHQ, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/judge of NeoLabHQ/context-engineering-kit.

Open the folder on GitHubat commit 23e2428

Compare with similar skills

Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Judge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Judge this skillNeoLabHQ/context-engineering-kit1.8k—~2kAutomated safety check: PassGPL-3.0
Claude Code Agent Developmentanthropics/claude-plugins-official38k7 repos~2.8kAutomated safety check: PassApache-2.0
Subagent Driven DevelopmentAsvarox/allkaraoke26138 repos~1.2kAutomated safety check: PassNone
Dispatching Parallel Agentsultralisp/ultralisp25841 repos~1.5kAutomated safety check: PassNone
Reflect on Session Learningscursor/plugins11k5 repos~1.2kAutomated safety check: PassNone
Paseo Advisor Second Opiniongetpaseo/paseo20k1 repos~756Automated safety check: PassCustom licence

Similar skills

  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed
  • Subagent Driven Development

    Asvarox/allkaraoke

    A skill your agent uses when executing implementation plans with independent tasks in the current session

    261 GitHub starsUsed in 38 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Dispatching Parallel Agents

    ultralisp/ultralisp

    A skill your agent uses when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies

    258 GitHub starsUsed in 41 repos~1.5k tokens
    Agent WorkflowsAuto-check passed
  • Official

    Starts three parallel reviewer subagents over the current conversation transcript, then turns their findings into concrete edits to existing skills.

    11k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • Launches one separate agent through Paseo to give a second opinion on the current task, with a self-contained briefing and no permission to edit files.

    20k GitHub starsUsed in 1 repo~756 tokens
    Agent WorkflowsAuto-check passed
  • Task Observer

    rebelytics/one-skill-to-rule-them-all

    Monitors task execution for skill improvement opportunities.

    3.2k GitHub starsUsed in 1 repo~11k tokens
    Agent WorkflowsAuto-check passed

More from NeoLabHQ/context-engineering-kit

All 57 skills in this repo
  • Git Notes

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when adding metadata to commits without changing history, tracking review status, test results, code quality annotations, or supplementing commit messages post-hoc - provides…

    1.8k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Load PR Comments

    NeoLabHQ/context-engineering-kit

    A skill your agent uses to load open/unresolved PR review comments then aggregate them as tasks in .specs/comments/.md for parallel agents to fix.

    1.8k GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Prompt Engineering

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when you writing commands, hooks, skills for Agent, or prompts for sub agents or any other LLM interaction, including optimizing prompts, improving LLM outputs, or designing…

    1.8k GitHub stars~4.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Multi Agent Patterns

    NeoLabHQ/context-engineering-kit

    Design multi-agent architectures for complex tasks. An agent skill from NeoLabHQ/context-engineering-kit.

    1.8k GitHub starsUsed in 6 repos~6k tokens
    Auto-check passed
  • Review PR

    NeoLabHQ/context-engineering-kit

    Review an existing GitHub pull request and post inline review comments on its diff.

    1.8k GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Subagent Driven Development

    NeoLabHQ/context-engineering-kit

    A skill your agent uses when executing implementation plans with independent tasks in the current session or facing 3+ independent issues that can be investigated without shared state or…

    1.8k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Judge

What does Judge do?

Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation. Judge is an agent skill from NeoLabHQ/context-engineering-kit.

When should I use Judge?

Judge fits situations like: tasks that involve Subagents.

How do I install Judge in Claude Code?

Run `npx skills add NeoLabHQ/context-engineering-kit --skill judge -a claude-code`. Or copy the skill folder (skills/judge in NeoLabHQ/context-engineering-kit) into .claude/skills/judge in your project. Claude Code loads it when a task matches its description.

How do I install Judge in Codex?

Run `npx skills add NeoLabHQ/context-engineering-kit --skill judge -a codex`. Or copy the skill folder (skills/judge in NeoLabHQ/context-engineering-kit) into .agents/skills/judge in your project. Codex loads it when a task matches its description.

Can I use Judge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeoLabHQ/context-engineering-kit --skill judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/judge, .gemini/skills/judge, .github/skills/judge and .opencode/skills/judge in your project.

What does Judge need to run?

SKILL.md names no scripts, command-line tools or credentials: Judge is instructions for the agent only.

Does Judge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Judge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Judge use?

Judge is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Judge use?

About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Judge?

Skills that share tags, products or a category with Judge: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Subagent Driven Development (Asvarox/allkaraoke, 261 stars), Dispatching Parallel Agents (ultralisp/ultralisp, 258 stars) and Reflect on Session Learnings (cursor/plugins, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Judge?

NeoLabHQ (a GitHub organization) maintains it in NeoLabHQ/context-engineering-kit, which has 1,750 GitHub stars. The repository holds 57 skills in this directory. The repository was last updated on August 26, 2026.

Source: NeoLabHQ/context-engineering-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.