Claude Code Agent Development
anthropics/claude-plugins-official
Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.
Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/judge .claude/skills/judge && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "judge" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge into .claude/skills/judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judgeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/judge .agents/skills/judge && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "judge" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge into .agents/skills/judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/judge .cursor/skills/judge && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "judge" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge into .cursor/skills/judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NeoLabHQ/context-engineering-kit.git --path skills/judge--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/judge .gemini/skills/judge && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "judge" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge into .gemini/skills/judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NeoLabHQ/context-engineering-kit judgeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/judge .github/skills/judge && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "judge" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge into .github/skills/judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/judge .opencode/skills/judge && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "judge" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge into .opencode/skills/judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
judgeLaunch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation
Judge is an agent skill from NeoLabHQ/context-engineering-kit. Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Subagents. The repository describes itself as: Hand-crafted Claude Code Skills focused on improving agent results quality. Compatible with OpenCode, Cursor, Antigravity, Gemini CLI, and others. Includes CodeRabbit open-source… The licence is GPL-3.0.
3 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 23e2428. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Judge loads about 2k tokens when it runs. Until then it costs about 26 tokens; SKILL.md has 391 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NeoLabHQ/context-engineering-kit at commit 23e2428, republished under its GPL-3.0 licence (© NeoLabHQ). 391 words, ~2,030 tokens.
.claude/skills/judge/SKILL.md (or your agent's skills folder).<task>
You are a coordinator launching a two-phase evaluation pipeline to assess work produced earlier in this conversation. First, a meta-judge generates tailored evaluation criteria. Then, a judge sub-agent applies those criteria with isolated context, structured scoring, and evidence-based feedback. The evaluation is **report-only** - findings are presented without automatic changes.
</task>
<context>
This command implements the **meta-judge -> LLM-as-Judge** pattern with context isolation:
- **Structured Evaluation**: Meta-judge produces tailored rubrics, checklists, and scoring criteria before judging
- **Context Isolation**: Judge operates with fresh context, preventing confirmation bias from accumulated session state
- **Evidence-Based**: Every score requires specific citations from the work (file locations, line numbers)
- **Multi-Dimensional Rubric**: Generated by meta-judge to match the specific artifact type and evaluation focus
- **Self-Verification**: Dynamic verification questions with documented adjustments
</context>
Before launching the evaluation pipeline, identify what needs evaluation:
Identify the work to evaluate:
Extract evaluation context:
Provide scope for user:
Evaluation Scope:
- Original request: [summary]
- Work produced: [description]
- Files involved: [list]
- Artifact type: [code | documentation | configuration | etc.]
- Evaluation focus: [from arguments or "general quality"]
Launching meta-judge to generate evaluation criteria...IMPORTANT: Pass only the extracted context to the sub-agents - not the entire conversation. This prevents context pollution and enables focused assessment.
Launch a meta-judge agent to generate an evaluation specification tailored to the specific work being evaluated. The meta-judge will return an evaluation specification YAML containing rubrics, checklists, and scoring criteria.
Meta-Judge Prompt:
## Task
Generate an evaluation specification yaml for the following evaluation task. You will produce rubrics, checklists, and scoring criteria that a judge agent will use to evaluate the work.
CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`
## User Prompt
{Original task or request that prompted the work}
## Context
{Any relevant context about the work being evaluated}
{Evaluation focus from arguments, or "General quality assessment"}
## Artifact Type
{code | documentation | configuration | etc.}
## Instructions
Return only the final evaluation specification YAML in your response.Dispatch:
Use Task tool:
- description: "Meta-judge: Generate evaluation criteria for {brief work summary}"
- prompt: {meta-judge prompt}
- model: opus
- subagent_type: "sadd:meta-judge"Wait for the meta-judge to complete before proceeding to Phase 3.
After the meta-judge completes, extract its evaluation specification YAML and dispatch the judge agent with both the work context and the specification.
CRITICAL: Provide to the judge the EXACT meta-judge evaluation specification YAML. Do not skip, add, modify, shorten, or summarize any text in it!
Judge Agent Prompt:
You are an Expert Judge evaluating the quality of work against an evaluation specification produced by the meta judge.
CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`
## Work Under Evaluation
[ORIGINAL TASK]
{paste the original request/task}
[/ORIGINAL TASK]
[WORK OUTPUT]
{summary of what was created/modified}
[/WORK OUTPUT]
[FILES INVOLVED]
{list of files with brief descriptions}
[/FILES INVOLVED]
## Evaluation Specification
```yaml
{meta-judge's evaluation specification YAML}Follow your full judge process as defined in your agent instructions!
CRITICAL: You must reply with this exact structured evaluation report format in YAML at the START of your response!
CRITICAL: NEVER provide score threshold to judges in any format. Judge MUST not know what threshold for score is, in order to not be biased!!!
**Dispatch:**
Use Task tool:
### Phase 4: Process and Present Results
After receiving the judge's evaluation:
1. **Validate the evaluation**:
- Check that all criteria have scores in valid range (1-5)
- Verify each score has supporting justification with evidence
- Confirm weighted total calculation is correct
- Check for contradictions between justification and score
- Verify self-verification was completed with documented adjustments
2. **If validation fails**:
- Note the specific issue
- Request clarification or re-evaluation if needed
3. **Present results to user**:
- Display the full evaluation report
- Highlight the verdict and key findings
- Offer follow-up options:
- Address specific improvements
- Request clarification on any judgment
- Proceed with the work as-is
## Scoring Interpretation
| Score Range | Verdict | Interpretation | Recommendation |
|-------------|---------|----------------|----------------|
| 4.50 - 5.00 | EXCELLENT | Exceptional quality, exceeds expectations | Ready as-is |
| 4.00 - 4.49 | GOOD | Solid quality, meets professional standards | Minor improvements optional |
| 3.50 - 3.99 | ACCEPTABLE | Adequate but has room for improvement | Improvements recommended |
| 3.00 - 3.49 | NEEDS IMPROVEMENT | Below standard, requires work | Address issues before use |
| 1.00 - 2.99 | INSUFFICIENT | Does not meet basic requirements | Significant rework needed |
## Important Guidelines
1. **Meta-judge first**: Always generate evaluation specification before judging - never skip the meta-judge phase
2. **Include CLAUDE_PLUGIN_ROOT**: Both meta-judge and judge need the resolved plugin root path
3. **Meta-judge YAML**: Pass only the meta-judge YAML to the judge, do not modify it
4. **Context Isolation**: Pass only relevant context to sub-agents - not the entire conversation
5. **Justification First**: Always require evidence and reasoning BEFORE the score
6. **Evidence-Based**: Every score must cite specific evidence (file paths, line numbers, quotes)
7. **Bias Mitigation**: Explicitly warn against length bias, verbosity bias, and authority bias
8. **Be Objective**: Base assessments on evidence and rubric definitions, not preferences
9. **Be Specific**: Cite exact locations, not vague observations
10. **Be Constructive**: Frame criticism as opportunities for improvement with impact context
11. **Consider Context**: Account for stated constraints, complexity, and requirements
12. **Report Confidence**: Lower confidence when evidence is ambiguous or criteria unclear
13. **Single Judge**: This command uses one focused judge for context isolation
## Notes
- This is a **report-only** command - it evaluates but does not modify work
- The meta-judge generates criteria tailored to the specific artifact type and evaluation focus
- The judge operates with fresh context for unbiased assessment
- Scores are calibrated to professional development standards
- Low scores indicate improvement opportunities, not failures
- Use the evaluation to inform next steps and iterations
- Low confidence evaluations may warrant human review© NeoLabHQ, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/judge of NeoLabHQ/context-engineering-kit.
Open the folder on GitHubat commit 23e2428
Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Judge this skillNeoLabHQ/context-engineering-kit | 1.8k | — | ~2k | Automated safety check: Pass | GPL-3.0 | |
| Claude Code Agent Developmentanthropics/claude-plugins-official | 38k | 7 repos | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| Subagent Driven DevelopmentAsvarox/allkaraoke | 261 | 38 repos | ~1.2k | Automated safety check: Pass | None | |
| Dispatching Parallel Agentsultralisp/ultralisp | 258 | 41 repos | ~1.5k | Automated safety check: Pass | None | |
| Reflect on Session Learningscursor/plugins | 11k | 5 repos | ~1.2k | Automated safety check: Pass | None | |
| Paseo Advisor Second Opiniongetpaseo/paseo | 20k | 1 repos | ~756 | Automated safety check: Pass | Custom licence |
anthropics/claude-plugins-official
Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.
Asvarox/allkaraoke
A skill your agent uses when executing implementation plans with independent tasks in the current session
ultralisp/ultralisp
A skill your agent uses when facing 2+ independent tasks that can be worked on without shared state or sequential dependencies
cursor/plugins
Starts three parallel reviewer subagents over the current conversation transcript, then turns their findings into concrete edits to existing skills.
getpaseo/paseo
Launches one separate agent through Paseo to give a second opinion on the current task, with a self-contained briefing and no permission to edit files.
rebelytics/one-skill-to-rule-them-all
Monitors task execution for skill improvement opportunities.
NeoLabHQ/context-engineering-kit
A skill your agent uses when adding metadata to commits without changing history, tracking review status, test results, code quality annotations, or supplementing commit messages post-hoc - provides…
NeoLabHQ/context-engineering-kit
A skill your agent uses to load open/unresolved PR review comments then aggregate them as tasks in .specs/comments/.md for parallel agents to fix.
NeoLabHQ/context-engineering-kit
A skill your agent uses when you writing commands, hooks, skills for Agent, or prompts for sub agents or any other LLM interaction, including optimizing prompts, improving LLM outputs, or designing…
NeoLabHQ/context-engineering-kit
Design multi-agent architectures for complex tasks. An agent skill from NeoLabHQ/context-engineering-kit.
NeoLabHQ/context-engineering-kit
Review an existing GitHub pull request and post inline review comments on its diff.
NeoLabHQ/context-engineering-kit
A skill your agent uses when executing implementation plans with independent tasks in the current session or facing 3+ independent issues that can be investigated without shared state or…
Categories
Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation. Judge is an agent skill from NeoLabHQ/context-engineering-kit.
Judge fits situations like: tasks that involve Subagents.
Run `npx skills add NeoLabHQ/context-engineering-kit --skill judge -a claude-code`. Or copy the skill folder (skills/judge in NeoLabHQ/context-engineering-kit) into .claude/skills/judge in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NeoLabHQ/context-engineering-kit --skill judge -a codex`. Or copy the skill folder (skills/judge in NeoLabHQ/context-engineering-kit) into .agents/skills/judge in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeoLabHQ/context-engineering-kit --skill judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/judge, .gemini/skills/judge, .github/skills/judge and .opencode/skills/judge in your project.
SKILL.md names no scripts, command-line tools or credentials: Judge is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Judge is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 8.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Judge: Claude Code Agent Development (anthropics/claude-plugins-official, 38k stars), Subagent Driven Development (Asvarox/allkaraoke, 261 stars), Dispatching Parallel Agents (ultralisp/ultralisp, 258 stars) and Reflect on Session Learnings (cursor/plugins, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NeoLabHQ (a GitHub organization) maintains it in NeoLabHQ/context-engineering-kit, which has 1,750 GitHub stars. The repository holds 57 skills in this directory. The repository was last updated on August 26, 2026.
Source: NeoLabHQ/context-engineering-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.