Arize Evaluator
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Evaluate solutions through multi-round debate between independent judges until consensus
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge-with-debate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/judge-with-debate .claude/skills/judge-with-debate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "judge-with-debate" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge-with-debate into .claude/skills/judge-with-debate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge-with-debate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge-with-debateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge-with-debate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/judge-with-debate .agents/skills/judge-with-debate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "judge-with-debate" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge-with-debate into .agents/skills/judge-with-debate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge-with-debate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge-with-debate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/judge-with-debate .cursor/skills/judge-with-debate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "judge-with-debate" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge-with-debate into .cursor/skills/judge-with-debate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge-with-debate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NeoLabHQ/context-engineering-kit.git --path skills/judge-with-debate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge-with-debate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/judge-with-debate .gemini/skills/judge-with-debate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "judge-with-debate" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge-with-debate into .gemini/skills/judge-with-debate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge-with-debate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NeoLabHQ/context-engineering-kit judge-with-debateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/judge-with-debate .github/skills/judge-with-debate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "judge-with-debate" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge-with-debate into .github/skills/judge-with-debate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge-with-debate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NeoLabHQ/context-engineering-kit judge-with-debate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NeoLabHQ/context-engineering-kit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/judge-with-debate .opencode/skills/judge-with-debate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "judge-with-debate" agent skill from https://github.com/NeoLabHQ/context-engineering-kit/tree/master/skills/judge-with-debate into .opencode/skills/judge-with-debate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "judge-with-debate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
judge-with-debateEvaluate solutions through multi-round debate between independent judges until consensus
Judge With Debate is an agent skill from NeoLabHQ/context-engineering-kit. Evaluate solutions through multi-round debate between independent judges until consensus
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Hand-crafted Claude Code Skills focused on improving agent results quality. Compatible with OpenCode, Cursor, Antigravity, Gemini CLI, and others. Includes CodeRabbit open-source… The licence is GPL-3.0.
2 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 23e2428. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Judge With Debate loads about 4.3k tokens when it runs. Until then it costs about 27 tokens; SKILL.md has 953 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from NeoLabHQ/context-engineering-kit at commit 23e2428, republished under its GPL-3.0 licence (© NeoLabHQ). 953 words, ~4,292 tokens.
.claude/skills/judge-with-debate/SKILL.md (or your agent's skills folder).<task>
Evaluate solutions through multi-agent debate where independent judges analyze, challenge each other's assessments, and iteratively refine their evaluations until reaching consensus or maximum rounds.
</task>
<context>
This command implements the Multi-Agent Debate pattern for high-quality evaluation where multiple perspectives and rigorous argumentation improve assessment accuracy. Unlike single-pass evaluation, debate forces judges to defend their positions with evidence and consider counter-arguments.
Key benefits:
</context>
This command implements iterative multi-judge debate:
Phase 0: Setup
mkdir -p .specs/reports
|
Phase 0.5: Dispatch Meta-Judge
Meta-Judge (Opus)
|
Evaluation Specification YAML
|
Phase 1: Independent Analysis (3 judges in parallel)
+- Judge 1 -> {name}.1.md -+
Solution +- Judge 2 -> {name}.2.md -+-+
+- Judge 3 -> {name}.3.md -+ |
|
Phase 2: Debate Round (iterative) |
Each judge reads others' reports |
| |
Argue + Defend + Challenge |
(grounded in eval specification) |
| |
Revise if convinced --------------+
| |
Check consensus |
+- Yes -> Final Report |
+- No -> Next Round ---------+Before starting evaluation, ensure the reports directory exists:
mkdir -p .specs/reportsReport naming convention: .specs/reports/{solution-name}-{YYYY-MM-DD}.[1|2|3].md
Where:
{solution-name} - Derived from solution filename (e.g., users-api from src/api/users.ts){YYYY-MM-DD} - Current date[1|2|3] - Judge numberBefore independent analysis, dispatch a meta-judge agent to generate a tailored evaluation specification. The meta-judge runs ONCE and produces rubrics, checklists, and scoring criteria that ALL judges will use across ALL rounds.
Meta-judge prompt template:
## Task
Generate an evaluation specification yaml for the following evaluation task. You will produce rubrics, checklists, and scoring criteria that multiple judge agents will use to evaluate the solution through independent analysis and multi-round debate.
CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`
## User Prompt
{task description - what the solution was supposed to accomplish}
## Context
{Any relevant context about the solution being evaluated}
## Artifact Type
{code | documentation | configuration | etc.}
## Evaluation Mode
Multi-judge debate with consensus-seeking across rounds
## Instructions
Return only the final evaluation specification YAML in your response.
The specification should support both independent analysis and debate-based refinement.Dispatch:
Use Task tool:
- description: "Meta-judge: generate evaluation specification for {solution-name}"
- prompt: {meta-judge prompt}
- model: opus
- subagent_type: "sadd:meta-judge"Wait for the meta-judge to complete and extract the evaluation specification YAML from its output before proceeding to Phase 1.
Launch 3 independent judge agents in parallel (Opus for rigor):
.specs/reports/{solution-name}-{date}.[1|2|3].mdKey principle: Independence in initial analysis prevents groupthink.
Prompt template for initial judges:
You are Judge {N} evaluating a solution independently against an evaluation specification produced by the meta judge.
CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`
## Solution
{path to solution file(s)}
## Task Description
{what the solution was supposed to accomplish}
## Evaluation Specification
```yaml
{meta-judge's evaluation specification YAML}.specs/reports/{solution-name}-{date}.{N}.md
Follow your full judge process as defined in your agent instructions!
Additional instructions:
Add to report beginning Done by Judge {N}
**Dispatch each judge:**
Use Task tool:
### Phase 2: Debate Rounds (Iterative)
For each debate round (max 3 rounds):
Launch **3 debate agents in parallel**:
1. Each judge agent receives:
- Path to their own previous report (`.specs/reports/{solution-name}-{date}.[1|2|3].md`)
- Paths to other judges' reports (`.specs/reports/{solution-name}-{date}.[1|2|3].md`)
- The original solution
- The meta-judge's evaluation specification YAML
2. Each judge:
- Identifies disagreements with other judges (>1 point score gap on any criterion)
- Defends their own ratings with evidence from the solution and evaluation specification
- Challenges other judges' ratings they disagree with
- Considers counter-arguments
- Revises their assessment if convinced
3. Updates their report file with new section: `## Debate Round {R}`
4. After they reply, if they reached agreement move to Phase 3: Consensus Report
**Key principle:** Judges communicate only through filesystem - orchestrator doesn't mediate and don't read reports files itself, it can overflow your context.
**Prompt template for debate judges:**
```markdown
You are Judge {N} in debate round {R}.
CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`
## Your Previous Report
{path to .specs/reports/{solution-name}-{date}.{N}.md}
## Other Judges' Reports
Judge 1: .specs/reports/{solution-name}-{date}.1.md
...
## Task Description
{what the solution was supposed to accomplish}
## Solution
{path to solution}
## Evaluation Specification
```yaml
{meta-judge's evaluation specification YAML}.specs/reports/{solution-name}-{date}.{N}.md (append to existing file)
Follow your full judge process as defined in your agent instructions!
Additional debate instructions:
CRITICAL:
**Dispatch each debate judge:**
Use Task tool:
### Consensus Check
After each debate round, check for consensus:
**Consensus achieved if:**
- All judges' overall scores within 0.5 points of each other
- No criterion has >1 point disagreement across any two judges
- All judges explicitly state they accept the consensus
**If no consensus after 3 rounds:**
- Report persistent disagreements
- Provide all judge reports for human review
- Flag that automated evaluation couldn't reach consensus
**Orchestration Instructions:**
**Step 1: Dispatch Meta-Judge (Phase 0.5)**
1. Launch meta-judge agent
2. Wait for meta-judge to complete
3. Extract the evaluation specification YAML from meta-judge output
**Step 2: Run Independent Analysis (Phase 1)**
1. Launch 3 judge agents in parallel (Judge 1, 2, 3) with the evaluation specification YAML
2. Each writes their independent assessment to `.specs/reports/{solution-name}-{date}.[1|2|3].md`
3. Wait for all 3 agents to complete
**Step 3: Check for Consensus**
Let's work through this systematically to ensure accurate consensus detection.
Read all three reports and extract:
- Each judge's overall weighted score
- Each judge's score for every criterion
Check consensus step by step:
1. First, extract all overall scores from each report and list them explicitly
2. Calculate the difference between the highest and lowest overall scores
- If difference <= 0.5 points -> overall consensus achieved
- If difference > 0.5 points -> no consensus yet
3. Next, for each criterion, list all three judges' scores side by side
4. For each criterion, calculate the difference between highest and lowest scores
- If any criterion has difference > 1.0 point -> no consensus on that criterion
5. Finally, verify consensus is achieved only if BOTH conditions are met:
- Overall scores within 0.5 points
- All criterion scores within 1.0 point
**Step 4: Decision Point**
- **If consensus achieved**: Go to Step 6 (Generate Consensus Report)
- **If no consensus AND round < 3**: Go to Step 5 (Run Debate Round)
- **If no consensus AND round = 3**: Go to Step 7 (Report No Consensus)
**Step 5: Run Debate Round**
1. Increment round counter (round = round + 1)
2. Launch 3 judge agents in parallel with the same evaluation specification YAML
3. Each agent reads:
- Their own previous report from filesystem
- Other judges' reports from filesystem
- Original solution
4. Each agent appends "Debate Round {R}" section to their own report file
5. Wait for all 3 agents to complete
6. Go back to Step 3 (Check for Consensus)
**Step 6: Reply with Report**
Let's synthesize the evaluation results step by step.
1. Read all final reports carefully
2. Before generating the report, analyze the following:
- What is the consensus status (achieved or not)?
- What were the key points of agreement across all judges?
- What were the main areas of disagreement, if any?
- How did the debate rounds change the evaluations?
3. Reply to user with a report that contains:
- If there is consensus:
- Consensus scores (average of all judges)
- Consensus strengths/weaknesses
- Number of rounds to reach consensus
- Final recommendation with clear justification
- If there is no consensus:
- All judges' final scores showing disagreements
- Specific criteria where consensus wasn't reached
- Analysis of why consensus couldn't be reached
- Flag for human review
4. Command complete
**Step 7: Report No Consensus**
- Report persistent disagreements
- Provide all judge reports for human review
- Flag that automated evaluation couldn't reach consensus
### Phase 3: Consensus Report
If consensus achieved, synthesize the final report by working through each section methodically:
```markdown
# Consensus Evaluation Report
Let's compile the final consensus by analyzing each component systematically.
## Consensus Scores
First, let's consolidate all judges' final scores:
| Criterion | Judge 1 | Judge 2 | Judge 3 | Final |
|-----------|---------|---------|---------|-------|
| {Name} | {X}/5 | {X}/5 | {X}/5 | {X}/5 |
...
**Consensus Overall Score**: {avg}/5.0
## Consensus Strengths
[Review each judge's identified strengths and extract the common themes that all judges agreed upon]
## Consensus Weaknesses
[Review each judge's identified weaknesses and extract the common themes that all judges agreed upon]
## Debate Summary
Let's trace how consensus was reached:
- Rounds to consensus: {N}
- Initial disagreements: {list with specific criteria and score gaps}
- How resolved: {for each disagreement, explain what evidence or argument led to resolution}
## Final Recommendation
Based on the consensus scores and the key strengths/weaknesses identified:
{Pass/Fail/Needs Revision with clear justification tied to the evidence}<output>
The command produces:
.specs/reports/ (created if not exists).specs/reports/{solution-name}-{date}.1.md, .specs/reports/{solution-name}-{date}.2.md, .specs/reports/{solution-name}-{date}.3.md</output>
/judge-with-debate Implement REST API for user management --solution "src/api/users.ts" Phase 0.5 - Meta-Judge (assuming date 2025-01-15):
Phase 1 - Independent Analysis (3 judges receive specification):
.specs/reports/users-api-2025-01-15.1.md - Judge 1 scores correctness 4/5, security 3/5.specs/reports/users-api-2025-01-15.2.md - Judge 2 scores correctness 4/5, security 5/5.specs/reports/users-api-2025-01-15.3.md - Judge 3 scores correctness 5/5, security 4/5Disagreement detected: Security scores range from 3-5
Phase 2 - Debate Round 1 (judges reference evaluation specification):
Debate Round 1 outputs:
Debate Round 2 (same evaluation specification):
Final consensus:
Correctness: 4.3/5
Design: 4.5/5
Security: 4.0/5 (2 debate rounds to consensus)
Performance: 4.7/5
Documentation: 4.0/5
Overall: 4.3/5 - PASS</output>
© NeoLabHQ, GPL-3.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/judge-with-debate of NeoLabHQ/context-engineering-kit.
Open the folder on GitHubat commit 23e2428
Judge With Debate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Judge With Debate this skillNeoLabHQ/context-engineering-kit | 1.8k | — | ~4.3k | Automated safety check: Pass | GPL-3.0 | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 1 repos | ~8.1k | Automated safety check: Notes | MIT | |
| EvaluatorsArize-ai/phoenix | 12k | — | ~1.7k | Automated safety check: Pass | Custom licence | |
| LLM Evaluationdavila7/claude-code-templates | 33k | 12 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Evaluationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Idea Evaluatorsickn33/agentic-awesome-skills | 47k | 1 repos | ~930 | Automated safety check: Pass | MIT |
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Arize-ai/phoenix
Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
sickn33/agentic-awesome-skills
Evaluates an idea by hosting a multi-turn debate between a Pro and Con agent, delivering a final verdict on whether it's worth pursuing.
sickn33/agentic-awesome-skills
A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.
NeoLabHQ/context-engineering-kit
A skill your agent uses when adding metadata to commits without changing history, tracking review status, test results, code quality annotations, or supplementing commit messages post-hoc - provides…
NeoLabHQ/context-engineering-kit
A skill your agent uses to load open/unresolved PR review comments then aggregate them as tasks in .specs/comments/.md for parallel agents to fix.
NeoLabHQ/context-engineering-kit
A skill your agent uses when you writing commands, hooks, skills for Agent, or prompts for sub agents or any other LLM interaction, including optimizing prompts, improving LLM outputs, or designing…
NeoLabHQ/context-engineering-kit
Design multi-agent architectures for complex tasks. An agent skill from NeoLabHQ/context-engineering-kit.
NeoLabHQ/context-engineering-kit
Review an existing GitHub pull request and post inline review comments on its diff.
NeoLabHQ/context-engineering-kit
A skill your agent uses when executing implementation plans with independent tasks in the current session or facing 3+ independent issues that can be investigated without shared state or…
Evaluate solutions through multi-round debate between independent judges until consensus. Judge With Debate is an agent skill from NeoLabHQ/context-engineering-kit.
Run `npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a claude-code`. Or copy the skill folder (skills/judge-with-debate in NeoLabHQ/context-engineering-kit) into .claude/skills/judge-with-debate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a codex`. Or copy the skill folder (skills/judge-with-debate in NeoLabHQ/context-engineering-kit) into .agents/skills/judge-with-debate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NeoLabHQ/context-engineering-kit --skill judge-with-debate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/judge-with-debate, .gemini/skills/judge-with-debate, .github/skills/judge-with-debate and .opencode/skills/judge-with-debate in your project.
SKILL.md names no scripts, command-line tools or credentials: Judge With Debate is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Judge With Debate is published under the GPL-3.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Judge With Debate: Arize Evaluator (github/awesome-copilot, 40k stars), Evaluators (Arize-ai/phoenix, 12k stars), LLM Evaluation (davila7/claude-code-templates, 33k stars) and Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NeoLabHQ (a GitHub organization) maintains it in NeoLabHQ/context-engineering-kit, which has 1,750 GitHub stars. The repository holds 57 skills in this directory. The repository was last updated on August 26, 2026.
Source: NeoLabHQ/context-engineering-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.