Agent skill

00 Arena Router

by agentscope-ai in agentscope-ai/OpenJudge

A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom…

Apache-2.0Auto-check passedResearch & Science

Install 00 Arena Router

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge 00-arena-router --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arena-eval/00-arena-router .claude/skills/00-arena-router && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
00-arena-router
GitHub stars
871
Token cost
~1k tokens
SKILL.md length
380 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom…

  • The user wants to compare
  • SKILL.md covers Diagnostic Question, Triage Table, Key distinction and Output
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • A benchmark specifically about reference/citation hallucination rate

What it does

00 Arena Router is an agent skill from agentscope-ai/OpenJudge. Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom task, or a benchmark specifically about reference/citation hallucination rate. Also use when the user mentions model arena, agent arena, pairwise model comparison, win-rate ranking, or comparing models on a task and hasn't specified whether that task is generic or about citation accuracy. This skill is the entry router for the arena-eval…

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Citation management. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user wants to compare
  • A benchmark specifically about reference/citation hallucination rate
  • The user mentions model arena
  • Pairwise model comparison

Example prompts

  • “/00-arena-router”

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

00 Arena Router loads about 1k tokens when it runs. Until then it costs about 164 tokens; SKILL.md has 380 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~164
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 380 words, ~1,032 tokens.

Download SKILL.mdSave it as .claude/skills/00-arena-router/SKILL.md (or your agent's skills folder).
name
00-arena-router
description
Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom task, or a benchmark specifically about reference/citation hallucination rate. Also use when the user mentions model arena, agent arena, pairwise model comparison, win-rate ranking, or comparing models on a task and hasn't specified whether that task is generic or about citation accuracy. This skill is the entry router for the arena-eval suite: it asks one diagnostic question when needed, then recommends the workflow or workflows needed to cover the request.

Arena Eval Router

Entry router for the arena-eval suite. You diagnose what the user wants to compare models on and route them to the appropriate sub-skill or both when the request spans both evaluation goals. You don't run comparisons yourself — you're the triage desk.

Each sub-skill is self-contained: it carries inline everything it needs, so it can be installed and used on its own.

Diagnostic Question

Ask (unless the user's request already makes the answer obvious):

To route you correctly: what are you comparing the models on?

a) A custom task of your own choosing (chatbot quality, summarization,
   coding, anything) — you'll get win-rate rankings from a judge model
b) Specifically how often each model fabricates or hallucinates references
   when asked to recommend citations

Shortcut rule: if the user already said "run an arena eval on my chatbot task" or "benchmark reference hallucination across these models", skip the question — the routing is already clear from their phrasing. Also skip the question when they explicitly ask for both general quality and citation accuracy; recommend both workflows.

Triage Table

User says / hasUse workflowWhat it does
"Compare/benchmark/rank these models on [any custom task]"01-auto-arenaGenerates queries from a task description, collects responses, auto-generates rubrics, runs pairwise judge comparisons, produces win-rate rankings
"Which model hallucinates citations least?" / "benchmark reference recommendation accuracy"02-ref-hallucination-arenaRuns reference-recommendation queries per model, verifies every returned citation against CrossRef/PubMed/arXiv/DBLP, ranks by verified accuracy
"Compare general helpfulness AND citation accuracy"01-auto-arena, then 02-ref-hallucination-arenaRuns separate evaluations for judge preference and verified citation accuracy, preserving both goals
"I want to review one paper's existing bibliography, not compare models"—Not this suite — see the academic-eval suite's 01-paper-review / 02-bib-verify instead
Show full SKILL.md (150 more words)Show less

Key distinction

Both workflows produce model rankings from head-to-head-style evaluation, but differ in what "correct" means:

  • 01-auto-arena: correctness is judge opinion — an LLM judge scores pairwise which response is better for an arbitrary task. Works for any task, needs no ground truth.
  • 02-ref-hallucination-arena: correctness is externally verifiable — every cited reference is checked against real bibliographic databases (CrossRef/PubMed/arXiv/DBLP), so the ranking reflects factual accuracy, not judge preference. Narrower scope (citation recommendation only) but higher ground-truth confidence.

If the user cares only about citation accuracy, prefer 02-ref-hallucination-arena over 01-auto-arena even if they phrase it as "which model is better."

Output

Recommended workflow: `[skill-name]`

Why: [one sentence tying the user's request to the triage table row]

Recommend one workflow when it covers the request. If the user asks for both general quality and citation accuracy, recommend 01-auto-arena followed by 02-ref-hallucination-arena as separate runs (or follow the user's requested order). Explain that the two runs measure different things and report their results separately; neither ranking substitutes for the other.

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/arena-eval/00-arena-router of agentscope-ai/OpenJudge.

Open the folder on GitHubat commit d1e0642

Compare with similar skills

00 Arena Router next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

00 Arena Router compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
00 Arena Router this skillagentscope-ai/OpenJudge871—~1kAutomated safety check: PassApache-2.0
Content Research Writerweapp-tailwindcss/weapp-tailwindcss1.9k25 repos~3.5kAutomated safety check: PassMIT
Systematic Review ScreenerImbad0202/academic-research-skills51k—~8.4kAutomated safety check: PassCustom licence
NetworkxzLanqing/codex-claude-academic-skills4.7k15 repos~3.2kAutomated safety check: PassBSD-3-Clause
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Openalex Databaseneflibata-feng/MyArxiv-Agent12612 repos~3kAutomated safety check: PassCustom licence

Similar skills

  • Content Research Writer

    weapp-tailwindcss/weapp-tailwindcss

    Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.

    1.9k GitHub starsUsed in 25 repos~3.5k tokens
    Research & ScienceAuto-check passed
  • Systematic Review Screener

    Imbad0202/academic-research-skills

    Screens records for systematic, scoping and rapid reviews against fixed eligibility rules, using two blinded AI reviewers and a third adjudicator, with traceable PRISMA counts.

    51k GitHub stars~8.4k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Networkx

    zLanqing/codex-claude-academic-skills

    Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python.

    4.7k GitHub starsUsed in 15 repos~3.2k tokens
    Research & ScienceAuto-check passed
  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Openalex Database

    neflibata-feng/MyArxiv-Agent

    Query and analyze scholarly literature using the OpenAlex database.

    126 GitHub starsUsed in 12 repos~3k tokens
    Research & ScienceAuto-check passed
  • Citation Verification Guide

    Galaxy-Dawn/claude-scholar

    Reference guidance for checking every citation in academic writing against canonical sources such as DOI, arXiv, CrossRef and Semantic Scholar, to catch fake or wrong references.

    5.7k GitHub starsUsed in 2 repos~1.9k tokens
    Research & ScienceAuto-check passed

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    871 GitHub stars~3.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    871 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    871 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Claude Authenticity

    agentscope-ai/OpenJudge

    Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

    871 GitHub starsUsed in 1 repo~5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    871 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: warnings
  • Find Skills Combo

    agentscope-ai/OpenJudge

    Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.

    871 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check: warnings

Questions about 00 Arena Router

What does 00 Arena Router do?

A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom…. 00 Arena Router is an agent skill from agentscope-ai/OpenJudge. Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom task, or a benchmark specifically about reference/citation hallucination rate.

When should I use 00 Arena Router?

00 Arena Router fits situations like: the user wants to compare; A benchmark specifically about reference/citation hallucination rate; the user mentions model arena; pairwise model comparison.

How do I install 00 Arena Router in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a claude-code`. Or copy the skill folder (skills/arena-eval/00-arena-router in agentscope-ai/OpenJudge) into .claude/skills/00-arena-router in your project. Claude Code loads it when a task matches its description.

How do I install 00 Arena Router in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a codex`. Or copy the skill folder (skills/arena-eval/00-arena-router in agentscope-ai/OpenJudge) into .agents/skills/00-arena-router in your project. Codex loads it when a task matches its description.

Can I use 00 Arena Router in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/00-arena-router, .gemini/skills/00-arena-router, .github/skills/00-arena-router and .opencode/skills/00-arena-router in your project.

What does 00 Arena Router need to run?

SKILL.md names no scripts, command-line tools or credentials: 00 Arena Router is instructions for the agent only.

Does 00 Arena Router access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is 00 Arena Router safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does 00 Arena Router use?

00 Arena Router is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does 00 Arena Router use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to 00 Arena Router?

Skills that share tags, products or a category with 00 Arena Router: Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Systematic Review Screener (Imbad0202/academic-research-skills, 51k stars), Networkx (zLanqing/codex-claude-academic-skills, 4.7k stars) and Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains 00 Arena Router?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.