Content Research Writer
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom…
$ npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentscope-ai/OpenJudge 00-arena-router --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arena-eval/00-arena-router .claude/skills/00-arena-router && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "00-arena-router" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/00-arena-router into .claude/skills/00-arena-router/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "00-arena-router", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/00-arena-routerType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentscope-ai/OpenJudge 00-arena-router --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/arena-eval/00-arena-router .agents/skills/00-arena-router && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "00-arena-router" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/00-arena-router into .agents/skills/00-arena-router/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "00-arena-router", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentscope-ai/OpenJudge 00-arena-router --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/arena-eval/00-arena-router .cursor/skills/00-arena-router && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "00-arena-router" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/00-arena-router into .cursor/skills/00-arena-router/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "00-arena-router", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentscope-ai/OpenJudge.git --path skills/arena-eval/00-arena-router--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentscope-ai/OpenJudge 00-arena-router --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/arena-eval/00-arena-router .gemini/skills/00-arena-router && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "00-arena-router" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/00-arena-router into .gemini/skills/00-arena-router/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "00-arena-router", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentscope-ai/OpenJudge 00-arena-routerInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/arena-eval/00-arena-router .github/skills/00-arena-router && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "00-arena-router" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/00-arena-router into .github/skills/00-arena-router/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "00-arena-router", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentscope-ai/OpenJudge 00-arena-router --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/arena-eval/00-arena-router .opencode/skills/00-arena-router && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "00-arena-router" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/00-arena-router into .opencode/skills/00-arena-router/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "00-arena-router", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
00-arena-routerA skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom…
00 Arena Router is an agent skill from agentscope-ai/OpenJudge. Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom task, or a benchmark specifically about reference/citation hallucination rate. Also use when the user mentions model arena, agent arena, pairwise model comparison, win-rate ranking, or comparing models on a task and hasn't specified whether that task is generic or about citation accuracy. This skill is the entry router for the arena-eval…
Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Citation management. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
00 Arena Router loads about 1k tokens when it runs. Until then it costs about 164 tokens; SKILL.md has 380 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 380 words, ~1,032 tokens.
.claude/skills/00-arena-router/SKILL.md (or your agent's skills folder).Entry router for the arena-eval suite. You diagnose what the user wants to
compare models on and route them to the appropriate sub-skill or both when
the request spans both evaluation goals. You don't run comparisons yourself
— you're the triage desk.
Each sub-skill is self-contained: it carries inline everything it needs, so it can be installed and used on its own.
Ask (unless the user's request already makes the answer obvious):
To route you correctly: what are you comparing the models on?
a) A custom task of your own choosing (chatbot quality, summarization,
coding, anything) — you'll get win-rate rankings from a judge model
b) Specifically how often each model fabricates or hallucinates references
when asked to recommend citationsShortcut rule: if the user already said "run an arena eval on my chatbot task" or "benchmark reference hallucination across these models", skip the question — the routing is already clear from their phrasing. Also skip the question when they explicitly ask for both general quality and citation accuracy; recommend both workflows.
| User says / has | Use workflow | What it does |
|---|---|---|
| "Compare/benchmark/rank these models on [any custom task]" | 01-auto-arena | Generates queries from a task description, collects responses, auto-generates rubrics, runs pairwise judge comparisons, produces win-rate rankings |
| "Which model hallucinates citations least?" / "benchmark reference recommendation accuracy" | 02-ref-hallucination-arena | Runs reference-recommendation queries per model, verifies every returned citation against CrossRef/PubMed/arXiv/DBLP, ranks by verified accuracy |
| "Compare general helpfulness AND citation accuracy" | 01-auto-arena, then 02-ref-hallucination-arena | Runs separate evaluations for judge preference and verified citation accuracy, preserving both goals |
| "I want to review one paper's existing bibliography, not compare models" | — | Not this suite — see the academic-eval suite's 01-paper-review / 02-bib-verify instead |
Both workflows produce model rankings from head-to-head-style evaluation, but differ in what "correct" means:
01-auto-arena: correctness is judge opinion — an LLM judge scores
pairwise which response is better for an arbitrary task. Works for any task,
needs no ground truth.02-ref-hallucination-arena: correctness is externally verifiable —
every cited reference is checked against real bibliographic databases
(CrossRef/PubMed/arXiv/DBLP), so the ranking reflects factual accuracy, not
judge preference. Narrower scope (citation recommendation only) but higher
ground-truth confidence.If the user cares only about citation accuracy, prefer
02-ref-hallucination-arena over 01-auto-arena even if they phrase it as
"which model is better."
Recommended workflow: `[skill-name]`
Why: [one sentence tying the user's request to the triage table row]Recommend one workflow when it covers the request. If the user asks for both
general quality and citation accuracy, recommend 01-auto-arena followed by
02-ref-hallucination-arena as separate runs (or follow the user's requested
order). Explain that the two runs measure different things and report their
results separately; neither ranking substitutes for the other.
© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/arena-eval/00-arena-router of agentscope-ai/OpenJudge.
Open the folder on GitHubat commit d1e0642
00 Arena Router next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| 00 Arena Router this skillagentscope-ai/OpenJudge | 871 | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Content Research Writerweapp-tailwindcss/weapp-tailwindcss | 1.9k | 25 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Systematic Review ScreenerImbad0202/academic-research-skills | 51k | — | ~8.4k | Automated safety check: Pass | Custom licence | |
| NetworkxzLanqing/codex-claude-academic-skills | 4.7k | 15 repos | ~3.2k | Automated safety check: Pass | BSD-3-Clause | |
| Literature Reviewneflibata-feng/MyArxiv-Agent | 126 | 20 repos | ~5.9k | Automated safety check: Notes | MIT | |
| Openalex Databaseneflibata-feng/MyArxiv-Agent | 126 | 12 repos | ~3k | Automated safety check: Pass | Custom licence |
weapp-tailwindcss/weapp-tailwindcss
Assists in writing high-quality content by conducting research, adding citations, improving hooks, iterating on outlines, and providing real-time feedback on each section.
Imbad0202/academic-research-skills
Screens records for systematic, scoping and rapid reviews against fixed eligibility rules, using two blinded AI reviewers and a third adjudicator, with traceable PRISMA counts.
zLanqing/codex-claude-academic-skills
Comprehensive toolkit for creating, analyzing, and visualizing complex networks and graphs in Python.
neflibata-feng/MyArxiv-Agent
Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).
neflibata-feng/MyArxiv-Agent
Query and analyze scholarly literature using the OpenAlex database.
Galaxy-Dawn/claude-scholar
Reference guidance for checking every citation in academic writing against canonical sources such as DOI, arXiv, CrossRef and Semantic Scholar, to catch fake or wrong references.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…
agentscope-ai/OpenJudge
A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.
agentscope-ai/OpenJudge
Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.
agentscope-ai/OpenJudge
A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…
agentscope-ai/OpenJudge
Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.
Categories
A skill your agent uses when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom…. 00 Arena Router is an agent skill from agentscope-ai/OpenJudge. Use when the user wants to compare or benchmark multiple LLMs/agents arena-style but it's unclear which specific workflow fits — a general-purpose win-rate comparison on a custom task, or a benchmark specifically about reference/citation hallucination rate.
00 Arena Router fits situations like: the user wants to compare; A benchmark specifically about reference/citation hallucination rate; the user mentions model arena; pairwise model comparison.
Run `npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a claude-code`. Or copy the skill folder (skills/arena-eval/00-arena-router in agentscope-ai/OpenJudge) into .claude/skills/00-arena-router in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a codex`. Or copy the skill folder (skills/arena-eval/00-arena-router in agentscope-ai/OpenJudge) into .agents/skills/00-arena-router in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill 00-arena-router -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/00-arena-router, .gemini/skills/00-arena-router, .github/skills/00-arena-router and .opencode/skills/00-arena-router in your project.
SKILL.md names no scripts, command-line tools or credentials: 00 Arena Router is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
00 Arena Router is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with 00 Arena Router: Content Research Writer (weapp-tailwindcss/weapp-tailwindcss, 1.9k stars), Systematic Review Screener (Imbad0202/academic-research-skills, 51k stars), Networkx (zLanqing/codex-claude-academic-skills, 4.7k stars) and Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.
Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.