Agent skill

01 Auto Arena

by agentscope-ai in agentscope-ai/OpenJudge

Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

Apache-2.0Auto-check passedEducation

Install 01 Auto Arena

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge 01-auto-arena --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arena-eval/01-auto-arena .claude/skills/01-auto-arena && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
01-auto-arena
GitHub stars
871
Token cost
~2.5k tokens
SKILL.md length
605 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

  • Works in 6 steps: Generate queries — LLM creates diverse… → Collect responses — query all target… → Generate rubrics — LLM produces… → …
  • The user asks to compare
  • SKILL.md covers Prerequisites, Gather from user before running, Quick start and CLI options, plus 7 more sections
  • Calls python and pip; reaches api.openai.com and dashscope.aliyuncs.com; needs OPENAI_API_KEY and DASHSCOPE_API_KEY

What it does

01 Auto Arena is an agent skill from agentscope-ai/OpenJudge. Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model hot-swap. Use when the user asks to compare, benchmark, or rank multiple models or agents on a custom task, or run an…

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments and Test data and fixtures. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user asks to compare
  • Rank multiple models
  • Agents on a custom task
  • Run an arena-style evaluation

Example prompts

  • “/01-auto-arena”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • A credential in DASHSCOPE_API_KEY

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Generate queries — LLM creates diverse test queries from task description
  2. Collect responses — query all target endpoints concurrently
  3. Generate rubrics — LLM produces evaluation criteria from task + sample queries
  4. Pairwise evaluation — judge model compares every model pair (with position-bias swap)
  5. Analyze & rank — compute win rates, win matrix, and rankings
  6. Report & charts — Markdown report + win-rate bar chart + optional matrix heatmap

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.openai.com
    • dashscope.aliyuncs.com

    Also links to:

    • agentscope-ai.github.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • DASHSCOPE_API_KEY
    • ANTHROPIC_API_KEY
    • DEEPSEEK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

01 Auto Arena loads about 2.5k tokens when it runs. Until then it costs about 139 tokens; SKILL.md has 605 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~139
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 605 words, ~2,467 tokens.

Download SKILL.mdSave it as .claude/skills/01-auto-arena/SKILL.md (or your agent's skills folder).
name
01-auto-arena
description
Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model hot-swap. Use when the user asks to compare, benchmark, or rank multiple models or agents on a custom task, or run an arena-style evaluation.

Auto Arena Skill

End-to-end automated model comparison using the OpenJudge AutoArenaPipeline:

  1. Generate queries — LLM creates diverse test queries from task description
  2. Collect responses — query all target endpoints concurrently
  3. Generate rubrics — LLM produces evaluation criteria from task + sample queries
  4. Pairwise evaluation — judge model compares every model pair (with position-bias swap)
  5. Analyze & rank — compute win rates, win matrix, and rankings
  6. Report & charts — Markdown report + win-rate bar chart + optional matrix heatmap

Prerequisites

bash
# Install OpenJudge
pip install py-openjudge

# Extra dependency for auto_arena (chart generation)
pip install matplotlib

Gather from user before running

InfoRequired?Notes
Task descriptionYesWhat the models/agents should do (set in config YAML)
Target endpointsYesAt least 2 OpenAI-compatible endpoints to compare
Judge endpointYesStrong model for pairwise evaluation (e.g. gpt-4, qwen-max)
API keysYesEnv vars: OPENAI_API_KEY, DASHSCOPE_API_KEY, etc.
Number of queriesNoDefault: 20
Seed queriesNoExample queries to guide generation style
System promptsNoPer-endpoint system prompts
Output directoryNoDefault: ./evaluation_results
Report languageNo"zh" (default) or "en"

Quick start

CLI
bash
# Run evaluation
python -m cookbooks.auto_arena --config config.yaml --save

# Use pre-generated queries
python -m cookbooks.auto_arena --config config.yaml \
  --queries_file queries.json --save

# Start fresh, ignore checkpoint
python -m cookbooks.auto_arena --config config.yaml --fresh --save

# Re-run only pairwise evaluation with new judge model
# (keeps queries, responses, and rubrics)
python -m cookbooks.auto_arena --config config.yaml --rerun-judge --save
Python API
python
import asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline

async def main():
    pipeline = AutoArenaPipeline.from_config("config.yaml")
    result = await pipeline.evaluate()

    print(f"Best model: {result.best_pipeline}")
    for rank, (model, win_rate) in enumerate(result.rankings, 1):
        print(f"{rank}. {model}: {win_rate:.1%}")

asyncio.run(main())
Minimal Python API (no config file)
python
import asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline
from cookbooks.auto_arena.schema import OpenAIEndpoint

async def main():
    pipeline = AutoArenaPipeline(
        task_description="Customer service chatbot for e-commerce",
        target_endpoints={
            "gpt4": OpenAIEndpoint(
                base_url="https://api.openai.com/v1",
                api_key="sk-...",
                model="gpt-4",
            ),
            "qwen": OpenAIEndpoint(
                base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
                api_key="sk-...",
                model="qwen-max",
            ),
        },
        judge_endpoint=OpenAIEndpoint(
            base_url="https://api.openai.com/v1",
            api_key="sk-...",
            model="gpt-4",
        ),
        num_queries=20,
    )
    result = await pipeline.evaluate()
    print(f"Best: {result.best_pipeline}")

asyncio.run(main())

CLI options

FlagDefaultDescription
--config—Path to YAML configuration file (required)
--output_dirconfig valueOverride output directory
--queries_file—Path to pre-generated queries JSON (skip generation)
--saveFalseSave results to file
--freshFalseStart fresh, ignore checkpoint
--rerun-judgeFalseRe-run pairwise evaluation only (keep queries/responses/rubrics)

Minimal config file

yaml
task:
  description: "Academic GPT assistant for research and writing tasks"

target_endpoints:
  model_v1:
    base_url: "https://api.openai.com/v1"
    api_key: "${OPENAI_API_KEY}"
    model: "gpt-4"
  model_v2:
    base_url: "https://api.openai.com/v1"
    api_key: "${OPENAI_API_KEY}"
    model: "gpt-3.5-turbo"

judge_endpoint:
  base_url: "https://api.openai.com/v1"
  api_key: "${OPENAI_API_KEY}"
  model: "gpt-4"

Full config reference

task
FieldRequiredDescription
descriptionYesClear description of the task models will be tested on
scenarioNoUsage scenario for additional context
target_endpoints.<name>
FieldDefaultDescription
base_url—API base URL (required)
api_key—API key, supports ${ENV_VAR} (required)
model—Model name (required)
system_prompt—System prompt for this endpoint
extra_params—Extra API params (e.g. temperature, max_tokens)
judge_endpoint

Same fields as target_endpoints.<name>. Use a strong model (e.g. gpt-4, qwen-max) with low temperature (~0.1) for consistent judgments.

query_generation
FieldDefaultDescription
num_queries20Total number of queries to generate
seed_queries—Example queries to guide generation
categories—Query categories with weights for stratified generation
endpointjudge endpointCustom endpoint for query generation
queries_per_call10Queries generated per API call (1–50)
num_parallel_batches3Parallel generation batches
temperature0.9Sampling temperature (0.0–2.0)
top_p0.95Top-p sampling (0.0–1.0)
max_similarity0.85Dedup similarity threshold (0.0–1.0)
enable_evolutionfalseEnable Evol-Instruct complexity evolution
evolution_rounds1Evolution rounds (0–3)
complexity_levels["constraints", "reasoning", "edge_cases"]Evolution strategies
Show full SKILL.md (231 more words)Show less
evaluation
FieldDefaultDescription
max_concurrency10Max concurrent API requests
timeout60Request timeout in seconds
retry_times3Retry attempts for failed requests
output
FieldDefaultDescription
output_dir./evaluation_resultsOutput directory
save_queriestrueSave generated queries
save_responsestrueSave model responses
save_detailstrueSave detailed results
report
FieldDefaultDescription
enabledfalseEnable Markdown report generation
language"zh"Report language: "zh" or "en"
include_examples3Examples per section (1–10)
chart.enabledtrueGenerate win-rate chart
chart.orientation"horizontal""horizontal" or "vertical"
chart.show_valuestrueShow values on bars
chart.highlight_besttrueHighlight best model
chart.matrix_enabledfalseGenerate win-rate matrix heatmap
chart.format"png"Chart format: "png", "svg", or "pdf"

Interpreting results

Win rate: percentage of pairwise comparisons a model wins. Each pair is evaluated in both orders (original + swapped) to eliminate position bias.

Rankings example:

  1. gpt4_baseline       [################----] 80.0%
  2. qwen_candidate      [############--------] 60.0%
  3. llama_finetuned      [##########----------] 50.0%

Win matrix: win_matrix[A][B] = how often model A beats model B across all queries.

Checkpoint & resume

The pipeline saves progress after each step. Interrupted runs resume automatically:

  • --fresh — ignore checkpoint, start from scratch
  • --rerun-judge — re-run only the pairwise evaluation step (useful when switching judge models); keeps queries, responses, and rubrics intact
  • Adding new endpoints to config triggers incremental response collection; existing responses are preserved

Output files

evaluation_results/
├── evaluation_results.json     # Rankings, win rates, win matrix
├── evaluation_report.md        # Detailed Markdown report (if enabled)
├── win_rate_chart.png          # Win-rate bar chart (if enabled)
├── win_rate_matrix.png         # Matrix heatmap (if matrix_enabled)
├── queries.json                # Generated test queries
├── responses.json              # All model responses
├── rubrics.json                # Generated evaluation rubrics
├── comparison_details.json     # Pairwise comparison details
└── checkpoint.json             # Pipeline checkpoint

API key by model

Model prefixEnvironment variable
gpt-*, o1-*, o3-*OPENAI_API_KEY
claude-*ANTHROPIC_API_KEY
qwen-*, dashscope/*DASHSCOPE_API_KEY
deepseek-*DEEPSEEK_API_KEY
Custom endpointset api_key + base_url in config

Additional resources

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/arena-eval/01-auto-arena of agentscope-ai/OpenJudge.

Open the folder on GitHubat commit d1e0642

Compare with similar skills

01 Auto Arena next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

01 Auto Arena compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
01 Auto Arena this skillagentscope-ai/OpenJudge871—~2.5kAutomated safety check: PassApache-2.0
Create Skill Testdotnet/skills5.6k1 repos~6.1kAutomated safety check: PassMIT
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch67k—~2kAutomated safety check: PassMIT
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch67k—~2.1kAutomated safety check: PassMIT
Evaluationguanyang/open-agent-hub9772 repos~4.2kAutomated safety check: PassMIT
Claude Certification Tutorrohitg00/ai-engineering-from-scratch67k—~3kAutomated safety check: PassMIT

Similar skills

  • Create Skill Test

    dotnet/skills

    Official

    Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.

    5.6k GitHub starsUsed in 1 repo~6.1k tokens
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    67k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    67k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed
  • Claude Certification Tutor

    rohitg00/ai-engineering-from-scratch

    Guides a learner through one of four independent Claude certification tracks with onboarding, lessons, practice labs, mock exams and remediation.

    67k GitHub stars~3k tokensUpdated today
    EducationAuto-check passed
  • Generate Verifiers Env

    adithya-s-k/FineEnvs

    Builds a Verifiers (PrimeIntellect) variant of an RL environment.

    461 GitHub starsUsed in 1 repo~2.3k tokens
    EducationAuto-check passed

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    871 GitHub stars~3.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    871 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    871 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Claude Authenticity

    agentscope-ai/OpenJudge

    Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

    871 GitHub starsUsed in 1 repo~5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    871 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: warnings
  • Find Skills Combo

    agentscope-ai/OpenJudge

    Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.

    871 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check: warnings

Questions about 01 Auto Arena

What does 01 Auto Arena do?

Automatically evaluate and compare multiple AI models or agents without pre-existing test data. 01 Auto Arena is an agent skill from agentscope-ai/OpenJudge. Automatically evaluate and compare multiple AI models or agents without pre-existing test data.

When should I use 01 Auto Arena?

01 Auto Arena fits situations like: the user asks to compare; rank multiple models; agents on a custom task; run an arena-style evaluation.

How do I install 01 Auto Arena in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a claude-code`. Or copy the skill folder (skills/arena-eval/01-auto-arena in agentscope-ai/OpenJudge) into .claude/skills/01-auto-arena in your project. Claude Code loads it when a task matches its description.

How do I install 01 Auto Arena in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a codex`. Or copy the skill folder (skills/arena-eval/01-auto-arena in agentscope-ai/OpenJudge) into .agents/skills/01-auto-arena in your project. Codex loads it when a task matches its description.

Can I use 01 Auto Arena in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/01-auto-arena, .gemini/skills/01-auto-arena, .github/skills/01-auto-arena and .opencode/skills/01-auto-arena in your project.

What does 01 Auto Arena need to run?

Going by SKILL.md and its folder, 01 Auto Arena needs the command-line tools its instructions call (python and pip) and credentials named OPENAI_API_KEY, DASHSCOPE_API_KEY, ANTHROPIC_API_KEY and DEEPSEEK_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in DASHSCOPE_API_KEY.

Does 01 Auto Arena access the network?

SKILL.md names 3 domains. In commands or code: api.openai.com and dashscope.aliyuncs.com; the agent is likely to contact these when it follows the instructions. As links in the text: agentscope-ai.github.io. This is read from the text; nothing was executed.

Is 01 Auto Arena safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does 01 Auto Arena use?

01 Auto Arena is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does 01 Auto Arena use?

About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to 01 Auto Arena?

Skills that share tags, products or a category with 01 Auto Arena: Create Skill Test (dotnet/skills, 5.6k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 67k stars), AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 67k stars) and Evaluation (guanyang/open-agent-hub, 977 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains 01 Auto Arena?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.