Create Skill Test
dotnet/skills
Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.
Automatically evaluate and compare multiple AI models or agents without pre-existing test data.
$ npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install agentscope-ai/OpenJudge 01-auto-arena --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arena-eval/01-auto-arena .claude/skills/01-auto-arena && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "01-auto-arena" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/01-auto-arena into .claude/skills/01-auto-arena/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "01-auto-arena", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/01-auto-arenaType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install agentscope-ai/OpenJudge 01-auto-arena --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/arena-eval/01-auto-arena .agents/skills/01-auto-arena && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "01-auto-arena" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/01-auto-arena into .agents/skills/01-auto-arena/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "01-auto-arena", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install agentscope-ai/OpenJudge 01-auto-arena --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/arena-eval/01-auto-arena .cursor/skills/01-auto-arena && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "01-auto-arena" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/01-auto-arena into .cursor/skills/01-auto-arena/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "01-auto-arena", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/agentscope-ai/OpenJudge.git --path skills/arena-eval/01-auto-arena--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install agentscope-ai/OpenJudge 01-auto-arena --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/arena-eval/01-auto-arena .gemini/skills/01-auto-arena && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "01-auto-arena" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/01-auto-arena into .gemini/skills/01-auto-arena/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "01-auto-arena", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install agentscope-ai/OpenJudge 01-auto-arenaInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/arena-eval/01-auto-arena .github/skills/01-auto-arena && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "01-auto-arena" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/01-auto-arena into .github/skills/01-auto-arena/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "01-auto-arena", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install agentscope-ai/OpenJudge 01-auto-arena --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/arena-eval/01-auto-arena .opencode/skills/01-auto-arena && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "01-auto-arena" agent skill from https://github.com/agentscope-ai/OpenJudge/tree/main/skills/arena-eval/01-auto-arena into .opencode/skills/01-auto-arena/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "01-auto-arena", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
01-auto-arenaAutomatically evaluate and compare multiple AI models or agents without pre-existing test data.
01 Auto Arena is an agent skill from agentscope-ai/OpenJudge. Automatically evaluate and compare multiple AI models or agents without pre-existing test data. Generates test queries from a task description, collects responses from all target endpoints, auto-generates evaluation rubrics, runs pairwise comparisons via a judge model, and produces win-rate rankings with reports and charts. Supports checkpoint resume, incremental endpoint addition, and judge model hot-swap. Use when the user asks to compare, benchmark, or rank multiple models or agents on a custom task, or run an…
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Education, covering Quizzes and assessments and Test data and fixtures. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.
6 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythonpipFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
api.openai.comdashscope.aliyuncs.comAlso links to:
agentscope-ai.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYDASHSCOPE_API_KEYANTHROPIC_API_KEYDEEPSEEK_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
01 Auto Arena loads about 2.5k tokens when it runs. Until then it costs about 139 tokens; SKILL.md has 605 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 605 words, ~2,467 tokens.
.claude/skills/01-auto-arena/SKILL.md (or your agent's skills folder).End-to-end automated model comparison using the OpenJudge AutoArenaPipeline:
# Install OpenJudge
pip install py-openjudge
# Extra dependency for auto_arena (chart generation)
pip install matplotlib| Info | Required? | Notes |
|---|---|---|
| Task description | Yes | What the models/agents should do (set in config YAML) |
| Target endpoints | Yes | At least 2 OpenAI-compatible endpoints to compare |
| Judge endpoint | Yes | Strong model for pairwise evaluation (e.g. gpt-4, qwen-max) |
| API keys | Yes | Env vars: OPENAI_API_KEY, DASHSCOPE_API_KEY, etc. |
| Number of queries | No | Default: 20 |
| Seed queries | No | Example queries to guide generation style |
| System prompts | No | Per-endpoint system prompts |
| Output directory | No | Default: ./evaluation_results |
| Report language | No | "zh" (default) or "en" |
# Run evaluation
python -m cookbooks.auto_arena --config config.yaml --save
# Use pre-generated queries
python -m cookbooks.auto_arena --config config.yaml \
--queries_file queries.json --save
# Start fresh, ignore checkpoint
python -m cookbooks.auto_arena --config config.yaml --fresh --save
# Re-run only pairwise evaluation with new judge model
# (keeps queries, responses, and rubrics)
python -m cookbooks.auto_arena --config config.yaml --rerun-judge --saveimport asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline
async def main():
pipeline = AutoArenaPipeline.from_config("config.yaml")
result = await pipeline.evaluate()
print(f"Best model: {result.best_pipeline}")
for rank, (model, win_rate) in enumerate(result.rankings, 1):
print(f"{rank}. {model}: {win_rate:.1%}")
asyncio.run(main())import asyncio
from cookbooks.auto_arena.auto_arena_pipeline import AutoArenaPipeline
from cookbooks.auto_arena.schema import OpenAIEndpoint
async def main():
pipeline = AutoArenaPipeline(
task_description="Customer service chatbot for e-commerce",
target_endpoints={
"gpt4": OpenAIEndpoint(
base_url="https://api.openai.com/v1",
api_key="sk-...",
model="gpt-4",
),
"qwen": OpenAIEndpoint(
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
api_key="sk-...",
model="qwen-max",
),
},
judge_endpoint=OpenAIEndpoint(
base_url="https://api.openai.com/v1",
api_key="sk-...",
model="gpt-4",
),
num_queries=20,
)
result = await pipeline.evaluate()
print(f"Best: {result.best_pipeline}")
asyncio.run(main())| Flag | Default | Description |
|---|---|---|
--config | — | Path to YAML configuration file (required) |
--output_dir | config value | Override output directory |
--queries_file | — | Path to pre-generated queries JSON (skip generation) |
--save | False | Save results to file |
--fresh | False | Start fresh, ignore checkpoint |
--rerun-judge | False | Re-run pairwise evaluation only (keep queries/responses/rubrics) |
task:
description: "Academic GPT assistant for research and writing tasks"
target_endpoints:
model_v1:
base_url: "https://api.openai.com/v1"
api_key: "${OPENAI_API_KEY}"
model: "gpt-4"
model_v2:
base_url: "https://api.openai.com/v1"
api_key: "${OPENAI_API_KEY}"
model: "gpt-3.5-turbo"
judge_endpoint:
base_url: "https://api.openai.com/v1"
api_key: "${OPENAI_API_KEY}"
model: "gpt-4"| Field | Required | Description |
|---|---|---|
description | Yes | Clear description of the task models will be tested on |
scenario | No | Usage scenario for additional context |
| Field | Default | Description |
|---|---|---|
base_url | — | API base URL (required) |
api_key | — | API key, supports ${ENV_VAR} (required) |
model | — | Model name (required) |
system_prompt | — | System prompt for this endpoint |
extra_params | — | Extra API params (e.g. temperature, max_tokens) |
Same fields as target_endpoints.<name>. Use a strong model (e.g. gpt-4, qwen-max) with low temperature (~0.1) for consistent judgments.
| Field | Default | Description |
|---|---|---|
num_queries | 20 | Total number of queries to generate |
seed_queries | — | Example queries to guide generation |
categories | — | Query categories with weights for stratified generation |
endpoint | judge endpoint | Custom endpoint for query generation |
queries_per_call | 10 | Queries generated per API call (1–50) |
num_parallel_batches | 3 | Parallel generation batches |
temperature | 0.9 | Sampling temperature (0.0–2.0) |
top_p | 0.95 | Top-p sampling (0.0–1.0) |
max_similarity | 0.85 | Dedup similarity threshold (0.0–1.0) |
enable_evolution | false | Enable Evol-Instruct complexity evolution |
evolution_rounds | 1 | Evolution rounds (0–3) |
complexity_levels | ["constraints", "reasoning", "edge_cases"] | Evolution strategies |
| Field | Default | Description |
|---|---|---|
max_concurrency | 10 | Max concurrent API requests |
timeout | 60 | Request timeout in seconds |
retry_times | 3 | Retry attempts for failed requests |
| Field | Default | Description |
|---|---|---|
output_dir | ./evaluation_results | Output directory |
save_queries | true | Save generated queries |
save_responses | true | Save model responses |
save_details | true | Save detailed results |
| Field | Default | Description |
|---|---|---|
enabled | false | Enable Markdown report generation |
language | "zh" | Report language: "zh" or "en" |
include_examples | 3 | Examples per section (1–10) |
chart.enabled | true | Generate win-rate chart |
chart.orientation | "horizontal" | "horizontal" or "vertical" |
chart.show_values | true | Show values on bars |
chart.highlight_best | true | Highlight best model |
chart.matrix_enabled | false | Generate win-rate matrix heatmap |
chart.format | "png" | Chart format: "png", "svg", or "pdf" |
Win rate: percentage of pairwise comparisons a model wins. Each pair is evaluated in both orders (original + swapped) to eliminate position bias.
Rankings example:
1. gpt4_baseline [################----] 80.0%
2. qwen_candidate [############--------] 60.0%
3. llama_finetuned [##########----------] 50.0%Win matrix: win_matrix[A][B] = how often model A beats model B across all queries.
The pipeline saves progress after each step. Interrupted runs resume automatically:
--fresh — ignore checkpoint, start from scratch--rerun-judge — re-run only the pairwise evaluation step (useful when switching judge models); keeps queries, responses, and rubrics intactevaluation_results/
├── evaluation_results.json # Rankings, win rates, win matrix
├── evaluation_report.md # Detailed Markdown report (if enabled)
├── win_rate_chart.png # Win-rate bar chart (if enabled)
├── win_rate_matrix.png # Matrix heatmap (if matrix_enabled)
├── queries.json # Generated test queries
├── responses.json # All model responses
├── rubrics.json # Generated evaluation rubrics
├── comparison_details.json # Pairwise comparison details
└── checkpoint.json # Pipeline checkpoint| Model prefix | Environment variable |
|---|---|
gpt-*, o1-*, o3-* | OPENAI_API_KEY |
claude-* | ANTHROPIC_API_KEY |
qwen-*, dashscope/* | DASHSCOPE_API_KEY |
deepseek-* | DEEPSEEK_API_KEY |
| Custom endpoint | set api_key + base_url in config |
© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/arena-eval/01-auto-arena of agentscope-ai/OpenJudge.
Open the folder on GitHubat commit d1e0642
01 Auto Arena next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| 01 Auto Arena this skillagentscope-ai/OpenJudge | 871 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| Create Skill Testdotnet/skills | 5.6k | 1 repos | ~6.1k | Automated safety check: Pass | MIT | |
| AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch | 67k | — | ~2k | Automated safety check: Pass | MIT | |
| AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch | 67k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Evaluationguanyang/open-agent-hub | 977 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | |
| Claude Certification Tutorrohitg00/ai-engineering-from-scratch | 67k | — | ~3k | Automated safety check: Pass | MIT |
dotnet/skills
Scaffolds eval.yaml evaluation specs for skills, custom agents, and redistributable gh-aw workflow packages in the dotnet/skills repository.
rohitg00/ai-engineering-from-scratch
Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.
rohitg00/ai-engineering-from-scratch
Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
rohitg00/ai-engineering-from-scratch
Guides a learner through one of four independent Claude certification tracks with onboarding, lessons, practice labs, mock exams and remediation.
adithya-s-k/FineEnvs
Builds a Verifiers (PrimeIntellect) variant of an RL environment.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…
agentscope-ai/OpenJudge
A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.
agentscope-ai/OpenJudge
A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.
agentscope-ai/OpenJudge
Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.
agentscope-ai/OpenJudge
A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…
agentscope-ai/OpenJudge
Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.
Automatically evaluate and compare multiple AI models or agents without pre-existing test data. 01 Auto Arena is an agent skill from agentscope-ai/OpenJudge. Automatically evaluate and compare multiple AI models or agents without pre-existing test data.
01 Auto Arena fits situations like: the user asks to compare; rank multiple models; agents on a custom task; run an arena-style evaluation.
Run `npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a claude-code`. Or copy the skill folder (skills/arena-eval/01-auto-arena in agentscope-ai/OpenJudge) into .claude/skills/01-auto-arena in your project. Claude Code loads it when a task matches its description.
Run `npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a codex`. Or copy the skill folder (skills/arena-eval/01-auto-arena in agentscope-ai/OpenJudge) into .agents/skills/01-auto-arena in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill 01-auto-arena -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/01-auto-arena, .gemini/skills/01-auto-arena, .github/skills/01-auto-arena and .opencode/skills/01-auto-arena in your project.
Going by SKILL.md and its folder, 01 Auto Arena needs the command-line tools its instructions call (python and pip) and credentials named OPENAI_API_KEY, DASHSCOPE_API_KEY, ANTHROPIC_API_KEY and DEEPSEEK_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in DASHSCOPE_API_KEY.
SKILL.md names 3 domains. In commands or code: api.openai.com and dashscope.aliyuncs.com; the agent is likely to contact these when it follows the instructions. As links in the text: agentscope-ai.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
01 Auto Arena is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 9.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with 01 Auto Arena: Create Skill Test (dotnet/skills, 5.6k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 67k stars), AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 67k stars) and Evaluation (guanyang/open-agent-hub, 977 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.
Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.