Agent skill

02 Ref Hallucination Arena

by agentscope-ai in agentscope-ai/OpenJudge

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP.

Apache-2.0Auto-check passedResearch & Science

Install 02 Ref Hallucination Arena

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill 02-ref-hallucination-arena -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge 02-ref-hallucination-arena --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/arena-eval/02-ref-hallucination-arena .claude/skills/02-ref-hallucination-arena && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
02-ref-hallucination-arena
GitHub stars
871
Token cost
~2.4k tokens
SKILL.md length
602 words
Files
1
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP.

  • Works in 6 steps: Load queries — from JSON/JSONL dataset → Collect responses — BibTeX-formatted… → Extract references — parse BibTeX… → …
  • The user asks to evaluate
  • SKILL.md covers Prerequisites, Gather from user before running, Quick start and CLI options, plus 7 more sections
  • Calls python and pip; reaches api.openai.com and dashscope.aliyuncs.com; needs OPENAI_API_KEY and DASHSCOPE_API_KEY

What it does

02 Ref Hallucination Arena is an agent skill from agentscope-ai/OpenJudge. Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mode. Use when the user asks to evaluate, benchmark, or compare models on academic reference hallucination, literature recommendation quality, or citation accuracy.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Academic paper search, Web search and Citation management. It works with PubMed, arXiv and React. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user asks to evaluate
  • Compare models on academic reference hallucination
  • Literature recommendation quality
  • Citation accuracy

Example prompts

  • “/02-ref-hallucination-arena”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • A credential in DASHSCOPE_API_KEY

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Load queries — from JSON/JSONL dataset
  2. Collect responses — BibTeX-formatted references from target models
  3. Extract references — parse BibTeX entries from model output
  4. Verify references — cross-check against Crossref / PubMed / arXiv / DBLP
  5. Score & rank — compute verification rate, per-field accuracy, discipline breakdown
  6. Generate report — Markdown report + visualization charts

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • api.openai.com
    • dashscope.aliyuncs.com

    Also links to:

    • huggingface.co
    • openjudge.me

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • DASHSCOPE_API_KEY
    • ANTHROPIC_API_KEY
    • DEEPSEEK_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

02 Ref Hallucination Arena loads about 2.4k tokens when it runs. Until then it costs about 122 tokens; SKILL.md has 602 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~122
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 602 words, ~2,424 tokens.

Download SKILL.mdSave it as .claude/skills/02-ref-hallucination-arena/SKILL.md (or your agent's skills folder).
name
02-ref-hallucination-arena
description
Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. Measures hallucination rate, per-field accuracy (title/author/year/DOI), discipline breakdown, and year constraint compliance. Supports tool-augmented (ReAct + web search) mode. Use when the user asks to evaluate, benchmark, or compare models on academic reference hallucination, literature recommendation quality, or citation accuracy.

Reference Hallucination Arena Skill

Evaluate how accurately LLMs recommend real academic references using the OpenJudge RefArenaPipeline:

  1. Load queries — from JSON/JSONL dataset
  2. Collect responses — BibTeX-formatted references from target models
  3. Extract references — parse BibTeX entries from model output
  4. Verify references — cross-check against Crossref / PubMed / arXiv / DBLP
  5. Score & rank — compute verification rate, per-field accuracy, discipline breakdown
  6. Generate report — Markdown report + visualization charts

Prerequisites

bash
# Install OpenJudge
pip install py-openjudge

# Extra dependency for ref_hallucination_arena (chart generation)
pip install matplotlib

Gather from user before running

InfoRequired?Notes
Config YAML pathYesDefines endpoints, dataset, verification settings
Dataset pathYesJSON/JSONL file with queries (can be set in config)
API keysYesEnv vars: OPENAI_API_KEY, DASHSCOPE_API_KEY, etc.
CrossRef emailNoImproves API rate limits for verification
PubMed API keyNoImproves PubMed rate limits
Output directoryNoDefault: ./evaluation_results/ref_hallucination_arena
Report languageNo"en" (default) or "zh"
Tavily API keyNoRequired only if using tool-augmented mode

Quick start

CLI
bash
# Run evaluation with config file
python -m cookbooks.ref_hallucination_arena --config config.yaml --save

# Resume from checkpoint (default behavior)
python -m cookbooks.ref_hallucination_arena --config config.yaml --save

# Start fresh, ignore checkpoint
python -m cookbooks.ref_hallucination_arena --config config.yaml --fresh --save

# Override output directory
python -m cookbooks.ref_hallucination_arena --config config.yaml \
  --output_dir ./my_results --save
Python API
python
import asyncio
from cookbooks.ref_hallucination_arena.pipeline import RefArenaPipeline

async def main():
    pipeline = RefArenaPipeline.from_config("config.yaml")
    result = await pipeline.evaluate()

    for rank, (model, score) in enumerate(result.rankings, 1):
        print(f"{rank}. {model}: {score:.1%}")

asyncio.run(main())

CLI options

FlagDefaultDescription
--config—Path to YAML configuration file (required)
--output_dirconfig valueOverride output directory
--saveFalseSave results to file
--freshFalseStart fresh, ignore checkpoint

Minimal config file

yaml
task:
  description: "Evaluate LLM reference recommendation capabilities"

dataset:
  path: "./data/queries.json"

target_endpoints:
  model_a:
    base_url: "https://api.openai.com/v1"
    api_key: "${OPENAI_API_KEY}"
    model: "gpt-4"
    system_prompt: "You are an academic literature recommendation expert. Recommend {num_refs} real papers in BibTeX format. Only recommend papers you are confident actually exist."

  model_b:
    base_url: "https://dashscope.aliyuncs.com/compatible-mode/v1"
    api_key: "${DASHSCOPE_API_KEY}"
    model: "qwen3-max"
    system_prompt: "You are an academic literature recommendation expert. Recommend {num_refs} real papers in BibTeX format. Only recommend papers you are confident actually exist."

Full config reference

task
FieldRequiredDescription
descriptionYesEvaluation task description
scenarioNoUsage scenario
dataset
FieldDefaultDescription
path—Path to JSON/JSONL dataset file (required)
shufflefalseShuffle queries before evaluation
max_queriesnullMax queries to use (null = all)
target_endpoints.<name>
FieldDefaultDescription
base_url—API base URL (required)
api_key—API key, supports ${ENV_VAR} (required)
model—Model name (required)
system_promptbuilt-inSystem prompt; use {num_refs} placeholder
max_concurrency5Max concurrent requests for this endpoint
extra_params—Extra API request params (e.g. temperature)
tool_config.enabledfalseEnable ReAct agent with Tavily web search
tool_config.tavily_api_keyenv varTavily API key
tool_config.max_iterations10Max ReAct iterations (1–30)
tool_config.search_depth"advanced""basic" or "advanced"
verification
FieldDefaultDescription
crossref_mailto—Email for Crossref polite pool
pubmed_api_key—PubMed API key
max_workers10Concurrent verification threads (1–50)
timeout30Per-request timeout in seconds
verified_threshold0.7Min composite score to count as VERIFIED
evaluation
FieldDefaultDescription
timeout120Model API request timeout in seconds
retry_times3Number of retry attempts
output
FieldDefaultDescription
output_dir./evaluation_results/ref_hallucination_arenaOutput directory
save_queriestrueSave loaded queries
save_responsestrueSave model responses
save_detailstrueSave verification details
Show full SKILL.md (238 more words)Show less
report
FieldDefaultDescription
enabledtrueEnable report generation
language"zh"Report language: "zh" or "en"
include_examples3Examples per section (1–10)
chart.enabledtrueGenerate charts
chart.orientation"vertical""horizontal" or "vertical"
chart.show_valuestrueShow values on bars
chart.highlight_besttrueHighlight best model

Dataset format

Each query in the JSON/JSONL dataset:

json
{
  "query": "Please recommend papers on Transformer architectures for NLP.",
  "discipline": "computer_science",
  "num_refs": 5,
  "language": "en",
  "year_constraint": {"min_year": 2020}
}
FieldRequiredDescription
queryYesPrompt for reference recommendation
disciplineNocomputer_science, biomedical, physics, chemistry, social_science, interdisciplinary, other
num_refsNoExpected number of references (default: 5)
languageNo"zh" or "en" (default: "zh")
year_constraintNo{"exact": 2023}, {"min_year": 2020}, {"max_year": 2015}, or {"min_year": 2020, "max_year": 2024}

Official dataset: OpenJudge/ref-hallucination-arena

Interpreting results

Overall accuracy (verification rate):

  • > 75% — Excellent: model rarely hallucinates references
  • 60–75% — Good: most references are real, some fabrication
  • 40–60% — Fair: significant hallucination, use with caution
  • < 40% — Poor: model frequently fabricates references

Per-field accuracy:

  • title_accuracy — % of titles matching real papers
  • author_accuracy — % of correct author lists
  • year_accuracy — % of correct publication years
  • doi_accuracy — % of valid DOIs

Verification status:

  • VERIFIED — title + author + year all exactly match a real paper
  • SUSPECT — partial match (e.g. title matches but authors differ)
  • NOT_FOUND — no match in any database
  • ERROR — API timeout or network failure

Ranking order: overall accuracy → year compliance rate → avg confidence → completeness

Output files

evaluation_results/ref_hallucination_arena/
├── evaluation_report.md          # Detailed Markdown report
├── evaluation_results.json       # Rankings, per-field accuracy, scores
├── verification_chart.png        # Per-field accuracy bar chart
├── discipline_chart.png          # Per-discipline accuracy chart
├── queries.json                  # Loaded evaluation queries
├── responses.json                # Raw model responses
├── extracted_refs.json           # Extracted BibTeX references
├── verification_results.json     # Per-reference verification details
└── checkpoint.json               # Pipeline checkpoint for resume

API key by model

Model prefixEnvironment variable
gpt-*, o1-*, o3-*OPENAI_API_KEY
claude-*ANTHROPIC_API_KEY
qwen-*, dashscope/*DASHSCOPE_API_KEY
deepseek-*DEEPSEEK_API_KEY
Custom endpointset api_key + base_url in config

Additional resources

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/arena-eval/02-ref-hallucination-arena of agentscope-ai/OpenJudge.

Open the folder on GitHubat commit d1e0642

Compare with similar skills

02 Ref Hallucination Arena next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

02 Ref Hallucination Arena compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
02 Ref Hallucination Arena this skillagentscope-ai/OpenJudge871—~2.4kAutomated safety check: PassApache-2.0
Literature Reviewneflibata-feng/MyArxiv-Agent12620 repos~5.9kAutomated safety check: NotesMIT
Citation ManagementK-Dense-AI/claude-scientific-writer2.4k2 repos~3.9kAutomated safety check: NotesMIT
Citation Managementneflibata-feng/MyArxiv-Agent12619 repos~8.1kAutomated safety check: NotesMIT
Literature Review AgentAr9av/PaperOrchestra6791 repos~5.2kAutomated safety check: PassCustom licence
Academic Search and Citation RouterYuan1z0825/nature-skills47k—~884Automated safety check: PassApache-2.0

Similar skills

  • Literature Review

    neflibata-feng/MyArxiv-Agent

    Conduct comprehensive, systematic literature reviews using multiple academic databases (PubMed, arXiv, bioRxiv, Semantic Scholar, etc.).

    126 GitHub starsUsed in 20 repos~5.9k tokens
    Research & ScienceAuto-check: notes
  • Citation Management

    K-Dense-AI/claude-scientific-writer

    Finds papers in OpenAlex, PubMed and Google Scholar, turns DOIs, PMIDs and arXiv IDs into clean BibTeX, and validates citations for a manuscript or thesis.

    2.4k GitHub starsUsed in 2 repos~3.9k tokens
    Research & ScienceAuto-check: notes
  • Citation Management

    neflibata-feng/MyArxiv-Agent

    Comprehensive citation management for academic research. An agent skill from neflibata-feng/MyArxiv-Agent.

    126 GitHub starsUsed in 19 repos~8.1k tokens
    Research & ScienceAuto-check: notes
  • Literature Review Agent

    Ar9av/PaperOrchestra

    Step 3 of the PaperOrchestra pipeline (arXiv:2604.05018). An agent skill from Ar9av/PaperOrchestra.

    679 GitHub starsUsed in 1 repo~5.2k tokens
    Research & ScienceAuto-check passed
  • Academic Search and Citation Router

    Yuan1z0825/nature-skills

    Finds papers across literature sources, verifies and converts citations, builds MeSH strategies and audits independent citations of a paper.

    47k GitHub stars~884 tokensUpdated today
    Research & ScienceAuto-check passed
  • Nature Academic Search

    jing1312/nature-figure-skill

    Multi-source literature search, citation verification, MeSH search strategy, citation file management (.nbib/.ris/.bib conversion), and reference management (BibTeX, related articles, ID conversion)…

    171 GitHub stars~1.3k tokensUpdated 1 mo ago
    Research & ScienceAuto-check: notes

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    871 GitHub stars~3.1k tokensUpdated 29 days ago
    Auto-check passed
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    871 GitHub stars~2.8k tokensUpdated 29 days ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    871 GitHub stars~2.4k tokensUpdated 29 days ago
    Auto-check passed
  • Claude Authenticity

    agentscope-ai/OpenJudge

    Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

    871 GitHub starsUsed in 1 repo~5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    871 GitHub stars~2.8k tokensUpdated 29 days ago
    Auto-check: warnings
  • Find Skills Combo

    agentscope-ai/OpenJudge

    Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.

    871 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check: warnings

Questions about 02 Ref Hallucination Arena

What does 02 Ref Hallucination Arena do?

Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP. 02 Ref Hallucination Arena is an agent skill from agentscope-ai/OpenJudge. Benchmark LLM reference recommendation capabilities by verifying every cited paper against Crossref, PubMed, arXiv, and DBLP.

When should I use 02 Ref Hallucination Arena?

02 Ref Hallucination Arena fits situations like: the user asks to evaluate; compare models on academic reference hallucination; literature recommendation quality; citation accuracy.

How do I install 02 Ref Hallucination Arena in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill 02-ref-hallucination-arena -a claude-code`. Or copy the skill folder (skills/arena-eval/02-ref-hallucination-arena in agentscope-ai/OpenJudge) into .claude/skills/02-ref-hallucination-arena in your project. Claude Code loads it when a task matches its description.

How do I install 02 Ref Hallucination Arena in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill 02-ref-hallucination-arena -a codex`. Or copy the skill folder (skills/arena-eval/02-ref-hallucination-arena in agentscope-ai/OpenJudge) into .agents/skills/02-ref-hallucination-arena in your project. Codex loads it when a task matches its description.

Can I use 02 Ref Hallucination Arena in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill 02-ref-hallucination-arena -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/02-ref-hallucination-arena, .gemini/skills/02-ref-hallucination-arena, .github/skills/02-ref-hallucination-arena and .opencode/skills/02-ref-hallucination-arena in your project.

What does 02 Ref Hallucination Arena need to run?

Going by SKILL.md and its folder, 02 Ref Hallucination Arena needs the command-line tools its instructions call (python and pip) and credentials named OPENAI_API_KEY, DASHSCOPE_API_KEY, ANTHROPIC_API_KEY and DEEPSEEK_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in DASHSCOPE_API_KEY.

Does 02 Ref Hallucination Arena access the network?

SKILL.md names 4 domains. In commands or code: api.openai.com and dashscope.aliyuncs.com; the agent is likely to contact these when it follows the instructions. As links in the text: huggingface.co and openjudge.me. This is read from the text; nothing was executed.

Is 02 Ref Hallucination Arena safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does 02 Ref Hallucination Arena use?

02 Ref Hallucination Arena is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does 02 Ref Hallucination Arena use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to 02 Ref Hallucination Arena?

Skills that share tags, products or a category with 02 Ref Hallucination Arena: Literature Review (neflibata-feng/MyArxiv-Agent, 126 stars), Citation Management (K-Dense-AI/claude-scientific-writer, 2.4k stars), Citation Management (neflibata-feng/MyArxiv-Agent, 126 stars) and Literature Review Agent (Ar9av/PaperOrchestra, 679 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains 02 Ref Hallucination Arena?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.