Agent skill

01 Paper Review

by agentscope-ai in agentscope-ai/OpenJudge

Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.

Apache-2.0Auto-check passedResearch & Science

Install 01 Paper Review

skills CLI
$ npx skills add agentscope-ai/OpenJudge --skill 01-paper-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install agentscope-ai/OpenJudge 01-paper-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/agentscope-ai/OpenJudge.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/academic-eval/01-paper-review .claude/skills/01-paper-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
01-paper-review
GitHub stars
871
Token cost
~2.4k tokens
SKILL.md length
929 words
Files
2
Skills in repo
19
Repo updated
First seen
Licence
Apache-2.0

At a glance

Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.

  • Works in 5 steps: Safety check — jailbreak detection +… → Correctness — objective errors (math,… → Review — quality, novelty, significance… → …
  • The user asks to review
  • SKILL.md covers Prerequisites, Gather from user before running, Quick start and All options, plus 4 more sections
  • Calls python, pip and curl; needs OPENAI_API_KEY and ANTHROPIC_API_KEY

What it does

01 Paper Review is an agent skill from agentscope-ai/OpenJudge. Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline. Supports PDF files and LaTeX source packages (.tar.gz/.zip). Covers 10 disciplines: cs, medicine, physics, chemistry, biology, economics, psychology, environmentalscience, mathematics, socialsciences. Use when the user asks to review, evaluate, critique, or assess a research paper, check references, or verify a BibTeX file.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `reference.md`).

It sits in Research & Science, covering Peer review, Citation management and LaTeX. It works with LaTeX and OpenAI. The repository describes itself as: OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards. The licence is Apache-2.0.

When your agent uses it

  • The user asks to review
  • Assess a research paper
  • Check references
  • Verify a BibTeX file

Example prompts

  • “/01-paper-review”

Requirements

  • Python 3
  • A credential in OPENAI_API_KEY
  • A credential in ANTHROPIC_API_KEY

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Safety check — jailbreak detection + format validation
  2. Correctness — objective errors (math, logic, data inconsistencies)
  3. Review — quality, novelty, significance (score 1–6)
  4. Criticality — severity of correctness issues
  5. BibTeX verification — cross-checks references against CrossRef/arXiv/DBLP

What it can do on your machine

Read from SKILL.md and the folder at commit d1e0642. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • pip
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.litellm.ai
    • platform.openai.com
    • docs.anthropic.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • OPENAI_API_KEY
    • ANTHROPIC_API_KEY
    • DASHSCOPE_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

01 Paper Review loads about 2.4k tokens when it runs. Until then it costs about 111 tokens; SKILL.md has 929 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~111
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from agentscope-ai/OpenJudge at commit d1e0642, republished under its Apache-2.0 licence (© agentscope-ai). 929 words, ~2,376 tokens.

Download SKILL.mdSave it as .claude/skills/01-paper-review/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
01-paper-review
description
Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline. Supports PDF files and LaTeX source packages (.tar.gz/.zip). Covers 10 disciplines: cs, medicine, physics, chemistry, biology, economics, psychology, environmental_science, mathematics, social_sciences. Use when the user asks to review, evaluate, critique, or assess a research paper, check references, or verify a BibTeX file.

Paper Review Skill

Multi-stage academic paper review using the OpenJudge PaperReviewPipeline:

  1. Safety check — jailbreak detection + format validation
  2. Correctness — objective errors (math, logic, data inconsistencies)
  3. Review — quality, novelty, significance (score 1–6)
  4. Criticality — severity of correctness issues
  5. BibTeX verification — cross-checks references against CrossRef/arXiv/DBLP

Prerequisites

bash
# Install OpenJudge
pip install py-openjudge

# Extra dependency for paper_review
pip install litellm
pip install pypdfium2  # only if using vision mode (use_vision_for_pdf=True)

Gather from user before running

InfoRequired?Notes
Paper file pathYesPDF or .tar.gz/.zip TeX package
API keyYesEnv var preferred: OPENAI_API_KEY, ANTHROPIC_API_KEY, etc.
Model nameNogpt-5.2, anthropic/claude-opus-4-6, dashscope/qwen-vl-plus. See Model selection below
DisciplineNoIf not given, uses general CS/ML-oriented prompts
VenueNoe.g. "NeurIPS 2025", "The Lancet"
InstructionsNoFree-form reviewer guidance, e.g. "Focus on experimental design"
LanguageNo"en" (default) or "zh" for Simplified Chinese output
BibTeX fileNoRequired only for reference verification
CrossRef emailNoImproves API rate limits for BibTeX verification

Quick start

File type is auto-detected: .pdf → PDF review, .tar.gz/.zip → TeX review, .bib → BibTeX verification.

bash
# Basic PDF review
python -m cookbooks.paper_review paper.pdf

# With discipline and venue
python -m cookbooks.paper_review paper.pdf \
  --discipline cs --venue "NeurIPS 2025"

# Chinese output
python -m cookbooks.paper_review paper.pdf --language zh

# Custom reviewer instructions
python -m cookbooks.paper_review paper.pdf \
  --instructions "Focus on experimental design and reproducibility"

# PDF + BibTeX verification
python -m cookbooks.paper_review paper.pdf \
  --bib references.bib --email your@email.com

# Vision mode (for models that prefer images over text extraction)
python -m cookbooks.paper_review paper.pdf \
  --vision --vision_max_pages 30 --format_vision_max_pages 10

# TeX source package
python -m cookbooks.paper_review paper_source.tar.gz \
  --discipline biology --email your@email.com

# TeX source package with Chinese output and custom instructions
python -m cookbooks.paper_review paper_source.tar.gz \
  --language zh --instructions "This is a short paper, be concise"

# Verify a standalone BibTeX file
python -m cookbooks.paper_review --bib_only references.bib --email your@email.com

All options

FlagDefaultDescription
input (positional)—Path to PDF, TeX package, or .bib file
--bib_only—Path to .bib file for standalone verification (no review)
--modelgpt-4oModel name
--api_keyenv varAPI key
--base_url—Custom API endpoint — must end at /v1, not /v1/chat/completions (litellm appends the path automatically)
--discipline—Academic discipline
--venue—Target conference/journal
--instructions—Free-form reviewer guidance
--languageenOutput language: en or zh
--bib—Path to .bib file (for PDF review + reference verification)
--email—CrossRef mailto for BibTeX check
--paper_namefilename stemPaper title in report
--outputautoOutput .md report path
--no_safetyoffSkip safety checks
--no_correctnessoffSkip correctness check
--no_criticalityoffSkip criticality verification
--no_biboffSkip BibTeX verification
--visiononUse vision mode (requires pypdfium2); enabled by default
--vision_max_pages30Max pages in vision mode (0 = all)
--format_vision_max_pages10Max pages for format check (0 = use --vision_max_pages)
--timeout7500API timeout in seconds

Interpreting results

Review score (1–6):

  • 1–2: Reject (major flaws or well-known results)
  • 3: Borderline reject
  • 4: Borderline accept
  • 5–6: Accept / Strong accept

Correctness score (1–3):

  • 1: No objective errors
  • 2: Minor errors (notation, arithmetic in non-critical parts)
  • 3: Major errors (wrong proofs, core algorithm flaws)

BibTeX verification:

  • verified: found in CrossRef/arXiv/DBLP
  • suspect: title/author mismatch or not found — manual check recommended

Model selection

This pipeline uses litellm for model calls. Provider prefixes are handled automatically by the pipeline — see the table below.

IMPORTANT: The model MUST support multimodal (vision) input. PDF review uses vision mode (--vision) to render pages as images, which requires a vision-capable model. Text-only models will fail or produce empty reviews.

The --model value uses a provider/model-name convention so the pipeline knows which API endpoint to call. The table below shows the exact string to pass:

Provider--model valueEnv varNotes
OpenAIgpt-5.2, gpt-5-mini, …OPENAI_API_KEYNo prefix needed; gpt-5.2 is the current flagship vision model; check OpenAI models for the latest
Anthropicanthropic/claude-opus-4-6, anthropic/claude-sonnet-4-6, …ANTHROPIC_API_KEYUse anthropic/ prefix; claude-opus-4-6 is the current flagship; check Anthropic models for the latest
DashScope (Qwen)dashscope/qwen-vl-plus, dashscope/qwen-vl-max, …DASHSCOPE_API_KEYUse dashscope/ prefix; the pipeline auto-routes to DashScope’s OpenAI-compatible endpoint
Custom endpointbare model name--api_key + --base_urlUse the model name your endpoint expects; no prefix needed when --base_url is set

Note on prefixes: The dashscope/ and anthropic/ prefixes are interpreted by the pipeline itself — do not add them to the actual API key or base URL. For OpenAI models the bare model name (e.g. gpt-5.2) is sufficient.

If the user does not specify a model, choose one based on available API keys:

  1. DASHSCOPE_API_KEY set → use dashscope/qwen-vl-plus (vision-capable)
  2. OPENAI_API_KEY set → search web for the latest vision-capable OpenAI model and use it (currently gpt-5.2)
  3. ANTHROPIC_API_KEY set → search web for the latest vision-capable Anthropic model and use it with anthropic/ prefix (currently anthropic/claude-opus-4-6)

Vision mode is enabled by default for PDF review. Pages are rendered as images, which preserves formatting, figures, and tables. To disable, pass --no_vision (not recommended). The model must support multimodal (vision) input.

Show full SKILL.md (283 more words)Show less

Additional resources

Troubleshooting API errors

CRITICAL: When the pipeline fails with an API error, you MUST diagnose and fix the root cause. Do NOT fall back to reading the PDF as plain text yourself and calling the API manually — this bypasses the entire review pipeline and produces incorrect, incomplete results.

Diagnose by reading the full error message, then follow the checklist below:

AuthenticationError / 401
  • The API key is wrong or not set.
  • Check the correct env var for the provider (see Model selection table).
  • For DashScope: echo $DASHSCOPE_API_KEY — must be non-empty.
  • Fix: export the correct key and re-run.
NotFoundError / 404 — model not found
  • The model name string is wrong.
  • Search the web for the provider's current model list and use the exact API ID.
  • Common mistakes: using a ChatGPT UI name instead of the API ID, outdated snapshot suffix.
  • Fix: correct --model and re-run.
BadRequestError / 400
  • Often caused by --base_url ending with /v1/chat/completions instead of /v1. litellm appends the path automatically — strip everything after /v1.
  • May also indicate the model does not support vision/image input. Use a vision-capable model (see Model selection) or omit --vision.
  • Fix: correct --base_url or switch to a vision-capable model and re-run.
Connection error / endpoint not reachable
  • --base_url points to the wrong host or port.
  • Test the endpoint first: curl <base_url>/models -H "Authorization: Bearer <key>"
  • Fix: correct --base_url to the reachable endpoint and re-run.
Timeout
  • The model is taking too long (common for long PDFs with vision mode).
  • Fix: increase --timeout (default 7500 s) or reduce --vision_max_pages.
After fixing, always re-run the full pipeline command.

Never summarise or interpret the paper yourself as a substitute for a failed pipeline run.

© agentscope-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/academic-eval/01-paper-review of agentscope-ai/OpenJudge.

  • SKILL.md
  • reference.md

Open the folder on GitHubat commit d1e0642

Compare with similar skills

01 Paper Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

01 Paper Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
01 Paper Review this skillagentscope-ai/OpenJudge871—~2.4kAutomated safety check: PassApache-2.0
Academic Paper Writing PipelineImbad0202/academic-research-skills51k—~16kAutomated safety check: PassCustom licence
Paper Reading ZhMrGeDiao/paper-reading-zh103—~1.3kAutomated safety check: PassMIT
Research Writing SkillzLanqing/codex-claude-academic-skills4.7k—~1.1kAutomated safety check: PassMIT
Check Review AlignmentInternScience/DrClaw172—~1.5kAutomated safety check: PassNone
Paper CovertGRIND-Lab-Core/night_owl_research_agent106—~2.1kAutomated safety check: NotesNone

Similar skills

  • Academic Paper Writing Pipeline

    Imbad0202/academic-research-skills

    Runs a 12-agent pipeline that plans, drafts, cites, reviews and formats academic papers, with modes for revision, rebuttals, abstracts and citation checks.

    51k GitHub stars~16k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Paper Reading Zh

    MrGeDiao/paper-reading-zh

    中文论文深读、总结/TL;DR、工程拆解、按论文实现/复现、比较与证据审计。用户给出论文 PDF、链接、标题、摘要、正文或图表,或明确要读论文时使用。只有论文锚点且无附言时澄清阅读目标;只有阅读意图时问哪篇。已有论文时,“看看这篇”直接深读。不用于仅翻译、解释单个术语、生成 BibTeX、找或下载论文。

    103 GitHub stars~1.3k tokensUpdated 7 days ago
    Research & ScienceAuto-check passed
  • Research Writing Skill

    zLanqing/codex-claude-academic-skills

    Chinese-first research paper writing, revision, polishing, section drafting, rebuttal, peer-review response, thesis prose improvement, and manuscript argument planning.

    4.7k GitHub stars~1.1k tokensUpdated 5 mo ago
    Research & ScienceAuto-check passed
  • Check Review Alignment

    InternScience/DrClaw

    当用户明确要求"核查/优化综述 {主题}review.tex 的正文引用"或"运行 check-review-alignment"时使用。通过宿主 AI 的语义理解逐条核查引用是否与文献内容吻合,只在发现致命性引用错误时对"包含引用的句子"做最小化改写,并复用 systematic-literature-review 的渲染脚本输出…

    172 GitHub stars~1.5k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Paper Covert

    GRIND-Lab-Core/night_owl_research_agent

    Converts the final Markdown manuscript from paper-draft / paper-review-loop into a submission package for the target venue — modular LaTeX (one file per section), compiled PDF, and Word .docx.

    106 GitHub stars~2.1k tokensUpdated 5 mo ago
    Documents & OfficeAuto-check: notes
  • Paper Audit

    bahayonghang/academic-writing-skills

    Reviewer-style audit and submission gate for academic papers in .tex, .typ, or .pdf.

    500 GitHub stars~5.1k tokensUpdated 13 days ago
    Documents & OfficeAuto-check passed

More from agentscope-ai/OpenJudge

All 19 skills in this repo
  • Align Human

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a judge/grader and human-labeled data, and wants to measure how well the judge agrees with humans, detect systematic biases, determine whether automatic…

    871 GitHub stars~3.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Prompt Regression

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has changed a prompt (system prompt, RAG template, agent instruction, etc.) and wants to know whether the candidate is better or worse than the baseline.

    871 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • RAG Eval

    agentscope-ai/OpenJudge

    A skill your agent uses when the user has a RAG (Retrieval-Augmented Generation) system and wants to evaluate its quality — separating retrieval issues from generation issues.

    871 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Claude Authenticity

    agentscope-ai/OpenJudge

    Detect whether an API endpoint is backed by genuine Claude (not a wrapper, proxy, or impersonator) using 9 weighted rule-based checks that mirror the claude-verify project.

    871 GitHub starsUsed in 1 repo~5k tokens
    Auto-check passed
  • Eval Design

    agentscope-ai/OpenJudge

    A skill your agent uses when the user needs to design evaluation datasets, create test cases, stratify samples, generate adversarial examples, extract eval dimensions from traces/specs, or build a…

    871 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: warnings
  • Find Skills Combo

    agentscope-ai/OpenJudge

    Discover and recommend combinations of agent skills to complete complex, multi-faceted tasks.

    871 GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check: warnings

Works with

Questions about 01 Paper Review

What does 01 Paper Review do?

Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline. 01 Paper Review is an agent skill from agentscope-ai/OpenJudge. Review academic papers for correctness, quality, and novelty using OpenJudge's multi-stage pipeline.

When should I use 01 Paper Review?

01 Paper Review fits situations like: the user asks to review; assess a research paper; check references; verify a BibTeX file.

How do I install 01 Paper Review in Claude Code?

Run `npx skills add agentscope-ai/OpenJudge --skill 01-paper-review -a claude-code`. Or copy the skill folder (skills/academic-eval/01-paper-review in agentscope-ai/OpenJudge) into .claude/skills/01-paper-review in your project. Claude Code loads it when a task matches its description.

How do I install 01 Paper Review in Codex?

Run `npx skills add agentscope-ai/OpenJudge --skill 01-paper-review -a codex`. Or copy the skill folder (skills/academic-eval/01-paper-review in agentscope-ai/OpenJudge) into .agents/skills/01-paper-review in your project. Codex loads it when a task matches its description.

Can I use 01 Paper Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add agentscope-ai/OpenJudge --skill 01-paper-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/01-paper-review, .gemini/skills/01-paper-review, .github/skills/01-paper-review and .opencode/skills/01-paper-review in your project.

What does 01 Paper Review need to run?

Going by SKILL.md and its folder, 01 Paper Review needs the command-line tools its instructions call (python, pip and curl) and credentials named OPENAI_API_KEY, ANTHROPIC_API_KEY and DASHSCOPE_API_KEY. Our summary lists: Python 3; A credential in OPENAI_API_KEY; A credential in ANTHROPIC_API_KEY.

Does 01 Paper Review access the network?

SKILL.md names 3 domains. As links in the text: docs.litellm.ai, platform.openai.com and docs.anthropic.com. This is read from the text; nothing was executed.

Is 01 Paper Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does 01 Paper Review use?

01 Paper Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does 01 Paper Review use?

About 2.4k tokens (SKILL.md is roughly 9.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to 01 Paper Review?

Skills that share tags, products or a category with 01 Paper Review: Academic Paper Writing Pipeline (Imbad0202/academic-research-skills, 51k stars), Paper Reading Zh (MrGeDiao/paper-reading-zh, 103 stars), Research Writing Skill (zLanqing/codex-claude-academic-skills, 4.7k stars) and Check Review Alignment (InternScience/DrClaw, 172 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains 01 Paper Review?

agentscope-ai (a GitHub organization) maintains it in agentscope-ai/OpenJudge, which has 871 GitHub stars. The repository holds 19 skills in this directory. The repository was last updated on September 11, 2026.

Source: agentscope-ai/OpenJudge on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.