Agent skill

Benchmark Paper Template

by HKUSTDial in HKUSTDial/Supervisor-Skills

Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

CC-BY-4.0Auto-check passedResearch & Science

Install Benchmark Paper Template

skills CLI
$ npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUSTDial/Supervisor-Skills benchmark-paper-template --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmark-paper-template .claude/skills/benchmark-paper-template && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-paper-template
GitHub stars
8.7k
Token cost
~2.8k tokens
SKILL.md length
970 words
Files
9 (incl. references)
Skills in repo
12
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

  • Works in 4 steps: Five-pillar completeness audit: is the… → Introduction six-part logic chain:… → Section skeleton for §2 to §7: Task and… → …
  • Writing a paper that introduces a new benchmark or evaluation dataset
  • SKILL.md covers Overview, Core capabilities, Benchmark paper vs technical… and The five pillars, plus 6 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The starting premise is that a benchmark paper wins by defining a new evaluation dimension and shipping a construction pipeline that makes the measurement high quality, scalable and reproducible, not by proposing an algorithm. The skill audits an idea against five pillars: the research gap, the construction pipeline, the evaluation framework, the empirical findings, and an optional companion method.

From there it produces a six-part Introduction chain (background with a running example, limits of existing benchmarks, research questions, design considerations, the proposal and contributions), a skeleton for Sections 2 to 7, and a reviewer-style checklist with Critical, Major and Minor severities. Reference files cover benchmark design, gap analysis, the construction pipeline, experiments and paper structure. Technical and position papers are sent to `tech-paper-template`, and a stand-alone Introduction outline to `intro-drafter`.

When your agent uses it

  • Writing a paper that introduces a new benchmark or evaluation dataset
  • Checking whether a benchmark idea is substantive before investing in it
  • Drafting the Introduction of a benchmark paper
  • Planning the data-construction pipeline and experiments

Example prompts

  • “I want to write a paper on a benchmark for long-document question answering. Audit my idea against the five pillars.”
  • “Draft the six-part Introduction chain for my code-review benchmark.”
  • “Give me a pre-submission checklist for the benchmark paper draft in paper/main.tex.”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Five-pillar completeness audit: is the Research Gap articulated? Is the Construction Pipeline principled? Is the Evaluation Framework…
  2. Introduction six-part logic chain: Background + Running Example, Existing-Benchmark Limitations (no more than three), Research Questions…
  3. Section skeleton for §2 to §7: Task and Design Goals, Construction Pipeline, Optional Companion Method, Experiments organized by RQ…
  4. Pre-submission self-check: four-category reviewer checklist (Introduction, Benchmark section, Experiments, Overall) with Critical, Major…

What it can do on your machine

Read from SKILL.md and the folder at commit 207bc6f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Paper Template loads about 2.8k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 970 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~135
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUSTDial/Supervisor-Skills at commit 207bc6f, republished under its CC-BY-4.0 licence (© HKUSTDial). 970 words, ~2,794 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-paper-template/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
benchmark-paper-template
description
Structures Benchmark and Evaluation papers using the five-pillar framework (Research Gap, Construction Pipeline, Evaluation Framework, Empirical Findings, optional Companion Method). Returns a completeness audit, a six-part Introduction logic chain, a Section 2-7 skeleton, and a pre-submission checklist. Use when writing a benchmark paper, structuring a benchmark paper, checking whether a benchmark idea is substantive, drafting a benchmark Introduction, or planning the data-construction pipeline or experiments.
license
CC-BY-4.0

Benchmark Paper Template

Overview

A Benchmark paper does not win by proposing a new algorithm. It wins by defining a new evaluation dimension and shipping a construction pipeline that makes the measurement high-quality, scalable, and reproducible. This skill scaffolds the five pillars a reviewer checks, then gives you a six-part Introduction chain, a Section 2-7 skeleton, and a pre-submission checklist. Stage-specific depth lives in seven reference files under references/.

Core capabilities

  1. Five-pillar completeness audit: is the Research Gap articulated? Is the Construction Pipeline principled? Is the Evaluation Framework fine-grained? Do the Empirical Findings reveal capability boundaries? Is a Companion Method warranted?
  2. Introduction six-part logic chain: Background + Running Example, Existing-Benchmark Limitations (no more than three), Research Questions, Design Considerations, Our Proposal, Contributions.
  3. Section skeleton for §2 to §7: Task and Design Goals, Construction Pipeline, Optional Companion Method, Experiments organized by RQ, Discussion and Research Opportunities, Related Work with benchmark comparison table, Conclusion.
  4. Pre-submission self-check: four-category reviewer checklist (Introduction, Benchmark section, Experiments, Overall) with Critical, Major, Minor severity.

Benchmark paper vs technical paper

DimensionTechnical paperBenchmark paper
Main contributionNovel algorithm or methodNovel evaluation dimension or dataset
Introduction axisKey Idea or MechanismEvaluation Gap and Benchmark Design Rationale
Problem definitionOne-sentence goalThe problem definition IS the contribution
Heaviest chapterMethodConstruction Pipeline + Evaluation Framework
Experiments purposeProve "my method beats baselines"Reveal "where model capability boundaries sit"
Canonical Figure 1Method framework diagramRunning example + pipeline diagram

For technical and position papers, use the tech-paper-template skill. For the Introduction outline in isolation, use intro-drafter.

The five pillars

  1. Research Gap. What dimension of evaluation does existing work miss? Ground the gap in a concrete failure case and cite at least three prior benchmarks whose limitations you are addressing (no more than three). Exemplars: StatQA highlights missing statistical-method appropriateness; nvBench 2.0 highlights query-ambiguity blindness; VisJudge-Bench highlights the fidelity-expressiveness-aesthetics trinity in visualization evaluation.
  2. Construction Pipeline. How do you build high-quality, scalable, reproducible data? Three common paradigms: Reverse Synthesis (seed knowledge then instantiate), Controlled Injection (seed queries then inject targeted ambiguity or error), Adaptive Generation with Expert Validation. Specify source selection, generation, annotation, quality control, split strategy, and statistical profile. Deep dive: references/construction-pipeline.md.
  3. Evaluation Framework. Beyond a single overall score: difficulty tiers, error taxonomy, per-dimension rubrics. Explain why this taxonomy diagnoses what the gap pointed at. Deep dive: references/benchmark-design.md.
  4. Empirical Findings. Multi-angle comparisons (Human vs LLM, architecture families, error distributions) condensed into bolded Finding X: sentences that read like lemmas. Each Finding must be actionable for future research. Deep dive: references/experiments.md.
  5. Companion Method (optional). A specialized model tuned for this benchmark signals that the community can act on the findings. Examples: Step-Text2Vis, VisJudge. Not mandatory, but strongly recommended for benchmarks targeting mature tasks.

Introduction six-part flowchart

  1. Research Background + Running Example (Figure 1). Establish the task, why it matters, and one concrete example that threads through the entire paper.
  2. Existing-Benchmark Limitations. At most three, each specific and traceable to an evaluation blind spot. Avoid vague "is limited" phrasing.
  3. Research Questions. Two or three RQs covering construction quality, capability boundaries, and the human-AI gap.
  4. Design Considerations. What should a good benchmark for this dimension have? Quality, scale, coverage, reproducibility, contamination resistance.
  5. Our Proposal. One paragraph: the benchmark plus the companion method if any.
  6. Contributions. Typically four items: benchmark + pipeline innovation + systematic evaluation + findings or companion method.
Show full SKILL.md (403 more words)Show less

Section skeleton

  • §2 Task + Design Goals: problem formulation, goals (G1 coverage, G2 fine-grained diagnostics, G3 reproducibility, G4 contamination resistance). See references/benchmark-design.md.
  • §3 Construction Pipeline: sources, generation, annotation protocol, QC, statistical profile. Figure 2 is the canonical pipeline diagram. See references/construction-pipeline.md.
  • §4 Companion Method (optional): a specialized model whose training set is this benchmark.
  • §5 Experiments: organized by RQ. Include the Overall Performance table (typically the largest table in the paper), fine-grained analysis, a human baseline when available, and bolded Finding X: summaries. See references/experiments.md.
  • §6 Discussion + Research Opportunities: what the findings reveal and what comes next.
  • §7 Related Work: a benchmark comparison table (often labelled Table 1) is essential, either here or at the end of §1.

The full section-by-section writing guide with page budgets and figure placement is in references/paper-structure.md.

Prompt template

Paste the block below into your AI assistant with the input slots filled.

markdown
# Role
You are a senior researcher who has published multiple Benchmark papers at top venues (NeurIPS Datasets and Benchmarks Track, SIGMOD, VLDB, ICML, ICLR). You know what reviewers look for in Benchmark submissions and how those criteria differ from Technical papers.

# Task
I will give you the core information about a Benchmark or Evaluation paper. Audit it against the five-pillar framework, then produce a complete logic skeleton for the paper.

# Five pillars (all must be addressed)
1. Research Gap: which dimension of evaluation does existing work miss?
2. Construction Pipeline: how is the data built at scale without losing quality?
3. Evaluation Framework: what is the fine-grained taxonomy?
4. Empirical Findings: what capability boundary does this reveal?
5. Companion Method (optional): a specialized model tuned for this benchmark.

# Input
- Research area: [e.g., Text-to-SQL, Text-to-Visualization, code generation]
- Benchmark name: [name]
- Research gap and motivation: [the evaluation blind spot you target]
- Construction approach: [how the data is built]
- Evaluation framework: [metrics and taxonomy]
- Data scale: [number of tasks, domains, difficulty tiers]
- Key findings or insights: [one to three]

# Output

## Step 1: Five-pillar completeness table

| Pillar | Covered? | Your content | Improvement suggestion |
|---|---|---|---|
| Research Gap | Y or N | ... | ... |
| Construction Pipeline | Y or N | ... | ... |
| Evaluation Framework | Y or N | ... | ... |
| Empirical Findings | Y or N | ... | ... |
| Companion Method | Y, N, or NA | ... | ... |

## Step 2: Introduction six-part logic chain

| Part | Your content |
|---|---|
| 1. Background + Running Example | ... |
| 2. Existing-benchmark limitations (up to 3) | Limitation 1: ... | Limitation 2: ... | Limitation 3: ... |
| 3. Research Questions | RQ1: ... | RQ2: ... | RQ3 (optional): ... |
| 4. Design Considerations | ... |
| 5. Our Proposal | ... |
| 6. Contributions | 1. ... | 2. ... | 3. ... | 4. ... |

## Step 3: Section outline for §2 to §7

For each section, produce a one-paragraph sketch naming the figure or table that carries its weight.

## Step 4: Pre-submission self-check

Load `references/checklist.md` and walk the four-category checklist. Report any Critical or Major items that are unresolved.

Reference exemplars

  • StatQA (NeurIPS 2024): gap is evaluation of statistical-method appropriateness; pipeline is reverse synthesis from textbooks; finding is that LLMs often pick the statistically wrong test even when the numeric answer is computed correctly.
  • nvBench 2.0 (NeurIPS 2025): gap is query-ambiguity blindness in Text-to-Visualization; pipeline is controlled ambiguity injection; finding is that LLM output quality swings dramatically with minor wording changes, while humans navigate via clarification dialogue.
  • VisJudge-Bench (ICLR 2026): gap is the fidelity-expressiveness-aesthetics trinity in visualization quality; pipeline is expert-curated with adaptive generation; companion method is VisJudge, a specialized judge model trained on this benchmark.

Usage tips

  • Use early, at scope lock. The cheapest fix for a missing pillar is before data construction starts.
  • When the user's answer to a pillar is "we have this but it is messy", point them to the specific file in references/ rather than trying to resolve it in one turn.
  • Do not confuse this Introduction flowchart with the technical-paper flowchart; they are structurally different. For technical papers, invoke tech-paper-template.
  • For pre-submission self-check, load references/checklist.md and walk it line by line with the user.

References

© HKUSTDial, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (references) in skills/benchmark-paper-template of HKUSTDial/Supervisor-Skills.

  • SKILL.md
  • references/benchmark-design.md
  • references/checklist.md
  • references/construction-pipeline.md
  • references/experiments.md
  • references/gap-analysis.md
  • references/instantiation-template.md
  • references/orchestrator-notes.md
  • references/paper-structure.md

Open the folder on GitHubat commit 207bc6f

Compare with similar skills

Benchmark Paper Template next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Paper Template compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Paper Template this skillHKUSTDial/Supervisor-Skills8.7k—~2.8kAutomated safety check: PassCC-BY-4.0
Scholar Evaluationjimmc414/Kosmos5951 repos~2.5kAutomated safety check: PassNone
Academic Researchvoidful/academic-skills134—~887Automated safety check: PassMIT
Scientific Workflow ToolsDrugClaw/DrugClaw125—~712Automated safety check: PassApache-2.0
Academic Paper Writing PipelineImbad0202/academic-research-skills51k—~16kAutomated safety check: PassCustom licence
Academic Paper ReviewerImbad0202/academic-research-skills51k—~11kAutomated safety check: PassCustom licence

Similar skills

  • Scholar Evaluation

    jimmc414/Kosmos

    Systematic framework for evaluating scholarly and research work based on the ScholarEval methodology.

    595 GitHub starsUsed in 1 repo~2.5k tokens
    Research & ScienceAuto-check passed
  • Academic Research

    voidful/academic-skills

    Complete academic research skill suite covering the full pipeline: paper reading (read/explain papers with storytelling), idea generation (brainstorm research directions), experiment design (plan…

    134 GitHub stars~887 tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Scientific Workflow Tools

    DrugClaw/DrugClaw

    Research-method workflow guide for hypothesis framing, peer-review style critique, reproducibility planning, study-design checks, and scientific-writing structure.

    125 GitHub stars~712 tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Academic Paper Writing Pipeline

    Imbad0202/academic-research-skills

    Runs a 12-agent pipeline that plans, drafts, cites, reviews and formats academic papers, with modes for revision, rebuttals, abstracts and citation checks.

    51k GitHub stars~16k tokensUpdated today
    Research & ScienceAuto-check passed
  • Academic Paper Reviewer

    Imbad0202/academic-research-skills

    Simulates a journal peer review of a manuscript with a five-seat reviewer panel, an editorial synthesizer and several review modes.

    51k GitHub stars~11k tokensUpdated today
    Research & ScienceAuto-check passed
  • Academic Research Pipeline

    Imbad0202/academic-research-skills

    Orchestrates a ten-stage academic workflow from research to finished manuscript, including integrity checks, two rounds of peer review and revision.

    51k GitHub stars~15k tokensUpdated today
    Research & ScienceAuto-check passed

More from HKUSTDial/Supervisor-Skills

All 12 skills in this repo
  • Draw.io Diagram Reconstruction

    HKUSTDial/Supervisor-Skills

    Rebuilds a reference diagram image as an editable, high-fidelity Draw.io file, mixing native elements, SVG icons and cropped PNGs, with a batch workflow for a folder of images.

    8.7k GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Paper Introduction Drafter

    HKUSTDial/Supervisor-Skills

    Drafts the Introduction of a technical paper as six paragraphs of flowing prose, positioning the work and matching contributions to challenges, with an outline on request.

    8.7k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Rebuttal Guidance

    HKUSTDial/Supervisor-Skills

    Turns pasted peer-review comments into a per-concern rebuttal plan, with reviewer mindset matching and strategy priorities, but not the final rebuttal text.

    8.7k GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Deep Research Literature Survey

    HKUSTDial/Supervisor-Skills

    Runs a survey-grade literature investigation: fixes the research questions, searches from adversarial angles, verifies citations and writes an evidence-first report.

    8.7k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Paper Figure Designer

    HKUSTDial/Supervisor-Skills

    Advises on designing the three core figures of a technical paper, then audits them against rules for format, fonts, color and captions.

    8.7k GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Research Idea Evaluator

    HKUSTDial/Supervisor-Skills

    Evaluates a draft research idea like a top-venue reviewer and advisor, scoring it on five dimensions, checking fit with your capacity and returning a clear verdict.

    8.7k GitHub stars~3.4k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Benchmark Paper Template

What does Benchmark Paper Template do?

Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist. The starting premise is that a benchmark paper wins by defining a new evaluation dimension and shipping a construction pipeline that makes the measurement high quality, scalable and reproducible, not by proposing an algorithm. The skill audits an idea against five pillars: the research gap, the construction pipeline, the evaluation framework, the empirical findings, and an optional companion method.

When should I use Benchmark Paper Template?

Benchmark Paper Template fits situations like: writing a paper that introduces a new benchmark or evaluation dataset; checking whether a benchmark idea is substantive before investing in it; drafting the Introduction of a benchmark paper; planning the data-construction pipeline and experiments.

How do I install Benchmark Paper Template in Claude Code?

Run `npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a claude-code`. Or copy the skill folder (skills/benchmark-paper-template in HKUSTDial/Supervisor-Skills) into .claude/skills/benchmark-paper-template in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Paper Template in Codex?

Run `npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a codex`. Or copy the skill folder (skills/benchmark-paper-template in HKUSTDial/Supervisor-Skills) into .agents/skills/benchmark-paper-template in your project. Codex loads it when a task matches its description.

Can I use Benchmark Paper Template in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-paper-template, .gemini/skills/benchmark-paper-template, .github/skills/benchmark-paper-template and .opencode/skills/benchmark-paper-template in your project.

What does Benchmark Paper Template need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark Paper Template is instructions for the agent only.

Does Benchmark Paper Template access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Paper Template safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Paper Template use?

Benchmark Paper Template is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Paper Template use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Benchmark Paper Template?

Skills that share tags, products or a category with Benchmark Paper Template: Scholar Evaluation (jimmc414/Kosmos, 595 stars), Academic Research (voidful/academic-skills, 134 stars), Scientific Workflow Tools (DrugClaw/DrugClaw, 125 stars) and Academic Paper Writing Pipeline (Imbad0202/academic-research-skills, 51k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Paper Template?

HKUSTDial (a GitHub organization) maintains it in HKUSTDial/Supervisor-Skills, which has 8,650 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 5, 2026.

Source: HKUSTDial/Supervisor-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.