Scholar Evaluation
jimmc414/Kosmos
Systematic framework for evaluating scholarly and research work based on the ScholarEval methodology.
Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.
$ npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install HKUSTDial/Supervisor-Skills benchmark-paper-template --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/benchmark-paper-template .claude/skills/benchmark-paper-template && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-paper-template" agent skill from https://github.com/HKUSTDial/Supervisor-Skills/tree/main/skills/benchmark-paper-template into .claude/skills/benchmark-paper-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-paper-template", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/HKUSTDial/Supervisor-Skills/tree/main/skills/benchmark-paper-templateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install HKUSTDial/Supervisor-Skills benchmark-paper-template --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/benchmark-paper-template .agents/skills/benchmark-paper-template && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-paper-template" agent skill from https://github.com/HKUSTDial/Supervisor-Skills/tree/main/skills/benchmark-paper-template into .agents/skills/benchmark-paper-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-paper-template", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install HKUSTDial/Supervisor-Skills benchmark-paper-template --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/benchmark-paper-template .cursor/skills/benchmark-paper-template && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-paper-template" agent skill from https://github.com/HKUSTDial/Supervisor-Skills/tree/main/skills/benchmark-paper-template into .cursor/skills/benchmark-paper-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-paper-template", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/HKUSTDial/Supervisor-Skills.git --path skills/benchmark-paper-template--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install HKUSTDial/Supervisor-Skills benchmark-paper-template --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/benchmark-paper-template .gemini/skills/benchmark-paper-template && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-paper-template" agent skill from https://github.com/HKUSTDial/Supervisor-Skills/tree/main/skills/benchmark-paper-template into .gemini/skills/benchmark-paper-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-paper-template", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install HKUSTDial/Supervisor-Skills benchmark-paper-templateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/benchmark-paper-template .github/skills/benchmark-paper-template && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-paper-template" agent skill from https://github.com/HKUSTDial/Supervisor-Skills/tree/main/skills/benchmark-paper-template into .github/skills/benchmark-paper-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-paper-template", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install HKUSTDial/Supervisor-Skills benchmark-paper-template --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/benchmark-paper-template .opencode/skills/benchmark-paper-template && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-paper-template" agent skill from https://github.com/HKUSTDial/Supervisor-Skills/tree/main/skills/benchmark-paper-template into .opencode/skills/benchmark-paper-template/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-paper-template", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-paper-templateStructures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.
The starting premise is that a benchmark paper wins by defining a new evaluation dimension and shipping a construction pipeline that makes the measurement high quality, scalable and reproducible, not by proposing an algorithm. The skill audits an idea against five pillars: the research gap, the construction pipeline, the evaluation framework, the empirical findings, and an optional companion method.
From there it produces a six-part Introduction chain (background with a running example, limits of existing benchmarks, research questions, design considerations, the proposal and contributions), a skeleton for Sections 2 to 7, and a reviewer-style checklist with Critical, Major and Minor severities. Reference files cover benchmark design, gap analysis, the construction pipeline, experiments and paper structure. Technical and position papers are sent to `tech-paper-template`, and a stand-alone Introduction outline to `intro-drafter`.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 207bc6f. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Paper Template loads about 2.8k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 135 tokens; SKILL.md has 970 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from HKUSTDial/Supervisor-Skills at commit 207bc6f, republished under its CC-BY-4.0 licence (© HKUSTDial). 970 words, ~2,794 tokens.
.claude/skills/benchmark-paper-template/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.A Benchmark paper does not win by proposing a new algorithm. It wins by defining a new evaluation dimension and shipping a construction pipeline that makes the measurement high-quality, scalable, and reproducible. This skill scaffolds the five pillars a reviewer checks, then gives you a six-part Introduction chain, a Section 2-7 skeleton, and a pre-submission checklist. Stage-specific depth lives in seven reference files under references/.
| Dimension | Technical paper | Benchmark paper |
|---|---|---|
| Main contribution | Novel algorithm or method | Novel evaluation dimension or dataset |
| Introduction axis | Key Idea or Mechanism | Evaluation Gap and Benchmark Design Rationale |
| Problem definition | One-sentence goal | The problem definition IS the contribution |
| Heaviest chapter | Method | Construction Pipeline + Evaluation Framework |
| Experiments purpose | Prove "my method beats baselines" | Reveal "where model capability boundaries sit" |
| Canonical Figure 1 | Method framework diagram | Running example + pipeline diagram |
For technical and position papers, use the tech-paper-template skill. For the Introduction outline in isolation, use intro-drafter.
references/construction-pipeline.md.references/benchmark-design.md.references/experiments.md.references/benchmark-design.md.references/construction-pipeline.md.references/experiments.md.The full section-by-section writing guide with page budgets and figure placement is in references/paper-structure.md.
Paste the block below into your AI assistant with the input slots filled.
# Role
You are a senior researcher who has published multiple Benchmark papers at top venues (NeurIPS Datasets and Benchmarks Track, SIGMOD, VLDB, ICML, ICLR). You know what reviewers look for in Benchmark submissions and how those criteria differ from Technical papers.
# Task
I will give you the core information about a Benchmark or Evaluation paper. Audit it against the five-pillar framework, then produce a complete logic skeleton for the paper.
# Five pillars (all must be addressed)
1. Research Gap: which dimension of evaluation does existing work miss?
2. Construction Pipeline: how is the data built at scale without losing quality?
3. Evaluation Framework: what is the fine-grained taxonomy?
4. Empirical Findings: what capability boundary does this reveal?
5. Companion Method (optional): a specialized model tuned for this benchmark.
# Input
- Research area: [e.g., Text-to-SQL, Text-to-Visualization, code generation]
- Benchmark name: [name]
- Research gap and motivation: [the evaluation blind spot you target]
- Construction approach: [how the data is built]
- Evaluation framework: [metrics and taxonomy]
- Data scale: [number of tasks, domains, difficulty tiers]
- Key findings or insights: [one to three]
# Output
## Step 1: Five-pillar completeness table
| Pillar | Covered? | Your content | Improvement suggestion |
|---|---|---|---|
| Research Gap | Y or N | ... | ... |
| Construction Pipeline | Y or N | ... | ... |
| Evaluation Framework | Y or N | ... | ... |
| Empirical Findings | Y or N | ... | ... |
| Companion Method | Y, N, or NA | ... | ... |
## Step 2: Introduction six-part logic chain
| Part | Your content |
|---|---|
| 1. Background + Running Example | ... |
| 2. Existing-benchmark limitations (up to 3) | Limitation 1: ... | Limitation 2: ... | Limitation 3: ... |
| 3. Research Questions | RQ1: ... | RQ2: ... | RQ3 (optional): ... |
| 4. Design Considerations | ... |
| 5. Our Proposal | ... |
| 6. Contributions | 1. ... | 2. ... | 3. ... | 4. ... |
## Step 3: Section outline for §2 to §7
For each section, produce a one-paragraph sketch naming the figure or table that carries its weight.
## Step 4: Pre-submission self-check
Load `references/checklist.md` and walk the four-category checklist. Report any Critical or Major items that are unresolved.references/ rather than trying to resolve it in one turn.tech-paper-template.references/checklist.md and walk it line by line with the user.references/gap-analysis.md: systematic identification of the evaluation blind spot.references/benchmark-design.md: design goals, task scope, taxonomy patterns, evaluation framework.references/construction-pipeline.md: the three construction paradigms, pipeline stages, quality control.references/experiments.md: baseline selection, RQ-driven analysis, Finding X pattern, case studies.references/paper-structure.md: section-by-section writing with page budgets and figure placement.references/checklist.md: four-category pre-submission checklist with severity classification.references/instantiation-template.md: fillable template for instantiating this thinking model on your paper.references/orchestrator-notes.md: historical notes from the earlier staged orchestrator architecture, kept for context.© HKUSTDial, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 8 other files (references) in skills/benchmark-paper-template of HKUSTDial/Supervisor-Skills.
Open the folder on GitHubat commit 207bc6f
Benchmark Paper Template next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Paper Template this skillHKUSTDial/Supervisor-Skills | 8.7k | — | ~2.8k | Automated safety check: Pass | CC-BY-4.0 | |
| Scholar Evaluationjimmc414/Kosmos | 595 | 1 repos | ~2.5k | Automated safety check: Pass | None | |
| Academic Researchvoidful/academic-skills | 134 | — | ~887 | Automated safety check: Pass | MIT | |
| Scientific Workflow ToolsDrugClaw/DrugClaw | 125 | — | ~712 | Automated safety check: Pass | Apache-2.0 | |
| Academic Paper Writing PipelineImbad0202/academic-research-skills | 51k | — | ~16k | Automated safety check: Pass | Custom licence | |
| Academic Paper ReviewerImbad0202/academic-research-skills | 51k | — | ~11k | Automated safety check: Pass | Custom licence |
jimmc414/Kosmos
Systematic framework for evaluating scholarly and research work based on the ScholarEval methodology.
voidful/academic-skills
Complete academic research skill suite covering the full pipeline: paper reading (read/explain papers with storytelling), idea generation (brainstorm research directions), experiment design (plan…
DrugClaw/DrugClaw
Research-method workflow guide for hypothesis framing, peer-review style critique, reproducibility planning, study-design checks, and scientific-writing structure.
Imbad0202/academic-research-skills
Runs a 12-agent pipeline that plans, drafts, cites, reviews and formats academic papers, with modes for revision, rebuttals, abstracts and citation checks.
Imbad0202/academic-research-skills
Simulates a journal peer review of a manuscript with a five-seat reviewer panel, an editorial synthesizer and several review modes.
Imbad0202/academic-research-skills
Orchestrates a ten-stage academic workflow from research to finished manuscript, including integrity checks, two rounds of peer review and revision.
HKUSTDial/Supervisor-Skills
Rebuilds a reference diagram image as an editable, high-fidelity Draw.io file, mixing native elements, SVG icons and cropped PNGs, with a batch workflow for a folder of images.
HKUSTDial/Supervisor-Skills
Drafts the Introduction of a technical paper as six paragraphs of flowing prose, positioning the work and matching contributions to challenges, with an outline on request.
HKUSTDial/Supervisor-Skills
Turns pasted peer-review comments into a per-concern rebuttal plan, with reviewer mindset matching and strategy priorities, but not the final rebuttal text.
HKUSTDial/Supervisor-Skills
Runs a survey-grade literature investigation: fixes the research questions, searches from adversarial angles, verifies citations and writes an evidence-first report.
HKUSTDial/Supervisor-Skills
Advises on designing the three core figures of a technical paper, then audits them against rules for format, fonts, color and captions.
HKUSTDial/Supervisor-Skills
Evaluates a draft research idea like a top-venue reviewer and advisor, scoring it on five dimensions, checking fit with your capacity and returning a clear verdict.
Categories
Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist. The starting premise is that a benchmark paper wins by defining a new evaluation dimension and shipping a construction pipeline that makes the measurement high quality, scalable and reproducible, not by proposing an algorithm. The skill audits an idea against five pillars: the research gap, the construction pipeline, the evaluation framework, the empirical findings, and an optional companion method.
Benchmark Paper Template fits situations like: writing a paper that introduces a new benchmark or evaluation dataset; checking whether a benchmark idea is substantive before investing in it; drafting the Introduction of a benchmark paper; planning the data-construction pipeline and experiments.
Run `npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a claude-code`. Or copy the skill folder (skills/benchmark-paper-template in HKUSTDial/Supervisor-Skills) into .claude/skills/benchmark-paper-template in your project. Claude Code loads it when a task matches its description.
Run `npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a codex`. Or copy the skill folder (skills/benchmark-paper-template in HKUSTDial/Supervisor-Skills) into .agents/skills/benchmark-paper-template in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUSTDial/Supervisor-Skills --skill benchmark-paper-template -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-paper-template, .gemini/skills/benchmark-paper-template, .github/skills/benchmark-paper-template and .opencode/skills/benchmark-paper-template in your project.
SKILL.md names no scripts, command-line tools or credentials: Benchmark Paper Template is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Benchmark Paper Template is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Benchmark Paper Template: Scholar Evaluation (jimmc414/Kosmos, 595 stars), Academic Research (voidful/academic-skills, 134 stars), Scientific Workflow Tools (DrugClaw/DrugClaw, 125 stars) and Academic Paper Writing Pipeline (Imbad0202/academic-research-skills, 51k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
HKUSTDial (a GitHub organization) maintains it in HKUSTDial/Supervisor-Skills, which has 8,650 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 5, 2026.
Source: HKUSTDial/Supervisor-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.