Agent skill

Research Idea Evaluator

by HKUSTDial in HKUSTDial/Supervisor-Skills

Evaluates a draft research idea like a top-venue reviewer and advisor, scoring it on five dimensions, checking fit with your capacity and returning a clear verdict.

CC-BY-4.0Auto-check passedResearch & Science

Install Research Idea Evaluator

skills CLI
$ npx skills add HKUSTDial/Supervisor-Skills --skill idea-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HKUSTDial/Supervisor-Skills idea-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HKUSTDial/Supervisor-Skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/idea-evaluator .claude/skills/idea-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
idea-evaluator
GitHub stars
8.5k
Token cost
~3.4k tokens
SKILL.md length
1,759 words
Files
12 (incl. references)
Skills in repo
12
Repo updated
First seen
Licence
CC-BY-4.0

At a glance

Evaluates a draft research idea like a top-venue reviewer and advisor, scoring it on five dimensions, checking fit with your capacity and returning a clear verdict.

  • Works in 12 steps: First impression and paper-type… → Fatal-flaws audit (early gate) → Lifecycle and capability matching → …
  • Deciding whether a draft research idea is worth pursuing
  • SKILL.md covers Overview, When to use this skill, When NOT to use this skill and Core procedure, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

The skill scores a preliminary idea on five improvement dimensions (Higher, Faster, Stronger, Cheaper, Broader), matches its lifecycle against your actual capability and weekly hours, probes for paradigm-shift potential and audits it for fatal flaws. It ends with one of three verdicts: Strong Accept, Accept with Revisions, or Reject and Pivot. The stated aim is to drop weak ideas before months are spent on them and to strengthen promising ones before writing starts.

The first step positions the idea as a novel problem, novel method or new setting and asks you to restate it if the story cannot be told in one sentence. The research paradigm is classified by method rather than department: experiment-based work continues with the five dimensions, while text-analysis or conceptual work routes to substitute frameworks. Reference files cover the dimensions, fatal flaws, lifecycle matching, paradigm probes and worked examples. It is not meant for finished work awaiting a paper, manuscript review or brainstorming from scratch.

When your agent uses it

  • Deciding whether a draft research idea is worth pursuing
  • Running a novelty or feasibility check before committing to a paper scope
  • Comparing two or three candidate ideas with a structured trade-off
  • Checking an idea for scope creep

Example prompts

  • “Evaluate this idea for my thesis: using contrastive learning to detect label noise in medical images.”
  • “Score this research idea and tell me whether I should pivot.”
  • “I have about ten hours a week. Is this a feasible research direction for a first-year student?”

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. First impression and paper-type positioning
  2. Fatal-flaws audit (early gate)
  3. Lifecycle and capability matching
  4. Five-dimension scoring
  5. Paradigm-shift probe
  6. Feasibility check
  7. Integrity gate
  8. Final verdict
  9. First impression
  10. Fatal-flaws audit (early gate)
  11. Lifecycle and capability match
  12. Five-dimension radar

What it can do on your machine

Read from SKILL.md and the folder at commit 207bc6f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Research Idea Evaluator loads about 3.4k tokens when it runs, and up to ~23k if it reads all its reference files. Until then it costs about 141 tokens; SKILL.md has 1,759 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~141
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~23k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from HKUSTDial/Supervisor-Skills at commit 207bc6f, republished under its CC-BY-4.0 licence (© HKUSTDial). 1,759 words, ~3,400 tokens.

Download SKILL.mdSave it as .claude/skills/idea-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 11 other files; get the full folder from GitHub.
name
idea-evaluator
description
Evaluates a preliminary research idea against a five-dimension framework (Higher, Faster, Stronger, Cheaper, Broader) plus idea-lifecycle and student-capability matching, paradigm-shift probing, and a fatal-flaws audit. Returns a reviewer-style verdict; non-STEM ideas route to substitute frameworks. Use when the user has a draft research idea and asks whether it is worth pursuing, asks to 'evaluate this idea', 'score this idea', 'assess feasibility', 'novelty check', 'is this a good research direction', or before committing to a paper scope.
license
CC-BY-4.0

Idea Evaluator

Overview

This skill evaluates a preliminary research idea from the combined perspective of a top-venue reviewer and an experienced advisor. It scores the idea against five improvement dimensions from the idea-generation guide (Higher, Faster, Stronger, Cheaper, Broader), matches the idea's lifecycle against the user's actual capability and available hours per week, probes whether the idea has paradigm-shift potential, flags fatal flaws, and returns one of three verdicts: Strong Accept, Accept with Revisions, or Reject and Pivot.

The goal is to kill weak ideas before the student invests months, and to shape promising-but-underdeveloped ideas into stronger forms before writing begins.

When to use this skill

  • The user has a draft idea and asks whether it is worth pursuing.
  • The user asks for novelty check, feasibility assessment, or scoring.
  • Before the user commits to a paper scope or starts implementation.
  • The user is comparing two or three candidate ideas and needs a structured trade-off.
  • The user suspects scope creep and wants an external check.
  • The user mentions 'evaluate this idea', 'score this idea', 'assess feasibility', or 'is this a good research direction'.

When NOT to use this skill

  • The user has already implemented the idea and is writing the paper. Use intro-drafter, tech-paper-template, or benchmark-paper-template (separate plugin) instead.
  • The user explicitly wants brainstorming of new ideas from scratch. Use plain conversation; see handbook 2.3 for a disruptive-innovation playbook.
  • The user asks for review of an existing manuscript. Use pre-submission-reviewer.
  • The user asks to evaluate a benchmark contribution specifically. Use benchmark-paper-template (separate plugin) in targeted mode.

Core procedure

Step 1: First impression and paper-type positioning

Read the user's idea description. In one paragraph, state whether the idea reads as Novel Problem, Novel Method, or New Setting. Is the story compelling in one sentence? If you cannot write that sentence, the idea itself is probably not yet clear enough for evaluation; ask the user to restate.

While restating, classify the research paradigm by method, not by department name: experiments, benchmarks, ablations, or model architectures mean STEM (continue with the five dimensions and fatal flaws below); text analysis, archives, or conceptual argument mean humanities; surveys, interviews, statistics, or fieldwork mean empirical social science; regressions, IV, DID, RDD, or panel data mean finance or economics; statutes, cases, or doctrine mean law.

See: references/domain-evaluation-frameworks.md for the substitute five-dimension frameworks and fatal-flaw substitutions used by the non-STEM paradigms. When the paradigm is unclear, ask one question: "is the core method experiments, surveys, text analysis, or theoretical derivation?"

Step 2: Fatal-flaws audit (early gate)

See: references/fatal-flaws.md for the ten canonical fatal flaws, each with a detection rule and a defense strategy.

Run the fatal-flaws audit before the scoring steps rather than after them. Identify at most two fatal flaws. For each, state the flaw, cite the detection rule, and recommend a concrete defense.

Ground the novelty flaw (F1) in real retrieval whenever the environment has a literature-search capability (a scholarly search tool, web search over scholarly indexes, or shell access to public APIs). Extract two or three keyword groups from the idea (core method plus domain; mechanism plus task; technique plus benchmark), search, and name the three to five closest published works with title, authors, and year. For each, state which axis actually differs: the object acted on, the mechanism, the input granularity, or the problem setting. A similar title alone never establishes duplication; duplication requires failing to find even one differing axis. Retrieval results support metadata-level judgments only (who did what, where); never quote numbers or method details from search snippets. And "not found" does not prove novelty: report it as "no directly overlapping work retrieved under these keywords". With no retrieval capability, label the novelty judgment "unverified; literature check required".

A data-refuted core mechanism is an automatic CRITICAL. If the user's own reported data or attachments already show the core mechanism matched or beaten by a baseline or simple control, lock the verdict to Reject and Pivot; write no defense and invent no optimistic threshold. See: references/fatal-flaws.md, the data-refuted section, for the exact boundary between "refuted by data" and "merely untested".

Short-circuit rule. If any fatal flaw is tagged CRITICAL in the severity taxonomy (single-handedly causes rejection, unfixable within the lifecycle), stop here and emit the verdict directly:

  • Verdict: Reject and Pivot.
  • Output sections 1 (First impression), 2 (Fatal flaws with the CRITICAL flaw), and 7 (Verdict with the flaw-driven rationale) only.
  • Do not run the five-dimension scoring, paradigm-shift probe, feasibility check, or integrity gate. Those would be decoration on a rejection.

If no CRITICAL flaw is found, continue to Step 3.

Step 3: Lifecycle and capability matching

See: references/lifecycle-capability-matching.md for the six-category lifecycle matrix, capability self-assessment rubric, and mismatch recovery strategies.

Map the idea onto one of six categories (Application, Foundational Theory, Cross-Disciplinary, Frontier Exploration, Data-Intensive, Innovative Technique). Match against the user's declared capability (effective hours per week, skill depth, theoretical versus applied strength). Output a mismatch flag if lifecycle is shorter than the user's realistic execution window.

Step 4: Five-dimension scoring

See: references/five-dimensions.md for each dimension's entry strategies, scoring rubric, and worked examples.

Score the idea on each of:

  • Higher: effectiveness and accuracy gains.
  • Faster: efficiency and cost reduction.
  • Stronger: robustness, noise tolerance, generalisation.
  • Cheaper: data, annotation, or solution cost reduction.
  • Broader: cross-domain transplantation or unification.

Score each 1-10 with explicit evidence from the user's stated contribution. Identify the two or three dimensions where the idea has the highest ceiling and recommend emphasising those in the paper.

Scoring discipline: start every dimension at 5 and justify movement. Two kinds of grounds move a score up, and both count: measured results the user reported (quote them), or a mechanism argument that holds up (label the score "mechanism-based, not yet confirmed by data"). A solid, untested mechanism can reach 8 or 9 with that label plus a named validation experiment; do not systematically cap untested ideas. A dimension with neither data nor mechanism stays at 5 with "no grounds given". Watch attribution: when an impressive gain plausibly comes from a peripheral factor (routing, post-processing, a stronger base model, favorable samples), cap that dimension until an ablation isolates the core mechanism. The scoring reference's final two sections cover both rules in detail.

For non-STEM paradigms, score the substitute dimensions from references/domain-evaluation-frameworks.md instead, under the same discipline.

Show full SKILL.md (719 more words)Show less
Step 5: Paradigm-shift probe

See: references/paradigm-shift-probe.md for the four probing principles (First Principles, Elephant in the Room, Technology Cycle, Hamming's Rule) and the cross-reference to handbook section 2.3 when deeper disruptive-innovation exploration is needed.

Test the idea against four questions:

  1. Does it challenge a hidden assumption the field takes for granted?
  2. Does it address an elephant-in-the-room problem everyone sees but nobody wants to touch?
  3. Does it ride a technology-cycle shift (for example, LLMs making a previously impractical approach now feasible)?
  4. If this problem solved itself, would the field change meaningfully? (Hamming's Rule)

Two or more yes answers means the idea has disruptive potential. Note that, and recommend reading handbook 2.3 to deepen the thinking on disruptive-innovation dimensions.

Step 6: Feasibility check

Against the user's stated resources (hardware, data access, team size, engineering skills, timeline), assess:

  • Compute risk: does the experiment fit on stated hardware?
  • Data risk: is the required data accessible, or does it need expensive annotation or private sources?
  • Engineering risk: does the implementation match the user's skill stack?
  • Timeline risk: does the estimated end-to-end duration (coding, experiments, writing, revision) fit within the idea's lifecycle?

If any risk is high, flag it explicitly with a suggested mitigation.

Step 7: Integrity gate

Before emitting the verdict, run the checks in the Integrity gate section below.

Step 8: Final verdict

Issue one of three verdicts:

  • Strong Accept: execute now. Two or more dimensions at 8+, no fatal flaws, capability match green, lifecycle fit.
  • Accept with Revisions: pivot the scope per recommendations before starting. Some dimensions weak, fixable flaws, or lifecycle mismatch that can be shortened.
  • Reject and Pivot: do not pursue this version. Dominated by a prior benchmark or method, unfixable capability mismatch, or more than one fatal flaw.

When the high scores are mechanism-based rather than data-based, qualify the verdict as "worth pursuing, pending the validation experiment", and name that experiment in the top-three actions.

Emit the evaluation in the Output format below.

Integrity gate

Each bullet is tagged with an enforceability class. [inspection] means the LLM can verify the bullet from the produced output alone. [attestation] means the LLM states it has done the check, but the user remains responsible for verification. [user-attest] means the bullet is a user-side rule the skill cannot confirm.

Before returning the verdict:

  1. [inspection] Every dimension score cites specific evidence from the user's stated contribution; no score is "gut feeling".
  2. [inspection] Feasibility claims reference the user's stated resources, not generic assumptions.
  3. [inspection] Novelty claims either cite specific prior work or are labelled "unverified; literature check required".
  4. [inspection] Fatal flaws are specific and actionable; "this might not work" is not a flaw statement.
  5. [inspection] Verdict is consistent with scoring: Strong Accept requires at least two dimensions at 8+ and zero CRITICAL flaws.
  6. [inspection] Paradigm-shift claim cites which probing question was answered positively.
  7. [attestation] Lifecycle prediction is reasoned from the field's recent pace; the user should sanity-check against their own knowledge of the subfield before acting on it.

If any [inspection] check fails, downgrade the verdict and mark the corresponding output section as "needs user attention". For [attestation] bullets, the skill states the check was run and the user confirms the result.

Run the gate silently. Do not print a per-gate pass or fail report; a failure surfaces as a concrete finding inside the affected output section, and the delivered evaluation stays free of internal checking rituals.

Output format

1. First impression
  • Paper type: <Novel Problem or Novel Method or New Setting>
  • One-sentence story: <...>
2. Fatal-flaws audit (early gate)
#FlawSeverityDefense
1...CRITICAL or MAJOR...

If any CRITICAL flaw is present, skip sections 3-6 and go to section 7 with verdict Reject and Pivot.

3. Lifecycle and capability match
AspectUser's inputAssessment
Idea category......
Lifecycle... months...
Weekly effective hours......
Fit...Green or Yellow or Red
4. Five-dimension radar
DimensionScore 1-10EvidenceLift suggestion
Higher.........
Faster.........
Stronger.........
Cheaper.........
Broader.........
5. Paradigm-shift probe
ProbeYes or NoRationale
First Principles......
Elephant in the Room......
Technology Cycle......
Hamming's Rule......

Disruptive potential: <none, possible, strong>.

6. Feasibility
RiskLevelMitigation
Compute......
Data......
Engineering......
Timeline......
7. Verdict

<Strong Accept or Accept with Revisions or Reject and Pivot> (mechanism-based high scores: append "worth pursuing, pending the validation experiment")

Top three actions to take first:

  1. ...
  2. ...
  3. ...

© HKUSTDial, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 11 other files (references) in skills/idea-evaluator of HKUSTDial/Supervisor-Skills.

  • SKILL.md
  • references/domain-evaluation-frameworks.md
  • references/fatal-flaws.md
  • references/five-dimensions.md
  • references/lifecycle-capability-matching.md
  • references/paradigm-elephant.md
  • references/paradigm-examples.md
  • references/paradigm-first-principles.md
  • references/paradigm-hamming.md
  • references/paradigm-shift-probe.md
  • references/paradigm-technology-cycle.md
  • references/worked-examples.md

Open the folder on GitHubat commit 207bc6f

Compare with similar skills

Research Idea Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Research Idea Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Research Idea Evaluator this skillHKUSTDial/Supervisor-Skills8.5k—~3.4kAutomated safety check: PassCC-BY-4.0
Scientific Workflow ToolsDrugClaw/DrugClaw125—~712Automated safety check: PassApache-2.0
Academic Paper Writing PipelineImbad0202/academic-research-skills51k—~16kAutomated safety check: PassCustom licence
Academic Research PipelineImbad0202/academic-research-skills51k—~15kAutomated safety check: PassCustom licence
Social Science Paper Writingfakerqwq/social-science-paper-writing-skill375—~7kAutomated safety check: PassNone
Paper GlanceDiaugeia/paper-glance-skill106—~433Automated safety check: PassNone

Similar skills

  • Scientific Workflow Tools

    DrugClaw/DrugClaw

    Research-method workflow guide for hypothesis framing, peer-review style critique, reproducibility planning, study-design checks, and scientific-writing structure.

    125 GitHub stars~712 tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Academic Paper Writing Pipeline

    Imbad0202/academic-research-skills

    Runs a 12-agent pipeline that plans, drafts, cites, reviews and formats academic papers, with modes for revision, rebuttals, abstracts and citation checks.

    51k GitHub stars~16k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • Academic Research Pipeline

    Imbad0202/academic-research-skills

    Orchestrates a ten-stage academic workflow from research to finished manuscript, including integrity checks, two rounds of peer review and revision.

    51k GitHub stars~15k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed
  • Social Science Paper Writing

    fakerqwq/social-science-paper-writing-skill

    Helps draft, diagnose, review and revise social science papers, from topic and research question to literature review, citation risks and pre-submission checks.

    375 GitHub stars~7k tokensUpdated 4 mo ago
    Research & ScienceAuto-check passed
  • Paper Glance

    Diaugeia/paper-glance-skill

    Reads an academic paper from a PDF, pasted text or abstract and offers five outputs: an analysis report, a mind map, a peer review, promotion scripts and a podcast.

    106 GitHub stars~433 tokensUpdated 3 mo ago
    Research & ScienceAuto-check passed
  • Canvas Week Plan

    vishalsachdev/canvas-mcp

    Student weekly assignment planner for Canvas LMS. An agent skill from vishalsachdev/canvas-mcp.

    284 GitHub stars~1.2k tokensUpdated today
    Research & ScienceAuto-check passed

More from HKUSTDial/Supervisor-Skills

All 12 skills in this repo
  • Benchmark Paper Template

    HKUSTDial/Supervisor-Skills

    Structures benchmark and evaluation papers around five pillars, with a completeness audit, an Introduction logic chain, a section skeleton and a pre-submission checklist.

    8.5k GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Draw.io Diagram Reconstruction

    HKUSTDial/Supervisor-Skills

    Rebuilds a reference diagram image as an editable, high-fidelity Draw.io file, mixing native elements, SVG icons and cropped PNGs, with a batch workflow for a folder of images.

    8.5k GitHub stars~5.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Paper Introduction Drafter

    HKUSTDial/Supervisor-Skills

    Drafts the Introduction of a technical paper as six paragraphs of flowing prose, positioning the work and matching contributions to challenges, with an outline on request.

    8.5k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Rebuttal Guidance

    HKUSTDial/Supervisor-Skills

    Turns pasted peer-review comments into a per-concern rebuttal plan, with reviewer mindset matching and strategy priorities, but not the final rebuttal text.

    8.5k GitHub stars~1.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Deep Research Literature Survey

    HKUSTDial/Supervisor-Skills

    Runs a survey-grade literature investigation: fixes the research questions, searches from adversarial angles, verifies citations and writes an evidence-first report.

    8.5k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Paper Figure Designer

    HKUSTDial/Supervisor-Skills

    Advises on designing the three core figures of a technical paper, then audits them against rules for format, fonts, color and captions.

    8.5k GitHub stars~1.9k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Research Idea Evaluator

What does Research Idea Evaluator do?

Evaluates a draft research idea like a top-venue reviewer and advisor, scoring it on five dimensions, checking fit with your capacity and returning a clear verdict. The skill scores a preliminary idea on five improvement dimensions (Higher, Faster, Stronger, Cheaper, Broader), matches its lifecycle against your actual capability and weekly hours, probes for paradigm-shift potential and audits it for fatal flaws. It ends with one of three verdicts: Strong Accept, Accept with Revisions, or Reject and Pivot.

When should I use Research Idea Evaluator?

Research Idea Evaluator fits situations like: deciding whether a draft research idea is worth pursuing; running a novelty or feasibility check before committing to a paper scope; comparing two or three candidate ideas with a structured trade-off; checking an idea for scope creep.

How do I install Research Idea Evaluator in Claude Code?

Run `npx skills add HKUSTDial/Supervisor-Skills --skill idea-evaluator -a claude-code`. Or copy the skill folder (skills/idea-evaluator in HKUSTDial/Supervisor-Skills) into .claude/skills/idea-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Research Idea Evaluator in Codex?

Run `npx skills add HKUSTDial/Supervisor-Skills --skill idea-evaluator -a codex`. Or copy the skill folder (skills/idea-evaluator in HKUSTDial/Supervisor-Skills) into .agents/skills/idea-evaluator in your project. Codex loads it when a task matches its description.

Can I use Research Idea Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HKUSTDial/Supervisor-Skills --skill idea-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/idea-evaluator, .gemini/skills/idea-evaluator, .github/skills/idea-evaluator and .opencode/skills/idea-evaluator in your project.

What does Research Idea Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Research Idea Evaluator is instructions for the agent only.

Does Research Idea Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Research Idea Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Research Idea Evaluator use?

Research Idea Evaluator is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Research Idea Evaluator use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 19k tokens, read only when the agent opens those files.

What are the alternatives to Research Idea Evaluator?

Skills that share tags, products or a category with Research Idea Evaluator: Scientific Workflow Tools (DrugClaw/DrugClaw, 125 stars), Academic Paper Writing Pipeline (Imbad0202/academic-research-skills, 51k stars), Academic Research Pipeline (Imbad0202/academic-research-skills, 51k stars) and Social Science Paper Writing (fakerqwq/social-science-paper-writing-skill, 375 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Research Idea Evaluator?

HKUSTDial (a GitHub organization) maintains it in HKUSTDial/Supervisor-Skills, which has 8,523 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 5, 2026.

Source: HKUSTDial/Supervisor-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.