Get a deep critical review of research from an external reviewer backend (Codex or manual).

MITAuto-check: notesResearch & Science

Install Research Review

skills CLI
$ npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wanshuiyin/Auto-claude-code-research-in-sleep research-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wanshuiyin/Auto-claude-code-research-in-sleep.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/research-review .claude/skills/research-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
research-review
GitHub stars
17k
Token cost
~3.1k tokens
SKILL.md length
1,115 words
Files
1
Skills in repo
26
Repo updated
First seen
Licence
MIT

At a glance

Get a deep critical review of research from an external reviewer backend (Codex or manual).

  • Works in 5 steps: Gather Research Context → Initial Review (Round 1) → Iterative Dialogue (Rounds 2-N) → …
  • User says review my research
  • SKILL.md covers Constants, Reviewer Calling Convention, Context: $ARGUMENTS and Prerequisites, plus 4 more sections
  • Calls claude

What it does

Research Review is an agent skill from wanshuiyin/Auto-claude-code-research-in-sleep. Get a deep critical review of research from an external reviewer backend (Codex or manual). Use when user says "review my research", "help me review", "get external review", or wants critical feedback on research ideas, papers, or experimental results.

Its SKILL.md is about 3.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Hypothesis generation. It works with Model Context Protocol and OpenAI. The repository describes itself as: ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework… The licence is MIT.

When your agent uses it

  • User says review my research
  • Get external review
  • Wants critical feedback on research ideas
  • Experimental results

Example prompts

  • “review my research”
  • “help me review”
  • “get external review”
  • “/research-review”

Requirements

  • Python 3
  • Pre-approved tools (allowed-tools): Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_review__review, mcp__manual_review__review_reply

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Gather Research Context
  2. Initial Review (Round 1)
  3. Iterative Dialogue (Rounds 2-N)
  4. Convergence
  5. Document Everything

What it can do on your machine

Read from SKILL.md and the folder at commit 26b95cf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash(*)
    • Read
    • Grep
    • Glob
    • Write
    • Edit
    • mcp__codex__codex
    • mcp__codex__codex-reply
    • mcp__manual_review__review
    • mcp__manual_review__review_reply

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Research Review loads about 3.1k tokens when it runs. Until then it costs about 67 tokens; SKILL.md has 1,115 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~67
When it runs · the whole SKILL.md, loaded when a task matches
~3.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_revi

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wanshuiyin/Auto-claude-code-research-in-sleep at commit 26b95cf, republished under its MIT licence (© wanshuiyin). 1,115 words, ~3,100 tokens.

Download SKILL.mdSave it as .claude/skills/research-review/SKILL.md (or your agent's skills folder).
name
research-review
description
Get a deep critical review of research from an external reviewer backend (Codex or manual). Use when user says "review my research", "help me review", "get external review", or wants critical feedback on research ideas, papers, or experimental results.
allowed-tools
Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_review__review, mcp__manual_review__review_reply
argument-hint
[topic-or-scope]

Research Review via External Reviewer Backend (ultra reasoning)

🔒 Do not wrap this skill in /loop, /schedule, or CronCreate. It is verdict-bearing — it produces a cross-model review verdict, multi-round with reviewer thread continuity. An external timer re-fires the verdict on wall-clock time and breaks the reviewer's round-to-round memory: zero new signal, full token cost. Schedule the external wait that precedes it (work ready → then review once), not the verdict. See shared-references/external-cadence.md.

Get a multi-round critical review of research work from the selected external reviewer backend with maximum reasoning depth.

Constants

  • REVIEWER_MODEL = gpt-6-astra — Default model for the Codex backend, reasoning effort ultra (deep-audit tier). Must be an OpenAI model (e.g., gpt-6-astra, gpt-5.5, o3). Manual backend uses a model the user chooses — it must be a recognized model from a different family (OpenAI, Anthropic, Google, DeepSeek, Moonshot/Kimi, Qwen).
  • REVIEWER_BACKEND = codex — Default: Codex MCP (ultra). Override with — reviewer: oracle-pro for Oracle MCP, or — reviewer: manual for Manual Review MCP. If manual-review MCP is unavailable, stop and print the install command; do not fall back to Codex. See shared-references/reviewer-routing.md.

Reviewer Calling Convention

When calling the reviewer, branch on REVIEWER_BACKEND:

If REVIEWER_BACKEND = codex: Use mcp__codex__codex for new review threads. Use mcp__codex__codex-reply for follow-up rounds (reuse threadId).

If REVIEWER_BACKEND = manual: Use mcp__manual_review__review for new review threads with: prompt: [exact same prompt that would go to Codex] config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true} Save the returned threadId. Use mcp__manual_review__review_reply for follow-up rounds with: threadId: [saved manual-review threadId] prompt: [follow-up prompt] config: {"model_reasoning_effort": "xhigh", "executor_model": "<actual executor model>", "require_reviewer_model": true}

Content fidelity: the manual reviewer should see the same substantive review brief Codex would read. If the manual UI supports file upload / attachment, reuse the same brief file; otherwise paste the brief contents inline because remote web UIs cannot read your local filesystem paths. Review tracing applies equally to both backends.

Context: $ARGUMENTS

Prerequisites

  • Codex MCP Server configured in Claude Code:
    bash
    claude mcp add codex -s user -- python3 "$HOME/aris_repo/mcp-servers/codex-exec/server.py"   # your ARIS clone's path
  • This gives Claude Code access to mcp__codex__codex and mcp__codex__codex-reply tools

Workflow

Step 1: Gather Research Context

Before calling the external reviewer, compile a comprehensive briefing:

  1. Read project narrative documents (e.g., STORY.md, README.md, paper drafts)
  2. Read any memory/notes files for key findings and experiment history
  3. Identify: core claims, methodology, key results, known weaknesses
Step 2: Initial Review (Round 1)

Send a detailed prompt with ultra reasoning, using the selected backend. For the codex backend, keep the MCP payload short: write the full briefing to RESEARCH_REVIEW_REQUEST.md, then point Codex at that file.

For codex backend:

mcp__codex__codex:
  model: gpt-6-astra
  config: {"model_reasoning_effort": "ultra"}
  prompt: |
    Read the review brief at <absolute path to RESEARCH_REVIEW_REQUEST.md>.
    Executor notes are not evidence beyond the files they cite, so verify the
    referenced artifacts before judging.
    Please act as a senior ML reviewer (NeurIPS/ICML level). Start from the
    assumption that the work is broken somewhere — your job is to find where.
    Be adversarial. Trust nothing the author tells you — verify everything
    yourself. Identify:
    1. Logical gaps or unjustified claims
    2. Missing experiments that would strengthen the story
    3. Narrative weaknesses
    4. Whether the contribution is sufficient for a top venue

    === SCOPE LIMITS (these bound what you PROPOSE, never what you look for) ===
    Report anything that is actually wrong here — including a rare-looking case, if
    this repo actually produces it. Then keep the fix in scope:
    1. This is a RESEARCH-WORKFLOW tool, not a security paper. Verification is
       welcome; over-defense is not. Assume a cooperating operator on their own
       machine — a malicious local user is NOT in the threat model.
    2. Do NOT propose SHA / hash / content-fingerprint / digest-binding schemes.
       Reporting a real defect in hashing code that already exists is fine.
    3. NO speculative machinery: do not add feature flags, migration frameworks,
       compat layers, wrappers, pins, or similar mechanisms unless evidence shows
       a current repo defect they fix or an explicit existing invariant they must
       preserve. "Load-bearing", "compatibility", and "not scaffolding" are labels,
       not evidence. Point to the failing path/artifact or invariant, and check the
       proposal's factual premises, such as whether a named package version exists.
    4. NO corner-case obsession: exotic encodings, symlink races, RTL text and
       millisecond races are out of scope unless you can show the case arises here.
    5. Where a rubric or checklist is genuinely needed, do not over-mechanize
       judgement. A clear sentence a human reads beats a scored table nobody
       maintains.
    Exception: code that runs remote commands, starts a network service, or installs
    an MCP server runs on the user's machine with their credentials — trust-boundary
    findings there are in scope and the default is strict.
    Say plainly when something is correct. Do not manufacture findings.
    Be brutally honest. If, after genuinely trying to break it, the work
    holds up and is ready, say so clearly.

The review brief should contain the full research context, the specific questions, and the primary artifact / raw-result paths the reviewer should inspect.

For manual backend: use mcp__manual_review__review with the same brief contents. If the manual-review UI supports attachments, attach RESEARCH_REVIEW_REQUEST.md; otherwise paste the brief inline. Save the returned threadId.

Step 3: Iterative Dialogue (Rounds 2-N)

For codex backend: use mcp__codex__codex-reply with the returned threadId. For manual backend: use mcp__manual_review__review_reply with the same threadId. Use the appropriate tool to continue the conversation. For Codex follow-up rounds, write an updated brief such as RESEARCH_REVIEW_ROUND_2.md and send only the path:

text
mcp__codex__codex-reply:
  threadId: [saved reviewer threadId from Step 2]
  # replies inherit the thread's model/effort (gpt-6-astra ultra)
  prompt: |
    Read the updated review brief at <absolute path to
    RESEARCH_REVIEW_ROUND_2.md>.
    Focus on unresolved weaknesses and whether the revision actually fixed them.

For manual follow-up rounds, attach that same updated brief if possible; otherwise paste it inline.

For each round:

  1. Respond to criticisms with evidence/counterarguments
  2. Ask targeted follow-ups on the most actionable points
  3. Request specific deliverables: experiment designs, paper outlines, claims matrices

Key follow-up patterns:

  • "If we reframe X as Y, does that change your assessment?"
  • "What's the minimum experiment to satisfy concern Z?"
  • "Please design the minimal additional experiment package (highest acceptance lift per GPU week)"
  • "Please write a mock NeurIPS/ICML review with scores"
  • "Give me a results-to-claims matrix for possible experimental outcomes"
Step 4: Convergence

Stop iterating when:

  • Both sides agree on the core claims and their evidence requirements
  • A concrete experiment plan is established
  • The narrative structure is settled
Show full SKILL.md (484 more words)Show less
Step 5: Document Everything

Save the full interaction and conclusions to a review document in the project root:

  • Round-by-round summary of criticisms and responses
  • Final consensus on claims, narrative, and experiments
  • Claims matrix (what claims are allowed under each possible outcome)
  • Prioritized TODO list with estimated compute costs
  • Paper outline if discussed

Update project memory/notes with key review conclusions.

Composed mode — if invoked with — composed: <canonical-report-path> (an orchestrator like /idea-discovery passes this), do not write a standalone review .md in the project root. The raw conversation is already persisted to .aris/traces/… (see Review Tracing below — that audit copy is kept in every mode); fold the review conclusions (consensus, claims matrix, prioritized TODOs) into the orchestrator's canonical report and cite the trace path there. Default (no — composed: directive): behave exactly as above — write the standalone review document. Never infer composed mode from a report file merely existing. Full rules: shared-references/output-composition.md.

Key Rules

  • ALWAYS pin model: gpt-6-astra + config: {"model_reasoning_effort": "ultra"} for reviews (deep-audit tier; capability fallback per reviewer-routing.md, never below xhigh)
  • That pin is the Codex backend's. For manual, use the identity-bearing config from the Reviewer Calling Convention above; model, sandbox and cwd are Codex-only
  • Put comprehensive context in the review brief. Codex can read local files when you pass an absolute path; manual reviewers usually cannot, so attach or paste the same brief there.
  • Be honest about weaknesses — hiding them leads to worse feedback
  • Push back on criticisms you disagree with, but accept valid ones
  • Focus on ACTIONABLE feedback — "what experiment would fix this?"
  • Document the threadId for potential future resumption
  • The review document should be self-contained (readable without the conversation)

Prompt Templates

For initial review:

"I'm going to present a complete ML research project for your critical review. Please act as a senior ML reviewer (NeurIPS/ICML level)..."

For experiment design:

"Please design the minimal additional experiment package that gives the highest acceptance lift per GPU week. Our compute: [describe]. Be very specific about configurations."

For paper structure:

"Please turn this into a concrete paper outline with section-by-section claims and figure plan."

For claims matrix:

"Please give me a results-to-claims matrix: what claim is allowed under each possible outcome of experiments X and Y?"

For mock review:

"Please write a mock NeurIPS review with: Summary, Strengths, Weaknesses, Questions for Authors, Score, Confidence, and What Would Move Toward Accept."

Review Tracing

After each reviewer call (mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_review__review, or mcp__manual_review__review_reply), save the trace following shared-references/review-tracing.md (Policy C — forensic; never silently skip). Use save_trace.sh (resolved per the chain in shared-references/integration-contract.md §2) or write files directly to .aris/traces/<skill>/<date>_run<NN>/. Respect the --- trace: parameter (default: full). A verdict-bearing manual response MUST begin with Reviewer-Model: <exact-model-id> — pass the model THIS session is actually running as in executor_model. Missing, unknown, or same-family identity cannot acquit; emit REVIEW_UNAVAILABLE rather than guessing. If the executor model cannot be named, manual review's cross-family claim is unprovable — say so in the report instead of asserting it.

© wanshuiyin, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/research-review of wanshuiyin/Auto-claude-code-research-in-sleep.

Open the folder on GitHubat commit 26b95cf

Compare with similar skills

Research Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Research Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Research Review this skillwanshuiyin/Auto-claude-code-research-in-sleep17k—~3.1kAutomated safety check: NotesMIT
Research ReviewGRIND-Lab-Core/night_owl_research_agent1065 repos~1.1kAutomated safety check: NotesNone
Research RefinezjYao36/Auto-Research-Refine1286 repos~6.9kAutomated safety check: NotesNone
Novelty CheckAI4Scientist/nano-scientist1284 repos~823Automated safety check: PassNone
Idea CreatorAI4Scientist/nano-scientist1284 repos~3.9kAutomated safety check: WarnNone
Novelty CheckGRIND-Lab-Core/night_owl_research_agent106—~1kAutomated safety check: PassNone

Similar skills

  • Research Review

    GRIND-Lab-Core/night_owl_research_agent

    Get a deep critical review of research idea from GPT via Codex MCP.

    106 GitHub starsUsed in 5 repos~1.1k tokens
    Research & ScienceAuto-check: notes
  • Research Refine

    zjYao36/Auto-Research-Refine

    Turns a vague research direction into a focused, problem-anchored method plan through up to five review rounds with a second model.

    128 GitHub starsUsed in 6 repos~6.9k tokens
    Research & ScienceAuto-check: notes
  • Novelty Check

    AI4Scientist/nano-scientist

    Verify research idea novelty against recent literature. An agent skill from AI4Scientist/nano-scientist.

    128 GitHub starsUsed in 4 repos~823 tokens
    Research & ScienceAuto-check passed
  • Idea Creator

    AI4Scientist/nano-scientist

    Generate and rank research ideas given a broad direction. An agent skill from AI4Scientist/nano-scientist.

    128 GitHub starsUsed in 4 repos~3.9k tokens
    Research & ScienceAuto-check: warnings
  • Novelty Check

    GRIND-Lab-Core/night_owl_research_agent

    Validates that a research idea is genuinely novel vs. An agent skill from GRIND-Lab-Core/night_owl_research_agent.

    106 GitHub stars~1k tokensUpdated 5 mo ago
    Research & ScienceAuto-check passed
  • Annotate Paper

    54yyyu/zotero-mcp

    Read the open paper and write study annotations into its PDF with zotero-cli - a context box on the title, a four-part summary on the abstract, role-coded abstract highlights, one box per figure…

    5.3k GitHub stars~1.5k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed

More from wanshuiyin/Auto-claude-code-research-in-sleep

All 26 skills in this repo
  • Academic Poster Builder

    wanshuiyin/Auto-claude-code-research-in-sleep

    Builds an academic conference poster as a single HTML and CSS file with measurement-based gates, real paper figures and a print-ready PDF rendered through headless Chromium.

    17k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check: notes
  • Proof Run Orchestrator

    wanshuiyin/Auto-claude-code-research-in-sleep

    Runs a mathematical proof project as a stateful pipeline of run directories: a local attempt first, then a manual GPT Pro handoff package, with an optional DeepSeek audit.

    17k GitHub starsUsed in 1 repo~4.7k tokens
    Auto-check passed
  • Render HTML

    wanshuiyin/Auto-claude-code-research-in-sleep

    Render an ARIS Markdown / JSON artifact (IDEAREPORT, AUTOREVIEW, KILLARGUMENT, PAPERPLAN, research-wiki state, etc.) into a single-file HTML view designed for human reading.

    17k GitHub starsUsed in 1 repo~5.4k tokens
    Auto-check: notes
  • Experiment Audit

    wanshuiyin/Auto-claude-code-research-in-sleep

    Audit experiment integrity before claiming results. An agent skill from wanshuiyin/Auto-claude-code-research-in-sleep.

    17k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check: notes
  • Integrity Forensics

    wanshuiyin/Auto-claude-code-research-in-sleep

    Run the Anti-Autoresearch integrity-forensics DETERMINISTIC slice (numeric core + rules-only reporter) against a paper via a SHA-pinned thin launcher, then convert the verdict into a typed policy…

    17k GitHub starsUsed in 1 repo~1.5k tokens
    Auto-check passed
  • Interview Cheatsheet

    wanshuiyin/Auto-claude-code-research-in-sleep

    Generate a long-form Chinese interview-prep cheat sheet on a specific ML/LLM topic — formulas with derivations, from-scratch PyTorch code, comparison tables, and 25 高频面试题 (L1 必会 / L2 进阶 / L3 顶级 lab).

    17k GitHub starsUsed in 1 repo~3.3k tokens
    Auto-check: notes

Questions about Research Review

What does Research Review do?

Get a deep critical review of research from an external reviewer backend (Codex or manual). Research Review is an agent skill from wanshuiyin/Auto-claude-code-research-in-sleep. Get a deep critical review of research from an external reviewer backend (Codex or manual).

When should I use Research Review?

Research Review fits situations like: user says review my research; get external review; wants critical feedback on research ideas; experimental results.

How do I install Research Review in Claude Code?

Run `npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-review -a claude-code`. Or copy the skill folder (skills/research-review in wanshuiyin/Auto-claude-code-research-in-sleep) into .claude/skills/research-review in your project. Claude Code loads it when a task matches its description.

How do I install Research Review in Codex?

Run `npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-review -a codex`. Or copy the skill folder (skills/research-review in wanshuiyin/Auto-claude-code-research-in-sleep) into .agents/skills/research-review in your project. Codex loads it when a task matches its description.

Can I use Research Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wanshuiyin/Auto-claude-code-research-in-sleep --skill research-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/research-review, .gemini/skills/research-review, .github/skills/research-review and .opencode/skills/research-review in your project.

What does Research Review need to run?

Going by SKILL.md and its folder, Research Review needs the command-line tools its instructions call (claude). Our summary lists: Python 3. Its frontmatter pre-approves these tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply, mcp__manual_review__review, mcp__manual_review__review_reply.

Does Research Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Research Review safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Research Review use?

Research Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Research Review use?

About 3.1k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Research Review?

Skills that share tags, products or a category with Research Review: Research Review (GRIND-Lab-Core/night_owl_research_agent, 106 stars), Research Refine (zjYao36/Auto-Research-Refine, 128 stars), Novelty Check (AI4Scientist/nano-scientist, 128 stars) and Idea Creator (AI4Scientist/nano-scientist, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Research Review?

wanshuiyin (a GitHub user) maintains it in wanshuiyin/Auto-claude-code-research-in-sleep, which has 17,205 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on October 7, 2026.

Source: wanshuiyin/Auto-claude-code-research-in-sleep on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.