Agent skill

Review

by raine in raine/consult-llm

Collect critical feedback from all registered LLMs on an artifact (architecture doc, implementation, plan).

MITAuto-check: notesAgent Workflows

Install Review

skills CLI
$ npx skills add raine/consult-llm --skill review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install raine/consult-llm review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/raine/consult-llm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/review .claude/skills/review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
review
GitHub stars
140
Token cost
~2.4k tokens
SKILL.md length
853 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Collect critical feedback from all registered LLMs on an artifact (architecture doc, implementation, plan).

  • Works in 5 steps: Load consult-llm skill → Read the Target → Independent Reviews → …
  • Tasks that involve Planning
  • SKILL.md covers Phase 0: Load consult-llm skill, Available Reviewers, Critical Rule: No Sycophancy and Phase 1: Read the Target, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Review is an agent skill from raine/consult-llm. Collect critical feedback from all registered LLMs on an artifact (architecture doc, implementation, plan). Intellectual debate with push-back — no sycophancy. Reports findings and unresolved disagreements.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Planning. It works with OpenAI and DeepSeek. The repository describes itself as: Get a second opinion from another AI model. The licence is MIT.

When your agent uses it

  • Tasks that involve Planning

Example prompts

  • “/review”

Requirements

  • Pre-approved tools (allowed-tools): Bash, Glob, Grep, Read, Write

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Load consult-llm skill
  2. Read the Target
  3. Independent Reviews
  4. Cross-Review and Push-Back
  5. Findings Report

What it can do on your machine

Read from SKILL.md and the folder at commit 69e3ecb. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Bash
    • Glob
    • Grep
    • Read
    • Write

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash and markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Review loads about 2.4k tokens when it runs. Until then it costs about 53 tokens; SKILL.md has 853 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~53
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Bash, Glob, Grep, Read, Write

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from raine/consult-llm at commit 69e3ecb, republished under its MIT licence (© raine). 853 words, ~2,421 tokens.

Download SKILL.mdSave it as .claude/skills/review/SKILL.md (or your agent's skills folder).
name
review
description
Collect critical feedback from all registered LLMs on an artifact (architecture doc, implementation, plan). Intellectual debate with push-back — no sycophancy. Reports findings and unresolved disagreements.
allowed-tools
Bash, Glob, Grep, Read, Write
disable-model-invocation
true

Collect critical, honest feedback from all LLMs on an artifact. Push back on weak arguments. Report both consensus findings and unresolved disagreements.

Phase 0: Load consult-llm skill

Load the consult-llm skill before proceeding — it defines the invocation contract (stdin heredoc, flags, output format, multi-turn). Do not call the CLI without loading it first.

Arguments: $ARGUMENTS

Check the arguments for flags:

Mode flags:

  • --rounds N → number of critique rounds (default: 2, max: 3)
  • --dry-run → skip the final synthesis, just show raw reviews
  • --models <list> → comma-separated selectors/model IDs to use as reviewers (default: gemini,openai,anthropic,deepseek)

Strip all flags from arguments to get the review target — a file path, directory, or topic description.

Set variables:

  • REVIEWERS: list of model selectors from --models flag, or ["gemini", "openai", "anthropic", "deepseek"] if omitted
  • Build the -m flags by repeating -m <selector> for each reviewer

Available Reviewers

Discover which selectors and models are available in this environment:

!`consult-llm models`

Default reviewers (used when no --models flag is given): gemini, openai, anthropic, deepseek — all four selectors that have a configured backend.

Override with --models flag: --models gemini,openai to review with only two, or --models gemini,openai,anthropic for three. Any selector or exact model ID from the list above is accepted.

Critical Rule: No Sycophancy

This skill exists to find problems, not to validate. Instruct every LLM call with:

  • Be critical. The goal is to find weaknesses, gaps, and risks — not to praise.
  • Disagree openly. If something looks wrong, say so directly. Do not soften criticism to be polite.
  • Push back. If another reviewer dismissed a concern too easily, challenge them.
  • Unresolved disagreements are fine. Not everything needs consensus. Flag genuine disagreements clearly rather than papering over them.

Phase 1: Read the Target

  1. Parse the arguments — determine what to review:

    • If it's a file path: read the file(s)
    • If it's a directory: explore and read key files
    • If it's a topic/description: gather relevant files from the codebase
  2. Gather context — use Glob, Grep, Read to understand:

    • The artifact itself (full content)
    • Surrounding code/docs it relates to
    • Existing patterns and conventions
  3. Prepare the review brief — a summary of:

    • What is being reviewed (the artifact and its purpose)
    • Relevant context from the codebase
    • Specific aspects to focus on (if the user mentioned any)

Phase 2: Independent Reviews

Have all four LLMs independently review the artifact in parallel using a single CLI call.

Review prompt:

You are a critical reviewer. Your job is to find problems, not to praise.

## What you are reviewing

[Review brief — artifact content and context]

## Your task

Provide a thorough, critical review:

1. **Problems found**: List concrete issues — bugs, logical errors, missing edge cases, architectural flaws, security concerns. Be specific with file paths and line numbers where applicable.
2. **Questionable decisions**: Decisions that might work but deserve scrutiny — are there better alternatives? What are the trade-offs not being considered?
3. **Missing considerations**: What's not addressed that should be? Gaps in error handling, testing, documentation, scalability, maintainability?
4. **Risks**: What could go wrong in production or during maintenance? What assumptions might not hold?
5. **What works well**: (Brief) What's genuinely solid and should be kept as-is?

Rules:
- Be direct and specific. "This could be improved" is useless. "The retry logic on line 45 silently swallows errors, which will make debugging impossible" is useful.
- Do NOT try to be balanced. If you find 10 problems and 1 good thing, report 10 problems and 1 good thing.
- Do NOT soften criticism. If something is bad, say it's bad and explain why.
- Prioritize your findings: critical issues first, minor nits last.

Invoke consult-llm with -m <selector> repeated for each reviewer in REVIEWERS, --task review, and -f <path> for each relevant file. Send the review prompt on stdin via quoted heredoc. All models are queried in parallel in a single call.

The response is in group format:

  • Line 1: [thread_id:group_xxx]
  • Each model section: ## Model: <id> header, then [model:<id>] [thread_id:<per-model-id>], then the response body

Extract thread IDs: Parse each model's thread_id from the per-model header lines. These are needed for Phase 3 since each model receives the other three's responses.

Present all reviews to the user.

Show full SKILL.md (369 more words)Show less

Phase 3: Cross-Review and Push-Back

For each round (default 2, configurable with --rounds N, max 3):

Share a combined summary of all other reviewers' findings with each reviewer and ask them to challenge, validate, or push back. Use -t <thread_id> to continue each LLM's conversation.

Cross-review prompt (for each reviewer, include the other reviewers' findings):

The other reviewers provided these assessments:

[Combined summary of the other reviewers' latest responses, labeled by provider name]

Respond critically:

1. **Agree**: Which of their findings are valid? Don't just agree to be agreeable — only agree if you genuinely think they're right.
2. **Disagree**: Which findings are wrong, exaggerated, or missing context? Explain why. If they dismissed one of YOUR concerns, push back if you still think it's valid.
3. **New findings**: Did their reviews make you notice anything you missed?
4. **Priority adjustment**: Given all reviews, what are the TOP 3 most critical issues?

Do NOT be diplomatic. If they're wrong, say they're wrong and explain why. If you change your mind, say so explicitly — don't quietly drop a previous point.

Each model receives a different prompt (the other reviewers' responses embedded). Invoke consult-llm once with one --run flag per reviewer, continuing each model's thread:

bash
consult-llm \
  --run "model=<selector>,thread=$THREAD,prompt-file=$PROMPT" \
  ...  # one --run per reviewer
  -f <path> ...

Write each model's cross-review prompt to a temp file with mktemp, using __CONSULT_LLM_END__ as the heredoc terminator and >| to overwrite.

Present all responses to the user after each round.

Phase 4: Findings Report

If --dry-run: Present the raw reviews without synthesis.

Analyze all rounds and produce a structured report:

1. Categorize findings

Go through every issue raised across all rounds and categorize:

  • Consensus findings: Majority of reviewers (3+) agree this is a problem
  • Partial consensus: Two reviewers agree, others disagree or are silent
  • Unresolved disagreements: Reviewers actively disagree — present all sides fairly
  • Dropped concerns: Issues raised then abandoned — note why
2. Write the report
markdown
## Review: [Artifact Name]

**Reviewed:** [What was reviewed — file paths or description]
**Reviewers:** Gemini, OpenAI, Anthropic, DeepSeek
**Rounds:** [N]

### Critical Issues (Consensus)

Issues where 3+ reviewers agree, ordered by severity:

1. **[Issue title]**
   - **What:** [Specific description]
   - **Where:** [File:line or section]
   - **Why it matters:** [Impact]
   - **Suggested fix:** [If one emerged from discussion]
   - **Raised by:** [Which reviewers]

### Disputed Issues

Issues where reviewers disagree — all positions presented:

1. **[Issue title]**
   - **For:** [Reviewers and their argument]
   - **Against:** [Reviewers and their argument]
   - **Moderator's take:** [Your assessment of who has the stronger argument]

### Minor Findings

Lower-severity issues and suggestions:
- [Finding 1]
- [Finding 2]

### What's Solid

Aspects reviewers consider well-done:
- [Strength 1]
- [Strength 2]

### Unresolved Questions

Open questions that need human judgment:
- [Question 1]
- [Question 2]
3. Moderator's assessment

Add your own honest assessment as moderator:

  • Which reviewer made the strongest arguments overall?
  • Are there issues NO reviewer caught that you noticed?
  • What's the single most important thing to address?

Save the report to history/review-<artifact-name>.md.

Critical Rules

  • Independence first. Phase 2 reviews are fully independent — all models receive the same prompt via parallel -m flags in a single call. Do not show one reviewer's output to another until Phase 3.
  • No sycophancy. Every prompt must instruct LLMs to be critical, disagree openly, and push back. This skill exists to find problems, not to validate.
  • Bash timeout 600000 on every consult-llm call — LLM responses routinely exceed the 2-minute default.
  • Unresolved disagreements are valid output. Do not force consensus. Flag genuine disagreements clearly rather than papering over them.
  • Verify before adopting. Findings are claims, not facts. If a reviewer cites a specific file/line, confirm it exists and says what they claim before including it in the report.
  • Phase discipline. Phase 1 is LLM-free (agent reads and prepares context). Phase 2 is independent LLM work. Phase 3 is adversarial cross-review. Phase 4 is agent synthesis.

© raine, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/review of raine/consult-llm.

Open the folder on GitHubat commit 69e3ecb

Compare with similar skills

Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Review this skillraine/consult-llm140—~2.4kAutomated safety check: NotesMIT
Codex with ChatGPT Planning LoopXiaoDuoYa/codex-with-chatgpt7.1k—~11kAutomated safety check: NotesMIT
LLM Councilgcpdev/llm-council-skill461—~1kAutomated safety check: NotesMIT
Remember Learningsdyad-sh/dyad22k—~1.1kAutomated safety check: PassCustom licence
Intent Debuggerbydtesla1609/intent-debugger126—~2.3kAutomated safety check: PassMIT
Test Ocas Openclawpwrdrvr/openclaw-codex-app-server265—~4kAutomated safety check: PassMIT

Similar skills

  • Codex with ChatGPT Planning Loop

    XiaoDuoYa/codex-with-chatgpt

    Uses ChatGPT in the browser as the planning and review brain for a Codex session, with Codex keeping all execution and ChatGPT reading the workspace through a bridge.

    7.1k GitHub stars~11k tokensUpdated 8 days ago
    Agent WorkflowsAuto-check: notes
  • LLM Council

    gcpdev/llm-council-skill

    Multi-LLM collaborative brainstorming and planning. An agent skill from gcpdev/llm-council-skill.

    461 GitHub stars~1k tokensUpdated 9 mo ago
    Agent WorkflowsAuto-check: notes
  • Remember Learnings

    dyad-sh/dyad

    Review the current session for errors, issues, snags, and hard-won knowledge, then update the rules/ files (or AGENTS.md if no suitable rule file exists) with actionable learnings.

    22k GitHub stars~1.1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Intent Debugger

    bydtesla1609/intent-debugger

    Interprets vague, conversational, or intuition-led product, software, AI, and feature ideas as a precise, checkable requirements draft, maps rough descriptions to useful professional terms, exposes…

    126 GitHub stars~2.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Test Ocas Openclaw

    pwrdrvr/openclaw-codex-app-server

    Regression test the OpenClaw Codex App Server plugin against a live local OpenClaw instance in Telegram or Discord.

    265 GitHub stars~4k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • DeepSeek V4 Thinking-Mode Rules

    codewhale-hq/Codewhale

    Three rules for multi-step work with DeepSeek V4 thinking models: verify references, use a verifier subagent before big edits, and write plans with exact path and line.

    41k GitHub stars~443 tokensUpdated today
    Agent WorkflowsAuto-check passed

More from raine/consult-llm

All 12 skills in this repo
  • Implement

    raine/consult-llm

    Explicit workflow for one bounded implementation using source-grounded discovery, a walking slice, evidence-gated review, validation, and commit.

    140 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check: notes
  • Collab

    raine/consult-llm

    Multiple LLMs collaboratively brainstorm solutions, building on each other's ideas across rounds.

    140 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Collab Vs

    raine/consult-llm

    The agent brainstorms with a partner LLM in alternating turns, building on each other's ideas.

    140 GitHub stars~1.8k tokensUpdated 2 days ago
    Auto-check passed
  • Consult

    raine/consult-llm

    Consult an external LLM with the user's query. An agent skill from raine/consult-llm.

    140 GitHub stars~963 tokensUpdated 2 days ago
    Auto-check: notes
  • Consult LLM

    raine/consult-llm

    How to invoke the consult-llm CLI. An agent skill from raine/consult-llm.

    140 GitHub stars~2.9k tokensUpdated 2 days ago
    Auto-check: notes
  • Debate

    raine/consult-llm

    LLMs propose and critique approaches, agent moderates the debate and synthesizes the best solution, then implements.

    140 GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed

Works with

Categories

Questions about Review

What does Review do?

Collect critical feedback from all registered LLMs on an artifact (architecture doc, implementation, plan). Review is an agent skill from raine/consult-llm. Collect critical feedback from all registered LLMs on an artifact (architecture doc, implementation, plan).

When should I use Review?

Review fits situations like: tasks that involve Planning.

How do I install Review in Claude Code?

Run `npx skills add raine/consult-llm --skill review -a claude-code`. Or copy the skill folder (skills/review in raine/consult-llm) into .claude/skills/review in your project. Claude Code loads it when a task matches its description.

How do I install Review in Codex?

Run `npx skills add raine/consult-llm --skill review -a codex`. Or copy the skill folder (skills/review in raine/consult-llm) into .agents/skills/review in your project. Codex loads it when a task matches its description.

Can I use Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add raine/consult-llm --skill review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review, .gemini/skills/review, .github/skills/review and .opencode/skills/review in your project.

What does Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Review is instructions for the agent only. Its frontmatter pre-approves these tools: Bash, Glob, Grep, Read, Write.

Does Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Review safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Review use?

Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Review use?

About 2.4k tokens (SKILL.md is roughly 9.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Review?

Skills that share tags, products or a category with Review: Codex with ChatGPT Planning Loop (XiaoDuoYa/codex-with-chatgpt, 7.1k stars), LLM Council (gcpdev/llm-council-skill, 461 stars), Remember Learnings (dyad-sh/dyad, 22k stars) and Intent Debugger (bydtesla1609/intent-debugger, 126 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Review?

raine (a GitHub user) maintains it in raine/consult-llm, which has 140 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on October 7, 2026.

Source: raine/consult-llm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.