Agent skill

Codex Review

by suyuan2022 in suyuan2022/suyuan-skill

Dual-model code review via Codex CLI. An agent skill from suyuan2022/suyuan-skill.

MITAuto-check: warningsSecurity

Install Codex Review

The automated check flagged lines worth reading first. See the safety section below.

skills CLI
$ npx skills add suyuan2022/suyuan-skill --skill codex-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install suyuan2022/suyuan-skill codex-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/suyuan2022/suyuan-skill.git skills-src && mkdir -p .claude/skills && cp -r skills-src/codex-review .claude/skills/codex-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
codex-review
GitHub stars
290
Token cost
~3.5k tokens
SKILL.md length
1,293 words
Files
6 (incl. references)
Skills in repo
6
Repo updated
First seen
Licence
MIT

At a glance

Dual-model code review via Codex CLI. An agent skill from suyuan2022/suyuan-skill.

  • Works in 3 steps: PRD documentation is the ultimate judge… → If the PRD doesn't cover it, check… → Only record items needing confirmation…
  • Let Codex check
  • Calls codex and npm
  • After complex coding tasks

What it does

Codex Review is an agent skill from suyuan2022/suyuan-skill. Dual-model code review via Codex CLI. Two GPT models review independently, Claude arbitrates disagreements. Supports review (read-only), fix (write), audit (full security scan), and autopilot (unattended) modes. Triggers on "review", "let Codex check", "security audit", "OWASP", "full scan". Also auto-triggers after complex coding tasks.

Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/code-security.md`, `references/experience-log.md` and `references/general.md`).

It sits in Security, covering Security review and Web application vulnerabilities. It works with OpenAI. The repository describes itself as: Claude Code Skills by Suyuan — AI-native productivity tools for real practitioners. The licence is MIT.

When your agent uses it

  • Let Codex check
  • After complex coding tasks

Example prompts

  • “review”
  • “let Codex check”
  • “security audit”
  • “/codex-review”

Requirements

  • Node.js

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. PRD documentation is the ultimate judge — if the PRD explicitly states something, follow the PRD
  2. If the PRD doesn't cover it, check existing code behavior before deciding
  3. Only record items needing confirmation when they involve product direction / business logic tradeoffs; batch-report at the end

What it can do on your machine

Read from SKILL.md and the folder at commit 423473d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • codex
    • npm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Codex Review loads about 3.5k tokens when it runs, and up to ~8.3k if it reads all its reference files. Until then it costs about 88 tokens; SKILL.md has 1,293 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~88
When it runs · the whole SKILL.md, loaded when a task matches
~3.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: warnings

The automated check found patterns that need a careful read before installing.

  • WarningTells the agent its actions are pre-authorized / not to stop for confirmationSKILL.md:30
    e arbitrates and decides whether to fix without asking the user. Decision criteria:

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from suyuan2022/suyuan-skill at commit 423473d, republished under its MIT licence (© suyuan2022). 1,293 words, ~3,548 tokens.

Download SKILL.mdSave it as .claude/skills/codex-review/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
codex-review
description
Dual-model code review via Codex CLI. Two GPT models review independently, Claude arbitrates disagreements. Supports review (read-only), fix (write), audit (full security scan), and autopilot (unattended) modes. Triggers on "review", "let Codex check", "security audit", "OWASP", "full scan". Also auto-triggers after complex coding tasks.

Prerequisites

  • Codex CLI installed and authenticated: npm install -g @openai/codex then codex login
  • An OpenAI account with access to GPT-5.4 and GPT-5.2 (or adjust model names below)

Roles
  • Codex (Model A + Model B): Two independent reviewers examining the same code simultaneously, unaware of each other
  • Claude: Controller + arbitrator. Constructs prompts, dispatches both models in parallel, compares findings, adjudicates disagreements, makes final decisions

Modes
SignalModeSandbox
review / check / look atreviewread-only
fix / modify / let Codex fixfixworkspace-write
security audit / OWASP / full scanauditread-only
autopilot / unattended / long-runningautopilotread-only
uncleardefault to reviewread-only

autopilot mode: When the user is away and Claude is executing code tasks autonomously. After Codex review results come back, Claude arbitrates and decides whether to fix without asking the user. Decision criteria:

  1. PRD documentation is the ultimate judge — if the PRD explicitly states something, follow the PRD
  2. If the PRD doesn't cover it, check existing code behavior before deciding
  3. Only record items needing confirmation when they involve product direction / business logic tradeoffs; batch-report at the end

Autopilot is not the default — requires explicit user request ("autopilot", "unattended", "run without me").


Post-Arbitration Action Classification

After arbitration, Claude handles each finding by these rules:

Fix directly (don't ask the user):

  • Code bugs, compile/runtime errors
  • Documentation wording / phrasing contradictions (not direction choices)
  • Inconsistencies with explicit PRD decisions
  • Security vulnerabilities

Must ask the user (don't skip any):

  • Product/business decisions needed (feature tradeoffs, scope changes)
  • User-provided information needed (credentials, domains, third-party service configs)
  • Choices affecting user experience not covered by PRD
  • Cost/budget decisions

In autopilot mode, "must ask" items are accumulated and batch-reported at the end without interrupting execution.


Dual-Model Parallel Review

Each review dispatches two parallel codex exec calls via Bash, one per model:

bash
# Model A (run in parallel with Model B)
codex exec -m "gpt-5.4" -s "read-only" -C "$REPO_DIR" \
  --skip-git-repo-check --ephemeral \
  -o /tmp/codex-review-a.md \
  - <<'PROMPT'
[constructed prompt here]
PROMPT

# Model B (run in parallel with Model A)
codex exec -m "gpt-5.2" -s "read-only" -C "$REPO_DIR" \
  --skip-git-repo-check --ephemeral \
  -o /tmp/codex-review-b.md \
  - <<'PROMPT'
[same prompt here]
PROMPT

Both calls use the same prompt but different models. Wait for both results before entering arbitration.

For fix mode, only one model executes the fix (default Model A), but the review phase still uses both models:

bash
codex exec -m "gpt-5.4" -s "workspace-write" -C "$REPO_DIR" \
  --skip-git-repo-check --full-auto \
  -o /tmp/codex-fix.md \
  - <<'PROMPT'
[fix prompt here]
PROMPT

To resume a session (for review-fix-review cycles):

bash
codex exec resume SESSION_ID "follow-up prompt here"

The session ID is printed in the codex exec output header (look for session id: xxx).


Arbitration: Claude as Judge

When both sets of findings arrive, Claude processes them in three zones:

Consensus zone — Both models flagged the same issue. Essentially confirmed; adopt directly. If fix suggestions differ, pick the one with harder evidence (file:line + code quote > vague description).

Single-source zone — Only one model flagged it. Claude reads the code to verify:

  • Hard evidence (specific line + reproducible logic chain) → adopt
  • Soft evidence ("might be a problem" without pinpointing) → downgrade to suggestion, don't force fix
  • Claude judges it's a false positive → discard, explain reasoning in output

Conflict zone — Two models disagree (A says problem, B says fine). Claude reads the code and rules. If unresolvable (e.g., business logic judgment), escalate to user.

Output format: Tag each finding with its source — [A+B] consensus, [A] or [B] single-source, [Conflict→Claude] conflict. User sees at a glance which findings are iron-clad vs. single opinions.


Invocation Parameters

-C must be set to the current working directory (Codex needs access to the same codebase).

Don't embed file contents in the prompt — Codex can read files itself. Only provide absolute paths in the prompt and let Codex read them. Save the token budget for review requirements, context, and focus areas.

Use --ephemeral for one-shot reviews (no session persistence needed). Omit it when you plan to do review-fix-review cycles.

Use --skip-git-repo-check when the target directory isn't a git repo.

Maximize initial review input — On the first call, give Codex the full review scope up front. Locating code is the most expensive phase of the first review; specific paths and ranges skip it entirely:

  • Required: Complete file path list (not vague like "the marketing-related files")
  • Required: Line ranges of interest (path:line-range) or precise diff range (<base>..<head> commit range / direct patch)
  • Required: Hard boundaries — which files/modules are context vs. review targets
  • Required: Internal checks already run (lint / typecheck / key grep results), so Codex doesn't repeat them
  • Principle: The more specific the file list in the first prompt, the shorter the review runs. Put the detail in the first prompt — don't wait for Codex to ask.

⚠️ Hard constraint — the path list is a locator, not a scope cap: The prompt must tell Codex explicitly — the file list/line ranges are locators to help you find entry points quickly, NOT a scope lock restricting your review to only these paths. Codex's review must be as comprehensive as possible:

  • Follow call chains; inspect files outside the list if they are affected
  • Check cross-file side effects, error paths, edge cases — don't constrain yourself to the listed paths
  • If you find issues outside the list, report them. Do NOT skip findings with "not in the given scope" reasoning

Include language like this in the prompt:

"The file list above is a locator to save you from scanning the repo from scratch. It is NOT a scope lock. Your review must be comprehensive — follow call chains, inspect files outside the list if they are affected, and report any issues you find there. Do not limit findings to the listed paths."


Show full SKILL.md (448 more words)Show less
Prompt Construction

Before constructing the prompt, assess the scenario and read the corresponding reference file:

ScenarioReference
Security audit (full scan)references/security-audit.md
Security-sensitive (auth/payment/data/crypto)references/code-security.md
Non-code (docs/config/prompts)references/non-code.md
Otherreferences/general.md
Before every promptScan references/experience-log.md for known pitfalls
XML Block Structure

All prompts sent to Codex use XML blocks, each with a fixed responsibility:

XML BlockPurposePrinciple
<task>Define the task and review objectiveOne run = one task, no mixing
<context>Code/content/diff/project backgroundFacts only, no judgments; paths for Codex to read
<instructions>Review dimensions + per-finding output formatRequire evidence: file:line + issue + reason + fix
<grounding_rules>Inference vs. confirmation boundaryInferences must be labeled "Inference:"; never stated as confirmed
<verification_loop>Anti-rubber-stampFinal check: each finding is material, supported, reproducible
<dig_deeper_nudge>Second-order check listGuide Codex to check cross-file impact, edge cases, error paths
<review_coverage>Review scope declarationExplicitly state what was and wasn't checked
<verdict>Final judgmentPass / Conditional pass / Fail
<default_follow_through_policy>Behavior on missing contextWrite "Cannot verify: [reason]" when uncertain; don't guess
<action_safety>Fix mode onlyRestrict change scope, no unrelated refactoring

Block selection by mode: review/audit must include grounding_rules + verification_loop + dig_deeper_nudge + review_coverage + verdict; fix mode adds action_safety, removes dig_deeper_nudge.

Anti-rubber-stamp: If either model returns "No issues found", check review_coverage scope adequacy. If scope is too narrow, do one more round. Both models say clean + adequate scope → pass.


Review-Fix Cycle
  1. Claude finishes implementation → constructs prompt → parallel dispatch to both models
  2. Both findings arrive → Claude arbitrates (consensus / single-source / conflict)
  3. Adopted findings → Claude fixes
  4. After fix → parallel dispatch both models again for re-review
  5. Both pass + adequate review scope → cycle ends
  6. Max 3 rounds. Beyond 3 → stop and report to user

Re-review prompt (resume session, or new session with diff):

Re-review's bottleneck isn't "finding the entry point" — it's verification. Codex has to confirm every fix actually landed and check for second-order problems. A natural-language summary ("fixed the AuthModal logic") forces Codex to locate code on its own, wasting its most expensive time.

You must provide a "Fix Evidence Index" structured at file:line granularity:

Fixed the previous round's findings. Fix evidence index below:

## Finding → Fix Mapping (finding → file:line → reason, all three required)
Every entry must include all three pieces: **what the finding is + where it landed (file:line) + why this fix**. A bare `file:line` without a reason forces Codex to reconstruct the "root cause → fix" logic chain on its own — defeating the purpose of the index.

- **[Finding-1 title]** → `<path>:<line-range>` [VERIFIED FIXED]
  - Reason: <what was done + why this approach; one line connecting root cause → fix>
- **[Finding-2 title]** → `<path>:<line-range>` [VERIFIED FIXED]
  - Reason: <...>
- **[Finding-3 title]** → [ACCEPTED — not fixing]
  - Reason: <why not fixing: risk/cost/unnecessary, or decided in earlier round>
- **[Finding-4 title]** → `<path>:<lines>` [PARTIALLY FIXED]
  - Reason: <what portion was fixed, why the rest is deferred, next-step plan>

(Reason is what Codex uses to judge whether the fix is on-target — not decoration. Without it, re-review either skips verification or spends time re-deriving the fix logic from code. Either way defeats the index.)

> **⚠️ Hard constraint — Reason is Claude's self-report, not Codex's conclusion**:
> The prompt must tell Codex explicitly — **the Reason field is Claude's own claim about the fix rationale (an unverified premise), NOT a fact you can cite as your conclusion**. Codex must:
> - Independently read the code at `<path>:<line-range>` and judge on its own whether the fix addresses the root cause
> - If agreeing with the Reason, produce **its own independent evidence from reading the code** (line number + code quote + reasoning for why this evidence supports the Reason). "Reason looks correct" or simply echoing Claude's wording is NOT acceptable.
> - If disagreeing, state exactly what's wrong and provide counter-evidence
> - ACCEPTED / PARTIALLY FIXED entries also require independent verification of whether the stated justification holds
>
> Include language like this in the prompt:
> > "The Reason field is Claude's self-reported fix rationale — treat it as an unverified claim, not a conclusion. For each VERIFIED FIXED entry, read the code at the stated path:line and produce your OWN independent evidence (line number + code quote + reasoning) for why the fix does or doesn't address the root cause. Echoing Claude's Reason without independent evidence is not acceptable."

## Files Changed This Round
- <path1>: <one-line summary of change>
- <path2>: <one-line summary of change>
(List only files actually touched this round. No vague summaries like "mainly updated marketing files".)

## Diff Range
Commit range: <base>..<head>  or  patch below
[Optional: inline diff]

## Verifications Already Run This Round
- `<command>`: <key result>
- `<command>`: <key result>
(e.g., `tsc --noEmit`: no errors; `rg "@gsap/react"`: 0 matches)

Please re-review based on this index:
1. For each [VERIFIED FIXED], confirm it actually landed (read the corresponding file:line)
2. Check for second-order issues (did changing A affect B?)
3. Evaluate whether ACCEPTED decisions are reasonable
4. Assess whether PARTIALLY FIXED remainders are critical

Why this format: Real Codex feedback shows ~70-80% of re-review time is spent on "locating and cross-referencing code", only 20-30% on actual thinking. A structured index front-loads the verification paths into the prompt so Codex can jump straight to file:line without translating natural-language descriptions into search targets.


Session Continuity

After each review, report both session IDs to the user:

Model A session: SESSION_ID_A
Model B session: SESSION_ID_B

When to Trigger

Review needed: New features (>=2 files or new functions), refactoring, complex bug fixes, user explicitly requests, high-risk scenarios.

Skip: Typos/naming/comments, single-line simple fixes, formatting adjustments, user asks for quick turnaround.


Interaction Language
  • To Codex: match the user's language (Chinese prompt → Chinese review)
  • To user: match the user's language

© suyuan2022, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in codex-review of suyuan2022/suyuan-skill.

  • SKILL.md
  • references/code-security.md
  • references/experience-log.md
  • references/general.md
  • references/non-code.md
  • references/security-audit.md

Open the folder on GitHubat commit 423473d

Compare with similar skills

Codex Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Codex Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Codex Review this skillsuyuan2022/suyuan-skill290—~3.5kAutomated safety check: WarnMIT
Security Auditoreigent-ai/eigent15k—~1.8kAutomated safety check: NotesApache-2.0
Security Reviewjewbetcha/opentrace11618 repos~3.1kAutomated safety check: NotesMIT
Code Audit3stoneBrother/code-audit8921 repos~2.7kAutomated safety check: PassNone
Security Audit Scannerruvnet/ruflo74k2 repos~823Automated safety check: PassMIT
Vibe Checkbenavlabs/vibe-check118—~1.1kAutomated safety check: NotesMIT

Similar skills

  • Security Auditor

    eigent-ai/eigent

    Audits source code, dependencies and config files for vulnerabilities and hardcoded secrets, using two bundled Python scanners and an OWASP Top 10 checklist.

    15k GitHub stars~1.8k tokensUpdated today
    SecurityAuto-check: notes
  • Security Review

    jewbetcha/opentrace

    A skill your agent uses when adding authentication, handling user input, working with secrets, creating API endpoints, or implementing payment/sensitive features.

    116 GitHub starsUsed in 18 repos~3.1k tokens
    SecurityAuto-check: notes
  • Code Audit

    3stoneBrother/code-audit

    Professional code security audit skill covering 55+ vulnerability types.

    892 GitHub starsUsed in 1 repo~2.7k tokens
    SecurityAuto-check passed
  • Runs claude-flow CLI security scans for input validation, path traversal, SQL injection, XSS, hardcoded secrets and known CVEs, and writes an audit report.

    74k GitHub starsUsed in 2 repos~823 tokens
    SecurityAuto-check passed
  • Vibe Check

    benavlabs/vibe-check

    Security audit for web apps, especially AI-built ("vibe coded") ones.

    118 GitHub stars~1.1k tokensUpdated 21 days ago
    SecurityAuto-check: notes
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    SecurityAuto-check: notes

More from suyuan2022/suyuan-skill

  • Claude Cleanup Audit

    suyuan2022/suyuan-skill

    Audit and safely modify local Claude Code and Claude Desktop cleanup/privacy settings on macOS.

    290 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Claude Cleanup

    suyuan2022/suyuan-skill

    在 macOS 上安全审计、完整备份、清理或重置 Claude Code 与 Claude Desktop 本机状态。适用于清理可再生缓存和日志、轮换本地 ID、重置桌面端登录态、移除应用、继承现有遥测状态、启用最小提示词模式或修改台北时区;不用于规避封禁或平台风控。

    290 GitHub stars~634 tokensUpdated 1 mo ago
    Auto-check: notes
  • Task Triage

    suyuan2022/suyuan-skill

    Task triage with position-based thinking. An agent skill from suyuan2022/suyuan-skill.

    290 GitHub stars~1.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Break AI Slop

    suyuan2022/suyuan-skill

    Break AI Slop — force the AI to think like a real expert before doing anything.

    290 GitHub stars~981 tokensUpdated 1 mo ago
    Auto-check passed
  • Whatis

    suyuan2022/suyuan-skill

    What Is This — explain anything as a picture-first HTML page.

    290 GitHub stars~606 tokensUpdated 1 mo ago
    Auto-check passed

Works with

Categories

Questions about Codex Review

What does Codex Review do?

Dual-model code review via Codex CLI. An agent skill from suyuan2022/suyuan-skill. Codex Review is an agent skill from suyuan2022/suyuan-skill. Dual-model code review via Codex CLI.

When should I use Codex Review?

Codex Review fits situations like: let Codex check; after complex coding tasks.

How do I install Codex Review in Claude Code?

Run `npx skills add suyuan2022/suyuan-skill --skill codex-review -a claude-code`. Or copy the skill folder (codex-review in suyuan2022/suyuan-skill) into .claude/skills/codex-review in your project. Claude Code loads it when a task matches its description.

How do I install Codex Review in Codex?

Run `npx skills add suyuan2022/suyuan-skill --skill codex-review -a codex`. Or copy the skill folder (codex-review in suyuan2022/suyuan-skill) into .agents/skills/codex-review in your project. Codex loads it when a task matches its description.

Can I use Codex Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add suyuan2022/suyuan-skill --skill codex-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/codex-review, .gemini/skills/codex-review, .github/skills/codex-review and .opencode/skills/codex-review in your project.

What does Codex Review need to run?

Going by SKILL.md and its folder, Codex Review needs the command-line tools its instructions call (codex and npm). Our summary lists: Node.js.

Does Codex Review access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Codex Review safe to install?

Our automated static check of SKILL.md flagged 1 warning(s): tells the agent its actions are pre-authorized / not to stop for confirmation. Read the flagged lines before installing; the check is not a guarantee either way.

What licence does Codex Review use?

Codex Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Codex Review use?

About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.7k tokens, read only when the agent opens those files.

What are the alternatives to Codex Review?

Skills that share tags, products or a category with Codex Review: Security Auditor (eigent-ai/eigent, 15k stars), Security Review (jewbetcha/opentrace, 116 stars), Code Audit (3stoneBrother/code-audit, 892 stars) and Security Audit Scanner (ruvnet/ruflo, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Codex Review?

suyuan2022 (a GitHub user) maintains it in suyuan2022/suyuan-skill, which has 290 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on August 24, 2026.

Source: suyuan2022/suyuan-skill on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.