Agent skill

Codex Review

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Independently validate the current analysis with a second model (OpenAI Codex).

MITAuto-check passedDatabases

Install Codex Review

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill codex-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst codex-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/codex-review .claude/skills/codex-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
codex-review
GitHub stars
304
Token cost
~3k tokens
SKILL.md length
1,535 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Independently validate the current analysis with a second model (OpenAI Codex).

  • Works in 7 steps: Preflight: is Codex usable? (decision… → Resolve what's being validated → Write the validation brief (blind to… → …
  • The user types /codex-review
  • SKILL.md covers Purpose, When to Use, Invocation and Instructions, plus 3 more sections
  • Calls codex, python3 and npm

What it does

Codex Review is an agent skill from ai-analyst-lab/ai-analyst. Independently validate the current analysis with a second model (OpenAI Codex). Codex re-derives the same answer from the same data — blind to Claude's SQL and numbers — and the skill reports AGREE / DISAGREE / PARTIAL per finding. Use when the user types "/codex-review", or says "validate with codex", "codex review", "second opinion from codex", "have the other model check this", "independently verify this analysis", "does codex agree", "cross-check this with gpt/codex", or wants a different model to confirm a…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Databases, covering SQL. It works with SQL. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • The user types /codex-review
  • Says validate with codex
  • Second opinion from codex
  • Have the other model check this

Example prompts

  • “/codex-review”
  • “validate with codex”
  • “codex review”
  • “/codex-review”

Requirements

  • Python 3
  • Node.js

Workflow steps

7 steps, taken from the step headings in SKILL.md.

  1. Preflight: is Codex usable? (decision matrix)
  2. Resolve what's being validated
  3. Write the validation brief (blind to Claude's numbers)
  4. Run Codex independently
  5. Compare (skeptical reconciliation)
  6. Record the run (tracked + auditable)
  7. Report (short, on-screen)

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • codex
    • python3
    • npm
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Codex Review loads about 3k tokens when it runs. Until then it costs about 191 tokens; SKILL.md has 1,535 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~191
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 1,535 words, ~3,011 tokens.

Download SKILL.mdSave it as .claude/skills/codex-review/SKILL.md (or your agent's skills folder).
name
codex-review
description
Independently validate the current analysis with a second model (OpenAI Codex). Codex re-derives the same answer from the same data — blind to Claude's SQL and numbers — and the skill reports AGREE / DISAGREE / PARTIAL per finding. Use when the user types "/codex-review", or says "validate with codex", "codex review", "second opinion from codex", "have the other model check this", "independently verify this analysis", "does codex agree", "cross-check this with gpt/codex", or wants a different model to confirm a result before acting on it. This is multi-model validation: a real independent re-analysis, not a critique of Claude's work. If the Codex plugin or CLI isn't installed, this skill detects that and walks the user through setup first.

Skill: Codex review

Purpose

Have a second model (OpenAI Codex) independently re-derive the current analysis from the same data and compare it to Claude's original. Codex gets the question and the metric definitions, but never sees Claude's SQL, numbers, or conclusions — it writes its own queries and computes its own results. The skill then reconciles the two: AGREE, DISAGREE, or PARTIAL per finding. Two models agreeing from independent derivations is strong evidence the analysis is sound; a disagreement points to exactly where to look.

This pairs with /reliability (same model, run N times — tests stability). /codex-review uses a different model once — it tests correctness by independent agreement.

When to Use

  • User says /codex-review, "validate with codex", "codex review", "second opinion from codex", "independently verify this", "does codex agree", "cross-check with the other model"
  • After producing a finding the user is about to act on and wants a second model to confirm
  • Routed here whenever multi-model validation of an analytical result is wanted

Invocation

/codex-review [finding or artifact path] — validate the most recent analysis by default, or scope to a single finding/file if given. Example: /codex-review after answering "What's our 30-day retention?"

Instructions

⛔ HARD GATE — read before anything else

This skill is worthless unless a different model (Codex) does the validation. If Codex is not ready, you (Claude) MUST NOT perform the validation yourself. Claude re-checking Claude's analysis is circular — it produces a confident "validated ✓" that means nothing and actively misleads the student.

The rule: if Step 1's preflight returns a non-empty missing list, your ONLY job this turn is to help the student set up Codex. You may not proceed to Steps 2–7, and you may not substitute any other model, your own reasoning, a re-run of the SQL, or an "approximate" check. There is no fallback that uses Claude. Setup is the task when Codex is missing — completing it is the helpful outcome, not skipping ahead to a verdict.

Step 1 — Preflight: is Codex usable? (decision matrix)

Run the deterministic check:

bash
python3 helpers/provenance/codex_validation.py --check

It returns JSON: {"codex_cli", "plugin", "auth", "missing": [...]}. Route on missing:

  • Empty missing → Codex is ready. Go to Step 2.
  • "codex_cli" present → the Codex CLI isn't installed. Tell the user to run:
    bash
    npm install -g @openai/codex
    (Requires Node.js 18.18+.)
  • "plugin" present → the Claude Code plugin isn't installed. Show these commands for the user to paste (the skill cannot run them — they're interactive):
    /plugin marketplace add openai/codex-plugin-cc
    /plugin install codex@openai-codex
    /reload-plugins
    /codex:setup
  • "auth" present → Codex is installed but not authenticated. Tell the user to run:
    bash
    codex login
    (Sign in with a ChatGPT account or an API key.)

If missing is non-empty, stop after giving the setup step and end the turn with "Once that's done, re-run /codex-review and I'll have Codex check it." The next invocation re-runs --check and proceeds only when missing is empty.

Restart gate. If the student just installed the plugin, also remind them the plugin's tools aren't loaded until they run /reload-plugins — so the sequence is install → /reload-plugins → re-run /codex-review. (auth is best-effort: if --check returns auth: null with the CLI and plugin present, proceed — the live Codex run is the real gate and will surface any login error.)

Keep this simple and one-step-at-a-time: name only the first missing piece, let the student fix it, then re-run the check. Setup may take two or three turns (CLI, then plugin + reload, then login); that is the expected, correct path — not a detour from the "real" work.

Step 2 — Resolve what's being validated

Identify, for the most recent analysis (or the scoped finding):

  • The question it answered.
  • The metric definitions / scope / time-window Claude used — pull from the metric dictionary (metrics/index.yaml), the analysis design spec, or the analysis itself.
  • The active dataset (.knowledge/active.yaml).
  • Claude's original result(s) — the headline number(s), the SQL, and the conclusion.

If it's ambiguous what to validate (no recent finding, multiple candidates), ask the user which finding or artifact to check, and offer a path.

Step 3 — Write the validation brief (blind to Claude's numbers)

Create a timestamped run directory: working/codex_validation/<UTC-timestamp>-<question-slug>/.

Write brief.md in it containing ONLY what Codex needs to answer the same question the same way, independently:

  • The question.
  • The metric definition(s), scope, and time-window (so Codex measures the same thing).
  • The active dataset id and how to reach the data: read .knowledge/active.yaml, the active dataset's .knowledge/datasets/{active}/schema.md and quirks.md; connect with from helpers.data.connection_manager import ConnectionManager (or the local DuckDB/CSV fallback in the dataset manifest's local_data if no warehouse is reachable).
  • An instruction to log its queries the way the repo expects.

Do NOT put Claude's SQL, result numbers, or conclusion in brief.md. That blindness is the whole point — it's what makes Codex's derivation independent.

Separately, stash Claude's original result in claude_original.md in the same run dir (headline number(s), SQL, conclusion). This file is for the Step 5 comparison only — it is not given to Codex.

Step 4 — Run Codex independently

Dispatch the codex:codex-rescue subagent (Agent tool) with brief.md and this output contract:

Independently answer the analytics question in this brief against the active dataset. Connect to the data and write your own SQL — do not ask for or assume anyone else's queries or numbers. Use the metric definition exactly as given. Log your queries. Then report ONLY:

  • headline: <the single number you'd report> (one per finding if multiple)
  • sql: <the query/queries you actually ran>
  • measured: <numerator, denominator, grain, window, filters>
  • conclusion: <one or two sentences>

Capture Codex's full response to codex_independent.md in the run dir.

(Fallback: if the subagent's output is unreliable or unavailable, run codex exec via Bash with the same brief and output contract, and save the result to the same file.)

Show full SKILL.md (620 more words)Show less
Step 5 — Compare (skeptical reconciliation)

Put Codex's numbers next to Claude's (claude_original.md) and assign a verdict per finding:

  • AGREE — numbers match within a sensible tolerance and the conclusions align.
  • DISAGREE — a material gap. Show both numbers, both SQL approaches, and the most likely cause (different filter, cohort, join grain, window). Investigate which derivation is right — do not average them.
  • PARTIAL — same direction, different magnitude, or agreement on some sub-results only.

Write a verdict.md (the human-readable comparison table) AND a verdict.json for the deterministic audit log, shaped: {"question": "<q>", "model": "codex", "findings": [{"name": "<finding>", "verdict": "AGREE|DISAGREE|PARTIAL"}, ...]}.

Step 6 — Record the run (tracked + auditable)

Append the run to the audit log:

bash
python3 helpers/provenance/codex_validation.py --log <that run directory>

It reads verdict.json, counts the verdicts deterministically, and appends one line to .knowledge/codex-review/log.jsonl. The run dir now holds the full provenance: brief.md, claude_original.md, codex_independent.md, verdict.md, verdict.json.

Step 7 — Report (short, on-screen)

Frame it as independent multi-model validation, then show the comparison:

  • Headline: e.g. "Claude found 38% 30-day retention; Codex independently derived 38% from its own query — AGREE." Or: "Codex got 31% vs Claude's 38% — DISAGREE: Codex filtered to activated users only; Claude counted all signups."
  • The per-finding Finding | Claude | Codex | Verdict | Why table from verdict.md.
  • Where it was saved (the run dir + .knowledge/codex-review/log.jsonl).

Then the honest framing:

  • All AGREE — "A second model independently reproduced this from its own queries. That's strong evidence the result is sound — not proof, but two independent derivations agreeing."
  • Any DISAGREE / PARTIAL — "The two models diverge here. That's the check earning its keep: one of these derivations is wrong, or the metric is under-defined. Resolve the gap before acting on the number."

On any DISAGREE, offer to re-run the relevant analysis step, define the metric via /metric-spec, or log the lesson via /log-correction.

Rules

  1. No Codex, no validation — and no Claude fallback. If preflight's missing is non-empty, stop at setup. Never validate with Claude, another model, your own reasoning, or a re-run of the SQL. A Claude-checks-Claude result is circular and must never be presented as a validation. This is the one rule that cannot be bent. (See the Hard Gate above.)
  2. Codex must be blind to Claude's numbers. Never include Claude's SQL, result numbers, or conclusions in brief.md. If you can't keep them out, the run isn't independent — say so rather than presenting a false validation.
  3. Counting is deterministic. Verdict tallies come from codex_validation.py --log reading verdict.json, never estimated in prose.
  4. One missing piece at a time in preflight. Don't dump every install step at once — name the first gap, let the user fix it, re-check.
  5. Respect the restart gate. After a plugin install, halt until /reload-plugins.
  6. Same definitions, independent derivation. Codex answers the same question with the same metric definition — only the SQL and numbers are its own.

Edge Cases

  • Codex not installed (the common student case) → help the student install/log in, then stop; validation happens on the next run once --check is clean.
  • No recent analysis to validate → ask the user what to check; offer a path or finding.
  • Codex can't reach the warehouse → the brief should hand it the local DuckDB/CSV fallback (manifest.local_data) so it can still derive independently.
  • Codex defines the metric differently anyway → flag it: the disagreement may be definitional, not an error. Surface both definitions and recommend /metric-spec.
  • auth: null from preflight → proceed; the live run surfaces any real login error.

Notes

  • The plugin (openai/codex-plugin-cc) is just the simple, supported path to an installed + authenticated Codex CLI. The validation itself runs Codex against the data, not a code diff — the plugin's /codex:review (diff review) is a different thing and isn't used here.
  • Complements /reliability: that re-runs the same model to test stability; this runs a different model once to test correctness by independent agreement.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/codex-review of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Codex Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Codex Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Codex Review this skillai-analyst-lab/ai-analyst304—~3kAutomated safety check: PassMIT
Evolving The Data ModelTriliumNext/Trilium38k—~2.1kAutomated safety check: PassAGPL-3.0
Orchardcore Data MigrationOrchardCMS/OrchardCore8.2k—~1.7kAutomated safety check: PassBSD-3-Clause
SQL Optimization Patternsynulihao/AgentSkillOS61711 repos~3.3kAutomated safety check: PassNone
SQL PortabilityHL7/sql-on-fhir150—~512Automated safety check: PassCustom licence
StmoSAP/project-foxhound1802 repos~1.8kAutomated safety check: PassGPL-3.0

Similar skills

  • Evolving The Data Model

    TriliumNext/Trilium

    A skill your agent uses when adding a DB migration or a new column/field to a Becca entity in Trilium ("add a migration", "new column on notes/attributes", "ALTER TABLE", "add a field to…

    38k GitHub stars~2.1k tokensUpdated today
    DatabasesAuto-check passed
  • Orchardcore Data Migration

    OrchardCMS/OrchardCore

    Creates and updates OrchardCore data migrations (DataMigration classes with CreateAsync/UpdateFromX).

    8.2k GitHub stars~1.7k tokensUpdated today
    DatabasesAuto-check passed
  • SQL Optimization Patterns

    ynulihao/AgentSkillOS

    Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries.

    617 GitHub starsUsed in 11 repos~3.3k tokens
    DatabasesAuto-check passed
  • SQL Portability

    HL7/sql-on-fhir

    Analyse whether a SQL query is portable across database implementations using sqlglot transpilation.

    150 GitHub stars~512 tokensUpdated 2 days ago
    DatabasesAuto-check passed
  • Stmo

    SAP/project-foxhound

    Official

    Manage Redash queries and dashboards on Mozilla's STMO (sql.telemetry.mozilla.org) using stmo-cli.

    180 GitHub starsUsed in 2 repos~1.8k tokens
    DatabasesAuto-check passed
  • DB Migrations

    kurealnum/dotfiles

    A skill your agent uses when generating or regenerating Drizzle migration files, changing database schema tables or columns, resolving migration sequence conflicts after rebase, reviewing migration…

    290 GitHub stars~820 tokensUpdated 5 mo ago
    DatabasesAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 8 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 8 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 8 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed

Works with

Categories

Questions about Codex Review

What does Codex Review do?

Independently validate the current analysis with a second model (OpenAI Codex). Codex Review is an agent skill from ai-analyst-lab/ai-analyst. Independently validate the current analysis with a second model (OpenAI Codex).

When should I use Codex Review?

Codex Review fits situations like: the user types /codex-review; says validate with codex; second opinion from codex; have the other model check this.

How do I install Codex Review in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill codex-review -a claude-code`. Or copy the skill folder (.claude/skills/codex-review in ai-analyst-lab/ai-analyst) into .claude/skills/codex-review in your project. Claude Code loads it when a task matches its description.

How do I install Codex Review in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill codex-review -a codex`. Or copy the skill folder (.claude/skills/codex-review in ai-analyst-lab/ai-analyst) into .agents/skills/codex-review in your project. Codex loads it when a task matches its description.

Can I use Codex Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill codex-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/codex-review, .gemini/skills/codex-review, .github/skills/codex-review and .opencode/skills/codex-review in your project.

What does Codex Review need to run?

Going by SKILL.md and its folder, Codex Review needs the command-line tools its instructions call (codex, python3, npm and claude). Our summary lists: Python 3; Node.js.

Does Codex Review access the network?

SKILL.md contains no URLs. Its commands use npm, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Codex Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Codex Review use?

Codex Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Codex Review use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Codex Review?

Skills that share tags, products or a category with Codex Review: Evolving The Data Model (TriliumNext/Trilium, 38k stars), Orchardcore Data Migration (OrchardCMS/OrchardCore, 8.2k stars), SQL Optimization Patterns (ynulihao/AgentSkillOS, 617 stars) and SQL Portability (HL7/sql-on-fhir, 150 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Codex Review?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.