---
name: codex-pr-review
description: >-
  Run an independent local Codex (gpt-6.1-sol, medium) review of a GitHub pull
  request and relay its verdict. Use for every PR an agent opens or updates,
  and whenever the user asks to review a PR "with Codex". The review is posted
  on the PR as approve / request-changes; it never edits, commits or merges.
---

# Local Codex PR review

One command does everything. Run it in the background and wait for it; do not
re-verify the Codex binary, the scripts, the token minter or the gh login first.
The script fails loudly if anything is missing.

```bash
nohup scripts/codex-review/codex-pr-review.sh <repo-dir> <pr-number> [extra-prompt-file] > <log> 2>&1 &
```

Then wait on the log (`until grep -q 'codex-pr-review: done\|FAILED' <log>; do sleep 15; done`)
and keep working. Every exit path prints exactly one of those markers, including preflight
failures. A run normally takes 5-20 minutes. The watchdog kills an attempt, with every
process it spawned, after 8 minutes without an event from the Codex process, or 30 minutes
total, and retries once unless the verdict was already posted. Never launch a bare
`codex exec` for a review.

## What the script does

- Fetches the PR head into the sibling worktree `<repo-dir>-pr-<n>`, so the main checkout is
  never touched. If GitHub has not advertised the PR's Git pull ref yet, it fetches the
  exact head from the PR API and rechecks that head before starting the review. It never
  substitutes a moving branch tip. The tool marks worktrees it created with a `<repo-dir>-pr-<n>.codex-review-owned`
  sentinel; on reruns it resets and cleans only those (ignored files such as `node_modules`
  are kept) and refuses a path at that location it did not create. A per-PR lock refuses a
  second concurrent run; a lock whose owner is dead is reclaimed through a short-lived
  `<lock>.reclaim` mutex, so two reclaimers cannot clobber each other.
  Cancelling the script (Ctrl-C, SIGTERM) stops the runner and every process of the
  attempt before the lock is released; the runner pid is kept in the lock so a runner that
  outlives a hard-killed wrapper still counts as live.
  A worktree created by an earlier version of the script has no sentinel: add the file by
  hand (any content) or remove the worktree with `git worktree remove`.
- Chooses the posting account. GitHub refuses a formal review from the PR author, so a PR
  authored by the logged-in gh user posts through a `mentra-release-coordinator` GitHub App
  token minted by `mentra-release-coordinator-token.mjs`; anyone else's PR posts from the
  logged-in account. Override with `GH_ACCOUNT=app|own`.
- Writes the standard prompt: read the diff against the base, read every existing comment
  (bots included) and judge each on its merits, favour low complexity and a coherent design over
  patchy fixes, no edits or pushes, and finish with `gh pr review <n> --approve` or
  `--request-changes` whose body starts with `Reviewed by local Codex (gpt-6.1-sol, medium).`
  and lists what was checked. If GitHub refuses the formal review it falls back to `--comment`
  with `Approve.` or `Request changes.` as the first line.
- Always adds an "Is this change needed?" step: Codex must judge whether the change is needed or
  desirable for the project at all (real problem, existing or planned alternative, surface area
  and maintenance cost versus benefit, smaller change possible) and let that weigh in the
  verdict, not only implementation quality.
- Runs `codex-review.sh` (watchdog + one retry) and prints Codex's final message. The runner
  starts `codex exec --json` by default, or the small stdio app-server adapter for a configured
  desktop project. Both use their own server event stream as the heartbeat, so other Codex sessions
  on the machine can never be mistaken for this one; every process of the attempt carries a
  unique `CODEX_REVIEW_ATTEMPT` marker so it can be terminated even if Codex exits first. Success means
  a review carrying the marker was found on the head commit after the run started, checked
  through the GitHub API by `review-receipt.sh`; the runner performs the same check before any
  retry so a verdict posted just before a crash is never duplicated. Artifacts:
  `$CODEX_REVIEW_HOME/<repo>-pr-<n>-<timestamp>/{prompt.txt,last-message.txt,runner.log,events-<attempt>.jsonl}`
  (default `~/.codex-reviews`).

## Configuration

### Use a desktop Review project

On the host that runs reviews, create a folder and add it as a local Codex project
named **Review**. Save its absolute path once:

```bash
mkdir -p "$HOME/dev/Codex-Reviews" "${CODEX_REVIEW_HOME:-$HOME/.codex-reviews}"
printf '%s\n' "$HOME/dev/Codex-Reviews" > "${CODEX_REVIEW_HOME:-$HOME/.codex-reviews}/project-directory"
```

Future skill invocations use `codex app-server --stdio` and the documented
`initialize` → `thread/start` → `turn/start` protocol, with that folder as the
saved session's starting directory. The server creates a persistent interactive
session and the adapter gives it a `<repo> #<PR> · review` title. This matters:
`codex exec` has a separate session source that default history lists exclude;
changing its cwd alone does not establish desktop visibility. `--thread-source`
is an analytics label and does not change that history source.

The reviewer receives the separate PR worktree path for all code inspection and tests;
worktree ownership, locks, watchdogs and verdict checks still apply. Existing sessions
are not moved. This is configured per host: a Mini review needs the folder registered
as a project on the Mini's connection. Do not use `--ephemeral` or a separate
`CODEX_HOME`, which would hide sessions from the usual desktop account.

For durable project assignment, initialize the same host's `codex app-server --stdio`
and call its documented `project/list` method. Follow pagination and match the
project's `roots[].path` to the exact configured directory. Save that returned ID in
`$CODEX_REVIEW_HOME/project-id` (default `~/.codex-reviews/project-id`) or set
`CODEX_REVIEW_PROJECT_ID`. Desktop-tool `list_projects` IDs can use a different
namespace: do not copy those IDs into this configuration. The adapter passes the
app-server ID as `thread/start.projectId` and verifies the response. It never reads
or edits Codex's database. Without an ID, grouping depends on the registered folder;
verify visibility on the next actual review rather than creating a duplicate review
as a probe.

Override the saved path with `CODEX_REVIEW_PROJECT_DIR=/absolute/path`, or set it to
an empty string for one invocation to start in the PR worktree instead. Without a
saved path or override, the existing `exec` behavior is unchanged. A configured
missing directory is an error, so reviews cannot silently appear somewhere else.
An explicit directory override ignores the saved project ID; also pass
`CODEX_REVIEW_PROJECT_ID` if that different folder needs a durable assignment.
Project reviews require Node.js and a Codex version supporting the app-server
protocol. Incompatible protocol, failed/interrupted turns and missing final text
fail the attempt; they never produce a successful review merely from progress.

Protocol reference: [Codex App Server](https://learn.chatgpt.com/docs/app-server),
including its `sourceKinds` history filter. The adapter uses the server's normal
session source; it does not relabel existing sessions or spoof a source.

| Variable                                                      | Default                                                | Purpose                                                                                       |
| ------------------------------------------------------------- | ------------------------------------------------------ | --------------------------------------------------------------------------------------------- |
| `CODEX_BIN`                                                   | `codex`                                                | Codex CLI binary. Point it at a specific install if the shim on PATH does not support `exec`. |
| `CODEX_REVIEW_MODEL` / `CODEX_REVIEW_EFFORT`                  | `gpt-6.1-sol` / `medium`                               | Model and reasoning effort.                                                                   |
| `CODEX_REVIEW_HOME`                                           | `~/.codex-reviews`                                     | Where prompts, final messages and runner logs are kept.                                       |
| `CODEX_REVIEW_PROJECT_DIR`                                    | Saved `project-directory` file, otherwise PR worktree | Starting directory for future sessions; register that folder in Codex to group reviews. |
| `CODEX_REVIEW_PROJECT_ID`                                     | Saved `project-id` file, otherwise unset             | Optional durable project assignment for the configured project on this host. |
| `MENTRA_RELEASE_COORDINATOR_KEY`                              | `~/.config/mentra-release-coordinator/private-key.pem` | App private key (0600) for posting on your own PRs.                                           |
| `GH_ACCOUNT`                                                  | auto                                                   | `app` or `own`, see above.                                                                    |
| `STALL_SECONDS` / `MAX_SECONDS` / `ATTEMPTS` / `POLL_SECONDS` | 480 / 1800 / 2 / 5                                     | Watchdog limits and poll interval.                                                            |
| `LOCK_STALE_SECONDS`                                          | 60                                                     | Age after which a lock directory without a pid file is reclaimed.                             |

The App must be installed on the target repository or the mint fails with HTTP 422 and the run
falls back to a comment review from the logged-in account.

With an already configured GitHub App installation credential, set `GH_ACCOUNT=own`
to retain that credential. An explicit account skips the user-only `/user` lookup;
it does not grant additional access or change the review receipt requirements.
The usual self-review comment fallback still applies. Controllers can inspect
`scripts/codex-review/codex-pr-review.sh --capabilities` before admitting this mode.

## Tests

`bun test scripts/codex-review` runs the lifecycle suite with fake `gh` and `codex` binaries:
markers on every exit path, worktree ownership, lock behaviour, receipt-before-retry,
and stall handling. Project transport tests cover the real JSON protocol, completion
ordering, final-text validation, cancellation and detached-child cleanup. They do
not claim a live desktop sidebar check. CI runs them on changes under `scripts/codex-review/`.

## Extra prompt file

Only when the review needs context Codex cannot get from the PR itself (a design decision made
in chat, a known-flaky test to ignore). Do not use it to steer the verdict on a specific
comment; let Codex weigh comments itself.

## After the run

- Relay the verdict and Codex's list to the user. Do not merge.
- If the PR is the agent's own and the verdict is request-changes: read the bot comments too,
  fix real defects with a coherent design change, say which comments were ignored and why, push,
  and rerun the script on the new head.
- If the PR belongs to someone else, report only; do not rework unless asked.
- On `codex-pr-review: FAILED`, read `runner.log`, report the failure, and do not claim a review
  was posted.
- Remove completed review worktrees first, before finished fixer worktrees, once
  the canonical outcome and needed reproduction material are preserved outside
  the checkout. Wait for the wrapper, runner and their children to exit; keep a
  checkout still needed by active or saved review follow-up.
- Before removal, check process use, incoming dependency symlinks and Git
  common-directory consumers. Preserve source commits/refs, review receipts,
  logs, evidence and shared dependencies. Use `git worktree remove`, never
  `--force`; if the checkout is still needed or removal refuses, keep it and
  report why. Store shared dependencies outside disposable review worktrees.
