---
name: github-pr-reviewer
description: >
  Create an automation that reviews GitHub pull requests when a configured
  reviewer is requested or a trigger label is applied. Starts one OpenHands
  review conversation per request with the pull request's exact head checked
  out, and publishes the review to GitHub.
triggers:
  - /pr-reviewer:setup
---

# GitHub PR Reviewer Automation

## Agent Canvas catalog

For new Agent Canvas installations, use the **GitHub code review** catalog
entry. Its deterministic `worker.py` delegates each requested exact head to a
stable conversation using the selected agent profile. It supports GitHub
reviewer-request events and scheduled label scans. The manual upload flow below
remains for existing deployments and is deprecated for new installations.

Create a cron automation that watches one or more GitHub repositories for pull
requests with a review trigger label, starts an OpenHands review conversation
once per label event, and publishes the AI review to GitHub.
Windows PowerShell equivalents for the setup, packaging, upload, and API-check shell snippets are in `references/windows.md`.

The automation script is deterministic: PR discovery, label-event tracking,
head eligibility, state persistence, stale-result suppression, the repository
checkout, and its removal are all handled in Python. The LLM is invoked only for
the review itself.

Before any review conversation is created, a scheduled scan evaluates the
current head's GitHub-required checks, read through the GraphQL `isRequired`
signal for the pull request. That signal is the merge policy's own source of
truth and is pull-request-scoped, so it stays correct for a stacked PR whose
symbolic base branch carries no branch rules of its own:

- Only checks GitHub reports as required decide the scheduled gate. An optional
  workflow that fails - including one that fails before creating any check run -
  does not block a head whose required checks pass.
- A completed required run on the exact head whose conclusion is `failure`,
  `cancelled`, or `timed_out` blocks the review. Any other unrecognized
  conclusion fails closed as a block rather than silently approving.
- A completed required run whose conclusion is `success`, `neutral`, or
  `skipped` does not block. A required `commit status` context is classified
  alongside required check runs.
- A `queued` or `in_progress` required run on the exact head, and a required
  context that has not reported a current-head check run at all, make the run
  exit with a **waiting on checks** outcome. Unfulfilled is never approval. The
  worker does not hold a slot polling; the next scheduled scan or explicit
  review request retries.
- When the required-check signal cannot be read or is empty, the gate falls back
  to every current-head check run and workflow run, so a red head still blocks.
  In that fallback, workflow runs are read as well as check runs, because a
  workflow can fail before creating any check run - a workflow-level error, or a
  `pull_request` run whose jobs never start. Such a run leaves a failed check
  suite with no check runs under it, so the commit's check-run rollup and
  `gh pr checks` both report success and only the workflow run reveals the red
  CI. A workflow run whose check suite already reported check runs is left to
  those runs, so a workflow is never counted twice.
- A scheduled scan considers the trigger label **and** every open PR, including drafts,
  that still holds an outstanding `all-hands-bot` review request. That is what
  resumes a request made while CI was running: the request is keyed by its own
  `review_requested` event, so repeated scans reuse one conversation and one
  review instead of creating duplicates, and a PR whose request was answered or
  withdrawn simply drops out of `requested_reviewers`.
- The scheduled scan also reviews open, non-draft PRs that **nobody requested**
  and that carry no trigger label, once their current head has no completed
  review by this account. That is what keeps reviewing PRs after the outstanding
  requests are exhausted. The delivery key is the stable
  `scan:{repository}:{number}:{head}`, so repeated scans reuse one conversation
  and one native review, and a changed head becomes eligible again under its new
  SHA.
- Each scheduled scan classifies the **whole open backlog**, including
  unrequested PRs. Blocked, pending, draft, or already-reviewed heads therefore
  cannot hide an eligible head behind an inspection window. The separate
  `max_new_per_run` quota limits only newly created review conversations, not the
  number of PRs inspected; explicit requests and trigger labels retain priority
  when eligible candidates are drained.
- A gate stop on an **unrequested** head posts no managed comment. A PR nobody
  asked about that is merely red or pending is simply skipped without starting an
  agent, which is what keeps a scan over a large backlog from posting a gate
  comment on every PR. The managed comment is the answer to an explicit request,
  so a requested or labeled head that is red or pending still gets its
  explanation.
- Workflow runs are read as well as check runs, because a workflow can fail
  before creating any check run - a workflow-level error, or a `pull_request`
  run whose jobs never start. Such a run leaves a failed check suite with no
  check runs under it, so the commit's check-run rollup and `gh pr checks` both
  report success and only the workflow run reveals the red CI. A workflow run
  whose check suite already reported check runs is left to those runs, so a
  workflow is never counted twice.
- Runs attributed to any other (obsolete) head SHA are ignored, so a stale
  failure cannot block the push that fixed it. This applies to workflow runs
  too.
- Only the latest run of each logical check or workflow counts. A logical check
  is its name plus the reporting app identity, and a logical workflow is its
  name plus workflow ID. The latest is chosen by the run ID (the reliable
  creation sequence) with the start time as a tie-break, so a re-run that fixed
  a check supersedes its earlier failure on the same SHA while a newer queued or
  in-progress re-run supersedes an earlier success and makes the head wait.
  Ordering by the run ID keeps a run whose start time is still absent in its
  true creation position in both directions.

An explicit `all-hands-bot` review request is the intake-policy exception: the
caller asked for that head by name, so the request dispatches even when required
CI is red or pending, and the gate leaves no explanatory comment. Draft, scope,
exact-head, and delivery-deduplication safeguards still apply. A blocked or
waiting scheduled stop consumes no trigger: once the head's required checks are
non-blocking, the scheduled scan, or a new `all-hands-bot` review request,
starts the normal review.

The gate needs no configured list of check names. When it stops a review it
leaves one concise explanation on the PR, identified by a hidden marker carrying
the head SHA and gate category, so a later run for a different head updates that
comment instead of posting another. Only a marker this reviewer account authored
counts as its own comment; every other PR comment is untrusted and can neither
suppress the explanation nor be edited. If the repository's own workflow already
posted a deterministic remediation comment that names the same current-head
checks, the gate adds nothing; a disclosure about some other check does not
suppress it. The gate applies to scheduled label scans; an explicit
`all-hands-bot` review request bypasses it by design.

Each explicit review request - a re-applied trigger label, or a new reviewer
request in event mode - is a fresh review. The conversation for a PR is reused,
but its earlier turns must not be trusted as current: before deciding a verdict
the reviewer re-fetches the mutable GitHub state (the exact head, the PR body,
review comments and threads, review requests, the linked issues' bodies and
labels, and the current-head Actions results) and ignores any earlier finding,
verdict, or label/priority claim that state no longer supports. Repository
analysis the conversation already did, such as reading `AGENTS.md`, stays
useful and is not repeated.

The waiting and blocked explanations name the retry the deployment actually has.
A scheduled run proves a scan is configured, so it says the next scan will
retry. An event-only run does not, so it tells the reader to remove the
outstanding `all-hands-bot` request and request `all-hands-bot` again instead of
promising a scan that does not exist: GitHub will not accept a second request
while the first is still outstanding.

The gate leaves only one explanation per head. When the same head keeps the same
hidden marker but the retry wording changes -- an event-only comment on a head
whose automation is later switched to a cron scan -- the scheduled run rewrites
that managed comment in place, so the comment always names the retry that is
actually deployed, and an unchanged body is left untouched.

## Bounded intake per scheduled scan

A scheduled scan drains outstanding reviewer requests fairly, but starting an
agent for every eligible pull request at once would overload the deployment, so
the scan starts at most `MAX_NEW_PER_RUN` new review conversations, counted
**across every configured repository** rather than per repository. The default
is `2`, and the rendered `config.json` overrides it through the
`max_new_per_run` key that `github-issue-to-pr` and `gitlab-issue-to-mr` already
use.

- Eligible candidates are ordered deterministically: an explicit `all-hands-bot`
  request (or a trigger label) is ordered before an unrequested PR, then by the
  oldest `all-hands-bot` `review_requested` event for a requested PR or the PR's
  own creation time (oldest first) for an unrequested one, then by repository,
  then by pull-request number, so the oldest work drains first and a later scan
  reaches the remainder.
- The bound counts the conversations a scan **starts**. A delivery that only
  deduplicates or reports an already-running conversation reuses a runtime and
  consumes no slot, so repeated scans make progress on the backlog instead of
  re-spending the bound on work already in flight.
- Reaching the bound never cuts the scan short. The scan still evaluates the
  exact-head checks of the remaining candidates, reconciles completed reviews,
  and runs the maintainer handoff. A candidate whose checks are pending or
  failing starts no conversation and consumes no slot, so it cannot block a later
  green candidate. Only an explicit request or a trigger label gets a waiting or
  blocked gate comment; an unrequested head that is merely red or pending is
  skipped silently.
- A dispatch that raises is reported and does not consume a slot or abort the
  scan, so the candidates behind it are still considered.
- The explicit `review_requested` event path and the trigger-label scan are
  unchanged: an explicit request still starts its conversation immediately, and
  only the scheduled scan's new conversations are bounded.

The review prompt starts with a scope gate: using the repository's own guidance
(its scope categories and ownership boundaries, not a list of individual PR
numbers), the reviewer decides whether the change belongs in this repository
and has the product/architecture direction it needs. When it does not, the review
stops with a single `event: COMMENT` review that says whether the change should
move repositories, close, or receive a maintainer decision, and ends with the
`🛑 MAINTAINER DECISION REQUIRED` verdict. That outcome is **neither an approval
nor a change request**: it does not approve or merge the PR. The completion
handler recognizes the verdict and requests one configured maintainer through the
same handoff used after an approval. Both paths require a successfully fetched
same-repository closing issue labeled `priority:medium` or `priority:high`; a
missing link or priority label, low/normal priority, unresolved reference, or PR
reference does not authorize handoff. One qualifying issue is sufficient even
with other low-priority or missing references. Existing manual review requests
are unchanged. An in-scope change continues the existing review unchanged.

Before approving a change to user-visible UI behavior, the prompt requires live
evidence from a real running application when the repository's guidance demands
it: a screenshot, screen recording, or equivalent capture of the running app
that exercises the production-facing path. Unit tests, CSS-token or contract
assertions, generated mockups, and reconstructed captures may support the review
but cannot substitute for that evidence. When the required evidence is missing,
the reviewer publishes one `event: COMMENT` review naming exactly what is
missing and ends with `🔄 CHANGES REQUESTED`, so approval is withheld and the
deterministic maintainer handoff does not fire. Non-UI changes keep the existing
requirement of the real command and its observed output, with tests insufficient
as sole proof.

The script prepares each review's workspace before the agent starts: the pull
request's head commit is downloaded as a tarball and extracted to a directory of
its own, which becomes the conversation's working directory. The agent is told
not to clone, fetch, check out, or delete anything, and the script removes the
checkout once the conversation has stopped. Nothing accumulates between runs.

---

The script imports shared GitHub transport from
`scripts/github_client.py`, installed with this skill. Include it beside
`main.py` when packaging manually, as shown below; catalog bundles include it
automatically.

## Prerequisites

### Required secret

Verify that the following secret is set in **OpenHands Settings -> Secrets**:

| Secret name | Token type | Minimum permissions |
|---|---|---|
| `GITHUB_PERSONAL_ACCESS_TOKEN` | Classic PAT | `repo` for private repos or `public_repo` for public repos |
| `GITHUB_PERSONAL_ACCESS_TOKEN` | Fine-grained PAT | Contents: Read, Metadata: Read, Pull requests: **Read and Write**, Issues: Read and Write |

Pull-request **write** access is required because the agent publishes a pull
request review, not just an issue comment. The Agent Canvas catalog worker may
also request a configured human reviewer after approval. A token with only Pull
requests: Read will poll happily and then fail at the point of publishing or
requesting the handoff.

When several repositories are monitored, the token must cover all of them.

Check with:
```bash
curl -s https://api.github.com/user \
  -H "Authorization: Bearer $GITHUB_PERSONAL_ACCESS_TOKEN" \
  | python3 -c "import json,sys; d=json.load(sys.stdin); print(d.get('login') or d.get('message'))"
```

If the token is missing or invalid, inform the user and stop.

---

## Setup Workflow

Follow these steps in order.

### Step 1 - Verify `GITHUB_PERSONAL_ACCESS_TOKEN`

Run the `curl` check above.

- If absent: *"GITHUB_PERSONAL_ACCESS_TOKEN is not set. Please add it in
  OpenHands Settings -> Secrets."* Stop.
- If the API returns `{"message": "Bad credentials"}`: tell the user the
  token is invalid and ask them to update it. Stop.

### Step 2 - Collect repositories

Ask: *"Which GitHub repositories should be monitored?
(Format: `owner/repo`, e.g. `myorg/backend`. List several separated by commas to
review them all from one automation.)"*

Validate access to **each** repository:
```bash
curl -s "https://api.github.com/repos/{owner}/{repo}" \
  -H "Authorization: Bearer $GITHUB_PERSONAL_ACCESS_TOKEN" \
  | python3 -c "
import json, sys
d = json.load(sys.stdin)
if 'message' in d:
    print('ERROR:', d['message'])
else:
    print(f\"Accessible. Private: {d.get('private')}. Permissions: {d.get('permissions')}\")
"
```

Record every accepted repository into `REPOS = ["{owner}/{repo}", ...]`. If one
repository fails the check, say which and ask whether to continue without it.

Each repository is polled independently and keeps its own state, so pull-request
numbers never collide between them. The trigger label, tone, and schedule are
shared by all of them; a repository needing different settings wants its own
automation.

### Step 3 - Collect trigger label

Ask: *"Which PR label should trigger a review?
(Press Enter for the default: `openhands-review`.)"*

Record the answer as `TRIGGER_LABEL`. If the label does not exist yet, tell the
user that GitHub will still record the event once the label is created and
applied to a PR.

The automation reviews a PR when it sees the latest matching `labeled` event for
that label. To request another review later, remove and re-apply the label.

### Step 4 - Collect review tone

Ask: *"What review tone should the reviewer use?
  1. Thorough (default) - comprehensive coverage of correctness, security, tests, style
  2. Concise - high-signal only, skips minor style feedback
  3. Friendly - constructive and encouraging
(Press Enter for Thorough, or type your choice or any custom style description)"*

Map the choice to `REVIEW_TONE`:

| Answer | `REVIEW_TONE` | `REVIEW_STYLE_INSTRUCTIONS` |
|---|---|---|
| 1 / Enter | `"thorough"` | `""` |
| 2 | `"concise"` | `""` |
| 3 | `"friendly"` | `""` |
| Custom text, e.g. `strict but kind` | `"thorough"` | the custom text verbatim |

### Step 5 - Collect cron schedule

Ask: *"How often should the automation poll for labeled PRs?
(Press Enter for the default: every 5 minutes.
Use a cron expression for a different interval, e.g. `0 * * * *` = hourly)"*

Default: `*/5 * * * *`.

Record as `CRON_SCHEDULE`.

### Step 6 - Generate the automation script

Read `scripts/main.py` from this skill's directory. Apply exactly seven constant
substitutions near the top of the file:

> The script also reads a `config.json` shipped beside it, if there is one, over
> these constants. That is how the catalog entry
> (`automations/catalog/github-pr-reviewer/`) configures an unmodified copy,
> since a declarative host cannot rewrite Python. This setup path substitutes the
> constants and ships no `config.json`, so the two never collide.

| Placeholder | Replace with |
|---|---|
| `REPOS = ["owner/repo"]` | `REPOS = ["{owner_repo}", ...]` - one entry per repository collected in Step 2 |
| `TRIGGER_LABEL = "openhands-review"` | `TRIGGER_LABEL = "{trigger_label}"` |
| `REVIEW_TONE = "thorough"` | `REVIEW_TONE = "{review_tone}"` |
| `REVIEW_STYLE_INSTRUCTIONS = ""` | `REVIEW_STYLE_INSTRUCTIONS = "{style_instructions}"` |
| `REPO_REVIEW_GUIDE_PATH = ".agents/skills/custom-codereview-guide.md"` | leave unchanged to auto-load a repo review guide at this path, or set to `""` to disable |
| `MAX_NEW_PER_RUN = 2` | leave unchanged to bound a scheduled scan to two new review conversations across all repositories, or raise it if the deployment can hold more agents at once |
| `DEFAULT_OPENHANDS_URL = "http://localhost:8000"` | leave unchanged unless the user has a preference |

Use a safe string writer such as `json.dumps(value)` when inserting user-provided
repository names, labels, or style instructions into Python string literals.
`json.dumps(list_of_repos)` produces the whole `REPOS` list safely in one step.

Run these commands from this skill's directory and write the customized script
to a temporary build directory:
```bash
mkdir -p /tmp/pr-reviewer-build
cp -L scripts/github_client.py /tmp/pr-reviewer-build/github_client.py
# write the customized main.py to /tmp/pr-reviewer-build/main.py
```

Validate syntax before packaging:
```bash
python3 -m py_compile /tmp/pr-reviewer-build/main.py && echo "Syntax OK"
```

Fix any syntax errors before proceeding.

### Step 7 - Package and upload

Determine the Automation backend URL and auth from the `<RUNTIME_SERVICES>`
block in your system context:
- **OPENHANDS_HOST**: the Automation backend `url_from_agent`
- **Auth**: `X-Session-API-Key: $OPENHANDS_AUTOMATION_API_KEY`

```bash
tar -czf /tmp/pr-reviewer.tar.gz -C /tmp/pr-reviewer-build .

TARBALL_PATH=$(curl -s -X POST \
  "${OPENHANDS_HOST}/api/automation/v1/uploads?name=github-pr-reviewer" \
  -H "X-Session-API-Key: $OPENHANDS_AUTOMATION_API_KEY" \
  -H "Content-Type: application/gzip" \
  --data-binary @/tmp/pr-reviewer.tar.gz \
  | python3 -c "import json,sys; print(json.load(sys.stdin)['tarball_path'])")

echo "Uploaded: $TARBALL_PATH"
```

### Step 8 - Register the automation

```bash
curl -s -X POST "${OPENHANDS_HOST}/api/automation/v1" \
  -H "X-Session-API-Key: $OPENHANDS_AUTOMATION_API_KEY" \
  -H "Content-Type: application/json" \
  -d "{
    \"name\": \"GitHub PR Reviewer: {repo_summary} label {trigger_label}\",
    \"trigger\": {\"type\": \"cron\", \"schedule\": \"{cron_schedule}\"},
    \"tarball_path\": \"$TARBALL_PATH\",
    \"entrypoint\": \"python3 main.py\",
    \"timeout\": 600
  }" | python3 -m json.tool
```

Use the single repository as `{repo_summary}` when there is one, and something
like `3 repos` when there are several. A poll now downloads a tarball per queued
review, so the timeout allows for that; a run never waits for a review to
finish, only for it to be started.

Record the returned `id`.

### Step 9 - Confirm

Tell the user:

> ✅ **GitHub PR Reviewer** is running!
>
> - Automation ID: `{id}`
> - Repositories: `{owner}/{repo}`, ... (one line each)
> - Trigger label: `{trigger_label}`
> - Review tone: `{tone}`
> - Polling schedule: `{cron_schedule}`
> - State file per repository:
>   `~/.openhands/workspaces/automation-state/github_pr_reviewer_label_event_{id}_{owner}__{repo}.json`
>
> Apply the `{trigger_label}` label to a pull request to queue a review. Each
> label event is processed once. To request another review, remove and re-apply
> the label.
>
> The review is published as a pull request review on the head commit, with
> inline comments where a finding maps to a changed line.

---

## Runtime Behaviour (per poll)

Each cron run executes `main.py`, which resolves and validates
`GITHUB_PERSONAL_ACCESS_TOKEN` once, then processes every repository in `REPOS`
independently. One repository failing does not stop the others; the run fails
only if every repository fails.

For each repository:

1. Loads that repository's state (see `references/state-schema.md`).
2. Verifies repository access.
3. Lists open PRs, newest-updated first.
4. For each open PR carrying `TRIGGER_LABEL`:
   - Refetches current PR metadata to avoid acting on stale list data.
   - Finds the latest matching GitHub `labeled` issue event.
   - Skips the event if it has already been tracked.
   - Downloads the PR's head commit as a tarball and extracts it to
     `{WORKSPACE_BASE}/repositories/{owner}__{repo}/pr-{number}-{sha12}`. The
     archive is checked as it is unpacked: a single root, no absolute or `..`
     paths, and symlinks skipped rather than materialised.
   - Starts an OpenHands conversation **whose working directory is that
     checkout**, with a review prompt carrying PR metadata, the exact head SHA,
     label event details, and the requirement to re-fetch the current mutable
     GitHub state before deciding a verdict.
   - Posts an acknowledgement comment with the label event, head SHA, and
     conversation link.
   - Records the review in state with `status: "active"` and the checkout path.
   - If the checkout or the conversation cannot be created, the checkout is
     removed and nothing is recorded, so the next poll retries the label event.
5. For each active review conversation:
   - Marks it closed without posting if the PR has closed or merged.
   - Suppresses stale results if the PR head SHA changed after the review was
     queued.
   - When the conversation reaches `idle`, `finished`, `error`, or `stuck`,
     asks GitHub whether a review by the token's own user exists for that head
     SHA. If it does, the review is complete. If it does not, the agent's final
     response is posted as a comment so the work is not lost.
   - Abandons a conversation that has not reached a terminal status within two
     hours, so its checkout can be reclaimed.
6. Removes the checkout of every finished review, but only after confirming the
   conversation has stopped - deleting it under a running agent would remove its
   working directory. When that cannot be confirmed the directory is left alone
   and the next poll tries again.
7. Saves that repository's state atomically.

The completion callback fires once for the whole run.

---

## Additional Resources

- **`references/state-schema.md`** - State JSON schema, field definitions, and
  review lifecycle diagram.
- **`scripts/main.py`** - The complete automation script. Customize the five
  constants at the top before packaging.
- **`tests/test_main.py`** - Unit tests for the checkout, its removal, and state
  handling. Run them from the skill root with `python -m pytest tests/` after
  editing the script.

---

## Troubleshooting

| Symptom | Likely cause | Fix |
|---|---|---|
| Bot never queues reviews | Trigger label not present or no matching `labeled` event | Apply the configured label to the PR |
| "Bad credentials" in run logs | Token expired | Rotate and update `GITHUB_PERSONAL_ACCESS_TOKEN` |
| 404 on repo access | Repo name wrong or no access | Re-check the entry in `REPOS` and the token's permissions |
| One repository is skipped, others work | That repository failed its access check | Read the `=== owner/repo ===` block in the run log |
| Same PR not reviewed after new commits | Label event was already processed | Remove and re-apply the trigger label |
| Review paused with a failing-check comment | A current-head required check reported `failure`, `cancelled`, or `timed_out` | Fix the named checks and push; the review starts on the new head, or request `all-hands-bot` to review immediately |
| Review reported waiting on checks | A current-head required check is `queued` or `in_progress`, or has not reported yet | No action; a later scan or a new review request retries |
| Optional workflow failed but no review was paused | The failed workflow is not required, so the required-only scheduled gate ignored it | No action; only GitHub-required checks gate scheduled discovery |
| Only a few reviews start on a large backlog | The per-scan `max_new_per_run` bound (default 2) reached | No action; later scans drain the remaining oldest eligible PRs, or raise `max_new_per_run` if the deployment can hold more agents |
| Review result never posts | Conversation still running or stuck | Open the conversation link from the acknowledgement comment |
| Stale review suppressed | PR head SHA changed while the agent was reviewing | Re-apply the trigger label after the latest commit |
| Review arrives as a plain comment, not a review | Publishing failed, so the script posted the text as a fallback | Check that the token has Pull requests: Read and Write |
| Agent reports it cannot clone the repo | Prompt asked it not to; the workspace is already the checkout | No action - the code is at the head SHA in its working directory |
| Checkouts remain under `repositories/` | Their conversations had not stopped yet | They are removed by a later poll once the conversation is terminal |
