PR Babysitter
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
Reviews a pull request like a strict senior engineer and posts one consolidated GitHub review, backing every comment with a verified file citation or an official source.
$ npx skills add tech-leads-club/agent-skills --skill the-judge -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tech-leads-club/agent-skills the-judge --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tech-leads-club/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'packages/skills-catalog/skills/(quality)/the-judge' .claude/skills/the-judge && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "the-judge" agent skill from https://github.com/tech-leads-club/agent-skills/tree/main/packages/skills-catalog/skills/(quality)/the-judge into .claude/skills/the-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "the-judge", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tech-leads-club/agent-skills/tree/main/packages/skills-catalog/skills/(quality)/the-judgeType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tech-leads-club/agent-skills --skill the-judge -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tech-leads-club/agent-skills the-judge --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tech-leads-club/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/'packages/skills-catalog/skills/(quality)/the-judge' .agents/skills/the-judge && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "the-judge" agent skill from https://github.com/tech-leads-club/agent-skills/tree/main/packages/skills-catalog/skills/(quality)/the-judge into .agents/skills/the-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "the-judge", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tech-leads-club/agent-skills --skill the-judge -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tech-leads-club/agent-skills the-judge --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tech-leads-club/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/'packages/skills-catalog/skills/(quality)/the-judge' .cursor/skills/the-judge && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "the-judge" agent skill from https://github.com/tech-leads-club/agent-skills/tree/main/packages/skills-catalog/skills/(quality)/the-judge into .cursor/skills/the-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "the-judge", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tech-leads-club/agent-skills.git --path 'packages/skills-catalog/skills/(quality)/the-judge'--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tech-leads-club/agent-skills --skill the-judge -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tech-leads-club/agent-skills the-judge --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tech-leads-club/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/'packages/skills-catalog/skills/(quality)/the-judge' .gemini/skills/the-judge && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "the-judge" agent skill from https://github.com/tech-leads-club/agent-skills/tree/main/packages/skills-catalog/skills/(quality)/the-judge into .gemini/skills/the-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "the-judge", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tech-leads-club/agent-skills the-judgeInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tech-leads-club/agent-skills --skill the-judge -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tech-leads-club/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/'packages/skills-catalog/skills/(quality)/the-judge' .github/skills/the-judge && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "the-judge" agent skill from https://github.com/tech-leads-club/agent-skills/tree/main/packages/skills-catalog/skills/(quality)/the-judge into .github/skills/the-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "the-judge", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tech-leads-club/agent-skills --skill the-judge -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tech-leads-club/agent-skills the-judge --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tech-leads-club/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/'packages/skills-catalog/skills/(quality)/the-judge' .opencode/skills/the-judge && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "the-judge" agent skill from https://github.com/tech-leads-club/agent-skills/tree/main/packages/skills-catalog/skills/(quality)/the-judge into .opencode/skills/the-judge/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "the-judge", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
the-judgeReviews a pull request like a strict senior engineer and posts one consolidated GitHub review, backing every comment with a verified file citation or an official source.
The Judge reviews with a high conviction bar, preferring three findings that matter over fifteen that waste the author's time. Its non-negotiable rules come first: any claim about the repository needs a file and line citation the reviewer actually read, and any claim about a library, framework or API needs a URL from official documentation fetched during that same review, never from memory. A finding without that evidence is simply not posted.
It runs the repository's own linters, type checks and focused tests before writing anything, so it never repeats what that tooling already flags. Inline comments are capped at 5 nits, with any overflow folded into a count in the summary rather than flooding the PR, and every comment and the summary must pass scripts/review_gate.py before anything is posted.
The review covers correctness, security, structural quality such as file-size limits and spaghetti growth, and AI-slop patterns like useless comments, and it ends in one verdict: APPROVE, COMMENT or REQUEST_CHANGES, weighted by what was actually found. Language follows the user's request, defaulting to English, while verdict tokens and code stay verbatim.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 6df68d5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 3 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
ghpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use gh, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
The Judge PR Reviewer loads about 5k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 234 tokens; SKILL.md has 2,481 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from tech-leads-club/agent-skills at commit 6df68d5, republished under its CC-BY-4.0 licence (© tech-leads-club). 2,481 words, ~5,002 tokens.
.claude/skills/the-judge/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.Review a pull request like a senior engineer with a high conviction bar: few comments, every one backed by evidence, posted as a single consolidated GitHub review. The Judge would rather post three findings that matter than fifteen observations that waste the author's time.
These rules override everything else in this skill. Read them before doing anything.
file:line citation you confirmed by reading the file. An external claim (about a library, API, framework, version, deprecation, vulnerability, or best practice) requires a URL from an official source fetched during this review. A finding without evidence is not posted. Period.scripts/review_gate.py with exit code 0 before posting. No exceptions, no manual overrides.scripts/scan_bypasses.py, not in prose. A finding that is deterministic by nature goes to the lint-rule flywheel so the next review costs less than this one.| Severity | Emoji | Definition | Verdict effect |
|---|---|---|---|
| blocker | 🔴 | Changes whether the PR should merge: data loss, exploitable security, incorrect money, broken auth, irreversible migration, PII in logs | REQUEST_CHANGES |
| should-fix | 🟠 | Real defect, but not a merge risk | COMMENT |
| nit | 🟡 | Minor. Capped at 5 inline; overflow counted in summary | No effect |
| pre-existing | 🟣 | Bug the PR did not introduce. Summary only, never inline | No effect |
Verdict mapping: any 🔴 present, REQUEST_CHANGES. Zero 🔴 and zero 🟠, APPROVE. Anything else, COMMENT.
Own-PR fallback: GitHub returns 422 when you APPROVE or REQUEST_CHANGES your own PR. scripts/post_review.py detects when the PR author equals the authenticated gh user, posts as COMMENT, and appends a one-line footer at the end of the summary stating the intended verdict. The TL;DR already carries the verdict token, so the footer only explains the mechanics. Do not fight this; it is API behavior.
gh pr view --json number,title,body,author,url,baseRefName,headRefName,additions,deletions,changedFiles
gh api user --jq .login
gh pr diff <number>
gh pr view <number> --json files --jq '.files[].path'Classify every changed file as core or mechanical (generated code, lockfiles, snapshots, vendored deps, build artifacts, migrations output). Mechanical files are skipped and listed in the summary. Detect moved code: 3+ consecutive lines deleted in one place and added identically elsewhere is a move, not new code; do not re-review it as new.
Determine the round. Fetch your own previous reviews on this PR:
gh api repos/{owner}/{repo}/pulls/{number}/reviews --jq '[.[] | select(.user.login=="<gh-login>") | {id, submitted_at, body}]'No previous review: this is round 1, run the full workflow. Previous review exists: this is round N, load the findings ledger (the ID table) from the latest previous summary and follow the Convergence Contract instead of a full re-run.
If the diff exceeds roughly 400 changed lines in core files and the harness supports subagents, run the Step 3 passes as parallel subagents. Otherwise run them sequentially. Never make subagent support a requirement.
Detect and run the repo's own checks: lint, typecheck, and tests focused on changed files (look at package.json scripts, Makefile, pyproject.toml, CI config). If the repo already runs security scanners (dependency, IaC, or SAST tools wired into its CI), run them or read their current output as deterministic input too. Record results. Findings these tools produce are excluded from your review scope; you judge only what they cannot. If the repo has no such tooling, note that in the summary and move on; do not install anything.
Then run the deterministic bypass scan:
gh pr diff <number> | python3 scripts/scan_bypasses.pyIt prints path:line category content for every bypass marker added by the diff (suppression directives, dodged tests, TLS and type-check bypasses, swallowed errors, sleep-as-synchronization). Scan hits are candidates for Pass F, not findings: a suppression carrying a justification and an issue link is acceptable; a naked one is not. The scan exists so zero LLM tokens are spent detecting what a regex detects.
Enumerate every external surface the diff touches: dependencies added or version-bumped (read the manifest/lockfile diff), APIs called, framework features used, language features that are version-sensitive. For each surface, search current official documentation, release notes, and security advisories. Log every consulted URL; the summary includes a research log. This step is not optional and not skippable, even when you feel confident. Confidence from memory is exactly the failure mode this step exists to kill. If the diff touches zero external surfaces, state that in the research log.
Completeness contract: round 1 covers all core files across all passes, in depth, in one shot. Nothing is deferred to "a later look". A finding you could have raised now and raise later is a broken contract with the author. One exception to volume: do not stack comments on code a structural finding will rewrite; if a 🔴 or 🟠 asks for a block to be restructured, withhold nits inside that block and note "nits withheld on lines the structural fix rewrites" in the summary.
Read references/review-standards.md now. Run six passes over core files:
Each pass produces candidate findings: claim, tentative severity, evidence pointer. If the harness supports choosing a model per subagent, use light, fast variants for mechanical work (file classification, dedupe, scan triage) and reserve the strongest model for the judgment passes and verification; burning the heavy model on cheap triage is waste, and burning the light model on judgment is false positives.
For each candidate: re-read the actual code at the cited location and confirm the claim holds. Re-apply the exclusion lists. Kill anything you cannot evidence. Deduplicate across passes. Assign final severity conservatively: a blocker you are not certain of is a should-fix phrased as a question. This pass exists because candidate generation is optimized for recall and posting is optimized for precision.
Two grounding rules:
repro evidence. The inverse binds too: when a reproduction was feasible and failed to reproduce the claim, the finding dies, whatever your reading of the code said.Read references/comment-voice.md now. Write the summary and every comment body under that spec. Produce findings.json:
{
"language": "en",
"round": 1,
"carryover": {"blocker": 0, "should-fix": 0},
"verdict": "REQUEST_CHANGES",
"summary": "## TL;DR\n...\n## Findings\n...\n## Promote to lint rule\n...\n## Research log\n...\n## Checks run\n...\n## Skipped files\n...",
"findings": [
{
"id": "F1",
"path": "src/billing/invoice.ts",
"line": 142,
"severity": "blocker",
"body": "🔴 `applyDiscount` divides by `items.length` with no empty-list guard (src/billing/invoice.ts:142)...",
"evidence": [
{"type": "internal", "ref": "src/billing/invoice.ts:142"},
{"type": "external", "ref": "https://official-docs.example/api#behavior"}
]
}
]
}round defaults to 1. carryover counts previous-round findings still unresolved (zeros in round 1); the gate uses it to compute the correct verdict when the new-findings list alone would understate open risk. id values are stable across rounds (F1 stays F1 forever) so the resolution ledger maps cleanly. Evidence entries accept three types: internal (verified path:line), external (https URL from an official source fetched this review), and repro (the local command that demonstrates the claim, with its observed result in the finding body).
Rules the gate enforces (write to them from the start): body starts with the severity emoji; blocker and should-fix require at least one evidence entry; external evidence must be an https URL; pre-existing findings have no line; the summary contains a TL;DR section and a lint-rule section.
python3 scripts/review_gate.py findings.jsonFix every reported violation and re-run until exit code 0. The gate is deterministic; arguing with it is arguing with a regex.
python3 scripts/post_review.py findings.jsonPosts everything as one review (single API call: summary body, verdict event, positioned inline comments), handles the own-PR fallback, prints the review URL. Inline comments only anchor on lines present in the diff; the script tells you which comment failed if GitHub rejects an anchor.
## TL;DR
One to three sentences: what the PR does and the verdict with its reason. Must contain the verdict token (APPROVE, COMMENT, or REQUEST_CHANGES) verbatim; the gate checks for it.
## Findings
| ID | Severity | Location | Summary |
|---|---|---|---|
## Resolution (round 2+)
| ID | Status | Note |
|---|---|---|
Every previous finding appears here with status resolved (with the commit), open, or declined-accepted (author's reason held). Omit this section in round 1.
## Promote to lint rule
Findings in this review that are deterministic by nature (banned pattern, naming convention, style rule). For each: the rule, and a one-line sketch of how to lint it. Reviews should get cheaper every cycle; this section is the flywheel. Write "none" if empty.
## Research log
URLs consulted in Step 2, one per line, with the claim each one supports or refutes. Sources that refuted a candidate finding belong here too: a documented non-finding is evidence of diligence and saves a future round. Write "no external surfaces touched" if that is the case.
## Checks run
Deterministic ladder results from Step 1, verbatim numbers with no interpretation: test counts (passed/failed/skipped), linter and typechecker outcome, bypass scan candidate count.
## Skipped files
Mechanical/generated files excluded from review.
## Nit overflow
"N additional nits not posted individually" when the cap was hit. Omit otherwise.The Judge is decisive: it converges reviews instead of stretching them. Round 1 is exhaustive; every later round shrinks. The failure mode this contract kills is the infinite ping-pong where each round discovers a new front.
Round N (N >= 2) scope is exactly two things:
Declined findings. When the author declines with a reason, first consider that they are closer to the code and may be right; if the reasoning holds from a code-health perspective, record declined-accepted and drop it permanently. If it does not hold, re-argue once with new evidence, at most once. Still unresolved after that: mark it open, state that the disagreement moves out of the review thread, and stop re-raising. Comment threads do not resolve philosophy.
Round cap. By round 3, every remaining open finding is either fixed, converted to a follow-up issue by agreement, or escalated out of the thread. The Judge never runs a fourth round on the same finding set.
Approval bar. Approve as soon as the PR definitely improves overall code health, even if imperfect; perfect code does not exist, only better code. Unresolved 🟡 never blocks an APPROVE (approve with nits). Withholding approval to extract polish beyond the severity table is scope creep by the reviewer.
The gate enforces the mechanical half of this contract: with "round": 2 or higher it rejects any non-blocker finding, requires the Resolution section, and requires carryover so the verdict reflects still-open previous findings.
User says: "revise esse PR" (on a branch with an open PR, 180 changed lines, TypeScript). Actions: Steps 0-2 (research: one dep bump, checked its changelog), sequential passes, 1 should-fix + 2 nits survive verification, gate passes on second run (first run caught an exclamation mark), post as COMMENT. Result: one consolidated review, three inline comments, summary with research log citing the changelog URL.
User says: "judge this PR before I merge".
Actions: Pass A finds an unguarded division in a money path with a traced failing input; verification confirms at src/billing/invoice.ts:142; verdict REQUEST_CHANGES; PR author equals gh user, so post_review posts as COMMENT with the TL;DR stating REQUEST_CHANGES and a footer explaining the fallback.
Result: review posted, blocker anchored inline with file:line evidence and the lazy fix (guard clause, three lines).
User says: "review this PR". PR description says "fixes the race condition in the worker pool". No test touches the pool. Actions: Pass E flags the claim; no repro or test exists; the finding becomes a 🟠 question asking for a failing-then-passing test, not an assertion that the fix is wrong. Result: comment asks for evidence instead of speculating.
User says: "julgue de novo" after pushing fixes.
Actions: Step 0 finds the previous review; round 2. Resolution check maps the ledger: F1 fixed in commit a1b2c3d, F2 declined with a reason that holds (recorded declined-accepted), F3 still open. Diff since last review scanned for new blockers only; none found. findings is empty, carryover is {"blocker": 1, "should-fix": 0} for F3, verdict REQUEST_CHANGES.
Result: a short review with only the Resolution table and TL;DR. No new comments, no new fronts, convergence preserved.
Cause: APPROVE/REQUEST_CHANGES on your own PR, or a comment anchored to a line not in the diff.
Solution: post_review.py auto-falls back to COMMENT for own PRs. For anchor failures, move the comment to a changed line shown in the patch, or drop line and fold it into the summary.
Cause: no gh auth login session.
Solution: ask the user to run gh auth status and authenticate. Do not attempt tokenless calls.
Cause: branch never pushed or PR not opened. Solution: report it and stop. Opening PRs is out of scope for The Judge (that is a different job).
Cause: rephrasing around a banned pattern instead of removing it. Solution: delete the sentence containing the violation. The gate output names the exact pattern and location.
© tech-leads-club, CC-BY-4.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 5 other files (scripts, references) in packages/skills-catalog/skills/(quality)/the-judge of tech-leads-club/agent-skills.
Open the folder on GitHubat commit 6df68d5
The Judge PR Reviewer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| The Judge PR Reviewer this skilltech-leads-club/agent-skills | 7k | — | ~5k | Automated safety check: Pass | CC-BY-4.0 | |
| PR Babysitteropeninterpreter/openinterpreter | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| GitHub Review Iterationprisma/orm | 48k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| PR Review State Fetchprisma/orm | 48k | — | ~767 | Automated safety check: Pass | Apache-2.0 | |
| PR Finalize Reviewmicrosoft/garnet | 12k | — | ~3.1k | Automated safety check: Pass | MIT | |
| Fastlane Pull Request Reviewfastlane/fastlane | 42k | — | ~550 | Automated safety check: Pass | MIT |
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
prisma/orm
Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.
prisma/orm
Fetches a pull request's canonical review state as JSON, validates it, and renders markdown, a text summary and triage target files from it using bundled scripts.
microsoft/garnet
Checks that a pull request's title and description match its implementation and reviews the code for Garnet best practices, reporting findings without posting them.
fastlane/fastlane
Reviews a fastlane pull request against its linked issue and the project guides, separating blocking from non-blocking findings and handling vulnerabilities privately.
thedotmack/claude-mem
Keeps watching a pull request, fixing real review and CI problems and resolving stale threads, until it is clean and ready to merge.
tech-leads-club/agent-skills
Guides design of modular-monolith platforms with DDD, flat-by-aggregate modules, anti-corruption layers, outbox events and resilience, plus an architecture document with SVG diagrams.
tech-leads-club/agent-skills
Generates Excalidraw diagram files from plain descriptions, choosing among flowcharts, mind maps, architecture, swimlane, class, sequence and ER diagrams.
tech-leads-club/agent-skills
Creates, validates and renders Mermaid diagrams to SVG, PNG or ASCII, including C4 and AWS architecture-beta, flowcharts, sequence diagrams and ERDs.
tech-leads-club/agent-skills
Answers AWS architecture, security and service-selection questions by searching AWS documentation through MCP tools first, then adapting advice to your stack and team.
tech-leads-club/agent-skills
Evaluates a repository's agent harness (AGENTS.md, rules, skills) for broken paths, redundant instructions and usefulness, and stops at reports.
tech-leads-club/agent-skills
Designs scalable NestJS modular monoliths with domain-driven design, Clean Architecture layers and optional CQRS, defining bounded contexts and strict module boundaries.
Works with
Categories
Reviews a pull request like a strict senior engineer and posts one consolidated GitHub review, backing every comment with a verified file citation or an official source. The Judge reviews with a high conviction bar, preferring three findings that matter over fifteen that waste the author's time. Its non-negotiable rules come first: any claim about the repository needs a file and line citation the reviewer actually read, and any claim about a library, framework or API needs a URL from official documentation fetched during that same review, never from memory.
The Judge PR Reviewer fits situations like: reviewing a pull request before merge with a strict evidence requirement; getting one consolidated GitHub review instead of many scattered comments; checking a PR for AI-generated filler comments and structural problems.
Run `npx skills add tech-leads-club/agent-skills --skill the-judge -a claude-code`. Or copy the skill folder (packages/skills-catalog/skills/(quality)/the-judge in tech-leads-club/agent-skills) into .claude/skills/the-judge in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tech-leads-club/agent-skills --skill the-judge -a codex`. Or copy the skill folder (packages/skills-catalog/skills/(quality)/the-judge in tech-leads-club/agent-skills) into .agents/skills/the-judge in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tech-leads-club/agent-skills --skill the-judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/the-judge, .gemini/skills/the-judge, .github/skills/the-judge and .opencode/skills/the-judge in your project.
Going by SKILL.md and its folder, The Judge PR Reviewer needs Python for the scripts in its folder and the command-line tools its instructions call (gh and python3). Our summary lists: The gh CLI with permission to post reviews; The repository's own linters and test runner.
SKILL.md contains no URLs. Its commands use gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
The Judge PR Reviewer is published under the CC-BY-4.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with The Judge PR Reviewer: PR Babysitter (openinterpreter/openinterpreter, 69k stars), GitHub Review Iteration (prisma/orm, 48k stars), PR Review State Fetch (prisma/orm, 48k stars) and PR Finalize Review (microsoft/garnet, 12k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tech-leads-club (a GitHub organization) maintains it in tech-leads-club/agent-skills, which has 7,045 GitHub stars. The repository holds 74 skills in this directory. The repository was last updated on October 9, 2026.
Source: tech-leads-club/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.