PR Babysitter
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification.
$ npx skills add tobihagemann/turbo --skill evaluate-findings -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install tobihagemann/turbo evaluate-findings --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/codex/skills/evaluate-findings .claude/skills/evaluate-findings && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluate-findings" agent skill from https://github.com/tobihagemann/turbo/tree/main/codex/skills/evaluate-findings into .claude/skills/evaluate-findings/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-findings", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/tobihagemann/turbo/tree/main/codex/skills/evaluate-findingsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add tobihagemann/turbo --skill evaluate-findings -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install tobihagemann/turbo evaluate-findings --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .agents/skills && cp -r skills-src/codex/skills/evaluate-findings .agents/skills/evaluate-findings && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluate-findings" agent skill from https://github.com/tobihagemann/turbo/tree/main/codex/skills/evaluate-findings into .agents/skills/evaluate-findings/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-findings", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tobihagemann/turbo --skill evaluate-findings -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install tobihagemann/turbo evaluate-findings --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/codex/skills/evaluate-findings .cursor/skills/evaluate-findings && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluate-findings" agent skill from https://github.com/tobihagemann/turbo/tree/main/codex/skills/evaluate-findings into .cursor/skills/evaluate-findings/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-findings", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/tobihagemann/turbo.git --path codex/skills/evaluate-findings--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add tobihagemann/turbo --skill evaluate-findings -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install tobihagemann/turbo evaluate-findings --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/codex/skills/evaluate-findings .gemini/skills/evaluate-findings && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluate-findings" agent skill from https://github.com/tobihagemann/turbo/tree/main/codex/skills/evaluate-findings into .gemini/skills/evaluate-findings/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-findings", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install tobihagemann/turbo evaluate-findingsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add tobihagemann/turbo --skill evaluate-findings -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .github/skills && cp -r skills-src/codex/skills/evaluate-findings .github/skills/evaluate-findings && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-findings" agent skill from https://github.com/tobihagemann/turbo/tree/main/codex/skills/evaluate-findings into .github/skills/evaluate-findings/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-findings", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add tobihagemann/turbo --skill evaluate-findings -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install tobihagemann/turbo evaluate-findings --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/codex/skills/evaluate-findings .opencode/skills/evaluate-findings && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluate-findings" agent skill from https://github.com/tobihagemann/turbo/tree/main/codex/skills/evaluate-findings into .opencode/skills/evaluate-findings/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluate-findings", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluate-findingsCritically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification.
Evaluate Findings is an agent skill from tobihagemann/turbo. Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification. Use when the user asks to "evaluate findings", "assess review comments", "triage review feedback", "evaluate review output", or "filter false positives".
Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Development, covering Code review. The repository describes itself as: Reusable workflows for planning, building, reviewing, and shipping with Claude Code and Codex. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 931eda5. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
gitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluate Findings loads about 5k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 3,004 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from tobihagemann/turbo at commit 931eda5, republished under its MIT licence (© tobihagemann). 3,004 words, ~5,013 tokens.
.claude/skills/evaluate-findings/SKILL.md (or your agent's skills folder).Assess external feedback (code reviews, AI suggestions, PR comments) with adversarial verification. Triage findings into actionable verdicts. Do not apply fixes.
If you already assessed a finding earlier in this session and recorded a verdict of Skip or Escalate — for example when an iterating loop re-runs review and the same finding resurfaces — do not re-adjudicate it from scratch. When a loop ledger path (.turbo/loops/<slug>.md) is in context, read it and treat its recorded verdicts the same way. When the re-reported finding matches one you already judged (same location and substance) and presents no new evidence beyond what your recorded reason already accounts for, keep that verdict and reason without re-reading the code, re-verifying, or routing it to the Devil's Advocate in Step 2. Assess fresh only when the finding raises materially new evidence, or when you have not judged it before in this session.
When several findings rest on a shared premise — for example a source-of-truth choice — verify that premise once before adjudicating them individually. Findings whose premise holds proceed through normal per-finding verification; when it fails, they are all Skip, citing the refuted premise.
When a plan governs the work, re-read the decisions it records before adjudicating. Having read it earlier in the session does not count: once it falls out of context, a recorded decision is indistinguishable from no decision at all.
When the repo keeps an improvements backlog (.turbo/improvements.md at the repo root or, inside a linked worktree, at the main checkout's root), search it before adjudicating for the paths and symbols the findings name and for entries on the same subject, leaving out any entry the work in hand sets out to implement. An entry that parks all or part of the change a finding asks for records a decision to defer it. When that finding holds and acting on it would make the change now, treat it as one that would reverse a decision the user made earlier, naming the entry as the original decision and quoting any condition it records for taking the change up. An entry that only shares the finding's area leaves the verdict to the finding's merits. In both cases, weigh what the entry records about that code when verifying the claim and assessing severity.
For each finding:
Read the referenced code at the mentioned location — include the full function or logical block, not just the flagged line
Check whether the code has diverged — if the finding references code that no longer exists or has since changed, skip it and note the divergence.
Determine scope — clarify whether the issue was introduced by the PR/changeset or is pre-existing.
Verify the claim against the actual code — does the issue genuinely exist?
Assess severity:
| Severity | Meaning |
|---|---|
| Critical | Drop everything. Blocking release or operations. |
| High | Urgent. Should be addressed in the next cycle. |
| Medium | Normal. To be fixed eventually. |
| Low | Nice to have. Minor improvement. |
If the upstream reviewer already assigned a priority (P0-P3), map it: P0→Critical, P1→High, P2→Medium, P3→Low. Then re-assess based on what the actual code reveals. The upstream level is a starting point, not a binding constraint. When the re-assessed severity differs from the upstream level, note the change and the reason.
If the finding has no upstream priority, assess severity from scratch.
Assign a verdict and confidence:
| Verdict | Criteria |
|---|---|
| Apply | The finding is real and in scope: clear bug, missing check, genuine improvement, style violation matching project conventions |
| Skip | False positive, subjective preference, reviewer is wrong, or the change's cost wildly dwarfs its benefit |
| Escalate | Needs the user's judgment: behavior might be intentional, involves product intent, requires domain knowledge the agent lacks, the finding is out of scope, or two findings present a genuine trade-off |
Also assign an internal confidence level — High, Medium, or Low — reflecting how certain you are about the verdict. Confidence is used solely to route findings to the Devil's Advocate in Step 2. It does not appear in the output.
Escalate guidance: When a finding questions whether behavior is intentional and neither docs, specs, nor code comments clarify the intent, assign Escalate. Do not autonomously accept or reject findings that hinge on product intent. If a counterpart implementation exists elsewhere, suggest checking it for consistency.
Conflict guidance: When two findings disagree about whether the code should change at all (one suggests a change to it, another the opposite change or none), treat the conflict as input, not a reason to skip. Verify each against the code and judge each on its merits as usual. If both are defensible and the choice is a genuine trade-off, assign Escalate to both, naming the opposing options so the user can decide.
An affirmation that something is correct is not a finding and carries no evidentiary weight; agreement among reviewers, or a reviewer's authority, does not settle whether a problem exists, nor whether a remedy the reviewers converged on works. When reviewers disagree on whether something is a problem at all — including one asserting it is fine while another flags it — treat the question as unresolved and verify it against the code, without letting the affirmation substitute for verification. When reviewers agree the code should change but propose opposing remedies, assign the verdict on the finding's own merits and name the opposing remedies in the Issue cell. Prefer a remedy whose measurement reports a result concrete enough to re-run over one resting on reading or on a reviewer's authority; a bare claim to have measured ranks no higher than reading. Where no remedy reports one, say so in that cell rather than choosing between them.
A reviewer's report that it could not verify something is a claim to check, not a fact to accept. Attempt the check independently, especially when the reported inability is what justifies skipping a verification step.
Verdict guidance:
After the initial assessment, challenge uncertain findings from a different angle.
Spawn when any finding has Medium or Low confidence. Send only those findings to the sub-agent. High-confidence findings pass through unchallenged. Skip this step entirely if all findings are High confidence.
Capture git status --short, git diff --cached | git hash-object --stdin, git diff | git hash-object --stdin, and git symbolic-ref --short -q HEAD before spawning.
Launch a single sub-agent (inherited model defaults). Provide the Medium/Low-confidence findings with their file locations, claims, and initial verdicts. Instruct the sub-agent to challenge each finding: try to prove it wrong, or confirm it with evidence.
Evidence standards: A refutation counts only when it rests on a defense, guarantee, or documented behavior the sub-agent located and read, or on behavior it observed by running the code; an expectation that a framework, caller, or type already handles the case returns Inconclusive and leaves the initial verdict standing. Confirmed applies to the claim the finding stands or falls on. Establishing the premise beneath that claim leaves it open: that code reads a value settles nothing about whether a test can control that value. When the sub-agent has established only the premise, it returns Inconclusive. Evidence that a test fails when its subject is changed settles only that the test pins the behavior; whether the pinned behavior is the required one stays open. Where the test was written alongside its implementation, their agreement is guaranteed by construction and carries no evidence about the requirement. A finding resting on that evidence is Confirmed only when the requirement itself has been checked against the code's production consumer, or against the governing plan; when neither is reachable, the sub-agent returns Inconclusive.
Protect the shared tree: The sub-agent's prompt must direct it to treat the shared working tree and its git index as read-only; an experiment that needs a scratch project runs in a temp directory outside the repo, or in an isolated git worktree created there and discarded afterward. HEAD stays where it is: read other refs with git show <ref>:<path> rather than git checkout or git switch. Refer to that directory or worktree by absolute path in every command and join chained steps with &&, so a failed step cannot leave the rest running in the shared checkout. Run teardown and verification as their own commands. Give that worktree its own dependency install rather than reaching the shared tree's install by any route: removing a worktree deletes through symlinks, and a redirected suite writes into the shared install. When its own install is not possible, the check is left unrun and reported as such. A check that runs in the shared checkout invokes an already-installed runner directly wherever a package-manager wrapper would front it, since such a wrapper reads as read-only while reconciling the shared install before it runs. Confine dependency installs and reconciliation to an isolated worktree. Every test runner the sub-agent starts, in a temp directory, a worktree, or the shared checkout, runs in its own process group under a timeout enforced from outside the runner. Before teardown, the sub-agent stops the process group of every runner it started, since stopping a runner can leave the processes it spawned alive. Afterward the sub-agent verifies that git worktree list no longer shows the worktree, that git status --short shows what it showed at the start, and that HEAD is still on the branch it started on. It also confirms that no process from those groups, and none whose command line names the temp directory or worktree path, if any, is still running, and reports by PID any process it could not stop. When it cannot list processes, it reports that check as unrun and names those process groups and the temp directory or worktree path, if any. After any check, in a worktree or in the shared checkout, it verifies that the shared tree's dependency directory still resolves (a destroyed install leaves git status unchanged, since it is gitignored). Damage the sub-agent cannot repair is reported with the exact repair command in place of findings.
Verify the tree: re-run all four commands when the sub-agent returns, including when it terminates early or reports incomplete results, and confirm alongside them that the shared tree's dependency directory still resolves, which the four cannot see. Delete what the sub-agent created, revert what it modified or staged, and return HEAD to the captured branch, leaving everything the pre-spawn capture already showed untouched.
The sub-agent picks research tools based on claim type:
| Claim Type | Tool |
|---|---|
| API deprecated/removed/changed | Documentation MCP tools or web search |
| Method doesn't exist / wrong signature | Documentation MCP tools, web search fallback |
| Code causes specific bug or behavior | Shell (isolated read-only test snippet) |
| Best practice or ecosystem claim | Web search |
| Migration or changelog lookup | Web search → web fetch |
Use whatever documentation tools are available. The specific tools vary by project setup.
Budget: max 2 research actions per finding. If the first action is conclusive, skip the second.
The sub-agent returns per finding:
Merge sub-agent results with the initial assessment:
Findings not investigated by the sub-agent keep their original verdict.
For Apply findings, document the issue and location. For Escalate findings, note what information would resolve the ambiguity. For Skip findings, document why.
Summarize the evaluated findings in a table:
| File | Issue | Source | Severity | Verdict |
|---|
When Step 2 ran (any finding was investigated by the Devil's Advocate sub-agent), add an Investigated column:
| File | Issue | Source | Severity | Verdict | Investigated |
|---|
Where Investigated shows:
For findings whose severity was re-assessed from the upstream level, append the change in the Severity cell (e.g., "High (was Medium)").
Carry each verdict into the Verdict column exactly as assessed, Escalate included.
For disputed findings, add a callout below the table showing both perspectives. For each finding, indicate scope in the Issue column (e.g., "Pre-existing:" prefix).
Then call update_plan to mark this step completed and continue with the next step of the active workflow.
© tobihagemann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in codex/skills/evaluate-findings of tobihagemann/turbo.
Open the folder on GitHubat commit 931eda5
Evaluate Findings next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluate Findings this skilltobihagemann/turbo | 408 | — | ~5k | Automated safety check: Pass | MIT | |
| PR Babysitteropeninterpreter/openinterpreter | 69k | 3 repos | ~4.2k | Automated safety check: Pass | Apache-2.0 | |
| Code Review ChecklistshareAI-lab/learn-claude-code | 78k | 4 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Backend Code Reviewlangflow-ai/langflow | 155k | — | ~3.5k | Automated safety check: Notes | MIT | |
| Mole Bug Patternstw93/Mole | 70k | — | ~2k | Automated safety check: Pass | GPL-3.0 | |
| Backend Code Reviewlanggenius/dify | 158k | — | ~676 | Automated safety check: Pass | Custom licence |
openinterpreter/openinterpreter
Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.
shareAI-lab/learn-claude-code
Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.
langflow-ai/langflow
Review backend code for quality, security, maintainability, and best practices based on established checklist rules.
tw93/Mole
A catalog of recurring bug shapes in the Mole Mac cleaner, used to review safety-sensitive diffs for deletion safety, unbounded commands, shell traps and weak tests.
langgenius/dify
Reviews backend code under api/ for concrete, reproducible defects, routes to rule packs for architecture, schema, repositories and SQLAlchemy, and ranks findings from P0 to P3.
woocommerce/woocommerce
Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.
tobihagemann/turbo
Consult ChatGPT Pro via ChatGPT browser automation for problems that resist standard approaches.
tobihagemann/turbo
Fetch and summarize review feedback and conversation from a GitHub PR (unresolved review threads, review bodies, and PR conversation comments) without making changes.
tobihagemann/turbo
Recall why a past change was made by locating the Claude Code transcript that produced it.
tobihagemann/turbo
Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.
tobihagemann/turbo
Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.
tobihagemann/turbo
Assess project-wide structural technical debt: complexity hotspots, deprecated API usage, duplication clusters, architecture rot, and low-value tests.
Categories
Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification. Evaluate Findings is an agent skill from tobihagemann/turbo. Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification.
Evaluate Findings fits situations like: the user asks to evaluate findings; assess review comments; triage review feedback; evaluate review output.
Run `npx skills add tobihagemann/turbo --skill evaluate-findings -a claude-code`. Or copy the skill folder (codex/skills/evaluate-findings in tobihagemann/turbo) into .claude/skills/evaluate-findings in your project. Claude Code loads it when a task matches its description.
Run `npx skills add tobihagemann/turbo --skill evaluate-findings -a codex`. Or copy the skill folder (codex/skills/evaluate-findings in tobihagemann/turbo) into .agents/skills/evaluate-findings in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tobihagemann/turbo --skill evaluate-findings -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-findings, .gemini/skills/evaluate-findings, .github/skills/evaluate-findings and .opencode/skills/evaluate-findings in your project.
Going by SKILL.md and its folder, Evaluate Findings needs the command-line tools its instructions call (git).
SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluate Findings is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluate Findings: PR Babysitter (openinterpreter/openinterpreter, 69k stars), Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), Backend Code Review (langflow-ai/langflow, 155k stars) and Mole Bug Patterns (tw93/Mole, 70k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
tobihagemann (a GitHub user) maintains it in tobihagemann/turbo, which has 408 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on October 9, 2026.
Source: tobihagemann/turbo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.