Agent skill

Evaluate Findings

by tobihagemann in tobihagemann/turbo

Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification.

MITAuto-check passedDevelopment

Install Evaluate Findings

skills CLI
$ npx skills add tobihagemann/turbo --skill evaluate-findings -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install tobihagemann/turbo evaluate-findings --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/tobihagemann/turbo.git skills-src && mkdir -p .claude/skills && cp -r skills-src/codex/skills/evaluate-findings .claude/skills/evaluate-findings && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-findings
GitHub stars
408
Token cost
~5k tokens
SKILL.md length
3,004 words
Files
1
Skills in repo
81
Repo updated
First seen
Licence
MIT

At a glance

Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification.

  • Works in 4 steps: Assess Each Finding → Devil's Advocate → Reconciliation → …
  • The user asks to evaluate findings
  • SKILL.md covers Step 1: Assess Each Finding, Step 2: Devil's Advocate, Step 3: Reconciliation and Step 4: Format Output
  • Calls git

What it does

Evaluate Findings is an agent skill from tobihagemann/turbo. Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification. Use when the user asks to "evaluate findings", "assess review comments", "triage review feedback", "evaluate review output", or "filter false positives".

Its SKILL.md is about 5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Code review. The repository describes itself as: Reusable workflows for planning, building, reviewing, and shipping with Claude Code and Codex. The licence is MIT.

When your agent uses it

  • The user asks to evaluate findings
  • Assess review comments
  • Triage review feedback
  • Evaluate review output

Example prompts

  • “evaluate findings”
  • “assess review comments”
  • “triage review feedback”
  • “/evaluate-findings”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Assess Each Finding
  2. Devil's Advocate
  3. Reconciliation
  4. Format Output

What it can do on your machine

Read from SKILL.md and the folder at commit 931eda5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate Findings loads about 5k tokens when it runs. Until then it costs about 80 tokens; SKILL.md has 3,004 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from tobihagemann/turbo at commit 931eda5, republished under its MIT licence (© tobihagemann). 3,004 words, ~5,013 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-findings/SKILL.md (or your agent's skills folder).
name
evaluate-findings
description
Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification. Use when the user asks to "evaluate findings", "assess review comments", "triage review feedback", "evaluate review output", or "filter false positives".

Evaluate Findings

Assess external feedback (code reviews, AI suggestions, PR comments) with adversarial verification. Triage findings into actionable verdicts. Do not apply fixes.

Step 1: Assess Each Finding

If you already assessed a finding earlier in this session and recorded a verdict of Skip or Escalate — for example when an iterating loop re-runs review and the same finding resurfaces — do not re-adjudicate it from scratch. When a loop ledger path (.turbo/loops/<slug>.md) is in context, read it and treat its recorded verdicts the same way. When the re-reported finding matches one you already judged (same location and substance) and presents no new evidence beyond what your recorded reason already accounts for, keep that verdict and reason without re-reading the code, re-verifying, or routing it to the Devil's Advocate in Step 2. Assess fresh only when the finding raises materially new evidence, or when you have not judged it before in this session.

When several findings rest on a shared premise — for example a source-of-truth choice — verify that premise once before adjudicating them individually. Findings whose premise holds proceed through normal per-finding verification; when it fails, they are all Skip, citing the refuted premise.

When a plan governs the work, re-read the decisions it records before adjudicating. Having read it earlier in the session does not count: once it falls out of context, a recorded decision is indistinguishable from no decision at all.

When the repo keeps an improvements backlog (.turbo/improvements.md at the repo root or, inside a linked worktree, at the main checkout's root), search it before adjudicating for the paths and symbols the findings name and for entries on the same subject, leaving out any entry the work in hand sets out to implement. An entry that parks all or part of the change a finding asks for records a decision to defer it. When that finding holds and acting on it would make the change now, treat it as one that would reverse a decision the user made earlier, naming the entry as the original decision and quoting any condition it records for taking the change up. An entry that only shares the finding's area leaves the verdict to the finding's merits. In both cases, weigh what the entry records about that code when verifying the claim and assessing severity.

For each finding:

  1. Read the referenced code at the mentioned location — include the full function or logical block, not just the flagged line

  2. Check whether the code has diverged — if the finding references code that no longer exists or has since changed, skip it and note the divergence.

  3. Determine scope — clarify whether the issue was introduced by the PR/changeset or is pre-existing.

    • Pre-existing issues in earlier commits on the same feature branch are in-scope by default — the entire branch is one coherent unit of work. Judge these on their merits like any in-scope finding.
    • Findings genuinely outside the branch's work are the user's call to include. Assign Escalate so the user decides whether to widen the changeset. Reserve Skip for changes whose cost wildly dwarfs the benefit.
  4. Verify the claim against the actual code — does the issue genuinely exist?

    • When the finding offers a concrete example as evidence — a claimed mishandled input, a claimed wrong output — verify that example independently: a finding can hold in substance while its example does not. Keep the finding and record the correction beside it; drop it only when the claim rests on that example alone.
    • When the finding asserts a compatibility property, establish two things before assigning Apply: what the existing check actually enforces, and what real counterparts produce today. A claim stronger than the check enforces is a premise error rather than a defect — Skip, citing what the check enforces, or narrow the finding to the property it does enforce and record the narrowing beside it. When neither can be established from the code, the artifacts, or authoritative documentation, keep the finding Escalate.
    • When the finding cites a rule or convention, read the cited text, then look for a place that already applied it before this changeset — the same file, or the nearest files the rule also governs. Where the text alone leaves the reading open, read the rule the way that application reads it; where no such application exists, judge on the text alone.
    • When the finding rests on a premise that reading the source cannot settle — what a platform API returns at runtime, or what a value measures once the system runs — establish that premise before assigning Apply or Escalate, using a targeted search or count over the source, or a measurement from a surface already running in this session. A premise of this kind reads as sound whether or not it holds, so confidence in the finding is no substitute. When nothing available settles it, assign Escalate on that ground, naming the premise as unverified in the Issue cell.
  5. Assess severity:

    SeverityMeaning
    CriticalDrop everything. Blocking release or operations.
    HighUrgent. Should be addressed in the next cycle.
    MediumNormal. To be fixed eventually.
    LowNice to have. Minor improvement.

    If the upstream reviewer already assigned a priority (P0-P3), map it: P0→Critical, P1→High, P2→Medium, P3→Low. Then re-assess based on what the actual code reveals. The upstream level is a starting point, not a binding constraint. When the re-assessed severity differs from the upstream level, note the change and the reason.

    If the finding has no upstream priority, assess severity from scratch.

  6. Assign a verdict and confidence:

VerdictCriteria
ApplyThe finding is real and in scope: clear bug, missing check, genuine improvement, style violation matching project conventions
SkipFalse positive, subjective preference, reviewer is wrong, or the change's cost wildly dwarfs its benefit
EscalateNeeds the user's judgment: behavior might be intentional, involves product intent, requires domain knowledge the agent lacks, the finding is out of scope, or two findings present a genuine trade-off

Also assign an internal confidence level — High, Medium, or Low — reflecting how certain you are about the verdict. Confidence is used solely to route findings to the Devil's Advocate in Step 2. It does not appear in the output.

Escalate guidance: When a finding questions whether behavior is intentional and neither docs, specs, nor code comments clarify the intent, assign Escalate. Do not autonomously accept or reject findings that hinge on product intent. If a counterpart implementation exists elsewhere, suggest checking it for consistency.

Conflict guidance: When two findings disagree about whether the code should change at all (one suggests a change to it, another the opposite change or none), treat the conflict as input, not a reason to skip. Verify each against the code and judge each on its merits as usual. If both are defensible and the choice is a genuine trade-off, assign Escalate to both, naming the opposing options so the user can decide.

An affirmation that something is correct is not a finding and carries no evidentiary weight; agreement among reviewers, or a reviewer's authority, does not settle whether a problem exists, nor whether a remedy the reviewers converged on works. When reviewers disagree on whether something is a problem at all — including one asserting it is fine while another flags it — treat the question as unresolved and verify it against the code, without letting the affirmation substitute for verification. When reviewers agree the code should change but propose opposing remedies, assign the verdict on the finding's own merits and name the opposing remedies in the Issue cell. Prefer a remedy whose measurement reports a result concrete enough to re-run over one resting on reading or on a reviewer's authority; a bare claim to have measured ranks no higher than reading. Where no remedy reports one, say so in that cell rather than choosing between them.

A reviewer's report that it could not verify something is a claim to check, not a fact to accept. Attempt the check independently, especially when the reported inability is what justifies skipping a verification step.

Verdict guidance:

  • A verdict records whether the finding is real: genuine defect, in scope, at what severity. A finding can hold while the remedy proposed for it does not, so a verdict never certifies the remedy. Judge the remedy's cost and scope where the bullets below call for that, and leave whether it works to be checked when it is applied. Where naming the likely direction helps the user, put it in the Issue cell flagged as unverified.
  • Never auto-dismiss findings about security defaults, permission escalation, or fail-open vs fail-closed behavior. Always surface these even if the behavior appears intentional.
  • Readability and clarity improvements that genuinely make code cleaner are valid. Do not auto-classify cosmetic changes as subjective.
  • Removing a comment that adds no information beyond the code is a valid Apply, not a subjective preference. Keep only comments that capture a constraint the code cannot express.
  • Be skeptical of "defensive coding" suggestions that wrap natural code in verbose guards without evidence of real-world failures. Apply a hardening finding only when it names a failure scenario reachable in this deployment, whatever severity the reviewer attached; when the plan's Context bounds the system (a single operator, no concurrent writers, a handful of invited users), a scenario that bound rules out is a Skip, citing the bound.
  • Machinery is scope. A finding whose fix adds a lease, lock, queue, versioning scheme, state machine, or new persistent entity expands the project even when the requirement count stays flat. Assign Escalate regardless of confidence; "making states explicit" or "staying within the approved plan" does not make the machinery proportionate. When a stated bound rules the machinery's failure scenario out entirely, Skip instead, citing the bound.
  • A finding that would reverse a decision the user made earlier — in discussion or recorded in the artifact — is Escalate, naming the original decision and the new evidence beside it. Judge by the outcome rather than the wording of the option the user chose: a reversal leaves the user with something materially different from what they chose. A finding that refutes only the factual premise the user's choice rested on, leaving the chosen outcome intact, is a premise correction: confirm the refutation against whichever of the code, the governing artifact, or authoritative documentation the premise turns on, and when none settles it, keep the finding Escalate. Otherwise assign Apply unless another bullet independently calls for Escalate, and add a callout below the table naming both the corrected premise and the chosen outcome it leaves standing.
  • In an iterating loop, a structural Apply triggers another full iteration; count that iteration in the change's cost when applying the Skip cost test. A finding that targets code or text introduced by an earlier iteration's accepted finding, and names no defect in it, is churn: Skip.
  • When a comment at the code in question records why a behavior is unobservable through every reachable path, verify that record against the code before honoring it. A coverage finding re-reporting the gap, naming no newly reachable path that would observe the behavior, is Skip, citing the record.
  • Weight reviewer authority. Feedback from trusted reviewers (repository maintainers or admins) should be treated with higher credibility even when phrased softly.
  • Plan deviation is not a verdict. Do not reject a finding on the grounds that it departs from a plan's prescribed shape. When the plan records a load-bearing reason for that shape, assign Escalate so the user can weigh the trade-off. When the plan is silent on why, or the recorded reason reads like "path of least deviation" or "minimal change", treat the shape as a default and judge the finding on its own merits.
Show full SKILL.md (1,070 more words)Show less

Step 2: Devil's Advocate

After the initial assessment, challenge uncertain findings from a different angle.

Spawn when any finding has Medium or Low confidence. Send only those findings to the sub-agent. High-confidence findings pass through unchallenged. Skip this step entirely if all findings are High confidence.

Capture git status --short, git diff --cached | git hash-object --stdin, git diff | git hash-object --stdin, and git symbolic-ref --short -q HEAD before spawning.

Launch a single sub-agent (inherited model defaults). Provide the Medium/Low-confidence findings with their file locations, claims, and initial verdicts. Instruct the sub-agent to challenge each finding: try to prove it wrong, or confirm it with evidence.

Evidence standards: A refutation counts only when it rests on a defense, guarantee, or documented behavior the sub-agent located and read, or on behavior it observed by running the code; an expectation that a framework, caller, or type already handles the case returns Inconclusive and leaves the initial verdict standing. Confirmed applies to the claim the finding stands or falls on. Establishing the premise beneath that claim leaves it open: that code reads a value settles nothing about whether a test can control that value. When the sub-agent has established only the premise, it returns Inconclusive. Evidence that a test fails when its subject is changed settles only that the test pins the behavior; whether the pinned behavior is the required one stays open. Where the test was written alongside its implementation, their agreement is guaranteed by construction and carries no evidence about the requirement. A finding resting on that evidence is Confirmed only when the requirement itself has been checked against the code's production consumer, or against the governing plan; when neither is reachable, the sub-agent returns Inconclusive.

Protect the shared tree: The sub-agent's prompt must direct it to treat the shared working tree and its git index as read-only; an experiment that needs a scratch project runs in a temp directory outside the repo, or in an isolated git worktree created there and discarded afterward. HEAD stays where it is: read other refs with git show <ref>:<path> rather than git checkout or git switch. Refer to that directory or worktree by absolute path in every command and join chained steps with &&, so a failed step cannot leave the rest running in the shared checkout. Run teardown and verification as their own commands. Give that worktree its own dependency install rather than reaching the shared tree's install by any route: removing a worktree deletes through symlinks, and a redirected suite writes into the shared install. When its own install is not possible, the check is left unrun and reported as such. A check that runs in the shared checkout invokes an already-installed runner directly wherever a package-manager wrapper would front it, since such a wrapper reads as read-only while reconciling the shared install before it runs. Confine dependency installs and reconciliation to an isolated worktree. Every test runner the sub-agent starts, in a temp directory, a worktree, or the shared checkout, runs in its own process group under a timeout enforced from outside the runner. Before teardown, the sub-agent stops the process group of every runner it started, since stopping a runner can leave the processes it spawned alive. Afterward the sub-agent verifies that git worktree list no longer shows the worktree, that git status --short shows what it showed at the start, and that HEAD is still on the branch it started on. It also confirms that no process from those groups, and none whose command line names the temp directory or worktree path, if any, is still running, and reports by PID any process it could not stop. When it cannot list processes, it reports that check as unrun and names those process groups and the temp directory or worktree path, if any. After any check, in a worktree or in the shared checkout, it verifies that the shared tree's dependency directory still resolves (a destroyed install leaves git status unchanged, since it is gitignored). Damage the sub-agent cannot repair is reported with the exact repair command in place of findings.

Verify the tree: re-run all four commands when the sub-agent returns, including when it terminates early or reports incomplete results, and confirm alongside them that the shared tree's dependency directory still resolves, which the four cannot see. Delete what the sub-agent created, revert what it modified or staged, and return HEAD to the captured branch, leaving everything the pre-spawn capture already showed untouched.

The sub-agent picks research tools based on claim type:

Claim TypeTool
API deprecated/removed/changedDocumentation MCP tools or web search
Method doesn't exist / wrong signatureDocumentation MCP tools, web search fallback
Code causes specific bug or behaviorShell (isolated read-only test snippet)
Best practice or ecosystem claimWeb search
Migration or changelog lookupWeb search → web fetch

Use whatever documentation tools are available. The specific tools vary by project setup.

Budget: max 2 research actions per finding. If the first action is conclusive, skip the second.

Sub-agent Verdicts

The sub-agent returns per finding:

  • Confirmed — found evidence supporting the claim (with source)
  • Disputed — found counter-evidence (with source and explanation)
  • Inconclusive — no definitive evidence either way

Step 3: Reconciliation

Merge sub-agent results with the initial assessment:

  • Confirmed: verdict and severity stand. Note the evidence source.
  • Disputed: if originally Apply, downgrade to Skip or Escalate. Re-assess severity if the evidence changes the impact picture. Show both perspectives.
  • Inconclusive: verdict and severity stand, note the uncertainty.

Findings not investigated by the sub-agent keep their original verdict.

For Apply findings, document the issue and location. For Escalate findings, note what information would resolve the ambiguity. For Skip findings, document why.

Step 4: Format Output

Summarize the evaluated findings in a table:

FileIssueSourceSeverityVerdict

When Step 2 ran (any finding was investigated by the Devil's Advocate sub-agent), add an Investigated column:

FileIssueSourceSeverityVerdictInvestigated

Where Investigated shows:

  • (empty) — not investigated by sub-agent
  • Confirmed (source) — sub-agent found supporting evidence
  • Disputed: [reason] — sub-agent found counter-evidence

For findings whose severity was re-assessed from the upstream level, append the change in the Severity cell (e.g., "High (was Medium)").

Carry each verdict into the Verdict column exactly as assessed, Escalate included.

For disputed findings, add a callout below the table showing both perspectives. For each finding, indicate scope in the Issue column (e.g., "Pre-existing:" prefix).

Then call update_plan to mark this step completed and continue with the next step of the active workflow.

© tobihagemann, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in codex/skills/evaluate-findings of tobihagemann/turbo.

Open the folder on GitHubat commit 931eda5

Compare with similar skills

Evaluate Findings next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate Findings compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate Findings this skilltobihagemann/turbo408—~5kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Code Review ChecklistshareAI-lab/learn-claude-code78k4 repos~1.1kAutomated safety check: PassMIT
Backend Code Reviewlangflow-ai/langflow155k—~3.5kAutomated safety check: NotesMIT
Mole Bug Patternstw93/Mole70k—~2kAutomated safety check: PassGPL-3.0
Backend Code Reviewlanggenius/dify158k—~676Automated safety check: PassCustom licence

Similar skills

  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Code Review Checklist

    shareAI-lab/learn-claude-code

    Reviews code against a five-part checklist covering security, correctness, performance, maintainability and testing, and reports findings in a fixed format.

    78k GitHub starsUsed in 4 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Backend Code Review

    langflow-ai/langflow

    Review backend code for quality, security, maintainability, and best practices based on established checklist rules.

    155k GitHub stars~3.5k tokensUpdated yesterday
    DevelopmentAuto-check: notes
  • A catalog of recurring bug shapes in the Mole Mac cleaner, used to review safety-sensitive diffs for deletion safety, unbounded commands, shell traps and weak tests.

    70k GitHub stars~2k tokensUpdated today
    DevelopmentAuto-check passed
  • Backend Code Review

    langgenius/dify

    Reviews backend code under api/ for concrete, reproducible defects, routes to rule packs for architecture, schema, repositories and SQLAlchemy, and ranks findings from P0 to P3.

    158k GitHub stars~676 tokensUpdated today
    DevelopmentAuto-check passed
  • WooCommerce Code Review

    woocommerce/woocommerce

    Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.

    11k GitHub starsUsed in 3 repos~1.1k tokens
    DevelopmentAuto-check passed

More from tobihagemann/turbo

All 81 skills in this repo
  • Consult Oracle

    tobihagemann/turbo

    Consult ChatGPT Pro via ChatGPT browser automation for problems that resist standard approaches.

    408 GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Fetch PR Comments

    tobihagemann/turbo

    Fetch and summarize review feedback and conversation from a GitHub PR (unresolved review threads, review bodies, and PR conversation comments) without making changes.

    408 GitHub stars~967 tokensUpdated 2 days ago
    Auto-check passed
  • Recall Rationale

    tobihagemann/turbo

    Recall why a past change was made by locating the Claude Code transcript that produced it.

    408 GitHub stars~1.3k tokensUpdated 2 days ago
    Auto-check passed
  • Resolve PR Comments

    tobihagemann/turbo

    Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.

    408 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Resolve PR Comments

    tobihagemann/turbo

    Evaluate, fix, answer, and reply to GitHub pull request review comments and conversation comments.

    408 GitHub stars~3.8k tokensUpdated 2 days ago
    Auto-check passed
  • Assess Technical Debt

    tobihagemann/turbo

    Assess project-wide structural technical debt: complexity hotspots, deprecated API usage, duplication clusters, architecture rot, and low-value tests.

    408 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed

Categories

Questions about Evaluate Findings

What does Evaluate Findings do?

Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification. Evaluate Findings is an agent skill from tobihagemann/turbo. Critically assess external feedback (code reviews, AI reviewers, PR comments) and decide which suggestions to apply using adversarial verification.

When should I use Evaluate Findings?

Evaluate Findings fits situations like: the user asks to evaluate findings; assess review comments; triage review feedback; evaluate review output.

How do I install Evaluate Findings in Claude Code?

Run `npx skills add tobihagemann/turbo --skill evaluate-findings -a claude-code`. Or copy the skill folder (codex/skills/evaluate-findings in tobihagemann/turbo) into .claude/skills/evaluate-findings in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate Findings in Codex?

Run `npx skills add tobihagemann/turbo --skill evaluate-findings -a codex`. Or copy the skill folder (codex/skills/evaluate-findings in tobihagemann/turbo) into .agents/skills/evaluate-findings in your project. Codex loads it when a task matches its description.

Can I use Evaluate Findings in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add tobihagemann/turbo --skill evaluate-findings -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-findings, .gemini/skills/evaluate-findings, .github/skills/evaluate-findings and .opencode/skills/evaluate-findings in your project.

What does Evaluate Findings need to run?

Going by SKILL.md and its folder, Evaluate Findings needs the command-line tools its instructions call (git).

Does Evaluate Findings access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evaluate Findings safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate Findings use?

Evaluate Findings is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate Findings use?

About 5k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate Findings?

Skills that share tags, products or a category with Evaluate Findings: PR Babysitter (openinterpreter/openinterpreter, 69k stars), Code Review Checklist (shareAI-lab/learn-claude-code, 78k stars), Backend Code Review (langflow-ai/langflow, 155k stars) and Mole Bug Patterns (tw93/Mole, 70k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate Findings?

tobihagemann (a GitHub user) maintains it in tobihagemann/turbo, which has 408 GitHub stars. The repository holds 81 skills in this directory. The repository was last updated on October 9, 2026.

Source: tobihagemann/turbo on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.