Agent skill

Meticulous Review

by FlintSH in FlintSH/Flare

Analyze a completed Meticulous test run — compare the diffs against the PR description to see what's expected, then focus on finding and flagging potential regressions.

MITAuto-check passedDevelopment

Install Meticulous Review

skills CLI
$ npx skills add FlintSH/Flare --skill meticulous-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install FlintSH/Flare meticulous-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/FlintSH/Flare.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/meticulous-review .claude/skills/meticulous-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
meticulous-review
GitHub stars
135
Token cost
~2.3k tokens
SKILL.md length
969 words
Files
1
Skills in repo
10
Repo updated
First seen
Licence
MIT

At a glance

Analyze a completed Meticulous test run — compare the diffs against the PR description to see what's expected, then focus on finding and flagging potential regressions.

  • Works in 9 steps: Establish what's expected → Get the replay diff summary → Get screenshot images → …
  • Asked to review Meticulous test results
  • SKILL.md covers Step 0 -- Establish what's…, Step 1 -- Get the replay diff…, Step 2 -- Get screenshot images and Step 3 -- Inspect the DOM diff…, plus 5 more sections
  • Calls gh, glab and git

What it does

Meticulous Review is an agent skill from FlintSH/Flare. Analyze a completed Meticulous test run — compare the diffs against the PR description to see what's expected, then focus on finding and flagging potential regressions. Resolves the test run from the local repo's current commit (the default), or from an explicit test-run ID or commit SHA. Use when asked to review Meticulous test results, when babysitting a pull/merge request's Meticulous Tests CI check, or right after implementing a frontend change yourself.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Pull requests. The repository describes itself as: A modern, lightning-fast file sharing platform built for self-hosting. Created with support for ShareX, KDE Spectacle, Flameshot, and easy to set up. The licence is MIT.

When your agent uses it

  • Asked to review Meticulous test results
  • Babysitting a pull/merge requests Meticulous Tests CI check
  • Right after implementing a frontend change yourself

Example prompts

  • “/meticulous-review”

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Establish what's expected
  2. Get the replay diff summary
  3. Get screenshot images
  4. Inspect the DOM diff (for structural detail)
  5. Get the replay timeline (optional, for diagnosing unexpected diffs)
  6. Decision guide
  7. Flag the diff
  8. Final report
  9. Report feedback to Meticulous

What it can do on your machine

Read from SKILL.md and the folder at commit c910523. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • glab
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh, glab and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Meticulous Review loads about 2.3k tokens when it runs. Until then it costs about 120 tokens; SKILL.md has 969 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~120
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from FlintSH/Flare at commit c910523, republished under its MIT licence (© FlintSH). 969 words, ~2,300 tokens.

Download SKILL.mdSave it as .claude/skills/meticulous-review/SKILL.md (or your agent's skills folder).
name
meticulous-review
description
Analyze a completed Meticulous test run — compare the diffs against the PR description to see what's expected, then focus on finding and flagging potential regressions. Resolves the test run from the local repo's current commit (the default), or from an explicit test-run ID or commit SHA. Use when asked to review Meticulous test results, when babysitting a pull/merge request's Meticulous Tests CI check, or right after implementing a frontend change yourself.
user-invocable
true

To review a Meticulous test run, follow the workflow below step by step, using the CLI or MCP commands as described.

Before starting, run the meticulous-cli-update skill to ensure the Meticulous CLI and skills are up to date — unless it has already run earlier in this conversation, in which case skip it.

This skill treats you as a reviewer, not the implementer — even if you did write the change earlier in this conversation. Its job is to catch regressions, not to iterate on the implementation (if you're mid-implementation and want to loop against Meticulous until things look right, see the meticulous-iterative-dev skill for feature work with intended visual changes, or the meticulous-zero-diff-task skill when the UI must not change; if diffs have already been reviewed and rejected and you're just here to fix what's flagged, see the meticulous-fix skill).

Step 0 -- Establish what's expected

Before looking at any diff, work out what visual change this PR is supposed to produce:

  • If you already have full context (same conversation that implemented the change, or the user just described the task), use that.
  • Otherwise — fetch the PR description (e.g. gh pr view <number> --json title,body for GitHub, glab mr view <id> --output json --jq '{title,description}' for GitLab, or GET /2.0/repositories/{workspace}/{repo_slug}/pullrequests/{id}?fields=title,description for Bitbucket) and pull out a brief bullet-point summary of any visual changes it calls out as expected.
  • If it doesn't mention any, note that explicitly — every diff below then gets extra scrutiny, since there's nothing on record it could be a known, accepted consequence of.

Step 1 -- Get the replay diff summary

Run from the local checkout to resolve the test run from the current commit's git HEAD — make sure HEAD matches the remote head CI ran on first (e.g. git pull), or you may review a stale or missing run:

bash
# CLI (infers the run from local git HEAD)
meticulous agent test-run-diffs

# MCP (git context is never inferred — resolve the testRunId from a commit first)
get_test_run_for_commit(commitSha="<sha>")
get_test_run_diffs(testRunId="<id>")

Returns a TSV of replayDiffId/screenshotName rows — a representative, priority-ordered subset of real visual differences; work through them top to bottom. To target a run explicitly instead of resolving from HEAD, pass --testRunId <id> or --commitSha <sha>. The CLI blocks until the run finishes by default (pass --dontWaitForTestRunToComplete to instead report an in-progress run and exit immediately); MCP never blocks, so keep polling until status is complete/failed.

Every returned row must be matched against Step 0 or flagged (see the Decision guide) before concluding the PR is good.

Step 2 -- Get screenshot images

For each representative screenshot:

bash
# CLI (downloads images to ~/.meticulous/agent-images/ and prints local paths)
meticulous agent image-files --replayDiffId <replayDiffId> --screenshotName <screenshotName>

# MCP (no download-to-disk tool — returns signed URLs instead; fetch them to view the images)
get_image_urls(replayDiffId="<replayDiffId>", screenshotName="<screenshotName>")

Open (or fetch) before, after, and diffImage to inspect the change — diffImage is usually the most informative, highlighting exactly which pixels changed. Always inspect the images, even when the DOM diff looks clear.

Step 3 -- Inspect the DOM diff (for structural detail)

bash
# CLI
meticulous agent dom-diff --replayDiffId <replayDiffId> --screenshotName <screenshotName>

# MCP
get_dom_diff(replayDiffId="<replayDiffId>", screenshotName="<screenshotName>")

Optional: --context <N|full> (CLI) controls how many context lines surround each hunk (default 3).

Output is a unified diff (+/-, indentation stripped), one [diff N]-headed block per independent change — for example (illustrative, not real output):

[diff 0]
 <div class="item">
-  <span class="label">old label</span>
+  <span class="label" data-flag="true">new label</span>
 </div>
[diff 1]
 <ul class="list">
+  <li>new item</li>
 </ul>

Step 4 -- Get the replay timeline (optional, for diagnosing unexpected diffs)

If a diff is unexpected and the images/DOM don't make it obvious why:

bash
# CLI
meticulous agent timeline-diff --replayDiffId <replayDiffId>

# MCP
get_timeline_diff(replayDiffId="<replayDiffId>")

TSV columns: diff ( identical, - removed, + added, ! changed), timeMs, event (user/screenshot/network/console/etc.), description. Look for failed network requests, unexpected redirects, or timing anomalies that could explain a visual change.

Show full SKILL.md (446 more words)Show less

Step 5 -- Decision guide

For each representative screenshot, compare the diff image and DOM diff against Step 0's expectations:

  • Expected — matches one of Step 0's expected changes (or, with full implementation context, is clearly a desired outcome). Check the diff actually looks like that change and nothing more — a diff can be expected in kind but still carry an extra, unrelated regression bundled into the same screenshot. Nothing to flag.
  • Unintended — not accounted for by Step 0. Use the timeline to rule out failed requests, redirects, or other anomalies, then flag it:
    • Potential regression (a real side effect, or otherwise clearly wrong) → reject.
    • Unrelated to the change under review — typically a flake, e.g. subpixel rendering noise or animation non-determinism → ignore.

Either way it's flagged, not silently dropped — a human still needs to see it.

This skill reviews and flags — it does not fix. Hand a rejected diff off to the meticulous-fix skill (or the person/skill implementing the change) — don't attempt code changes here.

Step 6 -- Flag the diff

bash
# CLI
meticulous agent reject-diff --replayDiffId=<id> --screenshotName=<name> --reason="<why>" --x=<0..1> --y=<0..1>
meticulous agent ignore-diff --replayDiffId=<id> --screenshotName=<name> --reason="<why>" --x=<0..1> --y=<0..1>
meticulous agent create-diff-comment --replayDiffId=<id> --screenshotName=<name> --text="<note>" --x=<0..1> --y=<0..1>

# MCP
reject_diff(replayDiffId="<id>", screenshotName="<name>", reason="<why>", x=<0..1>, y=<0..1>)
ignore_diff(replayDiffId="<id>", screenshotName="<name>", reason="<why>", x=<0..1>, y=<0..1>)
create_diff_comment(replayDiffId="<id>", screenshotName="<name>", text="<note>", x=<0..1>, y=<0..1>)

Call reject-diff or ignore-diff for every diff classified as unintended, in addition to including it in the final report. --reason is the succinct explanation from your classification above; --x/--y are the approximate normalized coordinates of the changed region, estimated from the diff image. create-diff-comment is the neutral option for anything you want on the record without a verdict.

Not symmetric: reject-diff writes a real, blocking decision, same as a human rejection. ignore-diff decides nothing — it's a comment only, so the diff stays unreviewed and the check stays pending either way. Only a human can clear a diff, so don't oversell an ignore-diff call in your final report as having resolved anything.

Step 7 -- Final report

Cover all significant visual changes.

  1. Expected changes — brief, a line or two each: what changed and which Step 0 expectation it matches.
  2. Flagged diffs (if any) — the main point of the review, so give these the most detail: replayDiffId/screenshotName (linked: https://app.meticulous.ai/test-runs/<testRunId>/replay-diff/<replayDiffId>?screenshot=<screenshotName>), whether you rejected or ignored it, the reason you gave when flagging it (Step 6), what the change looks like, and your best assessment of the cause.

The PR is only good when every diff has been matched or flagged. If any diff is flagged, the PR is not yet good: surface it clearly to the user in addition to the flag itself.

Step 8 -- Report feedback to Meticulous

Always do this as the last step — it's part of the review itself, not something the user has to ask for. Submit one brief note: did Meticulous catch a real problem, was anything confusing, what would have made the review easier. Positive feedback counts too — this isn't just for reporting friction.

bash
# CLI
meticulous agent submit-feedback --message="<one or two sentences>" --outcome=<helped|neutral|hindered> --testRunId=<id> --skill=meticulous-review

# MCP
submit_feedback(message="<one or two sentences>", outcome="<helped|neutral|hindered>", testRunId="<id>", skill="meticulous-review")

© FlintSH, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/meticulous-review of FlintSH/Flare.

Open the folder on GitHubat commit c910523

Compare with similar skills

Meticulous Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Meticulous Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Meticulous Review this skillFlintSH/Flare135—~2.3kAutomated safety check: PassMIT
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Check PRonyx-dot-app/onyx32k2 repos~2.3kAutomated safety check: PassMIT
PR Design DocOpenHands/OpenHands91k—~2.4kAutomated safety check: PassMIT
WooCommerce Code Reviewwoocommerce/woocommerce11k3 repos~1.1kAutomated safety check: PassCustom licence

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Check PR

    onyx-dot-app/onyx

    Checks a GitHub, GitLab, or Perforce (p4) pull request (or merge request, or shelved changelist) for unresolved review comments, failing status checks, and incomplete PR descriptions.

    32k GitHub starsUsed in 2 repos~2.3k tokens
    DevelopmentAuto-check passed
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    91k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • WooCommerce Code Review

    woocommerce/woocommerce

    Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.

    11k GitHub starsUsed in 3 repos~1.1k tokens
    DevelopmentAuto-check passed
  • Record PR Demo

    payloadcms/payload

    A skill your agent uses when a Payload pull request needs a concise visual walkthrough for reviewers.

    45k GitHub stars~1k tokensUpdated yesterday
    DevelopmentAuto-check passed

More from FlintSH/Flare

All 10 skills in this repo
  • Meticulous CLI

    FlintSH/Flare

    Overview of the Meticulous CLI tool and its global options. An agent skill from FlintSH/Flare.

    135 GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check passed
  • Meticulous Fix

    FlintSH/Flare

    Fix the visual diffs that have been reviewed and rejected on a Meticulous test run, following their review comments if given.

    135 GitHub stars~2.3k tokensUpdated 5 days ago
    Auto-check passed
  • Increase coverage for a Meticulous project by tracing specific under-covered files back to a real UI action in the codebase, driving that action with a real recorded browser session, and validating…

    135 GitHub stars~7k tokensUpdated 5 days ago
    Auto-check passed
  • Iterative frontend development loop using Meticulous for per-step visual validation.

    135 GitHub stars~1.3k tokensUpdated 5 days ago
    Auto-check passed
  • Run a Meticulous session simulation against a live URL and analyze the visual output — either by inspecting screenshots directly (quick-check mode) or by comparing pixel and HTML diffs against a…

    135 GitHub stars~1.8k tokensUpdated 5 days ago
    Auto-check passed
  • Meticulous Test

    FlintSH/Flare

    Run a Meticulous test run after implementing a frontend change, then hand off to the meticulous-review skill to classify each visual change as intended or unintended.

    135 GitHub stars~1.2k tokensUpdated 5 days ago
    Auto-check passed

Categories

Questions about Meticulous Review

What does Meticulous Review do?

Analyze a completed Meticulous test run — compare the diffs against the PR description to see what's expected, then focus on finding and flagging potential regressions. Meticulous Review is an agent skill from FlintSH/Flare. Analyze a completed Meticulous test run — compare the diffs against the PR description to see what's expected, then focus on finding and flagging potential regressions.

When should I use Meticulous Review?

Meticulous Review fits situations like: asked to review Meticulous test results; babysitting a pull/merge requests Meticulous Tests CI check; right after implementing a frontend change yourself.

How do I install Meticulous Review in Claude Code?

Run `npx skills add FlintSH/Flare --skill meticulous-review -a claude-code`. Or copy the skill folder (.agents/skills/meticulous-review in FlintSH/Flare) into .claude/skills/meticulous-review in your project. Claude Code loads it when a task matches its description.

How do I install Meticulous Review in Codex?

Run `npx skills add FlintSH/Flare --skill meticulous-review -a codex`. Or copy the skill folder (.agents/skills/meticulous-review in FlintSH/Flare) into .agents/skills/meticulous-review in your project. Codex loads it when a task matches its description.

Can I use Meticulous Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add FlintSH/Flare --skill meticulous-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/meticulous-review, .gemini/skills/meticulous-review, .github/skills/meticulous-review and .opencode/skills/meticulous-review in your project.

What does Meticulous Review need to run?

Going by SKILL.md and its folder, Meticulous Review needs the command-line tools its instructions call (gh, glab and git).

Does Meticulous Review access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Meticulous Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Meticulous Review use?

Meticulous Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Meticulous Review use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Meticulous Review?

Skills that share tags, products or a category with Meticulous Review: Finishing a Development Branch (obra/superpowers, 297k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars), Check PR (onyx-dot-app/onyx, 32k stars) and PR Design Doc (OpenHands/OpenHands, 91k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Meticulous Review?

FlintSH (a GitHub user) maintains it in FlintSH/Flare, which has 135 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on October 6, 2026.

Source: FlintSH/Flare on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.