Agent skill

System1-Agents Self-Review

by ThinkFlowLab in ThinkFlowLab/system1-agents

Prepares a local self-review report for a System1-Agents change before a PR is opened, checking the full diff, tests, docs, artifacts and claims against evidence.

Apache-2.0Auto-check passedDevelopment

Install System1-Agents Self-Review

skills CLI
$ npx skills add ThinkFlowLab/system1-agents --skill self-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ThinkFlowLab/system1-agents self-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ThinkFlowLab/system1-agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/self-review .claude/skills/self-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
self-review
GitHub stars
126
Token cost
~2.1k tokens
SKILL.md length
1,154 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
Apache-2.0

At a glance

Prepares a local self-review report for a System1-Agents change before a PR is opened, checking the full diff, tests, docs, artifacts and claims against evidence.

  • Self-reviewing a System1-Agents branch before opening a PR
  • SKILL.md covers Committed artifact hygiene, Task evidence, PR demo/evidence section and Report
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Checking a large change for readiness gaps against the contributing guide

What it does

This skill produces a contributor report for changes to the System1-Agents repository before a PR is opened or review is requested. The agent reads the checkout's CONTRIBUTING.md, PR template and local instructions, records the target branch and the base and head commits, reviews the full diff from the merge base plus uncommitted changes, and separates what is in the PR from local-only work. Above 3,000 changed code lines it expects the contributor's full self-review, split rationale, component map and validation before calling the PR ready, and size alone is not a correctness finding.

The checks include correctness, focused scope, decision-model and front contracts, fallback behavior, cancellation and timeouts, resource cleanup, tests for failure paths and a regression case for fixes, and docs that match the implementation. The agent runs the checks listed in CONTRIBUTING.md and reports exact commands, outcomes and skipped checks, and it inspects committed JSON, CSV, logs, reports and generated media for artifact hygiene. The skill does not authorize edits, commits, pushes, paid model calls or downloads.

When your agent uses it

  • Self-reviewing a System1-Agents branch before opening a PR
  • Checking a large change for readiness gaps against the contributing guide
  • Reviewing a use-case recipe against the recipe template

Example prompts

  • “Self-review my branch against CONTRIBUTING.md before I open the PR.”
  • “Check whether my diff is over the large-change threshold and what review preparation is missing.”
  • “Review the committed JSON and log files in this diff for artifact hygiene.”

Requirements

  • A System1-Agents checkout with CONTRIBUTING.md
  • Git

What it can do on your machine

Read from SKILL.md and the folder at commit 3a2c2c6. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

System1-Agents Self-Review loads about 2.1k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 1,154 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ThinkFlowLab/system1-agents at commit 3a2c2c6, republished under its Apache-2.0 licence (© ThinkFlowLab). 1,154 words, ~2,126 tokens.

Download SKILL.mdSave it as .claude/skills/self-review/SKILL.md (or your agent's skills folder).
name
self-review
description
Self-review System1-Agents changes before opening a PR or requesting review, with reproducible task outcomes, appropriate demos, and evidence-backed claims.

System1-Agents self-review

Read the target checkout's CONTRIBUTING.md, PR template, and applicable local instructions. Record the actual target branch and base/head commits, and review the full diff from their merge base plus relevant uncommitted changes. Distinguish what is in the PR from local-only work; disclose if the base could not be refreshed.

Apply the large-code-change requirements: report authored-code and total diff counts separately using the guide's counting convention. Above 3,000 changed code lines, verify the contributor's full self-review, split rationale, component map/review order, and validation across affected components and interfaces before recommending readiness. A quick precheck is insufficient; keep the PR draft until the contributor self-review is complete. Report missing preparation as a readiness gap; size alone is not a correctness finding or a reason to require expensive model/GPU runs.

Check correctness, focused scope, decision-model and front contracts, fallback behavior, cancellation/timeouts, and resource cleanup as relevant. Verify tests cover the changed behavior, including failure paths and a regression case for a fix. Match docs and commands to the implementation. Run the applicable checks in CONTRIBUTING.md; report exact commands, outcomes, and skipped checks with reasons. Documentation-only changes need link/example/claim checks, not unrelated model runs. For use-case recipes, check the recipe template: commands reach the named agent/backend, result checks can detect failure, and each profile's validation status matches its evidence.

This skill prepares a local contributor report. It does not itself authorize edits, commits, pushes, external posts, paid model calls, downloads, or changes to review status. Use only separately authorized execution resources and budgets.

Committed artifact hygiene

Apply this check in every review, including quick prechecks. Inspect added and changed artifacts in the complete diff, including JSON/JSONL, CSV, logs, reports, source/binary hash inventories, and generated media. Classify them by purpose and actual consumer, rather than rejecting a file extension.

  • Keep necessary configuration, request examples, maintained benchmark inputs, and small deterministic fixtures or reference oracles in the repository's intended locations. Identify the test, tool, or documented workflow that needs each retained artifact, such as labelled agent datasets or replay inputs.
  • Flag one-off run summaries, response dumps, cache statistics, profiler output, agent process notes, and duplicate historical results that have no maintained source-tree role. A link from PR prose or documentation alone does not justify committing generated run output. Report concrete paths and consumers.
  • Preserve raw measurements, failures, and provenance in a durable artifact archive or PR/CI evidence, and link the exact revision or run from the summary. Do not discard evidence to reduce the diff or hide it in a committed archive.
  • When removing redundant output, check its callers, links, and reproduction commands. Keep replay inputs and expected responses intact; verify their hashes and rerun the affected replay or documentation checks.

Task evidence

For changes to what an agent can accomplish, show a concrete task as input → actions → final result. Define success using an observable outcome, not just a DONE signal. Use a safe, reproducible scenario or fixture; include setup, commands, inputs/seed, model and configuration, environment, and exact source commits. Link the resulting logs or artifacts so a reviewer can trace the story.

When claiming improvement, compare baseline and head on the same tasks, inputs, success criteria, and budgets. Report completion counts/denominators, elapsed time, model calls and tool calls when measured. Include cost only if measured, with the accounting scope and pricing basis; do not infer total cost from latency or an incomplete token charge. Keep failures, retries, timeouts, and human intervention in the results. Report repetitions and variation; one successful demo is not a task-success rate. Mark missing metrics unmeasured, and remove or qualify unsupported claims rather than manufacturing a comparison.

Use CONTRIBUTING.md's video guide to classify the PR and prepare, record and attach its demo. Important PRs require a video of the application/task, System1-Agents decision-model agent and actual System1-Omni inference in the same run. Check all three parts against the linked trace; a terminal recording works for text agents and rails. Screenshots and logs support the clip. Use PR #35's recording linked in the guide as the example and choose a relevant README application candidate from the guide. Show the relevant input, action sequence, and result, with failures or human intervention visible. Label cuts, replay speed, and elapsed timing honestly; link a fuller trace when a clip omits context. Do not stage screens or present a replay as a live run. A replay must identify the source run/commit, workload, and speed; a historical or upstream model demo is not proof that the current agent integration works.

Choose figures that answer the review question: workflow screenshots, a short action timeline, or task-success comparisons backed by the run records. There is no asset quota. Only the guide's minor docs/formatting/test-only exemption permits N/A with a reason; a nonvisual task still needs a terminal video when the PR is important. If a required run cannot be made within the available authorization, resources, or budget, report the gap and its impact; keep the important PR draft until the video is supplied or a maintainer accepts the documented exception. Do not turn a missing run into a pass or require a production-scale demonstration.

Show full SKILL.md (309 more words)Show less

PR demo/evidence section

Prepare a Demo / evidence section for the PR containing what applies:

  • Required application + agents + Omni video and trace, or the documented exemption/blocker.
  • Task and observable result, or N/A for an exempt change with a concrete reason.
  • Reproduction command/fixture, configuration, and baseline/head commits.
  • Measured comparison and raw result links, including failures and limitations; distinguish personally run checks, author-reported results, and observed CI.
  • Demo/figure links with captions stating the workload, source revision, and whether each item is an actual run, recorded replay, or explanatory illustration.

Make evidence reusable for accurate reviews and public updates without implying permission to publish it elsewhere. Before attaching assets, check ownership, license/attribution, and permission to share. Use safe sample data and redact credentials, private URLs, personal/customer information, and sensitive screen or log content. Verify redaction in the final exported files, captions, and metadata. Keep useful measurement context after redaction. Clearly label diagrams, mockups, and generated artwork as illustrations; never fabricate screens, results, or performance claims. Link durable, reviewer-accessible artifacts rather than local paths. If rights or safe disclosure are unresolved, omit the asset and state why.

Check System1-Omni's current model, modality and hardware support as described in the video guide. Use a supported serving path for the required video when the branch has a compatible client; otherwise record the missing integration or configuration. Pin both repositories and verify the actual worker/frontend from run evidence. Keep agent-task evidence separate from serving/kernel measurements. Do not add an unrequested backend integration or run outside the authorized resources/budget merely to produce a demo.

Report

Lead with actionable findings and file/line references, then the reviewed scope, commands/results, demo/evidence summary, and remaining gaps. Say when there are no actionable findings without implying maintainer approval. Keep blocking gaps visible and recommend a draft while they remain. Do not check the contributor's boxes or publish on their behalf without separate authorization.

© ThinkFlowLab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/self-review of ThinkFlowLab/system1-agents.

Open the folder on GitHubat commit 3a2c2c6

Compare with similar skills

System1-Agents Self-Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

System1-Agents Self-Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
System1-Agents Self-Review this skillThinkFlowLab/system1-agents126—~2.1kAutomated safety check: PassApache-2.0
GitHub Review Iterationprisma/orm48k—~2.2kAutomated safety check: PassApache-2.0
Cherry Studio PR ReviewCherryHQ/cherry-studio52k—~3.9kAutomated safety check: PassAGPL-3.0
Review Triage Phaseprisma/orm48k—~995Automated safety check: PassApache-2.0
Deep Reviewdyad-sh/dyad22k—~1.4kAutomated safety check: PassCustom licence
PR Reviewjaemk/self_update961—~1.5kAutomated safety check: NotesMIT

Similar skills

  • Official

    Runs a loop on a GitHub pull request: fetch review state, triage comments into actions, implement them and resolve threads, repeating until nothing actionable is left.

    48k GitHub stars~2.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Cherry Studio PR Review

    CherryHQ/cherry-studio

    Reviews Cherry Studio branches, pull requests, commits, files and docs against the project's own architecture, naming, API-boundary and UI rules, report-only by default.

    52k GitHub stars~3.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Official

    Runs the triage step of the review-framework loop: reads fetched PR review state, builds `review-actions.json`, validates it and renders `review-actions.md`.

    48k GitHub stars~995 tokensUpdated today
    DevelopmentAuto-check passed
  • Deep Review

    dyad-sh/dyad

    Deep multi-agent code review run locally — a fleet of parallel finder agents reviews the diff from independent angles, then adversarial verifier agents reproduce each finding before it is reported.

    22k GitHub stars~1.4k tokensUpdated today
    DevelopmentAuto-check passed
  • PR Review

    jaemk/self_update

    Targeted, read-only review of a PR or checked-out branch. An agent skill from jaemk/self_update.

    961 GitHub stars~1.5k tokensUpdated 1 mo ago
    DevelopmentAuto-check: notes
  • Adopt PR Branch Context

    pydantic/pydantic-ai-harness

    Official

    Fills in issue-brief.md and pr-decisions.md for an existing pull request, so you can pick up a PR mid-flight with its linked issue and past review decisions summarized.

    952 GitHub stars~1.8k tokensUpdated 6 days ago
    DevelopmentAuto-check passed

More from ThinkFlowLab/system1-agents

  • System 1 Agent Builder

    ThinkFlowLab/system1-agents

    Scaffolds a new System 1 agent module for a named task in the system1-agents repo, after a fit probe, with its test and README row, verified model by model.

    126 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Review PR

    ThinkFlowLab/system1-agents

    Review pull requests for system1-agents with high-confidence, evidence-based feedback.

    126 GitHub stars~851 tokensUpdated today
    Auto-check passed
  • S1A Decision Agents

    ThinkFlowLab/system1-agents

    Delegates click-through web tasks, games and quizzes to S1A through its s1a command, or asks a fast decision model to pick one option from a list you provide.

    126 GitHub stars~1.1k tokensUpdated today
    Auto-check: notes
  • Add Agent Recipe

    ThinkFlowLab/system1-agents

    Add or update runnable use-case recipes for existing System1-Agents agents, with setup, inference-backend configuration, independent result checks, demo evidence and troubleshooting.

    126 GitHub stars~797 tokensUpdated today
    Auto-check passed

Questions about System1-Agents Self-Review

What does System1-Agents Self-Review do?

Prepares a local self-review report for a System1-Agents change before a PR is opened, checking the full diff, tests, docs, artifacts and claims against evidence. This skill produces a contributor report for changes to the System1-Agents repository before a PR is opened or review is requested.md, PR template and local instructions, records the target branch and the base and head commits, reviews the full diff from the merge base plus uncommitted changes, and separates what is in the PR from local-only work.

When should I use System1-Agents Self-Review?

System1-Agents Self-Review fits situations like: self-reviewing a System1-Agents branch before opening a PR; checking a large change for readiness gaps against the contributing guide; reviewing a use-case recipe against the recipe template.

How do I install System1-Agents Self-Review in Claude Code?

Run `npx skills add ThinkFlowLab/system1-agents --skill self-review -a claude-code`. Or copy the skill folder (.agents/skills/self-review in ThinkFlowLab/system1-agents) into .claude/skills/self-review in your project. Claude Code loads it when a task matches its description.

How do I install System1-Agents Self-Review in Codex?

Run `npx skills add ThinkFlowLab/system1-agents --skill self-review -a codex`. Or copy the skill folder (.agents/skills/self-review in ThinkFlowLab/system1-agents) into .agents/skills/self-review in your project. Codex loads it when a task matches its description.

Can I use System1-Agents Self-Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ThinkFlowLab/system1-agents --skill self-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/self-review, .gemini/skills/self-review, .github/skills/self-review and .opencode/skills/self-review in your project.

What does System1-Agents Self-Review need to run?

SKILL.md names no scripts, command-line tools or credentials: System1-Agents Self-Review is instructions for the agent only. Our summary lists: A System1-Agents checkout with CONTRIBUTING.md; Git.

Does System1-Agents Self-Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is System1-Agents Self-Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does System1-Agents Self-Review use?

System1-Agents Self-Review is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does System1-Agents Self-Review use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to System1-Agents Self-Review?

Skills that share tags, products or a category with System1-Agents Self-Review: GitHub Review Iteration (prisma/orm, 48k stars), Cherry Studio PR Review (CherryHQ/cherry-studio, 52k stars), Review Triage Phase (prisma/orm, 48k stars) and Deep Review (dyad-sh/dyad, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains System1-Agents Self-Review?

ThinkFlowLab (a GitHub organization) maintains it in ThinkFlowLab/system1-agents, which has 126 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on October 8, 2026.

Source: ThinkFlowLab/system1-agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.