Agent skill

Self Eval Bias

by Archive228 in Archive228/loopkit

Detect and interrupt the pattern where an agent confidently praises work it just produced instead of reviewing it critically.

MITAuto-check passed

Install Self Eval Bias

skills CLI
$ npx skills add Archive228/loopkit --skill self-eval-bias -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Archive228/loopkit self-eval-bias --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/self-eval-bias .claude/skills/self-eval-bias && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
self-eval-bias
GitHub stars
755
Token cost
~1k tokens
SKILL.md length
546 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Detect and interrupt the pattern where an agent confidently praises work it just produced instead of reviewing it critically.

  • Works in 6 steps: Notice the same-context tell. If the… → Force a fresh persona. Drop the… → Demand concrete evidence, not verdicts.… → …
  • SKILL.md covers When to apply, Procedure, Anti-patterns and When NOT to apply, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Self Eval Bias is an agent skill from Archive228/loopkit. Detect and interrupt the pattern where an agent confidently praises work it just produced instead of reviewing it critically. Same-context grading is not review — it's rationalization.

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 33 battle-tested skills + minimal .claude harness for any coding agent (Claude Code, Cursor, Codex, Gemini CLI). The licence is MIT.

Example prompts

  • “/self-eval-bias”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Notice the same-context tell. If the review verdict lands in under three sentences and contains "looks correct", "this should work", or…
  2. Force a fresh persona. Drop the generation context. Open a new subagent, or at minimum re-prompt with only the artifact (diff, plan…
  3. Demand concrete evidence, not verdicts. The reviewer must cite: the file:line it inspected, the input it ran, the observed output, and the…
  4. Adversarially probe. Ask the reviewer for the strongest case where the artifact fails. If it can't produce one, the review didn't happen…
  5. Run the artifact. For code, exercise it end-to-end (see [[broken-window-check]]). For a plan, walk the first two steps concretely…
  6. Rotate the reviewer periodically. In long multi-agent loops, re-prompt the evaluator from scratch every ~5 sprints — leniency drift…

What it can do on your machine

Read from SKILL.md and the folder at commit 5ae033e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • blog.anthropic.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Self Eval Bias loads about 1k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 546 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Archive228/loopkit at commit 5ae033e, republished under its MIT licence (© Archive228). 546 words, ~1,019 tokens.

Download SKILL.mdSave it as .claude/skills/self-eval-bias/SKILL.md (or your agent's skills folder).
name
self-eval-bias
description
Detect and interrupt the pattern where an agent confidently praises work it just produced instead of reviewing it critically. Same-context grading is not review — it's rationalization.
when_to_use
about to grade or accept output produced in the same context that generated it, reviewer verdict is high-confidence positive with no cited concrete evidence…

Self-Eval Bias

An agent that just produced a plan, a diff, or a report cannot fairly grade it in the same context. The reasoning that justified writing it is still loaded — every doubt was already resolved in favor of shipping. Asked to review, the same context reliably returns "looks good, ship it." This is not review. It is rationalization wearing a review's uniform.

The pattern shows up hardest in planner/generator/evaluator architectures where the evaluator drifts toward leniency over long runs — the prompts it reads fill up with the generator's reasoning, and skepticism erodes. (See Prithvi's March 2026 post on the three-agent harness: https://blog.anthropic.com/three-agent-harness-march-2026.)

When to apply

  • You just wrote code, a plan, or a claim, and the next step is "confirm it's correct".
  • A reviewer verdict comes back positive with no cited line numbers, no failing case explored, no counter-example attempted.
  • You're about to mark a feature passes: true, close an issue, or hand off to the next session.
  • The evaluator persona in a multi-agent loop has agreed with the last N generator outputs in a row.

Procedure

  1. Notice the same-context tell. If the review verdict lands in under three sentences and contains "looks correct", "this should work", or "no issues found" without a cited artifact — treat the verdict as unwritten.
  2. Force a fresh persona. Drop the generation context. Open a new subagent, or at minimum re-prompt with only the artifact (diff, plan, output) and the acceptance criteria — no reasoning trail, no self-justification.
  3. Demand concrete evidence, not verdicts. The reviewer must cite: the file:line it inspected, the input it ran, the observed output, and the criterion it matched against. "LGTM" without these is a null review — discard it.
  4. Adversarially probe. Ask the reviewer for the strongest case where the artifact fails. If it can't produce one, the review didn't happen — the reviewer just agreed.
  5. Run the artifact. For code, exercise it end-to-end (see [[broken-window-check]]). For a plan, walk the first two steps concretely. Same-context confidence collapses fast against a runtime.
  6. Rotate the reviewer periodically. In long multi-agent loops, re-prompt the evaluator from scratch every ~5 sprints — leniency drift compounds silently.
Show full SKILL.md (190 more words)Show less

Anti-patterns

  • Self-review in the same turn. "Let me double-check my work" followed by immediate approval. The doubt has to cost something to be real.
  • Praise as evidence. "This is a clean, well-structured implementation" is a vibe, not a finding. Findings cite lines.
  • Positive verdict, empty failure_scenario. If the reviewer can't describe what a failure would look like, they didn't look for one.
  • Rubber-stamping across a run. N consecutive "approved" verdicts from the same evaluator without a single rejection is a red flag, not a track record.
  • Fixing the criterion instead of the artifact. Reviewer notices a gap, then edits the spec to say the gap is out of scope. The gap is in the artifact. Fix that.

When NOT to apply

  • The output is trivial and cheap to redo if wrong (a one-line rename, a config toggle).
  • A separate reviewer with a fresh context already ran and cited concrete evidence — the check has been done, don't loop on it.
  • [[broken-window-check]] — the runtime-driven version of "don't trust the last claim".
  • [[adversarial-verify]] — the structural form of "find the strongest failure case".
  • [[shift-notes]] — where you record what the fresh-persona review actually found.

© Archive228, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/self-eval-bias of Archive228/loopkit.

Open the folder on GitHubat commit 5ae033e

Compare with similar skills

Self Eval Bias next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Self Eval Bias compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Self Eval Bias this skillArchive228/loopkit755—~1kAutomated safety check: PassMIT
Detecting Anomalous Authentication Patternsmukul975/Anthropic-Cybersecurity-Skills34k—~7.5kAutomated safety check: PassApache-2.0
Detecting Beaconing Patterns With Zeekmukul975/Anthropic-Cybersecurity-Skills34k—~662Automated safety check: PassApache-2.0
Golang Patternsaffaan-m/ECC276k—~1.1kAutomated safety check: PassMIT
Kotlin Exposed Patternsaffaan-m/ECC276k4 repos~5.5kAutomated safety check: PassMIT
Dotnet Patternsaffaan-m/ECC276k1 repos~2.3kAutomated safety check: PassMIT

Similar skills

  • Detecting Anomalous Authentication Patterns

    mukul975/Anthropic-Cybersecurity-Skills

    Detects anomalous authentication patterns using UEBA analytics, statistical baselines, and machine learning models to identify impossible travel, credential stuffing, brute force, password spraying…

    34k GitHub stars~7.5k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Detecting Beaconing Patterns With Zeek

    mukul975/Anthropic-Cybersecurity-Skills

    Performs statistical analysis of Zeek conn.log connection intervals to detect C2 beaconing patterns.

    34k GitHub stars~662 tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • Golang Patterns

    affaan-m/ECC

    Go-specific design patterns and best practices including functional options, small interfaces, dependency injection, concurrency patterns, error handling, and package organization.

    276k GitHub stars~1.1k tokensUpdated 4 days ago
    DevelopmentAuto-check passed
  • JetBrains Exposed ORM patterns including DSL queries, DAO pattern, transactions, HikariCP connection pooling, Flyway migrations, and repository pattern.

    276k GitHub starsUsed in 4 repos~5.5k tokens
    DatabasesAuto-check passed
  • Dotnet Patterns

    affaan-m/ECC

    Idiomatic C and .NET patterns, conventions, dependency injection, async/await, and best practices for building robust, maintainable .NET applications.

    276k GitHub starsUsed in 1 repo~2.3k tokens
    DevelopmentAuto-check passed
  • Fastapi Patterns

    affaan-m/ECC

    FastAPI patterns for async APIs, dependency injection, Pydantic request and response models, OpenAPI docs, tests, security, and production readiness.

    276k GitHub stars~2.3k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed

More from Archive228/loopkit

All 43 skills in this repo
  • Hitl Escalate

    Archive228/loopkit

    Escalate blocked runs to a human via configured channel or fallback to BLOCKED.md and exit the loop.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Structured Output

    Archive228/loopkit

    Get JSON out of the model reliably. An agent skill from Archive228/loopkit.

    755 GitHub stars~830 tokensUpdated 2 mo ago
    Auto-check passed
  • Using Loopkit

    Archive228/loopkit

    A skill your agent uses when starting any conversation in a loopkit-enabled project - establishes how to find and use loopkit's 49 skills, requiring skill invocation before ANY response including…

    755 GitHub stars~1.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Active Memory Reminder

    Archive228/loopkit

    Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable).

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Eval Harness

    Archive228/loopkit

    Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

    755 GitHub stars~876 tokensUpdated 2 mo ago
    Auto-check passed
  • Feature List JSON

    Archive228/loopkit

    Enumerate every end-to-end feature as strict JSON entries with passes:false, editable-passes-only discipline, and priority order.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Self Eval Bias

What does Self Eval Bias do?

Detect and interrupt the pattern where an agent confidently praises work it just produced instead of reviewing it critically. Self Eval Bias is an agent skill from Archive228/loopkit. Detect and interrupt the pattern where an agent confidently praises work it just produced instead of reviewing it critically.

How do I install Self Eval Bias in Claude Code?

Run `npx skills add Archive228/loopkit --skill self-eval-bias -a claude-code`. Or copy the skill folder (skills/self-eval-bias in Archive228/loopkit) into .claude/skills/self-eval-bias in your project. Claude Code loads it when a task matches its description.

How do I install Self Eval Bias in Codex?

Run `npx skills add Archive228/loopkit --skill self-eval-bias -a codex`. Or copy the skill folder (skills/self-eval-bias in Archive228/loopkit) into .agents/skills/self-eval-bias in your project. Codex loads it when a task matches its description.

Can I use Self Eval Bias in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Archive228/loopkit --skill self-eval-bias -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/self-eval-bias, .gemini/skills/self-eval-bias, .github/skills/self-eval-bias and .opencode/skills/self-eval-bias in your project.

What does Self Eval Bias need to run?

SKILL.md names no scripts, command-line tools or credentials: Self Eval Bias is instructions for the agent only.

Does Self Eval Bias access the network?

SKILL.md names 1 domain. As links in the text: blog.anthropic.com. This is read from the text; nothing was executed.

Is Self Eval Bias safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Self Eval Bias use?

Self Eval Bias is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Self Eval Bias use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Self Eval Bias?

Skills that share tags, products or a category with Self Eval Bias: Detecting Anomalous Authentication Patterns (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Detecting Beaconing Patterns With Zeek (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Golang Patterns (affaan-m/ECC, 276k stars) and Kotlin Exposed Patterns (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Self Eval Bias?

Archive228 (a GitHub user) maintains it in Archive228/loopkit, which has 755 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on July 14, 2026.

Source: Archive228/loopkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.