Agent skill

Evaluator Calibration

by Archive228 in Archive228/loopkit

Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.

MITAuto-check passedEducation

Install Evaluator Calibration

skills CLI
$ npx skills add Archive228/loopkit --skill evaluator-calibration -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Archive228/loopkit evaluator-calibration --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluator-calibration .claude/skills/evaluator-calibration && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluator-calibration
GitHub stars
755
Token cost
~1.1k tokens
SKILL.md length
566 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.

  • Works in 7 steps: Write the rubric as a scored checklist,… → Anchor every criterion with 2 concrete… → Forbid reading the generator's reasoning… → …
  • Tasks that involve Performance reviews
  • SKILL.md covers When to apply, Procedure, Anti-patterns and Related, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evaluator Calibration is an agent skill from Archive228/loopkit. Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Performance reviews, Quizzes and assessments and Prompt engineering. The repository describes itself as: 33 battle-tested skills + minimal .claude harness for any coding agent (Claude Code, Cursor, Codex, Gemini CLI). The licence is MIT.

When your agent uses it

  • Tasks that involve Performance reviews
  • Tasks that involve Quizzes and assessments
  • Tasks that involve Prompt engineering

Example prompts

  • “/evaluator-calibration”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Write the rubric as a scored checklist, not prose. Each criterion gets a name, a one-line definition, and a binary or 1-3 score. Prose…
  2. Anchor every criterion with 2 concrete examples — one pass, one fail. Real examples from prior runs, not invented ones. The evaluator…
  3. Forbid reading the generator's reasoning before scoring. The evaluator sees the artifact (code, diff, output) and the rubric. It does not…
  4. Require the evaluator to quote the artifact in every verdict. "Fails criterion 3 because " — not "fails criterion 3." Quoting forces…
  5. Re-prompt from scratch every N iterations. Empirically N=5 works. Kill the evaluator's context, reload the system prompt + rubric +…
  6. Log verdict distributions. Track pass rate per criterion per sprint. A criterion that goes from 40% pass to 90% pass without a spec change…
  7. Spot-check with a held-out fail. Every ~10 sprints, feed the evaluator an artifact from your example set that you know fails. If it…

What it can do on your machine

Read from SKILL.md and the folder at commit 5ae033e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluator Calibration loads about 1.1k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 566 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~38
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Archive228/loopkit at commit 5ae033e, republished under its MIT licence (© Archive228). 566 words, ~1,052 tokens.

Download SKILL.mdSave it as .claude/skills/evaluator-calibration/SKILL.md (or your agent's skills folder).
name
evaluator-calibration
description
Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.
when_to_use
standing up an evaluator/critic agent for a multi-agent harness, noticing evaluator scores drift upward across many iterations, grading skill/PR/diff output…

Evaluator Calibration

An evaluator agent that reads the generator's reasoning drifts lenient. The generator explains why the code is good; the evaluator, priming on that prose, starts nodding along. By sprint 8 the "skeptical critic" is a rubber stamp. Prithvi flagged this in the March 2026 planner/generator/evaluator writeup — evaluator leniency is the failure mode of the three-agent harness.

The fix is not "tell the evaluator to be stricter." That works for one iteration. The fix is anchoring the rubric with concrete pass/fail examples the evaluator re-reads every invocation, and re-prompting from scratch on a fixed cadence so drift can't accumulate.

When to apply

  • You're building a critic/evaluator/judge agent in a multi-agent loop.
  • You're using an LLM as a grader for skills, PRs, diffs, or agent output.
  • You've noticed pass rates creeping up while output quality hasn't changed — or worse, dropped.
  • You want two runs of the same evaluator on the same artifact to return the same verdict.

Procedure

  1. Write the rubric as a scored checklist, not prose. Each criterion gets a name, a one-line definition, and a binary or 1-3 score. Prose rubrics ("evaluate whether the code is well-designed") drift; checklists don't.

  2. Anchor every criterion with 2 concrete examples — one pass, one fail. Real examples from prior runs, not invented ones. The evaluator reads these every invocation. This is the calibration; without it you're just prompting hope.

  3. Forbid reading the generator's reasoning before scoring. The evaluator sees the artifact (code, diff, output) and the rubric. It does not see the generator's "here's why this is good" prose. Score first, then optionally read the reasoning to write the critique.

  4. Require the evaluator to quote the artifact in every verdict. "Fails criterion 3 because <quoted line>" — not "fails criterion 3." Quoting forces grounding and makes the verdict auditable.

  5. Re-prompt from scratch every N iterations. Empirically N=5 works. Kill the evaluator's context, reload the system prompt + rubric + examples fresh. Do not compact; compaction preserves the drift.

  6. Log verdict distributions. Track pass rate per criterion per sprint. A criterion that goes from 40% pass to 90% pass without a spec change is drift, not improvement.

  7. Spot-check with a held-out fail. Every ~10 sprints, feed the evaluator an artifact from your example set that you know fails. If it passes, the calibration has decayed — regenerate the example set from recent real runs.

Show full SKILL.md (172 more words)Show less

Anti-patterns

  • "Be skeptical" in the system prompt with no examples. Words don't calibrate. Examples calibrate.
  • Letting the evaluator read the planner's plan. Same drift mechanism as reading the generator's reasoning — priming on intent softens the critique.
  • One shared context across many gradings. Each grading should start from a clean rubric read. Batch-grading in one context is where leniency compounds fastest.
  • Invented examples. Fake pass/fail examples don't anchor to the artifact distribution the evaluator actually sees. Pull from real runs.
  • Scoring on a 1-10 scale. The evaluator will cluster at 7. Use binary or 1-3.
  • [[adversarial-verify]] — the single-shot form; evaluator-calibration is the standing-agent form.
  • [[shift-notes]] — evaluator verdicts belong in the ledger so drift is visible session-over-session.
  • [[broken-window-check]] — a mechanical version of "don't trust the last verdict"; pairs well when the evaluator is the thing being distrusted.

When NOT to apply

Single-shot grading with a fresh context every call — there's no drift to prevent, and the examples are overhead. Also skip for tasks under ~1 hour where the evaluator only runs 2-3 times.

© Archive228, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/evaluator-calibration of Archive228/loopkit.

Open the folder on GitHubat commit 5ae033e

Compare with similar skills

Evaluator Calibration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluator Calibration compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluator Calibration this skillArchive228/loopkit755—~1.1kAutomated safety check: PassMIT
Scoringxiaolai/nlpm150—~5.3kAutomated safety check: PassISC
Design AI BenchmarkingAperivue/medsci-skills333—~2.4kAutomated safety check: PassMIT
Interview System Designerborghei/Claude-Skills891—~1.7kAutomated safety check: PassMIT
Agent Prompt Quality Barmastra-ai/mastra29k—~2kAutomated safety check: PassCustom licence
Advanced Evaluationguanyang/open-agent-hub9772 repos~4.2kAutomated safety check: PassMIT

Similar skills

  • Scoring

    xiaolai/nlpm

    100-point NL artifact rubric: penalty tables per artifact type, calibration cases.

    150 GitHub stars~5.3k tokensUpdated today
    EducationAuto-check passed
  • Design AI Benchmarking

    Aperivue/medsci-skills

    A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.

    333 GitHub stars~2.4k tokensUpdated 4 days ago
    EducationAuto-check passed
  • Interview System Designer

    borghei/Claude-Skills

    Design calibrated interview loops, competency-based question banks, and hiring calibration.

    891 GitHub stars~1.7k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Agent Prompt Quality Bar

    mastra-ai/mastra

    Universal quality bar and final audit rubric for any agent system prompt.

    29k GitHub stars~2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Advanced Evaluation

    guanyang/open-agent-hub

    This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    AI & LLM EngineeringAuto-check passed
  • Prompt Lab

    Mathews-Tom/armory

    LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.

    329 GitHub stars~2.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed

More from Archive228/loopkit

All 43 skills in this repo
  • Hitl Escalate

    Archive228/loopkit

    Escalate blocked runs to a human via configured channel or fallback to BLOCKED.md and exit the loop.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Structured Output

    Archive228/loopkit

    Get JSON out of the model reliably. An agent skill from Archive228/loopkit.

    755 GitHub stars~830 tokensUpdated 2 mo ago
    Auto-check passed
  • Using Loopkit

    Archive228/loopkit

    A skill your agent uses when starting any conversation in a loopkit-enabled project - establishes how to find and use loopkit's 49 skills, requiring skill invocation before ANY response including…

    755 GitHub stars~1.4k tokensUpdated 2 mo ago
    Auto-check passed
  • Active Memory Reminder

    Archive228/loopkit

    Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable).

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Eval Harness

    Archive228/loopkit

    Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.

    755 GitHub stars~876 tokensUpdated 2 mo ago
    Auto-check passed
  • Feature List JSON

    Archive228/loopkit

    Enumerate every end-to-end feature as strict JSON entries with passes:false, editable-passes-only discipline, and priority order.

    755 GitHub stars~1.2k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Evaluator Calibration

What does Evaluator Calibration do?

Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs. Evaluator Calibration is an agent skill from Archive228/loopkit. Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.

When should I use Evaluator Calibration?

Evaluator Calibration fits situations like: tasks that involve Performance reviews; tasks that involve Quizzes and assessments; tasks that involve Prompt engineering.

How do I install Evaluator Calibration in Claude Code?

Run `npx skills add Archive228/loopkit --skill evaluator-calibration -a claude-code`. Or copy the skill folder (skills/evaluator-calibration in Archive228/loopkit) into .claude/skills/evaluator-calibration in your project. Claude Code loads it when a task matches its description.

How do I install Evaluator Calibration in Codex?

Run `npx skills add Archive228/loopkit --skill evaluator-calibration -a codex`. Or copy the skill folder (skills/evaluator-calibration in Archive228/loopkit) into .agents/skills/evaluator-calibration in your project. Codex loads it when a task matches its description.

Can I use Evaluator Calibration in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Archive228/loopkit --skill evaluator-calibration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluator-calibration, .gemini/skills/evaluator-calibration, .github/skills/evaluator-calibration and .opencode/skills/evaluator-calibration in your project.

What does Evaluator Calibration need to run?

SKILL.md names no scripts, command-line tools or credentials: Evaluator Calibration is instructions for the agent only.

Does Evaluator Calibration access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluator Calibration safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluator Calibration use?

Evaluator Calibration is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluator Calibration use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluator Calibration?

Skills that share tags, products or a category with Evaluator Calibration: Scoring (xiaolai/nlpm, 150 stars), Design AI Benchmarking (Aperivue/medsci-skills, 333 stars), Interview System Designer (borghei/Claude-Skills, 891 stars) and Agent Prompt Quality Bar (mastra-ai/mastra, 29k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluator Calibration?

Archive228 (a GitHub user) maintains it in Archive228/loopkit, which has 755 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on July 14, 2026.

Source: Archive228/loopkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.