Scoring
xiaolai/nlpm
100-point NL artifact rubric: penalty tables per artifact type, calibration cases.
Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.
$ npx skills add Archive228/loopkit --skill evaluator-calibration -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Archive228/loopkit evaluator-calibration --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/evaluator-calibration .claude/skills/evaluator-calibration && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluator-calibration" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/evaluator-calibration into .claude/skills/evaluator-calibration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluator-calibration", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Archive228/loopkit/tree/main/skills/evaluator-calibrationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Archive228/loopkit --skill evaluator-calibration -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Archive228/loopkit evaluator-calibration --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/evaluator-calibration .agents/skills/evaluator-calibration && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluator-calibration" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/evaluator-calibration into .agents/skills/evaluator-calibration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluator-calibration", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Archive228/loopkit --skill evaluator-calibration -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Archive228/loopkit evaluator-calibration --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/evaluator-calibration .cursor/skills/evaluator-calibration && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluator-calibration" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/evaluator-calibration into .cursor/skills/evaluator-calibration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluator-calibration", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Archive228/loopkit.git --path skills/evaluator-calibration--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Archive228/loopkit --skill evaluator-calibration -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Archive228/loopkit evaluator-calibration --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/evaluator-calibration .gemini/skills/evaluator-calibration && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluator-calibration" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/evaluator-calibration into .gemini/skills/evaluator-calibration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluator-calibration", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Archive228/loopkit evaluator-calibrationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Archive228/loopkit --skill evaluator-calibration -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/evaluator-calibration .github/skills/evaluator-calibration && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluator-calibration" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/evaluator-calibration into .github/skills/evaluator-calibration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluator-calibration", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Archive228/loopkit --skill evaluator-calibration -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Archive228/loopkit evaluator-calibration --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Archive228/loopkit.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/evaluator-calibration .opencode/skills/evaluator-calibration && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluator-calibration" agent skill from https://github.com/Archive228/loopkit/tree/main/skills/evaluator-calibration into .opencode/skills/evaluator-calibration/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluator-calibration", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluator-calibrationCalibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.
Evaluator Calibration is an agent skill from Archive228/loopkit. Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.
Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Education, covering Performance reviews, Quizzes and assessments and Prompt engineering. The repository describes itself as: 33 battle-tested skills + minimal .claude harness for any coding agent (Claude Code, Cursor, Codex, Gemini CLI). The licence is MIT.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 5ae033e. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluator Calibration loads about 1.1k tokens when it runs. Until then it costs about 38 tokens; SKILL.md has 566 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Archive228/loopkit at commit 5ae033e, republished under its MIT licence (© Archive228). 566 words, ~1,052 tokens.
.claude/skills/evaluator-calibration/SKILL.md (or your agent's skills folder).An evaluator agent that reads the generator's reasoning drifts lenient. The generator explains why the code is good; the evaluator, priming on that prose, starts nodding along. By sprint 8 the "skeptical critic" is a rubber stamp. Prithvi flagged this in the March 2026 planner/generator/evaluator writeup — evaluator leniency is the failure mode of the three-agent harness.
The fix is not "tell the evaluator to be stricter." That works for one iteration. The fix is anchoring the rubric with concrete pass/fail examples the evaluator re-reads every invocation, and re-prompting from scratch on a fixed cadence so drift can't accumulate.
Write the rubric as a scored checklist, not prose. Each criterion gets a name, a one-line definition, and a binary or 1-3 score. Prose rubrics ("evaluate whether the code is well-designed") drift; checklists don't.
Anchor every criterion with 2 concrete examples — one pass, one fail. Real examples from prior runs, not invented ones. The evaluator reads these every invocation. This is the calibration; without it you're just prompting hope.
Forbid reading the generator's reasoning before scoring. The evaluator sees the artifact (code, diff, output) and the rubric. It does not see the generator's "here's why this is good" prose. Score first, then optionally read the reasoning to write the critique.
Require the evaluator to quote the artifact in every verdict. "Fails criterion 3 because <quoted line>" — not "fails criterion 3." Quoting forces grounding and makes the verdict auditable.
Re-prompt from scratch every N iterations. Empirically N=5 works. Kill the evaluator's context, reload the system prompt + rubric + examples fresh. Do not compact; compaction preserves the drift.
Log verdict distributions. Track pass rate per criterion per sprint. A criterion that goes from 40% pass to 90% pass without a spec change is drift, not improvement.
Spot-check with a held-out fail. Every ~10 sprints, feed the evaluator an artifact from your example set that you know fails. If it passes, the calibration has decayed — regenerate the example set from recent real runs.
Single-shot grading with a fresh context every call — there's no drift to prevent, and the examples are overhead. Also skip for tasks under ~1 hour where the evaluator only runs 2-3 times.
© Archive228, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/evaluator-calibration of Archive228/loopkit.
Open the folder on GitHubat commit 5ae033e
Evaluator Calibration next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluator Calibration this skillArchive228/loopkit | 755 | — | ~1.1k | Automated safety check: Pass | MIT | |
| Scoringxiaolai/nlpm | 150 | — | ~5.3k | Automated safety check: Pass | ISC | |
| Design AI BenchmarkingAperivue/medsci-skills | 333 | — | ~2.4k | Automated safety check: Pass | MIT | |
| Interview System Designerborghei/Claude-Skills | 891 | — | ~1.7k | Automated safety check: Pass | MIT | |
| Agent Prompt Quality Barmastra-ai/mastra | 29k | — | ~2k | Automated safety check: Pass | Custom licence | |
| Advanced Evaluationguanyang/open-agent-hub | 977 | 2 repos | ~4.2k | Automated safety check: Pass | MIT |
xiaolai/nlpm
100-point NL artifact rubric: penalty tables per artifact type, calibration cases.
Aperivue/medsci-skills
A skill your agent uses when designing a study that benchmarks AI systems against a human-expert panel, before data collection.
borghei/Claude-Skills
Design calibrated interview loops, competency-based question banks, and hiring calibration.
mastra-ai/mastra
Universal quality bar and final audit rubric for any agent system prompt.
guanyang/open-agent-hub
This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated…
Mathews-Tom/armory
LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.
Archive228/loopkit
Escalate blocked runs to a human via configured channel or fallback to BLOCKED.md and exit the loop.
Archive228/loopkit
Get JSON out of the model reliably. An agent skill from Archive228/loopkit.
Archive228/loopkit
A skill your agent uses when starting any conversation in a loopkit-enabled project - establishes how to find and use loopkit's 49 skills, requiring skill invocation before ANY response including…
Archive228/loopkit
Before compaction Loopkit extracts decisions into claude-decisions.json (machine-readable).
Archive228/loopkit
Build a repeatable eval loop that grades agent output with an LLM judge, so prompt/skill changes get scored against a baseline instead of eyeballed.
Archive228/loopkit
Enumerate every end-to-end feature as strict JSON entries with passes:false, editable-passes-only discipline, and priority order.
Categories
Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs. Evaluator Calibration is an agent skill from Archive228/loopkit. Calibrate a reviewer persona with few-shot rubric examples so skepticism stays consistent and doesn't drift lenient over long runs.
Evaluator Calibration fits situations like: tasks that involve Performance reviews; tasks that involve Quizzes and assessments; tasks that involve Prompt engineering.
Run `npx skills add Archive228/loopkit --skill evaluator-calibration -a claude-code`. Or copy the skill folder (skills/evaluator-calibration in Archive228/loopkit) into .claude/skills/evaluator-calibration in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Archive228/loopkit --skill evaluator-calibration -a codex`. Or copy the skill folder (skills/evaluator-calibration in Archive228/loopkit) into .agents/skills/evaluator-calibration in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Archive228/loopkit --skill evaluator-calibration -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluator-calibration, .gemini/skills/evaluator-calibration, .github/skills/evaluator-calibration and .opencode/skills/evaluator-calibration in your project.
SKILL.md names no scripts, command-line tools or credentials: Evaluator Calibration is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluator Calibration is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluator Calibration: Scoring (xiaolai/nlpm, 150 stars), Design AI Benchmarking (Aperivue/medsci-skills, 333 stars), Interview System Designer (borghei/Claude-Skills, 891 stars) and Agent Prompt Quality Bar (mastra-ai/mastra, 29k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Archive228 (a GitHub user) maintains it in Archive228/loopkit, which has 755 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on July 14, 2026.
Source: Archive228/loopkit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.