Darwin Skill Optimizer
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
Review whether a repository's coding-agent harness can reliably carry work from intent through controlled execution, verification, delivery, and learning.
$ npx skills add majiayu000/spellbook --skill review-agent-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install majiayu000/spellbook review-agent-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/review-agent-harness .claude/skills/review-agent-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "review-agent-harness" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/review-agent-harness into .claude/skills/review-agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-agent-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/majiayu000/spellbook/tree/main/skills/review-agent-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add majiayu000/spellbook --skill review-agent-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install majiayu000/spellbook review-agent-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/review-agent-harness .agents/skills/review-agent-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "review-agent-harness" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/review-agent-harness into .agents/skills/review-agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-agent-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add majiayu000/spellbook --skill review-agent-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install majiayu000/spellbook review-agent-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/review-agent-harness .cursor/skills/review-agent-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "review-agent-harness" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/review-agent-harness into .cursor/skills/review-agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-agent-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/majiayu000/spellbook.git --path skills/review-agent-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add majiayu000/spellbook --skill review-agent-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install majiayu000/spellbook review-agent-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/review-agent-harness .gemini/skills/review-agent-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "review-agent-harness" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/review-agent-harness into .gemini/skills/review-agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-agent-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install majiayu000/spellbook review-agent-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add majiayu000/spellbook --skill review-agent-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/review-agent-harness .github/skills/review-agent-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "review-agent-harness" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/review-agent-harness into .github/skills/review-agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-agent-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add majiayu000/spellbook --skill review-agent-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install majiayu000/spellbook review-agent-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/majiayu000/spellbook.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/review-agent-harness .opencode/skills/review-agent-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "review-agent-harness" agent skill from https://github.com/majiayu000/spellbook/tree/main/skills/review-agent-harness into .opencode/skills/review-agent-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "review-agent-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
review-agent-harnessReview whether a repository's coding-agent harness can reliably carry work from intent through controlled execution, verification, delivery, and learning.
Review Agent Harness is an agent skill from majiayu000/spellbook. Review whether a repository's coding-agent harness can reliably carry work from intent through controlled execution, verification, delivery, and learning. Use when asked to assess agent readiness, repeated agent failures, Rules/Skills/Hooks/Memory effectiveness, missing validation or recovery loops, or whether a harness repair improved later outcomes. Do not use for code-only audits, AGENTS-only audits, individual skill reliability reviews, or executing the task itself.
Its SKILL.md is about 3.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 17 other files, including scripts and reference files (for example `agents/openai.yaml`, `evals/evals.json` and `references/finding-contract.md`).
It sits in Agent Workflows. It works with Git. The repository describes itself as: Cross-runtime skills for Claude Code, Codex, and multi-agent workflows. The licence is MIT.
5 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit ed52af7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
python3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Review Agent Harness loads about 3.5k tokens when it runs, and up to ~8.4k if it reads all its reference files. Until then it costs about 124 tokens; SKILL.md has 1,633 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from majiayu000/spellbook at commit ed52af7, republished under its MIT licence (© majiayu000). 1,633 words, ~3,515 tokens.
.claude/skills/review-agent-harness/SKILL.md (or your agent's skills folder). This skill also uses 13 other files; get the full folder from GitHub.Review the operating system around coding agents, not only its files. Separate declared assets, reachable routes, observed task use, and later outcomes.
Resolve the directory containing this SKILL.md before running its scripts.
Select one mode:
static: inspect the target repository only. Use by default.episode: add explicitly authorized Codex or Claude Code JSONL sources.longitudinal: compare a validated report with the existing ledger.Default to inline, read-only output. Write a durable report under the target only when the user explicitly requests an artifact or historical tracking. Never discover user-home Sessions, read Memory bodies, or inspect another provider merely because its files are available.
Use adjacent skills instead when their narrower owner is sufficient:
codebase-audit for code defects and architecture health;repo-agent-context-audit for AGENTS, Skills, and Specs alone;skill-lifeguard for one Skill's reliable contract;flowguard for running a long task;review-gate before landing an agent-generated diff.Record target, mode, provider, locale, decision, acceptance boundary, output
mode, included sources, excluded sources, and unavailable evidence. Treat a
missing required source as unobserved; do not substitute a broader directory,
another provider, or remembered results.
Resolve the target before interpreting assets or assigning scores. The collector
classifies it as exact_git_root, inside_git_worktree,
contains_nested_git_root, or non_git_directory. If the supplied directory
contains a nested Git root, stop and retarget that exact repository; do not
score the parent as though it were the project. For a Git target, the collector
uses Git's tracked and untracked inventory and excludes ignored worktrees and
prior review output from repository evidence.
Run static collection from the installed Skill directory:
python3 scripts/collect_evidence.py \
--target /absolute/target \
--mode static \
--locale zh-CN \
--decision "assess agent-harness readiness" \
--acceptance-boundary "resolve all five dimensions" \
--output-mode inline \
--output /temporary/evidence.jsonFor Session-informed review, require the user to authorize exact files or an exact root. Use one provider per evidence envelope:
python3 scripts/collect_evidence.py \
--target /absolute/target \
--mode episode \
--provider codex \
--session-file /explicit/session.jsonl \
--locale zh-CN \
--decision "explain the observed verification gap" \
--acceptance-boundary "separate configured and exercised routes" \
--output-mode inline \
--output /temporary/evidence.jsonUse --session-root only when that exact recursive scope was authorized. Add
--include-request-summaries only when sanitized request summaries are needed
for the decision. Read Privacy Boundary and
Session Adapters before Session-informed work.
Omit --output to stream evidence to stdout. Inline means no target writes;
environment-owned scratch remains allowed. validate_findings.py --input -
accepts findings JSON from stdin when the caller already has a stream.
Checkpoint: collection must return agent-harness-evidence; every stage must
be available, constrained, not_authorized, not_applicable, unavailable,
or unobserved. A depth-limited scan is constrained, never silently complete.
Stop on malformed output or an unexplained missing stage.
Copy the collector-owned scope.target_id and complete scope.snapshot
(baseline, target_relation, and id) into the findings document. Never
author these values manually. The renderer and ledger updater recompute the
binding from --target and reject a different local directory or any target
state that changed after collection. A previous report, ledger row, branch
name, or remembered result is a historical lead only. Recheck any retained
claim against the frozen current snapshot and label genuinely historical
evidence as such.
Keep the passes logically independent even when one agent runs them in sequence:
Do not launch parallel agents by default. If the user explicitly requests
threads, use threads with read-only lanes and bounded evidence packets. A
specialist proposes candidates; it does not assign final severity or claim
effectiveness.
Read Review Model before classifying the five
dimensions. Use present -> reachable -> exercised -> outcome_supported only
when each stronger state has direct evidence.
Resolve all 15 stable checks, three per dimension. Assign a score to each
dimension only after resolving its checks. The score is an evidence-bounded
summary, not a finding: present caps a dimension at 74, reachable at 84,
exercised at 94, and outcome_supported at 100; missing or unobserved
caps it at 59. Use the lowest applicable check ceiling and retain a short score
rationale. Do not compute an overall score.
Retain each distinct eligible candidate. Merge only when consequence, root cause, owner, and verifier are the same. The lead alone assigns severity, confidence, primary dimension, verification state, and priority.
Read Finding Contract. Every finding needs:
Counts, file absence without a requirement, similarity, theoretical risk,
score, or unavailable evidence never create a finding. Critical and High
findings require an adversarial check; retain an unavailable check as
unverified instead of presenting it as confirmed.
Record each executed verifier in verification_runs with a stable id, purpose,
result, exit code, final-state flag, and bounded summary. A confirmed Critical
or High finding must cite a final-state candidate_refutation or
targeted_reproduction run that supports the claim. Inspect aggregate exit
semantics: a child syntax error or failed subcheck paired with aggregate exit 0
is evidence of a false-green verifier, not a passing check.
Author one agent-harness-findings JSON object in environment-owned scratch
space, then validate it:
python3 scripts/validate_findings.py --input /temporary/findings.json --strict --jsonFix the findings data, not the validator. Stop if validation does not pass.
For inline review, render the overview, frozen snapshot, five-dimension scorecard, all 15 checks, structured verification runs, findings, evidence boundary, and at most three priority moves in the response. Do not write to the target.
When durable output is explicitly requested, render atomically:
python3 scripts/render_report.py \
--findings /temporary/findings.json \
--evidence /temporary/evidence.json \
--target /absolute/target \
--out /absolute/target/.agent-harness-review \
--jsonThe renderer refuses to replace an existing run and writes only validated
findings.json, privacy-safe evidence.json, and derived report.md.
For longitudinal mode, update the ledger after a fresh review:
python3 scripts/update_ledger.py \
--findings /temporary/findings.json \
--target /absolute/target \
--ledger /absolute/target/.agent-harness-review/ledger.json \
--jsonAn absent prior finding remains open with recheck_required until a targeted
spot-check produces an agent-harness-resolution-confirmations document. Each
confirmation must retain the finding id, verifier, and one bounded
evidence_ref. Pass it with
--resolution-confirmations /temporary/confirmations.json. Never resolve from
finder absence or an id-only assertion.
Read Repair Loop for follow-up. This review does
not authorize fixes. Route a selected finding to its owner in a separate task,
run its verifier on the final state, and update repair_state only.
Do not upgrade learning-retention from same-window repair evidence. For a
tool-backed route, collect the later Episode with --mechanism-category edit
or validation, --episode-role later, and an explicit --comparison-basis;
collect the baseline with the same basis and --episode-role baseline. The
adapter count shows only that coarse mechanism was exercised. Use bounded file
or policy evidence to map the category to the repaired route. Separately require
target-owned command or artifact references showing the result improved and
guardrails still passed. Adapter counts, collection time, or a request summary
alone never prove later effect.
When claiming outcome_supported, pass both collector envelopes to every gate:
validate_findings.py --evidence baseline.json --evidence later.json, and use
the same repeated --evidence flags with render_report.py or
update_ledger.py. The claim is rejected without exactly one bound baseline
and one bound later envelope.
Direct actions:
Escalate before:
Evidence-backed pushback: reject a requested score or conclusion when the target is not the exact repository, the relevant evidence is unavailable, an aggregate verifier hides a failed subcheck, or a historical claim was not rechecked on the frozen snapshot. State the concrete boundary and the smallest next command that could resolve it.
Feedback loop: replay the closest case in evals/evals.json after a miss or
false positive, add one focused regression test, and patch the smallest durable
owner in the collector, validator, renderer, or written contract.
unobserved and continue only with static mechanisms./absolute/target/.agent-harness-review, and the ledger updater accepts
only its ledger.json below that directory.Finish only when:
unobserved / not_applicable boundary;validate_findings.py --strict;Use evals/evals.json and the repository tests as the replay surface. Patch the
smallest durable owner when the Skill over-triggers, misses a primary request,
accepts private data, treats configuration as use, resolves from absence, or
claims later effectiveness from same-window checks.
scripts/collect_evidence.py: static collector and Session adapter facade.scripts/validate_findings.py: findings, evidence-state, and privacy gate.scripts/render_report.py: atomic durable Markdown renderer.scripts/update_ledger.py: conservative longitudinal ledger.references/review-model.md: dimensions and evidence semantics.references/finding-contract.md: authoring and reconciliation contract.references/privacy-boundary.md: authorization and redaction rules.references/session-adapters.md: Codex and Claude Code input boundaries.references/repair-loop.md: repair progress versus later effectiveness.© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 13 other files (scripts, references) in skills/review-agent-harness of majiayu000/spellbook.
Open the folder on GitHubat commit ed52af7
Review Agent Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Review Agent Harness this skillmajiayu000/spellbook | 287 | — | ~3.5k | Automated safety check: Pass | MIT | |
| Darwin Skill Optimizeralchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| CodeGraph Agent Evalcolbymchenry/codegraph | 74k | — | ~950 | Automated safety check: Pass | MIT | |
| Neat-Freak Knowledge CloseoutKKKKhazix/khazix-skills | 21k | — | ~1.9k | Automated safety check: Pass | MIT | |
| O2 Review Loopopenobserve/openobserve | 22k | — | ~3.7k | Automated safety check: Pass | AGPL-3.0 | |
| Beads Task Memorygastownhall/beads | 28k | — | ~1.2k | Automated safety check: Pass | MIT |
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
colbymchenry/codegraph
Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.
KKKKhazix/khazix-skills
Brings project docs, agent rule files, authorized memory and leftover workspace files back in line with what the code and runtime actually do at the end of a work session.
openobserve/openobserve
Splits a change into planner, coder and independent reviewer roles: you confirm a spec, a subagent implements it, and a separate reviewer checks each round's local WIP commit.
gastownhall/beads
Tracks multi-session work with dependencies in the bd issue tracker so the agent can find ready tasks and recover its context after conversation compaction.
slopus/happy
Searches past Claude Code, Codex and Cursor sessions and summarizes what was worked on, tried or decided, using extraction scripts instead of reading raw logs.
majiayu000/spellbook
Audits and repairs how coding-agent Skills are owned, copied and exposed across runtimes, from canonical sources to quarantine and retirement.
majiayu000/spellbook
Scans a repository for real evidence and proposes, or on request writes, a small stack of root and scoped AGENTS.md files with validation commands and generated-file boundaries.
majiayu000/spellbook
Plans, produces or diagnoses evidence-backed product demo videos: script, capture plan, pacing checks and verified final media built on real product behavior.
majiayu000/spellbook
Single entry point that routes long or ambiguous agent tasks, checks live state, bounds autonomous loops and leaves a resumable handoff.
majiayu000/spellbook
Scans a repository, its lockfiles and node_modules for known malicious npm package versions and install-time indicators, using a read-only Python scanner.
majiayu000/spellbook
Product management helpers: a RICE scoring script, an interview transcript analyzer and PRD templates for prioritizing features, synthesizing research and writing requirements.
Works with
Categories
Review whether a repository's coding-agent harness can reliably carry work from intent through controlled execution, verification, delivery, and learning. Review Agent Harness is an agent skill from majiayu000/spellbook. Review whether a repository's coding-agent harness can reliably carry work from intent through controlled execution, verification, delivery, and learning.
Review Agent Harness fits situations like: asked to assess agent readiness; repeated agent failures; rules/Skills/Hooks/Memory effectiveness; missing validation.
Run `npx skills add majiayu000/spellbook --skill review-agent-harness -a claude-code`. Or copy the skill folder (skills/review-agent-harness in majiayu000/spellbook) into .claude/skills/review-agent-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add majiayu000/spellbook --skill review-agent-harness -a codex`. Or copy the skill folder (skills/review-agent-harness in majiayu000/spellbook) into .agents/skills/review-agent-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/spellbook --skill review-agent-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/review-agent-harness, .gemini/skills/review-agent-harness, .github/skills/review-agent-harness and .opencode/skills/review-agent-harness in your project.
Going by SKILL.md and its folder, Review Agent Harness needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Review Agent Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.5k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Review Agent Harness: Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), CodeGraph Agent Eval (colbymchenry/codegraph, 74k stars), Neat-Freak Knowledge Closeout (KKKKhazix/khazix-skills, 21k stars) and O2 Review Loop (openobserve/openobserve, 22k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
majiayu000 (a GitHub user) maintains it in majiayu000/spellbook, which has 287 GitHub stars. The repository holds 97 skills in this directory. The repository was last updated on October 8, 2026.
Source: majiayu000/spellbook on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.