Agent skill

Coder Self Review

by ntorga in ntorga/agent-starter-kit

Deterministic self-evaluation rubric for Coder — scored every run using the GRASP framework.

MITAuto-check passedEducation

Install Coder Self Review

skills CLI
$ npx skills add ntorga/agent-starter-kit --skill coder-self-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ntorga/agent-starter-kit coder-self-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ntorga/agent-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/coder-self-review .claude/skills/coder-self-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
coder-self-review
GitHub stars
146
Token cost
~2.4k tokens
SKILL.md length
1,347 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Deterministic self-evaluation rubric for Coder — scored every run using the GRASP framework.

  • Works in 5 steps: Gather evidence. Honest self-review… → Score each criterion. Read the GRASP… → Output the Scorecard. Fill in the… → …
  • Tasks that involve Quizzes and assessments
  • SKILL.md covers Purpose, Procedure, Scorecard and GRASP Rubric, plus 1 more section
  • Calls rg

What it does

Coder Self Review is an agent skill from ntorga/agent-starter-kit. Deterministic self-evaluation rubric for Coder — scored every run using the GRASP framework.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments. The repository describes itself as: The scaffold for your multi-model, personalized Natural Language AI Harness (NLAH) . The licence is MIT.

When your agent uses it

  • Tasks that involve Quizzes and assessments

Example prompts

  • “/coder-self-review”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Gather evidence. Honest self-review makes verification efficient. Accurate scorecards confirm fast; dishonest ones fail and re-run…
  2. Score each criterion. Read the GRASP rubric below. For each letter, assign a score of 0, 1, or 2. You must
  3. Output the Scorecard. Fill in the scorecard below. This is not internal reasoning — this is your deliverable checkpoint.
  4. Apply the hard-fail rule. If any letter scores 0, do not deliver — go to step 5 immediately.
  5. Determine action by total score

What it can do on your machine

Read from SKILL.md and the folder at commit 851e942. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Coder Self Review loads about 2.4k tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 1,347 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ntorga/agent-starter-kit at commit 851e942, republished under its MIT licence (© ntorga). 1,347 words, ~2,440 tokens.

Download SKILL.mdSave it as .claude/skills/coder-self-review/SKILL.md (or your agent's skills folder).
name
coder-self-review
description
Deterministic self-evaluation rubric for Coder — scored every run using the GRASP framework.
usedBy
coder
version
0.2.0
lastUpdated
2026-09-12

Purpose

Before delivering a handoff, the Coder evaluates its own output against the GRASP rubric. Each letter is scored 0, 1, or 2 with evidence quoted from the rubric and cited from actual work. The total determines whether to deliver, rewrite, or abort.

Procedure

  1. Gather evidence. Honest self-review makes verification efficient. Accurate scorecards confirm fast; dishonest ones fail and re-run, wasting time and compute. Before scoring, run verification commands to gather proof. The examples below show common patterns — choose what provides the best evidence for your specific work.

    Examples:

    • Incomplete markers: rg -n 'TODO|FIXME|HACK|stub|placeholder' <changed-files> — zero results supports higher scores
    • Test suite: run the project's test command — all passing supports higher scores
    • Lint suppressions: rg -n 'nolint|eslint-disable|@ts-ignore|# noqa|disable-line' <changed-files> — zero results or justified suppressions support higher scores
    • Secrets: rg -n 'password|secret|api_key|token|PRIVATE.KEY' <changed-files> — zero results in source supports higher scores
    • Attack surface: identify every endpoint/handler accepting external input — verify each has explicit auth and input sanitization
    • Run the project's linter if one exists.

    These are examples, not mandates. Choose commands that provide the strongest proof for your work.

  2. Score each criterion. Read the GRASP rubric below. For each letter, assign a score of 0, 1, or 2. You must:

    • Quote the specific criterion level (0, 1, or 2) that your work matches.
    • Cite evidence from your actual work — file paths read, commands run, test output, patterns checked. Generic claims like "I followed all guidelines" are not evidence and score 0.
  3. Output the Scorecard. Fill in the scorecard below. This is not internal reasoning — this is your deliverable checkpoint.

  4. Apply the hard-fail rule. If any letter scores 0, do not deliver — go to step 5 immediately.

  5. Determine action by total score:

    • 9 – 10 — DELIVER — Implementation meets all criteria. Deliver to user.
    • 7 – 8 — FIX the scored < 2 criteria. a. Identify which letters scored below 2. b. Fix those gaps automatically (do NOT consult the user). c. Re-score, then deliver if 9-10. d. If still below 9, retry once more. e. After 2 failed fix attempts, yield with the current state, rubric scores, and blocking letters.
    • 0 – 6 — RESTART — The implementation is fundamentally broken. Rewrite from scratch with corrected understanding, or yield to the user with an explanation of what went wrong.

Scorecard

Complete this before delivering. Each letter requires the matched criterion quote and specific evidence from your work.

  • G — GUIDELINES — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [files read, commands run, test results]
  • R — REASONING — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [error paths verified, tests run, boundaries checked]
  • A — ARCHITECTURE — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [files checked, dependency directions verified]
  • S — STYLE — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [neighboring files read, linter results]
  • P — PROTECTION — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [attack surface analysis, rg results]

Total: X/10 → Action: [DELIVER/FIX/RESTART]

GRASP Rubric

G — GUIDELINES

Did I follow the playbook end-to-end with no scope creep?

  • 0 — Skipped todo creation or management. Did not review context files or FEATURE-MAP before touching code. Tests do not pass or were not run. Handoff format is wrong, missing sections, or absent. Scope expanded beyond the plan/brief.
  • 1 — Todo managed and context reviewed, but one or more procedural gaps: did not load relevant skills before implementing, did not update .context.md/FEATURE-MAP.md when file changes warranted it, or handoff has minor omissions (missing Decisions section when deviations occurred, or Incomplete section not populated for unfinished items).
  • 2 — All playbook steps followed: todo checked or created, plan type determined, context files reviewed, sibling files read for style, relevant skills loaded, tests written first and failed before implementation, all tests pass, .context.md and FEATURE-MAP.md updated where warranted, acceptance criteria verified, handoff delivered in exact format. Yield conditions evaluated honestly.
R — REASONING

Will this code survive real-world input and error paths?

  • 0 — Logic does not match the task brief's acceptance criteria. Error paths silently swallowed. Boundary conditions (nil, empty, zero, off-by-one) not handled. Resource leaks in error paths (unclosed connections, file handles, channels). Incomplete work markers present (TODO, FIXME, stub returns, skipped tests without justification).
  • 1 — Logic is sound and tests pass, but one or more edge cases are uncertain: retry logic missing backoff, external calls lack timeouts, backward compatibility of API changes not verified, or one error path logs but does not propagate.
  • 2 — All error paths handled or explicitly logged. Boundary conditions tested. No resource leaks. All external calls have timeouts. No incomplete markers. Backward compatibility verified or breaking changes documented in handoff Decisions. Tests cover Good, Bad, and Ugly lenses per the plan.
Show full SKILL.md (573 more words)Show less
A — ARCHITECTURE

Does this code respect the project's architectural boundaries?

  • 0 — Change violates layer boundaries (outer layer depends on inner, or vice versa). New dependency flows against the project's dependency direction. Duplicated logic where extraction exists, or premature abstraction where a simple function would do.
  • 1 — Boundaries respected, but one concern is unclear: a new cross-cutting dependency might create a cycle at scale, or an abstraction's necessity is questionable (not obviously wrong, but not obviously needed either).
  • 2 — Layer boundaries respected per .context.md definitions. All dependencies flow in the correct direction. No duplication — existing utilities reused where applicable. No premature abstractions — code is as simple as the problem requires. Structural coherence check passes.
S — STYLE

Does this code look like it belongs in this codebase?

  • 0 — Code does not match the local style of surrounding files (different naming conventions, structure, or patterns). MUST-level rule violations present. Cryptic one-liners or clever patterns that need comments to understand. Unjustified lint/type suppression markers added.
  • 1 — Style mostly matches, but one or two naming or formatting inconsistencies exist against the surrounding code. SHOULD-level deviations present without visible justification. One lint suppression added without an adjacent comment explaining why (a comment earned under rules/code/general.md § Comments — the tool contract is the external constraint).
  • 2 — Read two neighboring files before writing and matched their style exactly. All naming follows project conventions. Code is readable without comments — the structure explains itself. No new lint suppressions, or each has a clear adjacent justification. All rules checked and followed. MUST rules respected, SHOULD deviations justified.
P — PROTECTION

Did I map and secure every point where untrusted data enters?

  • 0 — Change introduces an endpoint, handler, or data flow accepting external input. Auth not enforced on mutating operations. Secrets visible in source or config. No sanitization on data reaching SQL, templates, file paths, or command sinks.
  • 1 — Attack surface identified and basic sanitization present, but one area is uncertain: auth enforcement relies on middleware ordering convention rather than explicit attachment, error messages may leak internals, or rate limiting missing on auth-adjacent endpoints.
  • 2 — No new attack surface (score 2). If surface exists: trace every untrusted data flow to its sink, use parameterized queries, attach auth to each endpoint, keep secrets out of source, keep TLS validation intact, use modern crypto with adequate key lengths.

Guardrails

  • Never deliver if any letter scores 0 — regardless of total. A zero is a hard fail.
  • Never skip scoring any letter — all 5 must be evaluated every run.
  • A letter without a specific evidence citation scores 0. "I followed all guidelines" is not evidence — cite what you did or did not do.
  • Every evidence citation must reference a specific action, file path, command output, or test result. Generic justifications are rejected.
  • The scorecard is not internal reasoning — it is a deliverable checkpoint. Output it.
  • The rubric is fixed — do not add or remove criteria. If a criterion proves inadequate, file a framework change request.
  • When fixing gaps (score 7-8 range), only address the letters that scored below 2. Do not rework letters that already scored 2. Fix automatically — do NOT stop to consult the user.
  • After 2 failed fix attempts, yield — do not keep looping. Present the current state, rubric scores, and blocking letters to the user.
  • The "restart" action (score 0-6) means: do not deliver the current output. Rewrite the implementation from scratch with corrected understanding, or yield to the user with a clear explanation of the failure mode.

© ntorga, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/coder-self-review of ntorga/agent-starter-kit.

Open the folder on GitHubat commit 851e942

Compare with similar skills

Coder Self Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Coder Self Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Coder Self Review this skillntorga/agent-starter-kit146—~2.4kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from ntorga/agent-starter-kit

All 21 skills in this repo
  • Agent Decision

    ntorga/agent-starter-kit

    Deterministic self-evaluation rubric for decision escalations — scored every run using the FRAME framework.

    146 GitHub stars~1.7k tokensUpdated 26 days ago
    Auto-check passed
  • Agent Memory

    ntorga/agent-starter-kit

    Long-term and session memory across sessions. An agent skill from ntorga/agent-starter-kit.

    146 GitHub stars~2.7k tokensUpdated 26 days ago
    Auto-check passed
  • Architect Design Tree

    ntorga/agent-starter-kit

    Builds the design tree for the grill — decisions mapped as nodes with dependencies, recommendations, and impact, pruned by path.

    146 GitHub stars~1.2k tokensUpdated 26 days ago
    Auto-check passed
  • Architect Impl Grounding

    ntorga/agent-starter-kit

    Grounds the grill's settled decisions in the codebase — annotates impl.md with file paths, signatures, reference files, test specs, and LOC; re-grounds the next epic after each landing.

    146 GitHub stars~944 tokensUpdated 26 days ago
    Auto-check passed
  • Boot

    ntorga/agent-starter-kit

    Session startup — gitignore, auto-update, memory, rules, context, CLI config, and greet.

    146 GitHub stars~937 tokensUpdated 26 days ago
    Auto-check passed
  • Browser Inspect

    ntorga/agent-starter-kit

    Browser inspection and interaction for verifying rendered web UI during development.

    146 GitHub stars~2k tokensUpdated 26 days ago
    Auto-check passed

Categories

Questions about Coder Self Review

What does Coder Self Review do?

Deterministic self-evaluation rubric for Coder — scored every run using the GRASP framework. Coder Self Review is an agent skill from ntorga/agent-starter-kit. Deterministic self-evaluation rubric for Coder — scored every run using the GRASP framework.

When should I use Coder Self Review?

Coder Self Review fits situations like: tasks that involve Quizzes and assessments.

How do I install Coder Self Review in Claude Code?

Run `npx skills add ntorga/agent-starter-kit --skill coder-self-review -a claude-code`. Or copy the skill folder (skills/coder-self-review in ntorga/agent-starter-kit) into .claude/skills/coder-self-review in your project. Claude Code loads it when a task matches its description.

How do I install Coder Self Review in Codex?

Run `npx skills add ntorga/agent-starter-kit --skill coder-self-review -a codex`. Or copy the skill folder (skills/coder-self-review in ntorga/agent-starter-kit) into .agents/skills/coder-self-review in your project. Codex loads it when a task matches its description.

Can I use Coder Self Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ntorga/agent-starter-kit --skill coder-self-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/coder-self-review, .gemini/skills/coder-self-review, .github/skills/coder-self-review and .opencode/skills/coder-self-review in your project.

What does Coder Self Review need to run?

Going by SKILL.md and its folder, Coder Self Review needs the command-line tools its instructions call (rg).

Does Coder Self Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Coder Self Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Coder Self Review use?

Coder Self Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Coder Self Review use?

About 2.4k tokens (SKILL.md is roughly 9.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Coder Self Review?

Skills that share tags, products or a category with Coder Self Review: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Coder Self Review?

ntorga (a GitHub user) maintains it in ntorga/agent-starter-kit, which has 146 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on September 12, 2026.

Source: ntorga/agent-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.