Agent skill

Contextualizer Self Review

by ntorga in ntorga/agent-starter-kit

Deterministic self-evaluation rubric for Contextualizer — scored every run using the TRACE framework.

MITAuto-check passedEducation

Install Contextualizer Self Review

skills CLI
$ npx skills add ntorga/agent-starter-kit --skill contextualizer-self-review -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ntorga/agent-starter-kit contextualizer-self-review --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ntorga/agent-starter-kit.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/contextualizer-self-review .claude/skills/contextualizer-self-review && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
contextualizer-self-review
GitHub stars
146
Token cost
~2.3k tokens
SKILL.md length
1,297 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Deterministic self-evaluation rubric for Contextualizer — scored every run using the TRACE framework.

  • Works in 5 steps: Gather evidence. Honest self-review… → Score each criterion. Read the TRACE… → Output the Scorecard. Fill in the… → …
  • Tasks that involve Quizzes and assessments
  • SKILL.md covers Purpose, Procedure, Scorecard and TRACE Rubric, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Contextualizer Self Review is an agent skill from ntorga/agent-starter-kit. Deterministic self-evaluation rubric for Contextualizer — scored every run using the TRACE framework.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments. The repository describes itself as: The scaffold for your multi-model, personalized Natural Language AI Harness (NLAH) . The licence is MIT.

When your agent uses it

  • Tasks that involve Quizzes and assessments

Example prompts

  • “/contextualizer-self-review”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Gather evidence. Honest self-review makes verification efficient. Accurate scorecards confirm fast; dishonest ones fail and re-run…
  2. Score each criterion. Read the TRACE rubric below. For each letter, assign a score of 0, 1, or 2. You must
  3. Output the Scorecard. Fill in the scorecard below. This is not internal reasoning — this is your deliverable checkpoint.
  4. Apply the hard-fail rule. If any letter scores 0, do not deliver — go to step 5 immediately.
  5. Determine action by total score

What it can do on your machine

Read from SKILL.md and the folder at commit 851e942. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Contextualizer Self Review loads about 2.3k tokens when it runs. Until then it costs about 32 tokens; SKILL.md has 1,297 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~32
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ntorga/agent-starter-kit at commit 851e942, republished under its MIT licence (© ntorga). 1,297 words, ~2,302 tokens.

Download SKILL.mdSave it as .claude/skills/contextualizer-self-review/SKILL.md (or your agent's skills folder).
name
contextualizer-self-review
description
Deterministic self-evaluation rubric for Contextualizer — scored every run using the TRACE framework.
usedBy
contextualizer
version
0.2.0
lastUpdated
2026-09-12

Purpose

Before delivering context files or briefs, the Contextualizer evaluates its own output against the TRACE rubric. Each letter is scored 0, 1, or 2 with evidence quoted from the rubric and cited from actual work. The total determines whether to deliver, rewrite, or abort.

Procedure

  1. Gather evidence. Honest self-review makes verification efficient. Accurate scorecards confirm fast; dishonest ones fail and re-run, wasting time and compute. Before scoring, run verification commands to gather proof. The examples below show common patterns — choose what provides the best evidence for your specific work.

    Examples:

    • Schema compliance: verify .context.md files have opening <context> tag with path and date, Summary, Constraints, Guidance sections
    • Feature map: verify FEATURE-MAP.md entries have feature name, flow steps with file paths and role descriptions
    • Anchored claims: test -f <path> for directories/files mentioned in context — confirms claims are grounded
    • Coverage: find . -type f | wc -l and find . -type d | wc -l to check project size against yield threshold
    • Brevity: wc -l <context-files> — each context file should be shorter than the directory it describes

    These are examples, not mandates. Choose commands that provide the strongest proof for your output.

  2. Score each criterion. Read the TRACE rubric below. For each letter, assign a score of 0, 1, or 2. You must:

    • Quote the specific criterion level (0, 1, or 2) that your work matches.
    • Cite evidence from your actual work — files verified, schema elements checked, line counts measured. Generic claims like "I verified everything" are not evidence and score 0.
  3. Output the Scorecard. Fill in the scorecard below. This is not internal reasoning — this is your deliverable checkpoint.

  4. Apply the hard-fail rule. If any letter scores 0, do not deliver — go to step 5 immediately.

  5. Determine action by total score:

    • 9 – 10 — DELIVER — Output meets all criteria. Deliver to user.
    • 7 – 8 — FIX the scored < 2 criteria. a. Identify which letters scored below 2. b. Fix those gaps automatically (do NOT consult the user). c. Re-score, then deliver if 9-10. d. If still below 9, retry once more. e. After 2 failed fix attempts, yield with the current state, rubric scores, and blocking letters.
    • 0 – 6 — RESTART — The output is fundamentally broken. Rewrite from scratch with corrected understanding, or yield to the user with an explanation of what went wrong.

Scorecard

Complete this before delivering. Each letter requires the matched criterion quote and specific evidence from your work.

  • T — TASK — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [task classification, deliverable type produced]
  • R — RUN — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [schema elements verified, skill followed]
  • A — ANCHORED — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [files verified with test -f, claims traced to code]
  • C — COVERAGE — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [project size checked, yield condition evaluated, incremental vs rewrite]
  • E — EVIDENT — Score: [0/1/2] — Matched: [quote the 0, 1, or 2 description] — Evidence: [line counts, brevity comparison]

Total: X/10 → Action: [DELIVER/FIX/RESTART]

TRACE Rubric

T — TASK

Did I classify the task correctly and produce the right deliverable?

  • 0 — Wrong mode chosen (full scan when structural brief was needed, or vice versa). Produced the wrong deliverable type entirely.
  • 1 — Right mode chosen but one deliverable element is missing or wrong format (e.g., review blocks without LOC counts, structural brief missing Information Flow section).
  • 2 — Correct mode, correct deliverable, correct format. Full scan produced .context.md files AND FEATURE-MAP.md. Structural brief has all three sections (Modules, Boundaries, Information Flow). Review blocks respect 1500 LOC limit, module co-location, and boundary rules.
R — RUN

Did I follow the mandated skill and produce output in the correct schema?

  • 0 — Did not follow skills/context-maintenance/SKILL.md for full scan. Used ad-hoc format instead of the prescribed .context.md or FEATURE-MAP.md schema. Did not run the directory scan script when producing context files. Structural brief does not use the required three-section format.
  • 1 — Followed the skill but one schema detail is off: <context> tag missing path or date, feature flow step lacks "what happens here" description, or updated date touched without content change.
  • 2 — Followed skills/context-maintenance/SKILL.md end-to-end. .context.md files use exact schema: opening <context> tag with path and date, Summary, Constraints, Guidance sections. FEATURE-MAP.md uses exact schema: feature name, flow steps in order with file paths and role descriptions. Structural brief uses exact three-section format. Updated only drifted features in the existing map — did not rewrite it.
Show full SKILL.md (562 more words)Show less
A — ANCHORED

Is every claim grounded in actual code? Nothing invented, nothing assumed.

  • 0 — Invented purpose for a directory without reading its contents. Added constraints or guidance to .context.md that cannot be verified from the code itself. Added a feature to the map without tracing its full path through the codebase.
  • 1 — Mostly grounded but one or more claims are uncertain: a directory's purpose is described vaguely rather than marked "unclear," or one feature flow step is inferred rather than confirmed by reading the actual code.
  • 2 — Every claim traceable to code. Directories with unclear purpose are labeled as such rather than guessed. All .context.md constraints verified from code itself. Every FEATURE-MAP.md entry traced end-to-end from entry point to output. No assumptions, no inference without evidence.
C — COVERAGE

Did I scope correctly — incremental update, yield when too big, respect boundaries?

  • 0 — Rewrote FEATURE-MAP.md from scratch when it already existed. Project exceeds 200 files / 50 directories and no yield or coverage report was produced. Review blocks cross major architectural boundaries without tight coupling.
  • 1 — Mostly scoped correctly but missed a drifted feature in the existing map, or did not report what was covered vs. what remains after a partial scan. One review block slightly exceeds 1500 LOC.
  • 2 — Incremental update only — changed entries in existing FEATURE-MAP.md, untouched stable ones. Yield condition evaluated: project size checked, coverage gap reported if exceeded. Review blocks each under 1500 LOC, module files co-located, boundaries respected. Nothing left silently unprocessed.
E — EVIDENT

Can someone arriving cold orient from this output alone? Is it brief enough?

  • 0 — Output is bloated — .context.md takes longer to read than the directory itself. Newcomer cannot determine what a directory does without reading the code. Feature map cannot be followed from entry point to output.
  • 1 — Output is mostly navigable but one .context.md summary is too detailed or one feature flow step is unclear without inspecting the code. Structure present but one section reads like prose instead of a quick-reference list.
  • 2 — Every .context.md is brief: one-to-two sentence description, one-line-per-file summary, constraints and guidance only when needed. Feature map follows a straight path from entry to output — a newcomer can trace the flow without opening code. Structure over prose. Brevity check passes: output is shorter than the directory it describes.

Guardrails

  • Never deliver if any letter scores 0 — regardless of total. A zero is a hard fail.
  • Never skip scoring any letter — all 5 must be evaluated every run.
  • A letter without a specific evidence citation scores 0. "I verified everything" is not evidence — cite what you verified.
  • Every evidence citation must reference a specific action, file path, command output, or verification result. Generic justifications are rejected.
  • The scorecard is not internal reasoning — it is a deliverable checkpoint. Output it.
  • The rubric is fixed — do not add or remove criteria. If a criterion proves inadequate, file a framework change request.
  • When fixing gaps (score 7-8 range), only address the letters that scored below 2. Do not rework letters that already scored 2. Fix automatically — do NOT stop to consult the user.
  • After 2 failed fix attempts, yield — do not keep looping. Present the current state, rubric scores, and blocking letters to the user.
  • The "restart" action (score 0-6) means: do not deliver the current output. Rewrite from scratch with corrected understanding, or yield to the user with a clear explanation of the failure mode.

© ntorga, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/contextualizer-self-review of ntorga/agent-starter-kit.

Open the folder on GitHubat commit 851e942

Compare with similar skills

Contextualizer Self Review next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Contextualizer Self Review compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Contextualizer Self Review this skillntorga/agent-starter-kit146—~2.3kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated yesterday
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from ntorga/agent-starter-kit

All 21 skills in this repo
  • Agent Decision

    ntorga/agent-starter-kit

    Deterministic self-evaluation rubric for decision escalations — scored every run using the FRAME framework.

    146 GitHub stars~1.7k tokensUpdated 26 days ago
    Auto-check passed
  • Agent Memory

    ntorga/agent-starter-kit

    Long-term and session memory across sessions. An agent skill from ntorga/agent-starter-kit.

    146 GitHub stars~2.7k tokensUpdated 26 days ago
    Auto-check passed
  • Architect Design Tree

    ntorga/agent-starter-kit

    Builds the design tree for the grill — decisions mapped as nodes with dependencies, recommendations, and impact, pruned by path.

    146 GitHub stars~1.2k tokensUpdated 26 days ago
    Auto-check passed
  • Architect Impl Grounding

    ntorga/agent-starter-kit

    Grounds the grill's settled decisions in the codebase — annotates impl.md with file paths, signatures, reference files, test specs, and LOC; re-grounds the next epic after each landing.

    146 GitHub stars~944 tokensUpdated 26 days ago
    Auto-check passed
  • Boot

    ntorga/agent-starter-kit

    Session startup — gitignore, auto-update, memory, rules, context, CLI config, and greet.

    146 GitHub stars~937 tokensUpdated 26 days ago
    Auto-check passed
  • Browser Inspect

    ntorga/agent-starter-kit

    Browser inspection and interaction for verifying rendered web UI during development.

    146 GitHub stars~2k tokensUpdated 26 days ago
    Auto-check passed

Categories

Questions about Contextualizer Self Review

What does Contextualizer Self Review do?

Deterministic self-evaluation rubric for Contextualizer — scored every run using the TRACE framework. Contextualizer Self Review is an agent skill from ntorga/agent-starter-kit. Deterministic self-evaluation rubric for Contextualizer — scored every run using the TRACE framework.

When should I use Contextualizer Self Review?

Contextualizer Self Review fits situations like: tasks that involve Quizzes and assessments.

How do I install Contextualizer Self Review in Claude Code?

Run `npx skills add ntorga/agent-starter-kit --skill contextualizer-self-review -a claude-code`. Or copy the skill folder (skills/contextualizer-self-review in ntorga/agent-starter-kit) into .claude/skills/contextualizer-self-review in your project. Claude Code loads it when a task matches its description.

How do I install Contextualizer Self Review in Codex?

Run `npx skills add ntorga/agent-starter-kit --skill contextualizer-self-review -a codex`. Or copy the skill folder (skills/contextualizer-self-review in ntorga/agent-starter-kit) into .agents/skills/contextualizer-self-review in your project. Codex loads it when a task matches its description.

Can I use Contextualizer Self Review in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ntorga/agent-starter-kit --skill contextualizer-self-review -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/contextualizer-self-review, .gemini/skills/contextualizer-self-review, .github/skills/contextualizer-self-review and .opencode/skills/contextualizer-self-review in your project.

What does Contextualizer Self Review need to run?

SKILL.md names no scripts, command-line tools or credentials: Contextualizer Self Review is instructions for the agent only.

Does Contextualizer Self Review access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Contextualizer Self Review safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Contextualizer Self Review use?

Contextualizer Self Review is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Contextualizer Self Review use?

About 2.3k tokens (SKILL.md is roughly 9.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Contextualizer Self Review?

Skills that share tags, products or a category with Contextualizer Self Review: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Contextualizer Self Review?

ntorga (a GitHub user) maintains it in ntorga/agent-starter-kit, which has 146 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on September 12, 2026.

Source: ntorga/agent-starter-kit on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.