Agent skill

Grade Iterate

by alirezarezvani in alirezarezvani/claude-skills

Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop.

MITAuto-check passedEducation

Install Grade Iterate

skills CLI
$ npx skills add alirezarezvani/claude-skills --skill grade-iterate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install alirezarezvani/claude-skills grade-iterate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/alirezarezvani/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-launcher/skills/grade-iterate .claude/skills/grade-iterate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
grade-iterate
GitHub stars
28k
Token cost
~1k tokens
SKILL.md length
324 words
Files
5 (incl. scripts)
Skills in repo
342
Repo updated
First seen
Licence
MIT

At a glance

Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop.

  • Works in 5 steps: Define the outcome. → Read every verdict first. → Loop invariant. Each iteration must move… → …
  • The user says grade my agent
  • SKILL.md covers Workflow, Hard rules, Forcing-question library… and Tools
  • Runs Python scripts from its folder; calls python3

What it does

Grade Iterate is an agent skill from alirezarezvani/claude-skills. Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop. Define a CMA outcome (a required markdown rubric graded by an isolated grader), read each verdict, decide the next move (sharpen / re-run / promote to schedule), and once a version passes, run held-back eval cases in parallel. Use when the user says "grade my agent", "make it pass the rubric", "iterate until it's good", "is it good enough", or when the orchestrator routes phase=grade-iterate. outcomebuilder.py builds the…

Its SKILL.md is about 1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts (for example `README.md`, `scripts/eval_scaffold.py` and `scripts/outcome_builder.py`).

It sits in Education, covering Quizzes and assessments. The repository describes itself as: 380 Claude Code skills & agent skills & plugins (30+ Agents, 70+ custom commands, 380+ skills, customizable references, scripts)for Claude Code, Codex, Gemini CLI, Cursor, and 8… The licence is MIT.

When your agent uses it

  • The user says grade my agent
  • Make it pass the rubric
  • Iterate until its good
  • Is it good enough

Example prompts

  • “grade my agent”
  • “make it pass the rubric”
  • “iterate until it”
  • “/grade-iterate”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Define the outcome.
  2. Read every verdict first.
  3. Loop invariant. Each iteration must move ≥1 rubric line fail→pass, or the
  4. Once a version passes, run held-back eval.
  5. Decide. SHIP as v0, or promote to a scheduled deployment (Phase 4). Record

What it can do on your machine

Read from SKILL.md and the folder at commit 19392f7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Grade Iterate loads about 1k tokens when it runs. Until then it costs about 216 tokens; SKILL.md has 324 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~216
When it runs · the whole SKILL.md, loaded when a task matches
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from alirezarezvani/claude-skills at commit 19392f7, republished under its MIT licence (© alirezarezvani). 324 words, ~1,032 tokens.

Download SKILL.mdSave it as .claude/skills/grade-iterate/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
grade-iterate
description
Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop. Define a CMA outcome (a required markdown rubric graded by an isolated grader), read each verdict, decide the next move (sharpen / re-run / promote to schedule), and once a version passes, run held-back eval cases in parallel. Use when the user says "grade my agent", "make it pass the rubric", "iterate until it's good", "is it good enough", or when the orchestrator routes phase=grade-iterate. outcome_builder.py builds the user.define_outcome payload (rubric required, max_iterations clamped 1..20 — never unbounded); verdict_reader.py reads the grader result and recommends the next move; eval_scaffold.py generates held-back cases + a parallel run plan (capped at the 25-thread CMA ceiling). Distinct from stage-launch (first launch) and run-without-you (scheduling).
version
2.11.2
author
Alireza Rezvani
license
MIT
tags
cma, outcome, rubric, grader, grade-iterate, loop, max-iterations, eval, held-back
compatible_tools
claude-code, codex-cli, cursor, antigravity, opencode, gemini-cli

Phase 3 — Grade → Iterate (the bounded loop)

This is the plugin's loop: CMA's outcome primitive self-grades the agent's work in an isolated context and feeds failing verdicts back for the next attempt. It is always bounded by max_iterations (1..20) — never "improve forever".

See ../../references/loops-and-workflows.md and the outcome section of ../../references/cma-primitives.md.

Workflow

  1. Define the outcome.
    bash
    python3 scripts/outcome_builder.py \
      --sheet ./my-agent/build-sheet.json --max-iterations 5 \
      --out ./my-agent/payloads/outcome.json
    The rubric is required; max_iterations is clamped to 1..20. Send the payload as a user.define_outcome event (append to the running session).
  2. Read every verdict first.
    bash
    python3 scripts/verdict_reader.py --result ./my-agent/last-verdict.json
    Tables the rubric outcome and recommends: SHIP (satisfied), SHARPEN then re-run (needs_revision), ESCALATE (max_iterations_reached / failed), RESUME (interrupted). With ≤1 iteration left it flips to "make the single highest-value fix or escalate now".
  3. Loop invariant. Each iteration must move ≥1 rubric line fail→pass, or the run halts at the cap and escalates. Don't burn the budget on cosmetic edits.
  4. Once a version passes, run held-back eval.
    bash
    python3 scripts/eval_scaffold.py \
      --sheet ./my-agent/build-sheet.json --out ./my-agent/eval.json --concurrency 5
    Held-back cases (never seen during iteration) run in parallel, capped at the 25-thread CMA ceiling, each graded against the same rubric.
  5. Decide. SHIP as v0, or promote to a scheduled deployment (Phase 4). Record the verdict on the goal: goal_state.py set --phase run-without-you.

Hard rules

  • Bounded, always. No outcome without a max_iterations cap.
  • Read the verdict before acting. The grader's explanation drives the next move.
  • Held-back cases are held back. Never grade generalization on cases the agent already iterated against.

Forcing-question library (recommend + cite)

  1. "What are the 3–5 rubric lines?" Recommend: grounded, checkable criteria. Cite: cma-primitives.md (rubric required).
  2. "How many iterations before you'd rather look yourself?" Recommend: 3–5. Cite: loops-and-workflows.md (bounded loop).
  3. "On a fail, sharpen the prompt or the tools?" Recommend: whichever rubric line failed points to. Cite: verdict_reader next-move table.
  4. "Which cases did the agent NOT see?" Recommend: hold back ≥3 for generalization. Cite: this SKILL (held-back eval).

Tools

  • scripts/outcome_builder.py — user.define_outcome payload (rubric required, cap 1..20).
  • scripts/verdict_reader.py — grader result → next move.
  • scripts/eval_scaffold.py — held-back cases + parallel run plan (≤25 threads).

© alirezarezvani, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts) in agent-launcher/skills/grade-iterate of alirezarezvani/claude-skills.

  • SKILL.md
  • README.md
  • scripts/eval_scaffold.py
  • scripts/outcome_builder.py
  • scripts/verdict_reader.py

Open the folder on GitHubat commit 19392f7

Compare with similar skills

Grade Iterate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Grade Iterate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Grade Iterate this skillalirezarezvani/claude-skills28k—~1kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated yesterday
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from alirezarezvani/claude-skills

All 342 skills in this repo
  • Agile Product Owner

    alirezarezvani/claude-skills

    Writes INVEST-checked user stories with acceptance criteria, splits epics, plans sprints from velocity and ranks the backlog with a weighted score.

    28k GitHub starsUsed in 3 repos~3.2k tokens
    Auto-check passed
  • Product Strategist

    alirezarezvani/claude-skills

    OKR cascade toolkit for product leaders: generates aligned company-to-team OKRs from five strategy types and scores how well they line up.

    28k GitHub starsUsed in 2 repos~1.8k tokens
    Auto-check passed
  • App Store Optimization

    alirezarezvani/claude-skills

    App Store Optimization (ASO) toolkit for researching keywords, analyzing competitor rankings, generating metadata suggestions, and improving app visibility on Apple App Store and Google Play Store.

    28k GitHub starsUsed in 1 repo~4.2k tokens
    Auto-check passed
  • AWS Solution Architect

    alirezarezvani/claude-skills

    Design AWS architectures for startups using serverless patterns and IaC templates.

    28k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Campaign Analytics

    alirezarezvani/claude-skills

    Calculates attribution, funnel and ROI figures for marketing campaigns with three Python scripts that need only the standard library.

    28k GitHub starsUsed in 1 repo~2.1k tokens
    Auto-check passed
  • Code to PRD

    alirezarezvani/claude-skills

    Reverse-engineers a frontend, backend or fullstack codebase into a product requirements document with per-page docs, an enum dictionary and an API inventory.

    28k GitHub starsUsed in 1 repo~4.9k tokens
    Auto-check passed

Categories

Questions about Grade Iterate

What does Grade Iterate do?

Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop. Grade Iterate is an agent skill from alirezarezvani/claude-skills. Phase 3 of building a Claude Managed Agent — the bounded grade→iterate loop.

When should I use Grade Iterate?

Grade Iterate fits situations like: the user says grade my agent; make it pass the rubric; iterate until its good; is it good enough.

How do I install Grade Iterate in Claude Code?

Run `npx skills add alirezarezvani/claude-skills --skill grade-iterate -a claude-code`. Or copy the skill folder (agent-launcher/skills/grade-iterate in alirezarezvani/claude-skills) into .claude/skills/grade-iterate in your project. Claude Code loads it when a task matches its description.

How do I install Grade Iterate in Codex?

Run `npx skills add alirezarezvani/claude-skills --skill grade-iterate -a codex`. Or copy the skill folder (agent-launcher/skills/grade-iterate in alirezarezvani/claude-skills) into .agents/skills/grade-iterate in your project. Codex loads it when a task matches its description.

Can I use Grade Iterate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add alirezarezvani/claude-skills --skill grade-iterate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/grade-iterate, .gemini/skills/grade-iterate, .github/skills/grade-iterate and .opencode/skills/grade-iterate in your project.

What does Grade Iterate need to run?

Going by SKILL.md and its folder, Grade Iterate needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Grade Iterate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Grade Iterate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Grade Iterate use?

Grade Iterate is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Grade Iterate use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Grade Iterate?

Skills that share tags, products or a category with Grade Iterate: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Grade Iterate?

alirezarezvani (a GitHub user) maintains it in alirezarezvani/claude-skills, which has 27,829 GitHub stars. The repository holds 342 skills in this directory. The repository was last updated on August 30, 2026.

Source: alirezarezvani/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.