Agent skill

Skillgrade Graders

by mgechev in mgechev/skillgrade

Authors deterministic and LLM rubric graders for skillgrade evaluations.

MITAuto-check passedEducation

Install Skillgrade Graders

skills CLI
$ npx skills add mgechev/skillgrade --skill skillgrade-graders -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mgechev/skillgrade skillgrade-graders --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mgechev/skillgrade.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skillgrade-graders .claude/skills/skillgrade-graders && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skillgrade-graders
GitHub stars
720
Token cost
~972 tokens
SKILL.md length
362 words
Files
2 (incl. references)
Skills in repo
2
Repo updated
First seen
Licence
MIT

At a glance

Authors deterministic and LLM rubric graders for skillgrade evaluations.

  • Works in 2 steps: Determine whether the task requires… → For most tasks, combine both:…
  • Creating scoring scripts
  • SKILL.md covers Procedures and Error Handling
  • Needs GEMINI_API_KEY and ANTHROPIC_API_KEY

What it does

Skillgrade Graders is an agent skill from mgechev/skillgrade. Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple graders with weighted scoring. Don't use for setting up eval pipelines, configuring eval.yaml defaults, or general test writing.

Its SKILL.md is about 970 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/grader-output-schema.md`).

It sits in Education, covering Quizzes and assessments. The repository describes itself as: "Unit tests" for your agent skills. The licence is MIT.

When your agent uses it

  • Creating scoring scripts
  • Writing evaluation rubrics
  • Combining multiple graders with weighted scoring
  • Setting up eval pipelines

Example prompts

  • “Use the skillgrade-graders skill to author deterministic and LLM rubric graders for skillgrade evaluations”
  • “/skillgrade-graders”

Requirements

  • A credential in GEMINI_API_KEY
  • A credential in ANTHROPIC_API_KEY

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. Determine whether the task requires objective verification (deterministic) or qualitative assessment (LLM rubric).
  2. For most tasks, combine both: deterministic graders verify outcomes (weight 0.7), LLM rubrics assess approach quality (weight 0.3).

What it can do on your machine

Read from SKILL.md and the folder at commit bb213d4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml and json).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GEMINI_API_KEY
    • ANTHROPIC_API_KEY
    • OPENAI_API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skillgrade Graders loads about 972 tokens when it runs, and up to ~1.3k if it reads all its reference files. Until then it costs about 76 tokens; SKILL.md has 362 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~76
When it runs · the whole SKILL.md, loaded when a task matches
~972
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mgechev/skillgrade at commit bb213d4, republished under its MIT licence (© mgechev). 362 words, ~972 tokens.

Download SKILL.mdSave it as .claude/skills/skillgrade-graders/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
skillgrade-graders
description
Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple graders with weighted scoring. Don't use for setting up eval pipelines, configuring eval.yaml defaults, or general test writing.

Skillgrade Grader Authoring

Procedures

Step 1: Identify the Grading Strategy

  1. Determine whether the task requires objective verification (deterministic) or qualitative assessment (LLM rubric).
  2. For most tasks, combine both: deterministic graders verify outcomes (weight 0.7), LLM rubrics assess approach quality (weight 0.3).

Step 2: Write a Deterministic Grader

  1. Create a script in the skill's graders/ directory (bash or TypeScript).
  2. The script must output a JSON object to stdout with the following structure:
    json
    {"score": 0.67, "details": "2/3 checks passed", "checks": [{"name": "check-name", "passed": true, "message": "Description"}]}
  3. score (0.0–1.0) and details are required. checks is optional but recommended.
  4. Read references/grader-output-schema.md for the full output specification.
  5. Use awk for arithmetic in bash scripts — bc is not available in node:20-slim.
  6. Reference the grader in eval.yaml:
    yaml
    - type: deterministic
      run: bash graders/check.sh
      weight: 0.7

Step 3: Write an LLM Rubric Grader

  1. Draft a rubric with explicit scoring criteria and point allocations.
  2. Structure the rubric into weighted sections that sum to 1.0:
    Workflow Compliance (0-0.5):
    - Did the agent follow the mandatory workflow steps?
    Efficiency (0-0.5):
    - Completed in ≤5 commands without trial-and-error?
  3. Reference the rubric in eval.yaml:
    yaml
    - type: llm_rubric
      rubric: |
        [rubric text or file path]
      weight: 0.3
      provider: gemini               # optional: gemini (default) | anthropic | openai
      model: gemini-3.5-flash        # optional model override (defaults to the latest dynamically resolved flash model)
  4. For long rubrics, store in a separate file and reference by path: rubric: rubrics/quality.md.

Step 4: Combine Multiple Graders

  1. Assign weights to each grader based on importance. Weights are normalized automatically.
  2. Final reward is calculated as: Σ (grader_score × weight) / Σ weight.
  3. Example configuration:
    yaml
    graders:
      - type: deterministic
        run: bash graders/check.sh
        weight: 0.7
      - type: llm_rubric
        rubric: rubrics/quality.md
        weight: 0.3
Show full SKILL.md (165 more words)Show less

Step 5: Validate Graders

  1. Create a reference solution script that produces the expected output.
  2. Run skillgrade --validate to verify graders score the reference solution correctly.
  3. Test only deterministic graders: skillgrade --grader=deterministic (skips LLM calls, faster iteration).
  4. Test only LLM rubric graders: skillgrade --grader=llm_rubric.
  5. Run a specific eval with a specific grader type: skillgrade --eval=my-eval --grader=deterministic.
  6. If a grader returns unexpected scores, inspect the script output and adjust scoring logic.

Error Handling

  • If a deterministic grader outputs non-JSON, ensure all echo/console.log statements except the final JSON result are redirected to stderr.
  • If an LLM rubric grader returns 0.00 with a missing API key message, set the appropriate key for your provider: GEMINI_API_KEY (provider: gemini), ANTHROPIC_API_KEY (provider: anthropic), or OPENAI_API_KEY (provider: openai).
  • To use a custom/self-hosted LLM endpoint, set ANTHROPIC_BASE_URL (for provider: anthropic) or OPENAI_BASE_URL (for provider: openai) — e.g. for Ollama or vLLM.
  • If scores are inconsistent across trials, reduce rubric ambiguity by adding concrete examples of passing and failing behavior.

© mgechev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/skillgrade-graders of mgechev/skillgrade.

  • SKILL.md
  • references/grader-output-schema.md

Open the folder on GitHubat commit bb213d4

Compare with similar skills

Skillgrade Graders next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skillgrade Graders compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skillgrade Graders this skillmgechev/skillgrade720—~972Automated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch67k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch67k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated 3 days ago
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    67k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    67k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from mgechev/skillgrade

  • Skillgrade Setup

    mgechev/skillgrade

    Sets up and runs skillgrade evaluation pipelines for Agent Skills.

    720 GitHub stars~915 tokensUpdated 5 days ago
    Auto-check passed

Categories

Questions about Skillgrade Graders

What does Skillgrade Graders do?

Authors deterministic and LLM rubric graders for skillgrade evaluations. Skillgrade Graders is an agent skill from mgechev/skillgrade. Authors deterministic and LLM rubric graders for skillgrade evaluations.

When should I use Skillgrade Graders?

Skillgrade Graders fits situations like: creating scoring scripts; writing evaluation rubrics; combining multiple graders with weighted scoring; setting up eval pipelines.

How do I install Skillgrade Graders in Claude Code?

Run `npx skills add mgechev/skillgrade --skill skillgrade-graders -a claude-code`. Or copy the skill folder (skills/skillgrade-graders in mgechev/skillgrade) into .claude/skills/skillgrade-graders in your project. Claude Code loads it when a task matches its description.

How do I install Skillgrade Graders in Codex?

Run `npx skills add mgechev/skillgrade --skill skillgrade-graders -a codex`. Or copy the skill folder (skills/skillgrade-graders in mgechev/skillgrade) into .agents/skills/skillgrade-graders in your project. Codex loads it when a task matches its description.

Can I use Skillgrade Graders in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mgechev/skillgrade --skill skillgrade-graders -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skillgrade-graders, .gemini/skills/skillgrade-graders, .github/skills/skillgrade-graders and .opencode/skills/skillgrade-graders in your project.

What does Skillgrade Graders need to run?

Going by SKILL.md and its folder, Skillgrade Graders needs credentials named GEMINI_API_KEY, ANTHROPIC_API_KEY and OPENAI_API_KEY. Our summary lists: A credential in GEMINI_API_KEY; A credential in ANTHROPIC_API_KEY.

Does Skillgrade Graders access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skillgrade Graders safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skillgrade Graders use?

Skillgrade Graders is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skillgrade Graders use?

About 972 tokens (SKILL.md is roughly 3.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 356 tokens, read only when the agent opens those files.

What are the alternatives to Skillgrade Graders?

Skills that share tags, products or a category with Skillgrade Graders: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 67k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 67k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skillgrade Graders?

mgechev (a GitHub user) maintains it in mgechev/skillgrade, which has 720 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 5, 2026.

Source: mgechev/skillgrade on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.