Agent skill

Evaluate Grader

by ai-analyst-lab in ai-analyst-lab/ai-analyst

Compare a narrow model grader with frozen human labels and inspect disagreement, bias probes, and repeated scoring stability.

MITAuto-check passedEducation

Install Evaluate Grader

skills CLI
$ npx skills add ai-analyst-lab/ai-analyst --skill evaluate-grader -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install ai-analyst-lab/ai-analyst evaluate-grader --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/ai-analyst-lab/ai-analyst.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/evaluate-grader .claude/skills/evaluate-grader && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluate-grader
GitHub stars
304
Token cost
~390 tokens
SKILL.md length
193 words
Files
1
Skills in repo
43
Repo updated
First seen
Licence
MIT

At a glance

Compare a narrow model grader with frozen human labels and inspect disagreement, bias probes, and repeated scoring stability.

  • Changing a model-based evaluator
  • Calls python3

What it does

Evaluate Grader is an agent skill from ai-analyst-lab/ai-analyst. Compare a narrow model grader with frozen human labels and inspect disagreement, bias probes, and repeated scoring stability. Use when building or changing a model-based evaluator.

Its SKILL.md is about 390 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education. The repository describes itself as: AI Product Analyst — Claude Code-powered data analysis toolkit. The licence is MIT.

When your agent uses it

  • Changing a model-based evaluator

Example prompts

  • “/evaluate-grader”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 52c0744. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluate Grader loads about 390 tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 193 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~390

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from ai-analyst-lab/ai-analyst at commit 52c0744, republished under its MIT licence (© ai-analyst-lab). 193 words, ~390 tokens.

Download SKILL.mdSave it as .claude/skills/evaluate-grader/SKILL.md (or your agent's skills folder).
name
evaluate-grader
description
Compare a narrow model grader with frozen human labels and inspect disagreement, bias probes, and repeated scoring stability. Use when building or changing a model-based evaluator.

Evaluate a grader

Freeze the human labels before running the grader. Use one narrow criterion with a written rubric and structured output. The grader must be able to return unknown or request human review.

Keep human labels outside the judge workspace. Use python3 -m helpers.evals.cli run-isolated-judge to create a fresh child workspace containing only the named examples and current rubric. Run every revised rubric in a different fresh child. A child must not receive human labels, prior verdicts, captured verdicts, or later rubric versions. Preserve the isolation record with the verdicts.

Use helpers.evals.judges.evaluate_alignment for the confusion table and disagreement set. Repeat at least one unchanged boundary example and use repeated_label_stability to measure scoring stability.

Inspect:

  • every human and grader disagreement;
  • label imbalance;
  • an answer-order reversal when judging pairs;
  • a verbosity trap where a longer answer is not the better answer;
  • an ambiguous example that should produce unknown; and
  • whether the generator and grader are actually independent contexts.

Revise one rubric criterion at a time and rerun only the working examples. Do not tune on the heldout judge set.

Call the classroom result an alignment check. A small agreeing sample is not completed calibration.

© ai-analyst-lab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/evaluate-grader of ai-analyst-lab/ai-analyst.

Open the folder on GitHubat commit 52c0744

Compare with similar skills

Evaluate Grader next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluate Grader compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluate Grader this skillai-analyst-lab/ai-analyst304—~390Automated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
Zhang Xuefeng Perspectivealchaincyf/zhangxuefeng-skill10k1 repos~2.6kAutomated safety check: PassMIT
Deep Reading Analystginobefun/deep-reading-analyst-skill3535 repos~3.6kAutomated safety check: PassMIT
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
OpenMAIC Setup and ExtensionTHU-MAIC/OpenMAIC40k—~1.7kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    EducationAuto-check passed
  • Zhang Xuefeng Perspective

    alchaincyf/zhangxuefeng-skill

    Answers education and career questions in the voice of Zhang Xuefeng, looking up current employment and admissions data before giving a direct verdict.

    10k GitHub starsUsed in 1 repo~2.6k tokens
    EducationAuto-check passed
  • Deep Reading Analyst

    ginobefun/deep-reading-analyst-skill

    Comprehensive framework for deep analysis of articles, papers, and long-form content using 10+ thinking models (SCQA, 5W2H, critical thinking, inversion, mental models, first principles, systems…

    353 GitHub starsUsed in 5 repos~3.6k tokens
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated 2 days ago
    EducationAuto-check passed
  • Guides setup, classroom generation and secondary development for OpenMAIC, the multi-agent interactive classroom, one confirmed phase at a time.

    40k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check: notes
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed

More from ai-analyst-lab/ai-analyst

All 43 skills in this repo
  • Always Compare

    ai-analyst-lab/ai-analyst

    Never present a metric or number in isolation; anchor every number to a comparison (prior period, benchmark, or another segment) or state that none is available.

    304 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed
  • Archaeology

    ai-analyst-lab/ai-analyst

    Retrieve proven SQL patterns, table cheatsheets, and join patterns from .knowledge/query-archaeology/ so past work gets reused.

    304 GitHub stars~1.3k tokensUpdated 8 days ago
    Auto-check passed
  • Archive Analysis

    ai-analyst-lab/ai-analyst

    Save completed analyses to the knowledge system's analysis archive for future reference.

    304 GitHub stars~2.7k tokensUpdated 8 days ago
    Auto-check passed
  • Auth Preflight

    ai-analyst-lab/ai-analyst

    Verify Google Workspace MCP authentication at the start of any session that needs Google APIs (Docs, Slides, Drive).

    304 GitHub stars~3.1k tokensUpdated 8 days ago
    Auto-check passed
  • Causal

    ai-analyst-lab/ai-analyst

    Causal inference toolkit for when experiments are not possible: estimate treatment effects from observational data with assumption checks and mandatory caveats.

    304 GitHub stars~1.8k tokensUpdated 8 days ago
    Auto-check passed
  • Chart To Drive

    ai-analyst-lab/ai-analyst

    Standardized workflow for uploading local chart PNGs to Google Drive and making them available for insertion into Google Docs and Slides.

    304 GitHub stars~1.4k tokensUpdated 8 days ago
    Auto-check passed

Categories

Questions about Evaluate Grader

What does Evaluate Grader do?

Compare a narrow model grader with frozen human labels and inspect disagreement, bias probes, and repeated scoring stability. Evaluate Grader is an agent skill from ai-analyst-lab/ai-analyst. Compare a narrow model grader with frozen human labels and inspect disagreement, bias probes, and repeated scoring stability.

When should I use Evaluate Grader?

Evaluate Grader fits situations like: changing a model-based evaluator.

How do I install Evaluate Grader in Claude Code?

Run `npx skills add ai-analyst-lab/ai-analyst --skill evaluate-grader -a claude-code`. Or copy the skill folder (.claude/skills/evaluate-grader in ai-analyst-lab/ai-analyst) into .claude/skills/evaluate-grader in your project. Claude Code loads it when a task matches its description.

How do I install Evaluate Grader in Codex?

Run `npx skills add ai-analyst-lab/ai-analyst --skill evaluate-grader -a codex`. Or copy the skill folder (.claude/skills/evaluate-grader in ai-analyst-lab/ai-analyst) into .agents/skills/evaluate-grader in your project. Codex loads it when a task matches its description.

Can I use Evaluate Grader in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ai-analyst-lab/ai-analyst --skill evaluate-grader -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluate-grader, .gemini/skills/evaluate-grader, .github/skills/evaluate-grader and .opencode/skills/evaluate-grader in your project.

What does Evaluate Grader need to run?

Going by SKILL.md and its folder, Evaluate Grader needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Evaluate Grader access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluate Grader safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluate Grader use?

Evaluate Grader is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluate Grader use?

About 390 tokens (SKILL.md is roughly 1.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluate Grader?

Skills that share tags, products or a category with Evaluate Grader: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), Zhang Xuefeng Perspective (alchaincyf/zhangxuefeng-skill, 10k stars), Deep Reading Analyst (ginobefun/deep-reading-analyst-skill, 353 stars) and AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluate Grader?

ai-analyst-lab (a GitHub organization) maintains it in ai-analyst-lab/ai-analyst, which has 304 GitHub stars. The repository holds 43 skills in this directory. The repository was last updated on September 30, 2026.

Source: ai-analyst-lab/ai-analyst on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.