Agent skill

Evaluation Lab

by aircrushin in aircrushin/promptMinder

Build compact, meaningful prompt evaluations with representative cases, explicit rubrics, and regression thresholds.

No licenceAuto-check passedEducation

Install Evaluation Lab

skills CLI
$ npx skills add aircrushin/promptMinder --skill evaluation-lab -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aircrushin/promptMinder evaluation-lab --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aircrushin/promptMinder.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/premium/evaluation-lab .claude/skills/evaluation-lab && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation-lab
GitHub stars
111
Token cost
~501 tokens
SKILL.md length
274 words
Files
1
Skills in repo
2
Repo updated
First seen
Licence
None found

At a glance

Build compact, meaningful prompt evaluations with representative cases, explicit rubrics, and regression thresholds.

  • Works in 7 steps: Write the evaluation question as a… → Collect a small case set that covers… → Remove identifying information and avoid… → …
  • Tasks that involve Quizzes and assessments
  • SKILL.md covers Working contract, Method, Output format and Quality rules
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Evaluation Lab is an agent skill from aircrushin/promptMinder. Build compact, meaningful prompt evaluations with representative cases, explicit rubrics, and regression thresholds.

Its SKILL.md is about 500 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Education, covering Quizzes and assessments. The repository describes itself as: 一个开源的,专注于提示词管理的平台 / An open-source platform focused on prompt management.

When your agent uses it

  • Tasks that involve Quizzes and assessments

Example prompts

  • “/evaluation-lab”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Write the evaluation question as a falsifiable statement.
  2. Collect a small case set that covers common, boundary, adversarial, and refusal scenarios.
  3. Remove identifying information and avoid using examples that reveal the expected answer through wording.
  4. Create a rubric with three to five dimensions. Each dimension needs observable anchors for fail, acceptable, and strong.
  5. Run the baseline and the candidate under the same inputs, context, and tool permissions.
  6. Record per-case results and inspect disagreements. A mean score cannot hide a critical failure.
  7. Set a release rule: minimum overall score, zero tolerance failures, and an allowed regression budget.

What it can do on your machine

Read from SKILL.md and the folder at commit 4c5134f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation Lab loads about 501 tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 274 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~501

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 274 words (~501 tokens).

“Use this skill when you are comparing prompt versions, models, or workflow changes and need evidence that survives beyond a single impressive example.”

— opening of SKILL.md by aircrushin
name
evaluation-lab

Read the full SKILL.md on GitHub

Files

Just SKILL.md in skills/premium/evaluation-lab of aircrushin/promptMinder.

Open the folder on GitHubat commit 4c5134f

Compare with similar skills

Evaluation Lab next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation Lab compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation Lab this skillaircrushin/promptMinder111—~501Automated safety check: PassNone
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch67k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch67k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated 3 days ago
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    67k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    67k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from aircrushin/promptMinder

  • Promptminder CLI

    aircrushin/promptMinder

    A skill your agent uses when running promptminder or promptminder-agent commands, setting PROMPTMINDERTOKEN, passing --team for workspace scoping, handling JSON stderr errors like "Missing token" or…

    111 GitHub stars~1.5k tokensUpdated today
    Auto-check passed
  • Promptminder CLI

    aircrushin/promptMinder

    A skill your agent uses when running PromptMinder CLI or agent-wrapper commands, configuring its token, or debugging team-scoped JSON I/O errors.

    111 GitHub stars~1k tokensUpdated today
    Auto-check passed
  • Agent Launch Checklist

    aircrushin/promptMinder

    Review an AI agent before release across scope, tools, state, safety, observability, human handoff, and rollback.

    111 GitHub stars~499 tokensUpdated today
    Auto-check passed
  • Content Repurposer

    aircrushin/promptMinder

    Reframe one source into channel-ready content while preserving claims, voice, disclosure requirements, and a clear review trail.

    111 GitHub stars~510 tokensUpdated today
    Auto-check passed
  • Prompt Architect

    aircrushin/promptMinder

    Design production-grade prompts with explicit task contracts, variable schemas, failure handling, and reusable quality gates.

    111 GitHub stars~602 tokensUpdated today
    Auto-check passed

Categories

Questions about Evaluation Lab

What does Evaluation Lab do?

Build compact, meaningful prompt evaluations with representative cases, explicit rubrics, and regression thresholds. Evaluation Lab is an agent skill from aircrushin/promptMinder. Build compact, meaningful prompt evaluations with representative cases, explicit rubrics, and regression thresholds.

When should I use Evaluation Lab?

Evaluation Lab fits situations like: tasks that involve Quizzes and assessments.

How do I install Evaluation Lab in Claude Code?

Run `npx skills add aircrushin/promptMinder --skill evaluation-lab -a claude-code`. Or copy the skill folder (skills/premium/evaluation-lab in aircrushin/promptMinder) into .claude/skills/evaluation-lab in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation Lab in Codex?

Run `npx skills add aircrushin/promptMinder --skill evaluation-lab -a codex`. Or copy the skill folder (skills/premium/evaluation-lab in aircrushin/promptMinder) into .agents/skills/evaluation-lab in your project. Codex loads it when a task matches its description.

Can I use Evaluation Lab in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aircrushin/promptMinder --skill evaluation-lab -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation-lab, .gemini/skills/evaluation-lab, .github/skills/evaluation-lab and .opencode/skills/evaluation-lab in your project.

What does Evaluation Lab need to run?

SKILL.md names no scripts, command-line tools or credentials: Evaluation Lab is instructions for the agent only.

Does Evaluation Lab access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluation Lab safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluation Lab use?

No licence was found for Evaluation Lab or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Evaluation Lab use?

About 501 tokens (SKILL.md is roughly 2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluation Lab?

Skills that share tags, products or a category with Evaluation Lab: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 67k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 67k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation Lab?

aircrushin (a GitHub user) maintains it in aircrushin/promptMinder, which has 111 GitHub stars. The repository holds 2 skills in this directory. The repository was last updated on October 11, 2026.

Source: aircrushin/promptMinder on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.