Agent skill

Skill Evaluator

by hamzafarooq in hamzafarooq/claude-code-starter

Evaluate any Skill by scoring its output against ground truth.

MITAuto-check passed

Install Skill Evaluator

skills CLI
$ npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install hamzafarooq/claude-code-starter skill-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/hamzafarooq/claude-code-starter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/skill-evaluator .claude/skills/skill-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-evaluator
GitHub stars
145
Token cost
~498 tokens
SKILL.md length
265 words
Files
1
Skills in repo
29
Repo updated
First seen
Licence
MIT

At a glance

Evaluate any Skill by scoring its output against ground truth.

  • Works in 2 steps: The Skill's system prompt (or the path… → The ground truth table (or path to…
  • Checking if a skill is ready to ship
  • SKILL.md covers When given a Skill to evaluate, Scoring rubric (per test case), Output format and Confidence score interpretation
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Skill Evaluator is an agent skill from hamzafarooq/claude-code-starter. Evaluate any Skill by scoring its output against ground truth. Use when asked to eval, test, or score a skill, or when checking if a skill is ready to ship.

Its SKILL.md is about 500 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The licence is MIT.

When your agent uses it

  • Checking if a skill is ready to ship
  • Tasks that involve Test generation

Example prompts

  • “/skill-evaluator”

Workflow steps

2 steps, taken from the first numbered list in SKILL.md.

  1. The Skill's system prompt (or the path to its SKILL.md)
  2. The ground truth table (or path to docs/eval-ground-truth.md)

What it can do on your machine

Read from SKILL.md and the folder at commit 172c531. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Evaluator loads about 498 tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 265 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~43
When it runs · the whole SKILL.md, loaded when a task matches
~498

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from hamzafarooq/claude-code-starter at commit 172c531, republished under its MIT licence (© hamzafarooq). 265 words, ~498 tokens.

Download SKILL.mdSave it as .claude/skills/skill-evaluator/SKILL.md (or your agent's skills folder).
name
skill-evaluator
description
Evaluate any Skill by scoring its output against ground truth. Use when asked to eval, test, or score a skill, or when checking if a skill is ready to ship.
model
claude-opus-4-7
tools
Read

You are an evaluator for Claude Code Skills.

Your job is to score a Skill's actual output against expected ground truth and identify what to fix in the system prompt.

When given a Skill to evaluate

Ask the user for:

  1. The Skill's system prompt (or the path to its SKILL.md)
  2. The ground truth table (or path to docs/eval-ground-truth.md)

If a ground truth file is provided, read it. If not, ask for at least 3 input/output pairs to work with.

Scoring rubric (per test case)

Score each output 0–2:

ScoreMeaning
2Matches ground truth — correct structure, correct content
1Partially correct — right structure, wrong or missing detail
0Wrong, missing, or hallucinated

Output format

Return this exact format:


Skill Eval Report

Skill: [name] Test cases run: [N] Pass (score ≥ 2): [N] Partial (score = 1): [N] Fail (score = 0): [N] Confidence score: [X / 10]

Results by test case:

Test 1 — Score: [0/1/2] Input: [what was passed in] Expected: [ground truth] Actual: [what the skill produced] Reason: [one line — why this score]

[repeat for each test case]

Failure pattern: [If multiple failures share a root cause, name it here. e.g. "The skill always drops the Risks section when the PRD is under 500 words." If no pattern, write "No consistent failure pattern."]

Fix to make: [One specific change to the system prompt that would address the most failures. Quote the exact line to add or change.]


Confidence score interpretation

ScoreRecommendation
9–10Ship it
7–8Fix failures, rerun
5–6Find root cause, rewrite prompt
< 5Rethink task definition

Do not summarize. Return the report only.

© hamzafarooq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/skill-evaluator of hamzafarooq/claude-code-starter.

Open the folder on GitHubat commit 172c531

Compare with similar skills

Skill Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Evaluator this skillhamzafarooq/claude-code-starter145—~498Automated safety check: PassMIT
Write and Verify Playwright Testsappsmithorg/appsmith41k—~2.9kAutomated safety check: NotesApache-2.0
Adk Verify Snippetsgoogle/adk-python22k—~1.4kAutomated safety check: PassApache-2.0
Engine E2Ewix/react-native-navigation13k—~1.1kAutomated safety check: PassMIT
Hermetic Python Unit TestsdimensionalOS/dimos4.6k—~1.4kAutomated safety check: PassCustom licence
Emcaklofas/kicad-happy1.4k1 repos~2.8kAutomated safety check: PassMIT

Similar skills

  • Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.

    41k GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check: notes
  • Adk Verify Snippets

    google/adk-python

    Official

    Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…

    22k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Engine E2E

    wix/react-native-navigation

    Official

    Run Wix Engine (mobile-apps-engine) iOS E2E tests locally to validate RNN changes.

    13k GitHub stars~1.1k tokensUpdated 2 days ago
    Testing & QAAuto-check passed
  • Hermetic Python Unit Tests

    dimensionalOS/dimos

    Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.

    4.6k GitHub stars~1.4k tokensUpdated today
    Testing & QAAuto-check passed
  • Emc

    aklofas/kicad-happy

    EMC pre-compliance risk analysis for KiCad PCB designs — 18 check categories, 44 rule IDs covering ground planes, decoupling, I/O filtering, switching harmonics, clock routing, differential pair…

    1.4k GitHub starsUsed in 1 repo~2.8k tokens
    Testing & QAAuto-check passed
  • Swig Test

    swig/swig

    Run SWIG test suite for specific languages. An agent skill from swig/swig.

    6.3k GitHub stars~2.3k tokensUpdated yesterday
    Testing & QAAuto-check passed

More from hamzafarooq/claude-code-starter

All 29 skills in this repo
  • Beautiful HTML

    hamzafarooq/claude-code-starter

    Turn any document, proposal, report, or outline into a stunning single-file HTML presentation using 34 pre-built professional templates.

    145 GitHub stars~1.6k tokensUpdated 24 days ago
    Auto-check passed
  • Claude Code Deck

    hamzafarooq/claude-code-starter

    Generate a complete, ready-to-share HTML presentation explaining Claude Code, Skills, Sub Agents, Hooks, and Multi-Agent Systems — designed for product managers.

    145 GitHub stars~972 tokensUpdated 24 days ago
    Auto-check passed
  • Competitor Research

    hamzafarooq/claude-code-starter

    Research 3–5 competitors for any product or feature. An agent skill from hamzafarooq/claude-code-starter.

    145 GitHub stars~1.6k tokensUpdated 24 days ago
    Auto-check passed
  • Frontend Slides

    hamzafarooq/claude-code-starter

    Create stunning, animation-rich single-file HTML presentations from scratch or by converting a PowerPoint (.pptx) file.

    145 GitHub stars~1.4k tokensUpdated 24 days ago
    Auto-check passed
  • Md To PDF

    hamzafarooq/claude-code-starter

    Convert any markdown file to a clean PDF. An agent skill from hamzafarooq/claude-code-starter.

    145 GitHub stars~598 tokensUpdated 24 days ago
    Auto-check passed
  • User Story Writer

    hamzafarooq/claude-code-starter

    Converts feature ideas or rough bullets into user stories with acceptance criteria.

    145 GitHub stars~531 tokensUpdated 24 days ago
    Auto-check passed

Questions about Skill Evaluator

What does Skill Evaluator do?

Evaluate any Skill by scoring its output against ground truth. Skill Evaluator is an agent skill from hamzafarooq/claude-code-starter. Evaluate any Skill by scoring its output against ground truth.

When should I use Skill Evaluator?

Skill Evaluator fits situations like: checking if a skill is ready to ship; tasks that involve Test generation.

How do I install Skill Evaluator in Claude Code?

Run `npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a claude-code`. Or copy the skill folder (.claude/skills/skill-evaluator in hamzafarooq/claude-code-starter) into .claude/skills/skill-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Skill Evaluator in Codex?

Run `npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a codex`. Or copy the skill folder (.claude/skills/skill-evaluator in hamzafarooq/claude-code-starter) into .agents/skills/skill-evaluator in your project. Codex loads it when a task matches its description.

Can I use Skill Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-evaluator, .gemini/skills/skill-evaluator, .github/skills/skill-evaluator and .opencode/skills/skill-evaluator in your project.

What does Skill Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Skill Evaluator is instructions for the agent only.

Does Skill Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Evaluator use?

Skill Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Evaluator use?

About 498 tokens (SKILL.md is roughly 2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Evaluator?

Skills that share tags, products or a category with Skill Evaluator: Write and Verify Playwright Tests (appsmithorg/appsmith, 41k stars), Adk Verify Snippets (google/adk-python, 22k stars), Engine E2E (wix/react-native-navigation, 13k stars) and Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Evaluator?

hamzafarooq (a GitHub user) maintains it in hamzafarooq/claude-code-starter, which has 145 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on September 13, 2026.

Source: hamzafarooq/claude-code-starter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.