Write and Verify Playwright Tests
appsmithorg/appsmith
Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.
Evaluate any Skill by scoring its output against ground truth.
$ npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install hamzafarooq/claude-code-starter skill-evaluator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/hamzafarooq/claude-code-starter.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/skill-evaluator .claude/skills/skill-evaluator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-evaluator" agent skill from https://github.com/hamzafarooq/claude-code-starter/tree/main/.claude/skills/skill-evaluator into .claude/skills/skill-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/hamzafarooq/claude-code-starter/tree/main/.claude/skills/skill-evaluatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install hamzafarooq/claude-code-starter skill-evaluator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hamzafarooq/claude-code-starter.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/skill-evaluator .agents/skills/skill-evaluator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-evaluator" agent skill from https://github.com/hamzafarooq/claude-code-starter/tree/main/.claude/skills/skill-evaluator into .agents/skills/skill-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install hamzafarooq/claude-code-starter skill-evaluator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hamzafarooq/claude-code-starter.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/skill-evaluator .cursor/skills/skill-evaluator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-evaluator" agent skill from https://github.com/hamzafarooq/claude-code-starter/tree/main/.claude/skills/skill-evaluator into .cursor/skills/skill-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/hamzafarooq/claude-code-starter.git --path .claude/skills/skill-evaluator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install hamzafarooq/claude-code-starter skill-evaluator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hamzafarooq/claude-code-starter.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/skill-evaluator .gemini/skills/skill-evaluator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-evaluator" agent skill from https://github.com/hamzafarooq/claude-code-starter/tree/main/.claude/skills/skill-evaluator into .gemini/skills/skill-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install hamzafarooq/claude-code-starter skill-evaluatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/hamzafarooq/claude-code-starter.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/skill-evaluator .github/skills/skill-evaluator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-evaluator" agent skill from https://github.com/hamzafarooq/claude-code-starter/tree/main/.claude/skills/skill-evaluator into .github/skills/skill-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install hamzafarooq/claude-code-starter skill-evaluator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/hamzafarooq/claude-code-starter.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/skill-evaluator .opencode/skills/skill-evaluator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-evaluator" agent skill from https://github.com/hamzafarooq/claude-code-starter/tree/main/.claude/skills/skill-evaluator into .opencode/skills/skill-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-evaluatorEvaluate any Skill by scoring its output against ground truth.
Skill Evaluator is an agent skill from hamzafarooq/claude-code-starter. Evaluate any Skill by scoring its output against ground truth. Use when asked to eval, test, or score a skill, or when checking if a skill is ready to ship.
Its SKILL.md is about 500 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The licence is MIT.
2 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 172c531. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Evaluator loads about 498 tokens when it runs. Until then it costs about 43 tokens; SKILL.md has 265 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from hamzafarooq/claude-code-starter at commit 172c531, republished under its MIT licence (© hamzafarooq). 265 words, ~498 tokens.
.claude/skills/skill-evaluator/SKILL.md (or your agent's skills folder).You are an evaluator for Claude Code Skills.
Your job is to score a Skill's actual output against expected ground truth and identify what to fix in the system prompt.
Ask the user for:
If a ground truth file is provided, read it. If not, ask for at least 3 input/output pairs to work with.
Score each output 0–2:
| Score | Meaning |
|---|---|
| 2 | Matches ground truth — correct structure, correct content |
| 1 | Partially correct — right structure, wrong or missing detail |
| 0 | Wrong, missing, or hallucinated |
Return this exact format:
Skill Eval Report
Skill: [name] Test cases run: [N] Pass (score ≥ 2): [N] Partial (score = 1): [N] Fail (score = 0): [N] Confidence score: [X / 10]
Results by test case:
Test 1 — Score: [0/1/2] Input: [what was passed in] Expected: [ground truth] Actual: [what the skill produced] Reason: [one line — why this score]
[repeat for each test case]
Failure pattern: [If multiple failures share a root cause, name it here. e.g. "The skill always drops the Risks section when the PRD is under 500 words." If no pattern, write "No consistent failure pattern."]
Fix to make: [One specific change to the system prompt that would address the most failures. Quote the exact line to add or change.]
| Score | Recommendation |
|---|---|
| 9–10 | Ship it |
| 7–8 | Fix failures, rerun |
| 5–6 | Find root cause, rewrite prompt |
| < 5 | Rethink task definition |
Do not summarize. Return the report only.
© hamzafarooq, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/skills/skill-evaluator of hamzafarooq/claude-code-starter.
Open the folder on GitHubat commit 172c531
Skill Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Evaluator this skillhamzafarooq/claude-code-starter | 145 | — | ~498 | Automated safety check: Pass | MIT | |
| Write and Verify Playwright Testsappsmithorg/appsmith | 41k | — | ~2.9k | Automated safety check: Notes | Apache-2.0 | |
| Adk Verify Snippetsgoogle/adk-python | 22k | — | ~1.4k | Automated safety check: Pass | Apache-2.0 | |
| Engine E2Ewix/react-native-navigation | 13k | — | ~1.1k | Automated safety check: Pass | MIT | |
| Hermetic Python Unit TestsdimensionalOS/dimos | 4.6k | — | ~1.4k | Automated safety check: Pass | Custom licence | |
| Emcaklofas/kicad-happy | 1.4k | 1 repos | ~2.8k | Automated safety check: Pass | MIT |
appsmithorg/appsmith
Writes a Playwright end-to-end test from a prompt, runs it against a live Appsmith deployment and retries with fixes up to three times until it passes.
google/adk-python
Checks that every Python code block in a Markdown file actually compiles and runs, by extracting each block to a temporary file, executing it in an isolated subprocess, and writing a pass/fail…
wix/react-native-navigation
Run Wix Engine (mobile-apps-engine) iOS E2E tests locally to validate RNN changes.
dimensionalOS/dimos
Rules for writing, fixing and reviewing pytest unit tests that are hermetic: behavior-focused, deterministic, isolated and cheap to run.
aklofas/kicad-happy
EMC pre-compliance risk analysis for KiCad PCB designs — 18 check categories, 44 rule IDs covering ground planes, decoupling, I/O filtering, switching harmonics, clock routing, differential pair…
swig/swig
Run SWIG test suite for specific languages. An agent skill from swig/swig.
hamzafarooq/claude-code-starter
Turn any document, proposal, report, or outline into a stunning single-file HTML presentation using 34 pre-built professional templates.
hamzafarooq/claude-code-starter
Generate a complete, ready-to-share HTML presentation explaining Claude Code, Skills, Sub Agents, Hooks, and Multi-Agent Systems — designed for product managers.
hamzafarooq/claude-code-starter
Research 3–5 competitors for any product or feature. An agent skill from hamzafarooq/claude-code-starter.
hamzafarooq/claude-code-starter
Create stunning, animation-rich single-file HTML presentations from scratch or by converting a PowerPoint (.pptx) file.
hamzafarooq/claude-code-starter
Convert any markdown file to a clean PDF. An agent skill from hamzafarooq/claude-code-starter.
hamzafarooq/claude-code-starter
Converts feature ideas or rough bullets into user stories with acceptance criteria.
Evaluate any Skill by scoring its output against ground truth. Skill Evaluator is an agent skill from hamzafarooq/claude-code-starter. Evaluate any Skill by scoring its output against ground truth.
Skill Evaluator fits situations like: checking if a skill is ready to ship; tasks that involve Test generation.
Run `npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a claude-code`. Or copy the skill folder (.claude/skills/skill-evaluator in hamzafarooq/claude-code-starter) into .claude/skills/skill-evaluator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a codex`. Or copy the skill folder (.claude/skills/skill-evaluator in hamzafarooq/claude-code-starter) into .agents/skills/skill-evaluator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add hamzafarooq/claude-code-starter --skill skill-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-evaluator, .gemini/skills/skill-evaluator, .github/skills/skill-evaluator and .opencode/skills/skill-evaluator in your project.
SKILL.md names no scripts, command-line tools or credentials: Skill Evaluator is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Skill Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 498 tokens (SKILL.md is roughly 2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Skill Evaluator: Write and Verify Playwright Tests (appsmithorg/appsmith, 41k stars), Adk Verify Snippets (google/adk-python, 22k stars), Engine E2E (wix/react-native-navigation, 13k stars) and Hermetic Python Unit Tests (dimensionalOS/dimos, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
hamzafarooq (a GitHub user) maintains it in hamzafarooq/claude-code-starter, which has 145 GitHub stars. The repository holds 29 skills in this directory. The repository was last updated on September 13, 2026.
Source: hamzafarooq/claude-code-starter on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.