Eval-Driven Development Harness
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.
$ npx skills add athola/claude-night-market --skill skills-eval -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install athola/claude-night-market skills-eval --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/abstract/skills/skills-eval .claude/skills/skills-eval && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skills-eval" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-eval into .claude/skills/skills-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skills-eval", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-evalType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add athola/claude-night-market --skill skills-eval -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install athola/claude-night-market skills-eval --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/abstract/skills/skills-eval .agents/skills/skills-eval && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skills-eval" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-eval into .agents/skills/skills-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skills-eval", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add athola/claude-night-market --skill skills-eval -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install athola/claude-night-market skills-eval --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/abstract/skills/skills-eval .cursor/skills/skills-eval && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skills-eval" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-eval into .cursor/skills/skills-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skills-eval", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/athola/claude-night-market.git --path plugins/abstract/skills/skills-eval--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add athola/claude-night-market --skill skills-eval -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install athola/claude-night-market skills-eval --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/abstract/skills/skills-eval .gemini/skills/skills-eval && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skills-eval" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-eval into .gemini/skills/skills-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skills-eval", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install athola/claude-night-market skills-evalInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add athola/claude-night-market --skill skills-eval -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/abstract/skills/skills-eval .github/skills/skills-eval && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skills-eval" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-eval into .github/skills/skills-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skills-eval", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add athola/claude-night-market --skill skills-eval -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install athola/claude-night-market skills-eval --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/abstract/skills/skills-eval .opencode/skills/skills-eval && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skills-eval" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/abstract/skills/skills-eval into .opencode/skills/skills-eval/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skills-eval", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills-evalEvaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.
Skills Eval is an agent skill from athola/claude-night-market. Evaluate Claude skill quality through auditing. Use when reviewing or auditing skills.
Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including scripts (for example `README.md`, `modules/advanced-tool-use-analysis.md` and `modules/authoring-checklist.md`).
The repository describes itself as: 23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context… The licence is MIT.
Read from SKILL.md and the folder at commit 9f3eb00. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/, which the agent can run.
Shell commands in SKILL.md call:
makeFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skills Eval loads about 1.6k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 489 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from athola/claude-night-market at commit 9f3eb00, republished under its MIT licence (© athola). 489 words, ~1,646 tokens.
.claude/skills/skills-eval/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.abstract:skill-authoring)abstract:hooks-eval)abstract:rules-eval)This framework audits Claude skills against quality standards to improve performance and reduce token consumption. Automated tools analyze skill structure, measure context usage, and identify specific technical improvements. Run verification commands after each audit to confirm fixes work correctly.
The skills-auditor provides structural analysis, while the improvement-suggester ranks fixes by impact. Compliance is verified through the compliance-checker. Runtime efficiency is monitored by tool-performance-analyzer and token-usage-tracker.
Run a full audit of all skills or target a specific file to identify structural issues.
# Audit all skills
make audit-all
# Audit specific skill
make audit-skill TARGET=path/to/skill/SKILL.mdUse skill_analyzer.py for complexity checks and token_estimator.py to verify the context budget.
make analyze-skill TARGET=path/to/skill/SKILL.md
make estimate-tokens TARGET=path/to/skill/SKILL.mdGenerate a prioritized plan and verify standards compliance using improvement_suggester.py and compliance_checker.py.
make improve-skill TARGET=path/to/skill/SKILL.md
make check-compliance TARGET=path/to/skill/SKILL.mdStart with make audit-all to inventory skills and identify high-priority targets. For each skill requiring attention, run analysis with analyze-skill to map complexity. Generate an improvement plan, apply fixes, and run check-compliance to verify the skill meets project standards. Finalize by checking the token budget for efficiency.
Quality assessments use the skills-auditor and improvement-suggester to generate detailed reports. Performance analysis focuses on token efficiency through the token-usage-tracker and tool performance via tool-performance-analyzer. For standards compliance, the compliance-checker automates common fixes for structural issues.
We evaluate skills across five dimensions: structure compliance, content quality, token efficiency, activation reliability, and tool integration. Scores above 90 represent production-ready skills, while scores below 50 indicate critical issues requiring immediate attention.
Improvements are prioritized by impact. Critical issues include security vulnerabilities or broken functionality. High-priority items cover structural flaws that hinder discoverability. Medium and low priorities focus on best practices and minor optimizations.
Deprecated: skills/shared/modules/ directories. Shared modules must be relocated into the consuming skill's own modules/ directory. The evaluator flags any remaining skills/shared/ as a structural warning.
Current: Each skill owns its modules at skills/<skill-name>/modules/. Cross-skill references use relative paths (e.g., ../skill-authoring/modules/description-writing.md).
modules/trigger-isolation-analysis.mdmodules/authoring-checklist.mdmodules/evaluation-workflows.mdmodules/advanced-tool-use-analysis.mdmodules/evaluation-framework.mdmodules/integration.mdmodules/troubleshooting.mdmodules/pressure-testing.mdmodules/integration-testing.mdmodules/performance-benchmarking.mdscripts/ directory.scripts/automation/.skills/shared/ module reference is listed as a structural
warning in the audit output.make check-compliance TARGET=<skill> exits 0 for skills reported as passing, confirming
the compliance-checker agrees with the audit score.© athola, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 14 other files (scripts) in plugins/abstract/skills/skills-eval of athola/claude-night-market.
Open the folder on GitHubat commit 9f3eb00
Skills Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skills Eval this skillathola/claude-night-market | 341 | — | ~1.6k | Automated safety check: Pass | MIT | |
| Eval-Driven Development Harnessaffaan-m/ECC | 276k | — | ~1.5k | Automated safety check: Pass | MIT | |
| Evalalirezarezvani/claude-skills | 28k | — | ~618 | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 276k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 276k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| LLM Eval Pipeline Auditai-evals-course/evals-skills | 1.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 |
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
alirezarezvani/claude-skills
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
affaan-m/ECC
Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi
ai-evals-course/evals-skills
Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.
diegosouzapw/OmniRoute
Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.
athola/claude-night-market
Run and interpret repo diagnostic scripts (ratchets, validators, token stats).
athola/claude-night-market
Coordinates Claude agent teams via filesystem protocol. An agent skill from athola/claude-night-market.
athola/claude-night-market
Delegates execution to eight CLIs (Gemini, Qwen, MiniMax, GLM, Muse, Codex, OpenCode, Glimmer).
athola/claude-night-market
Guide minimal code via a decision ladder with full safety, edge, and negative-case coverage.
athola/claude-night-market
Build a project skill library in .claude/skills/ via discovery, parallel authoring, and review.
athola/claude-night-market
Orchestrates full project lifecycle by auto-detecting state and routing to the correct phase.
Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market. Skills Eval is an agent skill from athola/claude-night-market. Evaluate Claude skill quality through auditing.
Skills Eval fits situations like: auditing skills.
Run `npx skills add athola/claude-night-market --skill skills-eval -a claude-code`. Or copy the skill folder (plugins/abstract/skills/skills-eval in athola/claude-night-market) into .claude/skills/skills-eval in your project. Claude Code loads it when a task matches its description.
Run `npx skills add athola/claude-night-market --skill skills-eval -a codex`. Or copy the skill folder (plugins/abstract/skills/skills-eval in athola/claude-night-market) into .agents/skills/skills-eval in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add athola/claude-night-market --skill skills-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skills-eval, .gemini/skills/skills-eval, .github/skills/skills-eval and .opencode/skills/skills-eval in your project.
Going by SKILL.md and its folder, Skills Eval needs the command-line tools its instructions call (make).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Skills Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Skills Eval: Eval-Driven Development Harness (affaan-m/ECC, 276k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 276k stars) and Eval Harness (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
athola (a GitHub user) maintains it in athola/claude-night-market, which has 341 GitHub stars. The repository holds 154 skills in this directory. The repository was last updated on October 6, 2026.
Source: athola/claude-night-market on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.