Agent skill

Skills Eval

by athola in athola/claude-night-market

Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.

MITAuto-check passed

Install Skills Eval

skills CLI
$ npx skills add athola/claude-night-market --skill skills-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install athola/claude-night-market skills-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/abstract/skills/skills-eval .claude/skills/skills-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skills-eval
GitHub stars
341
Token cost
~1.6k tokens
SKILL.md length
489 words
Files
15 (incl. scripts)
Skills in repo
154
Repo updated
First seen
Licence
MIT

At a glance

Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.

  • Auditing skills
  • SKILL.md covers When NOT To Use, Overview, Quick Start and Evaluation Workflow, plus 3 more sections
  • Calls make

What it does

Skills Eval is an agent skill from athola/claude-night-market. Evaluate Claude skill quality through auditing. Use when reviewing or auditing skills.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 16 other files, including scripts (for example `README.md`, `modules/advanced-tool-use-analysis.md` and `modules/authoring-checklist.md`).

The repository describes itself as: 23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context… The licence is MIT.

When your agent uses it

  • Auditing skills

Example prompts

  • “/skills-eval”

What it can do on your machine

Read from SKILL.md and the folder at commit 9f3eb00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/, which the agent can run.

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skills Eval loads about 1.6k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 489 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from athola/claude-night-market at commit 9f3eb00, republished under its MIT licence (© athola). 489 words, ~1,646 tokens.

Download SKILL.mdSave it as .claude/skills/skills-eval/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
skills-eval
description
Evaluate Claude skill quality through auditing. Use when reviewing or auditing skills.
alwaysApply
false
category
skill-management
tags
evaluation, improvement, skills, optimization, quality-assurance, tool-use, performance-metrics
dependencies
modular-skills
provides.infrastructure
evaluation-framework, quality-assurance, improvement-planning
provides.patterns
skill-analysis, token-optimization, modular-design
provides.sdk_features
agent-sdk-compatibility, advanced-metrics, dynamic-discovery
estimated_tokens
1800
usage_patterns
skill-audit, quality-assessment, improvement-planning, skills-inventory, tool-performance-evaluation, dynamic-discovery-optimization…
complexity
advanced

Skills Evaluation and Improvement

When NOT To Use

  • Writing a new skill (use abstract:skill-authoring)
  • Evaluating hooks (use abstract:hooks-eval)
  • Evaluating rules (use abstract:rules-eval)

Overview

This framework audits Claude skills against quality standards to improve performance and reduce token consumption. Automated tools analyze skill structure, measure context usage, and identify specific technical improvements. Run verification commands after each audit to confirm fixes work correctly.

The skills-auditor provides structural analysis, while the improvement-suggester ranks fixes by impact. Compliance is verified through the compliance-checker. Runtime efficiency is monitored by tool-performance-analyzer and token-usage-tracker.

Quick Start

Basic Audit

Run a full audit of all skills or target a specific file to identify structural issues.

bash
# Audit all skills
make audit-all

# Audit specific skill
make audit-skill TARGET=path/to/skill/SKILL.md
Analysis and Optimization

Use skill_analyzer.py for complexity checks and token_estimator.py to verify the context budget.

bash
make analyze-skill TARGET=path/to/skill/SKILL.md
make estimate-tokens TARGET=path/to/skill/SKILL.md
Improvements

Generate a prioritized plan and verify standards compliance using improvement_suggester.py and compliance_checker.py.

bash
make improve-skill TARGET=path/to/skill/SKILL.md
make check-compliance TARGET=path/to/skill/SKILL.md

Evaluation Workflow

Start with make audit-all to inventory skills and identify high-priority targets. For each skill requiring attention, run analysis with analyze-skill to map complexity. Generate an improvement plan, apply fixes, and run check-compliance to verify the skill meets project standards. Finalize by checking the token budget for efficiency.

Evaluation and Optimization

Quality assessments use the skills-auditor and improvement-suggester to generate detailed reports. Performance analysis focuses on token efficiency through the token-usage-tracker and tool performance via tool-performance-analyzer. For standards compliance, the compliance-checker automates common fixes for structural issues.

Scoring and Prioritization

We evaluate skills across five dimensions: structure compliance, content quality, token efficiency, activation reliability, and tool integration. Scores above 90 represent production-ready skills, while scores below 50 indicate critical issues requiring immediate attention.

Improvements are prioritized by impact. Critical issues include security vulnerabilities or broken functionality. High-priority items cover structural flaws that hinder discoverability. Medium and low priorities focus on best practices and minor optimizations.

Show full SKILL.md (197 more words)Show less
Structural Patterns

Deprecated: skills/shared/modules/ directories. Shared modules must be relocated into the consuming skill's own modules/ directory. The evaluator flags any remaining skills/shared/ as a structural warning.

Current: Each skill owns its modules at skills/<skill-name>/modules/. Cross-skill references use relative paths (e.g., ../skill-authoring/modules/description-writing.md).

Resources

Shared Modules: Cross-Skill Patterns
Skill-Specific Modules
  • Trigger Isolation Analysis: See modules/trigger-isolation-analysis.md
  • Authoring Checklist: See modules/authoring-checklist.md
  • Evaluation Workflows: See modules/evaluation-workflows.md
  • Advanced Tool Use Analysis: See modules/advanced-tool-use-analysis.md
  • Evaluation Framework: See modules/evaluation-framework.md
  • Integration Patterns: See modules/integration.md
  • Troubleshooting: See modules/troubleshooting.md
  • Pressure Testing: See modules/pressure-testing.md
  • Integration Testing: See modules/integration-testing.md
  • Performance Benchmarking: See modules/performance-benchmarking.md
Tools and Automation
  • Tools: Executable analysis utilities in scripts/ directory.
  • Automation: Setup and validation scripts in scripts/automation/.

Exit Criteria

  • Every audited skill receives a score across all five dimensions (structure compliance, content quality, token efficiency, activation reliability, tool integration) summing to 100.
  • Any skill with a deprecated skills/shared/ module reference is listed as a structural warning in the audit output.
  • make check-compliance TARGET=<skill> exits 0 for skills reported as passing, confirming the compliance-checker agrees with the audit score.
  • An improvement plan is produced for any skill scoring below 75, with findings ordered by priority: critical > high > medium > low.

© athola, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts) in plugins/abstract/skills/skills-eval of athola/claude-night-market.

  • SKILL.md
  • README.md
  • modules/advanced-tool-use-analysis.md
  • modules/authoring-checklist.md
  • modules/evaluation-criteria.md
  • modules/evaluation-framework.md
  • modules/evaluation-workflows.md
  • modules/integration-testing.md
  • modules/integration.md
  • modules/performance-benchmarking.md
  • modules/pressure-testing.md
  • modules/skill-authoring-best-practices.md
  • modules/trigger-isolation-analysis.md
  • modules/troubleshooting.md
  • scripts/README.md

Open the folder on GitHubat commit 9f3eb00

Compare with similar skills

Skills Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skills Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skills Eval this skillathola/claude-night-market341—~1.6kAutomated safety check: PassMIT
Eval-Driven Development Harnessaffaan-m/ECC276k—~1.5kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k—~618Automated safety check: PassMIT
Eval Harnessaffaan-m/ECC276k—~2.2kAutomated safety check: PassMIT
Eval Harnessaffaan-m/ECC276k1 repos~1.7kAutomated safety check: PassMIT
LLM Eval Pipeline Auditai-evals-course/evals-skills1.5k—~2.5kAutomated safety check: PassApache-2.0

Similar skills

  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    276k GitHub stars~1.5k tokensUpdated 5 days ago
    Agent WorkflowsAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub stars~618 tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    276k GitHub stars~2.2k tokensUpdated 5 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    276k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • LLM Eval Pipeline Audit

    ai-evals-course/evals-skills

    Inspects an LLM evaluation setup for missing error analysis, unvalidated judges and vanity metrics, and ranks the problems by impact with fixes.

    1.5k GitHub stars~2.5k tokensUpdated 15 days ago
    AI & LLM EngineeringAuto-check passed
  • OmniRoute CLI Evals

    diegosouzapw/OmniRoute

    Creates and runs LLM evaluation suites from the omniroute CLI, follows live runs, shows scorecards, compares models and ties eval runs into CI.

    74k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from athola/claude-night-market

All 154 skills in this repo
  • Night Market Diagnostics Toolkit

    athola/claude-night-market

    Run and interpret repo diagnostic scripts (ratchets, validators, token stats).

    341 GitHub stars~3.4k tokensUpdated 4 days ago
    Auto-check passed
  • Agent Teams

    athola/claude-night-market

    Coordinates Claude agent teams via filesystem protocol. An agent skill from athola/claude-night-market.

    341 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Delegation Core

    athola/claude-night-market

    Delegates execution to eight CLIs (Gemini, Qwen, MiniMax, GLM, Muse, Codex, OpenCode, Glimmer).

    341 GitHub stars~2.5k tokensUpdated 4 days ago
    Auto-check passed
  • Elegant Code

    athola/claude-night-market

    Guide minimal code via a decision ladder with full safety, edge, and negative-case coverage.

    341 GitHub stars~2.1k tokensUpdated 4 days ago
    Auto-check passed
  • Skill Library Mission

    athola/claude-night-market

    Build a project skill library in .claude/skills/ via discovery, parallel authoring, and review.

    341 GitHub stars~1.6k tokensUpdated 4 days ago
    Auto-check passed
  • Mission Orchestrator

    athola/claude-night-market

    Orchestrates full project lifecycle by auto-detecting state and routing to the correct phase.

    341 GitHub stars~3.4k tokensUpdated 4 days ago
    Auto-check: warnings

Questions about Skills Eval

What does Skills Eval do?

Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market. Skills Eval is an agent skill from athola/claude-night-market. Evaluate Claude skill quality through auditing.

When should I use Skills Eval?

Skills Eval fits situations like: auditing skills.

How do I install Skills Eval in Claude Code?

Run `npx skills add athola/claude-night-market --skill skills-eval -a claude-code`. Or copy the skill folder (plugins/abstract/skills/skills-eval in athola/claude-night-market) into .claude/skills/skills-eval in your project. Claude Code loads it when a task matches its description.

How do I install Skills Eval in Codex?

Run `npx skills add athola/claude-night-market --skill skills-eval -a codex`. Or copy the skill folder (plugins/abstract/skills/skills-eval in athola/claude-night-market) into .agents/skills/skills-eval in your project. Codex loads it when a task matches its description.

Can I use Skills Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add athola/claude-night-market --skill skills-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skills-eval, .gemini/skills/skills-eval, .github/skills/skills-eval and .opencode/skills/skills-eval in your project.

What does Skills Eval need to run?

Going by SKILL.md and its folder, Skills Eval needs the command-line tools its instructions call (make).

Does Skills Eval access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skills Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skills Eval use?

Skills Eval is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skills Eval use?

About 1.6k tokens (SKILL.md is roughly 6.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skills Eval?

Skills that share tags, products or a category with Skills Eval: Eval-Driven Development Harness (affaan-m/ECC, 276k stars), Eval (alirezarezvani/claude-skills, 28k stars), Eval Harness (affaan-m/ECC, 276k stars) and Eval Harness (affaan-m/ECC, 276k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skills Eval?

athola (a GitHub user) maintains it in athola/claude-night-market, which has 341 GitHub stars. The repository holds 154 skills in this directory. The repository was last updated on October 6, 2026.

Source: athola/claude-night-market on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.