Agent skill

Evaluation Framework

by athola in athola/claude-night-market

Provides weighted scoring, rubrics, and decision-threshold patterns.

MITAuto-check passedTesting & QA

Install Evaluation Framework

skills CLI
$ npx skills add athola/claude-night-market --skill evaluation-framework -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install athola/claude-night-market evaluation-framework --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/leyline/skills/evaluation-framework .claude/skills/evaluation-framework && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation-framework
GitHub stars
342
Token cost
~1.3k tokens
SKILL.md length
322 words
Files
7
Skills in repo
159
Repo updated
First seen
Licence
MIT

At a glance

Provides weighted scoring, rubrics, and decision-threshold patterns.

  • Works in 4 steps: Define Criteria → Score Each Criterion → Calculate Weighted Total → …
  • Designing quality gates
  • SKILL.md covers Overview, When To Use, When NOT To Use and Core Pattern, plus 5 more sections
  • Calls pytest

What it does

Evaluation Framework is an agent skill from athola/claude-night-market. Provides weighted scoring, rubrics, and decision-threshold patterns. Use when designing quality gates, evaluation systems, or decision frameworks.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `README.md`, `modules/decision-thresholds.md` and `modules/evaluation-rubric.md`).

It sits in Testing & QA, covering Quality gates and Quizzes and assessments. The repository describes itself as: 23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context… The licence is MIT.

When your agent uses it

  • Designing quality gates
  • Evaluation systems
  • Decision frameworks

Example prompts

  • “Use the evaluation-framework skill to provide weighted scoring, rubrics, and decision-threshold patterns”
  • “/evaluation-framework”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Define Criteria
  2. Score Each Criterion
  3. Calculate Weighted Total
  4. Apply Decision Thresholds

What it can do on your machine

Read from SKILL.md and the folder at commit 9f3eb00. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • pytest

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation Framework loads about 1.3k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 322 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~42
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from athola/claude-night-market at commit 9f3eb00, republished under its MIT licence (© athola). 322 words, ~1,251 tokens.

Download SKILL.mdSave it as .claude/skills/evaluation-framework/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
evaluation-framework
description
Provides weighted scoring, rubrics, and decision-threshold patterns. Use when designing quality gates, evaluation systems, or decision frameworks.
alwaysApply
false
category
infrastructure
tags
evaluation, scoring, decision-making, metrics, quality
provides.infrastructure
weighted-scoring, threshold-decisions, evaluation-patterns
provides.patterns
criteria-definition, scoring-methodology, decision-logic
usage_patterns
quality-evaluation, scoring-systems, decision-frameworks, rubric-design
complexity
beginner
model_hint
fast
estimated_tokens
550
progressive_loading
true

Evaluation Framework

Overview

A generic framework for weighted scoring and threshold-based decision making. Provides reusable patterns for evaluating any artifact against configurable criteria with consistent scoring methodology.

This framework abstracts the common pattern of: define criteria → assign weights → score against criteria → apply thresholds → make decisions.

When To Use

  • Implementing quality gates or evaluation rubrics
  • Building scoring systems for artifacts, proposals, or submissions
  • Need consistent evaluation methodology across different domains
  • Want threshold-based automated decision making
  • Creating assessment tools with weighted criteria

When NOT To Use

  • Simple pass/fail without scoring needs

Core Pattern

1. Define Criteria
yaml
criteria:
  - name: criterion_name
    weight: 0.30          # 30% of total score
    description: What this measures
    scoring_guide:
      90-100: Exceptional
      70-89: Strong
      50-69: Acceptable
      30-49: Weak
      0-29: Poor

Verification: Run the command with --help flag to verify availability.

2. Score Each Criterion
python
scores = {
    "criterion_1": 85,  # Out of 100
    "criterion_2": 92,
    "criterion_3": 78,
}

Verification: Run the command with --help flag to verify availability.

3. Calculate Weighted Total
python
total = sum(score * weights[criterion] for criterion, score in scores.items())
# Example: (85 × 0.30) + (92 × 0.40) + (78 × 0.30) = 85.5

Verification: Run the command with --help flag to verify availability.

4. Apply Decision Thresholds
yaml
thresholds:
  80-100: Accept with priority
  60-79: Accept with conditions
  40-59: Review required
  20-39: Reject with feedback
  0-19: Reject

Verification: Run the command with --help flag to verify availability.

Quick Start

Define Your Evaluation
  1. Identify criteria: What aspects matter for your domain?
  2. Assign weights: Which criteria are most important? (sum to 1.0)
  3. Create scoring guides: What does each score range mean?
  4. Set thresholds: What total scores trigger which decisions?
Example: Code Review Evaluation
yaml
criteria:
  correctness: {weight: 0.40, description: Does code work as intended?}
  maintainability: {weight: 0.25, description: Is it readable?}
  performance: {weight: 0.20, description: Meets performance needs?}
  testing: {weight: 0.15, description: Tests detailed?}

thresholds:
  85-100: Approve immediately
  70-84: Approve with minor feedback
  50-69: Request changes
  0-49: Reject, major issues

Verification: Run pytest -v to verify tests pass.

Evaluation Workflow
text
**Verification:** Run the command with `--help` flag to verify availability.
1. Review artifact against each criterion
2. Assign 0-100 score for each criterion
3. Calculate: total = Σ(score × weight)
4. Compare total to thresholds
5. Take action based on threshold range

Verification: Run the command with --help flag to verify availability.

Common Use Cases

Quality Gates: Code review, PR approval, release readiness Content Evaluation: Document quality, knowledge intake, skill assessment Resource Allocation: Backlog prioritization, investment decisions, triage

Integration Pattern

yaml
# In your skill's frontmatter
dependencies: [leyline:evaluation-framework]

Verification: Run the command with --help flag to verify availability.

Then customize the framework for your domain:

  • Define domain-specific criteria
  • Set appropriate weights for your context
  • Establish meaningful thresholds
  • Document what each score range means

Detailed Resources

  • Scoring Patterns: See modules/scoring-patterns.md for detailed methodology
  • Decision Thresholds: See modules/decision-thresholds.md for threshold design

Exit Criteria

  • Criteria defined with clear descriptions
  • Weights assigned and sum to 1.0
  • Scoring guides documented for each criterion
  • Thresholds mapped to specific actions
  • Evaluation process documented and reproducible

© athola, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files in plugins/leyline/skills/evaluation-framework of athola/claude-night-market.

  • SKILL.md
  • README.md
  • modules/decision-thresholds.md
  • modules/evaluation-rubric.md
  • modules/multi-metric-evaluation-methodology.md
  • modules/quality-metrics.md
  • modules/scoring-patterns.md

Open the folder on GitHubat commit 9f3eb00

Compare with similar skills

Evaluation Framework next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation Framework compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation Framework this skillathola/claude-night-market342—~1.3kAutomated safety check: PassMIT
Evaluationguanyang/open-agent-hub9752 repos~4.2kAutomated safety check: PassMIT
Creating A Coral TaskHuman-Agent-Society/CORAL1k—~2.2kAutomated safety check: PassApache-2.0
Feature Plannerserendipity1004/cc-feature-implementer176—~2.4kAutomated safety check: PassNone
Ccg Workflowfengshao1227/ccg-workflow5.9k—~2.3kAutomated safety check: PassMIT
Conducty Checkpointrobertbarclayy/conducty176—~1.5kAutomated safety check: PassMIT

Similar skills

  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed
  • Creating A Coral Task

    Human-Agent-Society/CORAL

    Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…

    1k GitHub stars~2.2k tokensUpdated 29 days ago
    Testing & QAAuto-check passed
  • Feature Planner

    serendipity1004/cc-feature-implementer

    Creates phase-based feature plans with quality gates and incremental delivery structure.

    176 GitHub stars~2.4k tokensUpdated 9 mo ago
    Testing & QAAuto-check passed
  • Ccg Workflow

    fengshao1227/ccg-workflow

    How to run a non-trivial change end to end with the CCG role tools (ccganalyze / ccgdesign / ccgbuild / ccgdebug / ccgoptimize / ccgreview / ccgtest) and the verify- quality gates.

    5.9k GitHub stars~2.3k tokensUpdated 22 days ago
    Testing & QAAuto-check passed
  • Conducty Checkpoint

    robertbarclayy/conducty

    Quality gate between parallelization groups. An agent skill from robertbarclayy/conducty.

    176 GitHub stars~1.5k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed
  • Mission Planner

    jdforsythe/forge

    Decomposes goals into team blueprints using evidence-based scaling laws, topology selection, and role design.

    151 GitHub stars~3.5k tokensUpdated 3 mo ago
    Testing & QAAuto-check passed

More from athola/claude-night-market

All 159 skills in this repo
  • Night Market Diagnostics Toolkit

    athola/claude-night-market

    Run and interpret repo diagnostic scripts (ratchets, validators, token stats).

    342 GitHub stars~3.4k tokensUpdated 2 days ago
    Auto-check passed
  • Skills Eval

    athola/claude-night-market

    Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.

    342 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed
  • Agent Teams

    athola/claude-night-market

    Coordinates Claude agent teams via filesystem protocol. An agent skill from athola/claude-night-market.

    342 GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed
  • Delegation Core

    athola/claude-night-market

    Delegates execution to eight CLIs (Gemini, Qwen, MiniMax, GLM, Muse, Codex, OpenCode, Glimmer).

    342 GitHub stars~2.5k tokensUpdated 2 days ago
    Auto-check passed
  • Elegant Code

    athola/claude-night-market

    Guide minimal code via a decision ladder with full safety, edge, and negative-case coverage.

    342 GitHub stars~2.1k tokensUpdated 2 days ago
    Auto-check passed
  • Skill Library Mission

    athola/claude-night-market

    Build a project skill library in .claude/skills/ via discovery, parallel authoring, and review.

    342 GitHub stars~1.6k tokensUpdated 2 days ago
    Auto-check passed

Questions about Evaluation Framework

What does Evaluation Framework do?

Provides weighted scoring, rubrics, and decision-threshold patterns. Evaluation Framework is an agent skill from athola/claude-night-market. Provides weighted scoring, rubrics, and decision-threshold patterns.

When should I use Evaluation Framework?

Evaluation Framework fits situations like: designing quality gates; evaluation systems; decision frameworks.

How do I install Evaluation Framework in Claude Code?

Run `npx skills add athola/claude-night-market --skill evaluation-framework -a claude-code`. Or copy the skill folder (plugins/leyline/skills/evaluation-framework in athola/claude-night-market) into .claude/skills/evaluation-framework in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation Framework in Codex?

Run `npx skills add athola/claude-night-market --skill evaluation-framework -a codex`. Or copy the skill folder (plugins/leyline/skills/evaluation-framework in athola/claude-night-market) into .agents/skills/evaluation-framework in your project. Codex loads it when a task matches its description.

Can I use Evaluation Framework in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add athola/claude-night-market --skill evaluation-framework -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation-framework, .gemini/skills/evaluation-framework, .github/skills/evaluation-framework and .opencode/skills/evaluation-framework in your project.

What does Evaluation Framework need to run?

Going by SKILL.md and its folder, Evaluation Framework needs the command-line tools its instructions call (pytest). Our summary lists: Python 3.

Does Evaluation Framework access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Evaluation Framework safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluation Framework use?

Evaluation Framework is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluation Framework use?

About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Evaluation Framework?

Skills that share tags, products or a category with Evaluation Framework: Evaluation (guanyang/open-agent-hub, 975 stars), Creating A Coral Task (Human-Agent-Society/CORAL, 1k stars), Feature Planner (serendipity1004/cc-feature-implementer, 176 stars) and Ccg Workflow (fengshao1227/ccg-workflow, 5.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation Framework?

athola (a GitHub user) maintains it in athola/claude-night-market, which has 342 GitHub stars. The repository holds 159 skills in this directory. The repository was last updated on October 6, 2026.

Source: athola/claude-night-market on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.