Evaluation
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
Provides weighted scoring, rubrics, and decision-threshold patterns.
$ npx skills add athola/claude-night-market --skill evaluation-framework -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install athola/claude-night-market evaluation-framework --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/leyline/skills/evaluation-framework .claude/skills/evaluation-framework && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluation-framework" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/leyline/skills/evaluation-framework into .claude/skills/evaluation-framework/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-framework", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/athola/claude-night-market/tree/master/plugins/leyline/skills/evaluation-frameworkType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add athola/claude-night-market --skill evaluation-framework -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install athola/claude-night-market evaluation-framework --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/leyline/skills/evaluation-framework .agents/skills/evaluation-framework && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluation-framework" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/leyline/skills/evaluation-framework into .agents/skills/evaluation-framework/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-framework", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add athola/claude-night-market --skill evaluation-framework -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install athola/claude-night-market evaluation-framework --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/leyline/skills/evaluation-framework .cursor/skills/evaluation-framework && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluation-framework" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/leyline/skills/evaluation-framework into .cursor/skills/evaluation-framework/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-framework", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/athola/claude-night-market.git --path plugins/leyline/skills/evaluation-framework--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add athola/claude-night-market --skill evaluation-framework -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install athola/claude-night-market evaluation-framework --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/leyline/skills/evaluation-framework .gemini/skills/evaluation-framework && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluation-framework" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/leyline/skills/evaluation-framework into .gemini/skills/evaluation-framework/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-framework", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install athola/claude-night-market evaluation-frameworkInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add athola/claude-night-market --skill evaluation-framework -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/leyline/skills/evaluation-framework .github/skills/evaluation-framework && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluation-framework" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/leyline/skills/evaluation-framework into .github/skills/evaluation-framework/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-framework", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add athola/claude-night-market --skill evaluation-framework -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install athola/claude-night-market evaluation-framework --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/athola/claude-night-market.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/leyline/skills/evaluation-framework .opencode/skills/evaluation-framework && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluation-framework" agent skill from https://github.com/athola/claude-night-market/tree/master/plugins/leyline/skills/evaluation-framework into .opencode/skills/evaluation-framework/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-framework", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluation-frameworkProvides weighted scoring, rubrics, and decision-threshold patterns.
Evaluation Framework is an agent skill from athola/claude-night-market. Provides weighted scoring, rubrics, and decision-threshold patterns. Use when designing quality gates, evaluation systems, or decision frameworks.
Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files (for example `README.md`, `modules/decision-thresholds.md` and `modules/evaluation-rubric.md`).
It sits in Testing & QA, covering Quality gates and Quizzes and assessments. The repository describes itself as: 23 Claude Code plugins: TDD enforcement hooks, git/PR workflows, spec-driven development, code review, project lifecycle, fix-from-error, maintenance automation, context… The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 9f3eb00. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pytestFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluation Framework loads about 1.3k tokens when it runs. Until then it costs about 42 tokens; SKILL.md has 322 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from athola/claude-night-market at commit 9f3eb00, republished under its MIT licence (© athola). 322 words, ~1,251 tokens.
.claude/skills/evaluation-framework/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.A generic framework for weighted scoring and threshold-based decision making. Provides reusable patterns for evaluating any artifact against configurable criteria with consistent scoring methodology.
This framework abstracts the common pattern of: define criteria → assign weights → score against criteria → apply thresholds → make decisions.
criteria:
- name: criterion_name
weight: 0.30 # 30% of total score
description: What this measures
scoring_guide:
90-100: Exceptional
70-89: Strong
50-69: Acceptable
30-49: Weak
0-29: PoorVerification: Run the command with --help flag to verify availability.
scores = {
"criterion_1": 85, # Out of 100
"criterion_2": 92,
"criterion_3": 78,
}Verification: Run the command with --help flag to verify availability.
total = sum(score * weights[criterion] for criterion, score in scores.items())
# Example: (85 × 0.30) + (92 × 0.40) + (78 × 0.30) = 85.5Verification: Run the command with --help flag to verify availability.
thresholds:
80-100: Accept with priority
60-79: Accept with conditions
40-59: Review required
20-39: Reject with feedback
0-19: RejectVerification: Run the command with --help flag to verify availability.
criteria:
correctness: {weight: 0.40, description: Does code work as intended?}
maintainability: {weight: 0.25, description: Is it readable?}
performance: {weight: 0.20, description: Meets performance needs?}
testing: {weight: 0.15, description: Tests detailed?}
thresholds:
85-100: Approve immediately
70-84: Approve with minor feedback
50-69: Request changes
0-49: Reject, major issuesVerification: Run pytest -v to verify tests pass.
**Verification:** Run the command with `--help` flag to verify availability.
1. Review artifact against each criterion
2. Assign 0-100 score for each criterion
3. Calculate: total = Σ(score × weight)
4. Compare total to thresholds
5. Take action based on threshold rangeVerification: Run the command with --help flag to verify availability.
Quality Gates: Code review, PR approval, release readiness Content Evaluation: Document quality, knowledge intake, skill assessment Resource Allocation: Backlog prioritization, investment decisions, triage
# In your skill's frontmatter
dependencies: [leyline:evaluation-framework]Verification: Run the command with --help flag to verify availability.
Then customize the framework for your domain:
modules/scoring-patterns.md for detailed methodologymodules/decision-thresholds.md for threshold design© athola, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files in plugins/leyline/skills/evaluation-framework of athola/claude-night-market.
Open the folder on GitHubat commit 9f3eb00
Evaluation Framework next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluation Framework this skillathola/claude-night-market | 342 | — | ~1.3k | Automated safety check: Pass | MIT | |
| Evaluationguanyang/open-agent-hub | 975 | 2 repos | ~4.2k | Automated safety check: Pass | MIT | |
| Creating A Coral TaskHuman-Agent-Society/CORAL | 1k | — | ~2.2k | Automated safety check: Pass | Apache-2.0 | |
| Feature Plannerserendipity1004/cc-feature-implementer | 176 | — | ~2.4k | Automated safety check: Pass | None | |
| Ccg Workflowfengshao1227/ccg-workflow | 5.9k | — | ~2.3k | Automated safety check: Pass | MIT | |
| Conducty Checkpointrobertbarclayy/conducty | 176 | — | ~1.5k | Automated safety check: Pass | MIT |
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
Human-Agent-Society/CORAL
Author a new CORAL task — the three pieces that must line up (task.yaml, seed/, a packaged grader/), the coral init → coral validate → smoke-test loop, and how to pick a grader pattern (stdout…
serendipity1004/cc-feature-implementer
Creates phase-based feature plans with quality gates and incremental delivery structure.
fengshao1227/ccg-workflow
How to run a non-trivial change end to end with the CCG role tools (ccganalyze / ccgdesign / ccgbuild / ccgdebug / ccgoptimize / ccgreview / ccgtest) and the verify- quality gates.
robertbarclayy/conducty
Quality gate between parallelization groups. An agent skill from robertbarclayy/conducty.
jdforsythe/forge
Decomposes goals into team blueprints using evidence-based scaling laws, topology selection, and role design.
athola/claude-night-market
Run and interpret repo diagnostic scripts (ratchets, validators, token stats).
athola/claude-night-market
Evaluate Claude skill quality through auditing. An agent skill from athola/claude-night-market.
athola/claude-night-market
Coordinates Claude agent teams via filesystem protocol. An agent skill from athola/claude-night-market.
athola/claude-night-market
Delegates execution to eight CLIs (Gemini, Qwen, MiniMax, GLM, Muse, Codex, OpenCode, Glimmer).
athola/claude-night-market
Guide minimal code via a decision ladder with full safety, edge, and negative-case coverage.
athola/claude-night-market
Build a project skill library in .claude/skills/ via discovery, parallel authoring, and review.
Categories
Provides weighted scoring, rubrics, and decision-threshold patterns. Evaluation Framework is an agent skill from athola/claude-night-market. Provides weighted scoring, rubrics, and decision-threshold patterns.
Evaluation Framework fits situations like: designing quality gates; evaluation systems; decision frameworks.
Run `npx skills add athola/claude-night-market --skill evaluation-framework -a claude-code`. Or copy the skill folder (plugins/leyline/skills/evaluation-framework in athola/claude-night-market) into .claude/skills/evaluation-framework in your project. Claude Code loads it when a task matches its description.
Run `npx skills add athola/claude-night-market --skill evaluation-framework -a codex`. Or copy the skill folder (plugins/leyline/skills/evaluation-framework in athola/claude-night-market) into .agents/skills/evaluation-framework in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add athola/claude-night-market --skill evaluation-framework -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation-framework, .gemini/skills/evaluation-framework, .github/skills/evaluation-framework and .opencode/skills/evaluation-framework in your project.
Going by SKILL.md and its folder, Evaluation Framework needs the command-line tools its instructions call (pytest). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluation Framework is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.3k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Evaluation Framework: Evaluation (guanyang/open-agent-hub, 975 stars), Creating A Coral Task (Human-Agent-Society/CORAL, 1k stars), Feature Planner (serendipity1004/cc-feature-implementer, 176 stars) and Ccg Workflow (fengshao1227/ccg-workflow, 5.9k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
athola (a GitHub user) maintains it in athola/claude-night-market, which has 342 GitHub stars. The repository holds 159 skills in this directory. The repository was last updated on October 6, 2026.
Source: athola/claude-night-market on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.