Agent skill

Evaluation Methodology

by wshobson in wshobson/agents

PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas.

MITAuto-check passedEducation

Install Evaluation Methodology

skills CLI
$ npx skills add wshobson/agents --skill evaluation-methodology -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install wshobson/agents evaluation-methodology --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/plugin-eval/skills/evaluation-methodology .claude/skills/evaluation-methodology && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
evaluation-methodology
GitHub stars
40k
Token cost
~2k tokens
SKILL.md length
963 words
Files
3 (incl. references)
Skills in repo
142
Repo updated
First seen
Licence
MIT

At a glance

PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas.

  • Understanding how plugin quality is measured
  • SKILL.md covers Evaluation depths, Static layer (lint), LLM judge layer (experimental) and Monte Carlo layer (experimental), plus 6 more sections
  • Calls uv
  • Interpreting a low score on a specific dimension

What it does

Evaluation Methodology is an agent skill from wshobson/agents. PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas. Use this skill when understanding how plugin quality is measured, when interpreting a low score on a specific dimension, when deciding how to improve a skill's triggering accuracy or orchestration fitness, when setting score thresholds for your marketplace, or when explaining quality badges to external partners like Neon.

Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/improving-scores.md` and `references/rubrics.md`).

It sits in Education, covering Quizzes and assessments. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.

When your agent uses it

  • Understanding how plugin quality is measured
  • Interpreting a low score on a specific dimension
  • Deciding how to improve a skills triggering accuracy
  • Orchestration fitness

Example prompts

  • “/evaluation-methodology”

What it can do on your machine

Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Evaluation Methodology loads about 2k tokens when it runs, and up to ~9.1k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 963 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~108
When it runs · the whole SKILL.md, loaded when a task matches
~2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 963 words, ~1,977 tokens.

Download SKILL.mdSave it as .claude/skills/evaluation-methodology/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
evaluation-methodology
description
PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas. Use this skill when understanding how plugin quality is measured, when interpreting a low score on a specific dimension, when deciding how to improve a skill's triggering accuracy or orchestration fitness, when setting score thresholds for your marketplace, or when explaining quality badges to external partners like Neon.

Evaluation methodology

PluginEval scores a skill or a plugin from 0 to 100 by combining up to three layers. The static layer is a lint. It's fast and deterministic, and it's useful for checking structure. The LLM judge and Monte Carlo layers are experimental, not validated against human labels, so treat their numbers as rough signals. For the trace-based eval program, see evals/README.md at the repository root.

The judge rubric anchors are in references/rubrics.md. Fixes for each anti-pattern and tips for each dimension are in references/improving-scores.md.

Evaluation depths

DepthLayersConfidence label
quickstaticEstimated
standard (default for score)static and judgeAssessed
deep (used by certify)static, judge, and Monte Carlo with 50 runsCertified
thoroughstatic, judge, and Monte Carlo with 100 runsCertified+

The label names the depth that ran, not a check against human judgment. A plugin directory gets the static layer only, whatever depth you ask for.

Static layer (lint)

The static layer reads SKILL.md and makes no model calls. It computes seven sub-scores, and the first six feed composite dimensions:

  • frontmatter_quality (feeds triggering_accuracy)
  • orchestration_wiring (feeds orchestration_fitness)
  • progressive_disclosure, structural_completeness, token_efficiency, ecosystem_coherence
  • harness_portability

harness_portability maps to no dimension, so it doesn't change a skill's composite score. It does carry 6% of the static layer's own score. A plugin's score is built from that layer score, so portability findings can lower a plugin's score a little. Its findings are not counted as anti-patterns.

The static layer flags six anti-patterns: OVER_CONSTRAINED, EMPTY_DESCRIPTION, MISSING_TRIGGER, BLOATED_SKILL, ORPHAN_REFERENCE, and DEAD_CROSS_REF. Each flag cuts the score by 5%, down to a floor of 50%:

text
penalty = max(0.5, 1.0 - 0.05 * anti_pattern_count)

The report prints a severity for each flag, but the penalty counts flags and ignores severity.

LLM judge layer (experimental)

The judge layer makes four model calls and returns four holistic scores from 0 to 1:

  • triggering_accuracy: Haiku reads the description and writes 10 test prompts, 5 that should trigger and 5 that should not. It predicts the outcome for each prompt and reports its own F1. Nothing checks those predictions against real triggering.
  • orchestration_fitness and scope_calibration: Sonnet rates the skill on a five-point rubric.
  • output_quality: Sonnet imagines three tasks and rates the output it expects.

The three Sonnet calls see only the first 3,000 characters of SKILL.md. Only one judge runs, because nothing reads the judges setting.

Monte Carlo layer (experimental)

Haiku writes 15 prompts that should trigger the skill, and the layer repeats them to reach 50 runs (100 at thorough). Each run sends the SKILL.md text and one prompt to the model, and the layer records four measures:

  • Activation rate is the share of runs with any non-empty reply, so it shows whether the model answered, not whether the skill should have fired.
  • Output quality is reply length divided by 500, capped at 1.0.
  • Failure rate is the share of runs that errored.
  • Token efficiency is 1 - median_tokens / 8000.

Every prompt is one that should trigger, so the layer never checks that the skill stays out of unrelated requests. The layer's JSON includes Wilson, bootstrap, and Clopper-Pearson intervals for its own measures. The composite ci_lower and ci_upper fields are always null.

Show full SKILL.md (447 more words)Show less

Composite score

First, for each dimension, the engine blends the layer scores that exist, and it renormalizes the blend weights over those layers. Second, it sums the weighted dimension scores, and it renormalizes the dimension weights over the dimensions that have a score. Third, it multiplies the sum by 100 and by the anti-pattern penalty.

DimensionWeightStaticJudgeMonte Carlo
triggering_accuracy0.250.150.250.60
orchestration_fitness0.200.100.70none
output_quality0.15none0.400.60
scope_calibration0.12none0.55none
progressive_disclosure0.100.80nonenone
token_efficiency0.060.40none0.50
robustness0.05nonenone0.80
structural_completeness0.030.90nonenone
code_template_quality0.02nonenonenone
ecosystem_coherence0.020.85nonenone

A cell reads "none" when its layer produces no score for the dimension, even where LAYER_BLENDS lists a weight. No layer produces code_template_quality, so it's always unmeasured. For a plugin directory, the composite is the static layer's mean score across the plugin's skills and agents, times 100, times the penalty for the plugin's anti-pattern count.

Badges and grades

Badges come from the composite score alone. Platinum needs at least 90, Gold at least 80, Silver at least 70, and Bronze at least 60. Badge.from_scores accepts an Elo rating, but no command computes one. Plugin-level badges, including the ones in the weekly CI report, come from the static layer alone. Skill-level badges at standard depth or deeper also include the experimental layers.

Each measured dimension gets a letter grade on the 0 to 100 scale, from A+ at 97 down to D- at 60, and F below 60.

Usage

bash
plugin-eval score ./path/to/skill --depth quick     # static lint only
plugin-eval score ./path/to/skill                   # static and judge
plugin-eval certify ./path/to/skill                 # deep depth
plugin-eval compare ./skill-a ./skill-b             # quick depth by default
plugin-eval score ./path/to/skill --depth quick --output json --threshold 70

At standard depth or deeper, score, certify, and compare print a note on stderr that the judge and Monte Carlo layers are experimental. For a plugin directory, the CLI prints a warning that only the static layer runs instead. With --threshold, the command exits with code 1 when the composite is below the value. plugin-eval init writes a corpus index, but no other command reads it.

Examples

An abridged example of the JSON output follows. Scripts can read composite.score from it:

json
{
  "layers": [{"layer": "static", "score": 0.75, "sub_scores": {}, "anti_patterns": []}],
  "composite": {"score": 76.9, "ci_lower": null, "ci_upper": null, "badge": "silver",
                "confidence_label": "Estimated", "dimensions": []},
  "elo": null
}

Troubleshooting

  • When a score drops after you add content, check layers[0].anti_patterns in the JSON.
  • If triggering_accuracy is low at quick depth, add a trigger phrase such as "Use this skill when" to the description, followed by several comma-separated contexts.
  • Judge scores change between runs, because the model writes new test prompts and tasks each time. Use the static layer for comparisons you want to repeat.
  • If stderr says the judge could not measure some dimensions, install the LLM extra with uv sync --extra llm.

The eval-judge agent scores the four judge dimensions inside Claude Code, and the eval-orchestrator agent runs the CLI and merges the results.

© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in plugins/plugin-eval/skills/evaluation-methodology of wshobson/agents.

  • SKILL.md
  • references/improving-scores.md
  • references/rubrics.md

Open the folder on GitHubat commit 46891e7

Compare with similar skills

Evaluation Methodology next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Evaluation Methodology compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Evaluation Methodology this skillwshobson/agents40k—~2kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated today
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated yesterday
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated yesterday
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    975 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from wshobson/agents

All 142 skills in this repo
  • Billing Automation

    wshobson/agents

    Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.

    40k GitHub starsUsed in 14 repos~473 tokens
    Auto-check passed
  • Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.

    40k GitHub starsUsed in 14 repos~1.7k tokens
    Auto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    Auto-check passed
  • Portfolio Risk Metrics

    wshobson/agents

    Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.

    40k GitHub starsUsed in 13 repos~502 tokens
    Auto-check passed
  • Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.

    40k GitHub starsUsed in 12 repos~1.3k tokens
    Auto-check passed
  • Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.

    40k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed

Categories

Questions about Evaluation Methodology

What does Evaluation Methodology do?

PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas. Evaluation Methodology is an agent skill from wshobson/agents. PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas.

When should I use Evaluation Methodology?

Evaluation Methodology fits situations like: understanding how plugin quality is measured; interpreting a low score on a specific dimension; deciding how to improve a skills triggering accuracy; orchestration fitness.

How do I install Evaluation Methodology in Claude Code?

Run `npx skills add wshobson/agents --skill evaluation-methodology -a claude-code`. Or copy the skill folder (plugins/plugin-eval/skills/evaluation-methodology in wshobson/agents) into .claude/skills/evaluation-methodology in your project. Claude Code loads it when a task matches its description.

How do I install Evaluation Methodology in Codex?

Run `npx skills add wshobson/agents --skill evaluation-methodology -a codex`. Or copy the skill folder (plugins/plugin-eval/skills/evaluation-methodology in wshobson/agents) into .agents/skills/evaluation-methodology in your project. Codex loads it when a task matches its description.

Can I use Evaluation Methodology in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill evaluation-methodology -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation-methodology, .gemini/skills/evaluation-methodology, .github/skills/evaluation-methodology and .opencode/skills/evaluation-methodology in your project.

What does Evaluation Methodology need to run?

Going by SKILL.md and its folder, Evaluation Methodology needs the command-line tools its instructions call (uv).

Does Evaluation Methodology access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Evaluation Methodology safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Evaluation Methodology use?

Evaluation Methodology is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Evaluation Methodology use?

About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.1k tokens, read only when the agent opens those files.

What are the alternatives to Evaluation Methodology?

Skills that share tags, products or a category with Evaluation Methodology: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Evaluation Methodology?

wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.

Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.