DeepTutor CLI
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas.
$ npx skills add wshobson/agents --skill evaluation-methodology -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wshobson/agents evaluation-methodology --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/plugin-eval/skills/evaluation-methodology .claude/skills/evaluation-methodology && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "evaluation-methodology" agent skill from https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology into .claude/skills/evaluation-methodology/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-methodology", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodologyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wshobson/agents --skill evaluation-methodology -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wshobson/agents evaluation-methodology --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/plugin-eval/skills/evaluation-methodology .agents/skills/evaluation-methodology && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "evaluation-methodology" agent skill from https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology into .agents/skills/evaluation-methodology/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-methodology", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill evaluation-methodology -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wshobson/agents evaluation-methodology --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/plugin-eval/skills/evaluation-methodology .cursor/skills/evaluation-methodology && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "evaluation-methodology" agent skill from https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology into .cursor/skills/evaluation-methodology/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-methodology", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wshobson/agents.git --path plugins/plugin-eval/skills/evaluation-methodology--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wshobson/agents --skill evaluation-methodology -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wshobson/agents evaluation-methodology --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/plugin-eval/skills/evaluation-methodology .gemini/skills/evaluation-methodology && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "evaluation-methodology" agent skill from https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology into .gemini/skills/evaluation-methodology/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-methodology", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wshobson/agents evaluation-methodologyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wshobson/agents --skill evaluation-methodology -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/plugin-eval/skills/evaluation-methodology .github/skills/evaluation-methodology && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "evaluation-methodology" agent skill from https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology into .github/skills/evaluation-methodology/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-methodology", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wshobson/agents --skill evaluation-methodology -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wshobson/agents evaluation-methodology --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wshobson/agents.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/plugin-eval/skills/evaluation-methodology .opencode/skills/evaluation-methodology && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "evaluation-methodology" agent skill from https://github.com/wshobson/agents/tree/main/plugins/plugin-eval/skills/evaluation-methodology into .opencode/skills/evaluation-methodology/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "evaluation-methodology", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
evaluation-methodologyPluginEval quality methodology, covering dimensions, rubrics, and scoring formulas.
Evaluation Methodology is an agent skill from wshobson/agents. PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas. Use this skill when understanding how plugin quality is measured, when interpreting a low score on a specific dimension, when deciding how to improve a skill's triggering accuracy or orchestration fitness, when setting score thresholds for your marketplace, or when explaining quality badges to external partners like Neon.
Its SKILL.md is about 2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `references/improving-scores.md` and `references/rubrics.md`).
It sits in Education, covering Quizzes and assessments. The repository describes itself as: Multi-harness agentic plugin marketplace for Claude Code, Codex, Cursor, OpenCode, GitHub Copilot, Google Antigravity, and Pi. The licence is MIT.
Read from SKILL.md and the folder at commit 46891e7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
uvFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Evaluation Methodology loads about 2k tokens when it runs, and up to ~9.1k if it reads all its reference files. Until then it costs about 108 tokens; SKILL.md has 963 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wshobson/agents at commit 46891e7, republished under its MIT licence (© wshobson). 963 words, ~1,977 tokens.
.claude/skills/evaluation-methodology/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.PluginEval scores a skill or a plugin from 0 to 100 by combining up to three layers. The static
layer is a lint. It's fast and deterministic, and it's useful for checking structure. The LLM
judge and Monte Carlo layers are experimental, not validated against human labels, so treat their
numbers as rough signals. For the trace-based eval program, see evals/README.md at the
repository root.
The judge rubric anchors are in references/rubrics.md. Fixes for each anti-pattern and tips for each dimension are in references/improving-scores.md.
| Depth | Layers | Confidence label |
|---|---|---|
quick | static | Estimated |
standard (default for score) | static and judge | Assessed |
deep (used by certify) | static, judge, and Monte Carlo with 50 runs | Certified |
thorough | static, judge, and Monte Carlo with 100 runs | Certified+ |
The label names the depth that ran, not a check against human judgment. A plugin directory gets the static layer only, whatever depth you ask for.
The static layer reads SKILL.md and makes no model calls. It computes seven sub-scores, and the first six feed composite dimensions:
frontmatter_quality (feeds triggering_accuracy)orchestration_wiring (feeds orchestration_fitness)progressive_disclosure, structural_completeness, token_efficiency, ecosystem_coherenceharness_portabilityharness_portability maps to no dimension, so it doesn't change a skill's composite score. It
does carry 6% of the static layer's own score. A plugin's score is built from that layer score,
so portability findings can lower a plugin's score a little. Its findings are not counted as
anti-patterns.
The static layer flags six anti-patterns: OVER_CONSTRAINED, EMPTY_DESCRIPTION,
MISSING_TRIGGER, BLOATED_SKILL, ORPHAN_REFERENCE, and DEAD_CROSS_REF. Each flag cuts the
score by 5%, down to a floor of 50%:
penalty = max(0.5, 1.0 - 0.05 * anti_pattern_count)The report prints a severity for each flag, but the penalty counts flags and ignores severity.
The judge layer makes four model calls and returns four holistic scores from 0 to 1:
triggering_accuracy: Haiku reads the description and writes 10 test prompts, 5 that should
trigger and 5 that should not. It predicts the outcome for each prompt and reports its own F1.
Nothing checks those predictions against real triggering.orchestration_fitness and scope_calibration: Sonnet rates the skill on a five-point rubric.output_quality: Sonnet imagines three tasks and rates the output it expects.The three Sonnet calls see only the first 3,000 characters of SKILL.md. Only one judge runs,
because nothing reads the judges setting.
Haiku writes 15 prompts that should trigger the skill, and the layer repeats them to reach 50
runs (100 at thorough). Each run sends the SKILL.md text and one prompt to the model, and the
layer records four measures:
1 - median_tokens / 8000.Every prompt is one that should trigger, so the layer never checks that the skill stays out of
unrelated requests. The layer's JSON includes Wilson, bootstrap, and Clopper-Pearson intervals
for its own measures. The composite ci_lower and ci_upper fields are always null.
First, for each dimension, the engine blends the layer scores that exist, and it renormalizes the blend weights over those layers. Second, it sums the weighted dimension scores, and it renormalizes the dimension weights over the dimensions that have a score. Third, it multiplies the sum by 100 and by the anti-pattern penalty.
| Dimension | Weight | Static | Judge | Monte Carlo |
|---|---|---|---|---|
triggering_accuracy | 0.25 | 0.15 | 0.25 | 0.60 |
orchestration_fitness | 0.20 | 0.10 | 0.70 | none |
output_quality | 0.15 | none | 0.40 | 0.60 |
scope_calibration | 0.12 | none | 0.55 | none |
progressive_disclosure | 0.10 | 0.80 | none | none |
token_efficiency | 0.06 | 0.40 | none | 0.50 |
robustness | 0.05 | none | none | 0.80 |
structural_completeness | 0.03 | 0.90 | none | none |
code_template_quality | 0.02 | none | none | none |
ecosystem_coherence | 0.02 | 0.85 | none | none |
A cell reads "none" when its layer produces no score for the dimension, even where
LAYER_BLENDS lists a weight. No layer produces code_template_quality, so it's always
unmeasured. For a plugin directory, the composite is the static layer's mean score across the
plugin's skills and agents, times 100, times the penalty for the plugin's anti-pattern count.
Badges come from the composite score alone. Platinum needs at least 90, Gold at least 80, Silver
at least 70, and Bronze at least 60. Badge.from_scores accepts an Elo rating, but no command
computes one. Plugin-level badges, including the ones in the weekly CI report, come from the
static layer alone. Skill-level badges at standard depth or deeper also include the
experimental layers.
Each measured dimension gets a letter grade on the 0 to 100 scale, from A+ at 97 down to D- at 60, and F below 60.
plugin-eval score ./path/to/skill --depth quick # static lint only
plugin-eval score ./path/to/skill # static and judge
plugin-eval certify ./path/to/skill # deep depth
plugin-eval compare ./skill-a ./skill-b # quick depth by default
plugin-eval score ./path/to/skill --depth quick --output json --threshold 70At standard depth or deeper, score, certify, and compare print a note on stderr that the
judge and Monte Carlo layers are experimental. For a plugin directory, the CLI prints a warning
that only the static layer runs instead. With --threshold, the command exits with code 1 when the
composite is below the value. plugin-eval init writes a corpus index, but no other command
reads it.
An abridged example of the JSON output follows. Scripts can read composite.score from it:
{
"layers": [{"layer": "static", "score": 0.75, "sub_scores": {}, "anti_patterns": []}],
"composite": {"score": 76.9, "ci_lower": null, "ci_upper": null, "badge": "silver",
"confidence_label": "Estimated", "dimensions": []},
"elo": null
}layers[0].anti_patterns in the JSON.triggering_accuracy is low at quick depth, add a trigger phrase such as "Use this skill
when" to the description, followed by several comma-separated contexts.uv sync --extra llm.The eval-judge agent scores the four judge dimensions inside Claude Code, and the
eval-orchestrator agent runs the CLI and merges the results.
© wshobson, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in plugins/plugin-eval/skills/evaluation-methodology of wshobson/agents.
Open the folder on GitHubat commit 46891e7
Evaluation Methodology next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Evaluation Methodology this skillwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| DeepTutor CLIHKUDS/DeepTutor | 41k | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | |
| AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2k | Automated safety check: Pass | MIT | |
| Codebase to Coursezarazhangrui/codebase-to-course | 5.7k | — | ~4.4k | Automated safety check: Pass | None | |
| AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch | 66k | — | ~2.1k | Automated safety check: Pass | MIT | |
| Scholar EvaluationK-Dense-AI/claude-scientific-writer | 2.4k | 2 repos | ~2.9k | Automated safety check: Notes | MIT |
HKUDS/DeepTutor
Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.
rohitg00/ai-engineering-from-scratch
Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.
zarazhangrui/codebase-to-course
Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.
rohitg00/ai-engineering-from-scratch
Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.
K-Dense-AI/claude-scientific-writer
Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.
guanyang/open-agent-hub
This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…
wshobson/agents
Covers building subscription billing: billing cycles, subscription states, invoice generation, proration, tax handling and dunning for failed payments.
wshobson/agents
Cuts cloud spend across AWS, Azure, GCP and OCI with cost tagging, rightsizing, commitment and spot pricing models, and architecture changes.
wshobson/agents
Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.
wshobson/agents
Covers portfolio risk measurement with VaR, CVaR, Sharpe, Sortino and drawdown, plus guidance on limits, stress tests and tail risk.
wshobson/agents
Writes unit tests for shell scripts with Bats: error-condition tests, fixtures and mocks, cross-shell checks, parallel runs, helper files and CI integration.
wshobson/agents
Plans memory headroom, works through out-of-memory failures and watches temperature and power during long ML training jobs on NVIDIA DGX Spark.
Categories
PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas. Evaluation Methodology is an agent skill from wshobson/agents. PluginEval quality methodology, covering dimensions, rubrics, and scoring formulas.
Evaluation Methodology fits situations like: understanding how plugin quality is measured; interpreting a low score on a specific dimension; deciding how to improve a skills triggering accuracy; orchestration fitness.
Run `npx skills add wshobson/agents --skill evaluation-methodology -a claude-code`. Or copy the skill folder (plugins/plugin-eval/skills/evaluation-methodology in wshobson/agents) into .claude/skills/evaluation-methodology in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wshobson/agents --skill evaluation-methodology -a codex`. Or copy the skill folder (plugins/plugin-eval/skills/evaluation-methodology in wshobson/agents) into .agents/skills/evaluation-methodology in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wshobson/agents --skill evaluation-methodology -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/evaluation-methodology, .gemini/skills/evaluation-methodology, .github/skills/evaluation-methodology and .opencode/skills/evaluation-methodology in your project.
Going by SKILL.md and its folder, Evaluation Methodology needs the command-line tools its instructions call (uv).
SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Evaluation Methodology is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2k tokens (SKILL.md is roughly 7.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.1k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Evaluation Methodology: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wshobson (a GitHub user) maintains it in wshobson/agents, which has 40,287 GitHub stars. The repository holds 142 skills in this directory. The repository was last updated on October 5, 2026.
Source: wshobson/agents on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.