Agent skill

Package Evaluator

by Mathews-Tom in Mathews-Tom/armory

Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports.

MITAuto-check passed

Install Package Evaluator

skills CLI
$ npx skills add Mathews-Tom/armory --skill package-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Mathews-Tom/armory package-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/package-evaluator .claude/skills/package-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
package-evaluator
GitHub stars
328
Token cost
~5.1k tokens
SKILL.md length
2,006 words
Files
3 (incl. references)
Skills in repo
80
Repo updated
First seen
Licence
MIT

At a glance

Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports.

  • Works in 4 steps: Input → Analysis → Scoring → …
  • : evaluate package
  • SKILL.md covers Reference Files, Audit Modes, Evaluation Dimensions and Type-Specific Dimensions, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Package Evaluator is an agent skill from Mathews-Tom/armory. Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports. Triggers on: "evaluate package", "audit agent quality", "score this hook", "package audit", "skill quality check". NOT for LLM prompts, use prompt-lab.

Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `evals/cases.yaml` and `references/evaluation-rubric.md`).

The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.

When your agent uses it

  • : evaluate package
  • Audit agent quality
  • Score this hook
  • Skill quality check

Example prompts

  • “evaluate package”
  • “audit agent quality”
  • “score this hook”
  • “/package-evaluator”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Input
  2. Analysis
  3. Scoring
  4. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Package Evaluator loads about 5.1k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 2,006 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~5.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~10k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 2,006 words, ~5,097 tokens.

Download SKILL.mdSave it as .claude/skills/package-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
package-evaluator
description
Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports. Triggers on: "evaluate package", "audit agent quality", "score this hook", "package audit", "skill quality check". NOT for LLM prompts, use prompt-lab.
metadata.version
1.3.1
metadata.category
review
metadata.tags
quality, audit, scoring, frontmatter
metadata.difficulty
intermediate
metadata.phase
review

Package Evaluator

Packages that do not activate on relevant queries waste the entire investment in writing them. A skill can have deep, well-structured content and still deliver zero value if its frontmatter description lacks the trigger phrases users actually type. An agent without a decision tree produces inconsistent results. A hook without a handler script is inert. Quality evaluation catches trigger gaps, missing sections, structural deficiencies, and shallow content before deployment — turning a package from a static document into a reliable tool.

Reference Files

FileContents
references/evaluation-rubric.mdDetailed 1-5 scoring criteria per dimension, weight justifications, type-specific criteria, worked examples for calibration

Audit Modes

Two modes, selected by input:

  • Quick Audit: Evaluate a single package. Produces a full per-dimension scored report with findings, severity classifications, and recommendations.
  • Full Audit: Evaluate all packages in the repository. Produces a comparative ranking table sorted by overall score, plus condensed per-package summaries. Optionally filtered to a single package type.
Mode Selection
InputMode
Path to a specific package directory or definition fileQuick Audit
"all", "every package", no path specifiedFull Audit
"--type agents" or type filterFull Audit filtered to one type
Multiple specific pathsQuick Audit for each, then comparative summary

Evaluation Dimensions

Six dimensions, each scored 1-5. Weighted sum determines overall percentage.

D1: Frontmatter Quality (20%)

Evaluates the YAML frontmatter block for completeness and discoverability.

Signals:

  • name field present and non-empty
  • description field present and non-empty
  • Description length between 200-800 characters (sweet spot for keyword density without bloat)
  • Description contains explicit trigger phrases users would type
  • Description includes a "Use this skill when..." clause or equivalent
  • Description is keyword-dense, not generic filler

Scoring constraints: A description under 100 characters caps this dimension at 2/5. A missing name or description field caps at 1/5.

D2: Trigger Coverage (18%)

Evaluates whether the package activates on the queries users actually type.

Signals:

  • Synonym breadth — multiple phrasings for the same intent (e.g., "review", "audit", "critique", "evaluate", "assess", "check")
  • Implied contexts — situations where the package applies even without explicit keywords (e.g., "user provides a design doc and asks for feedback")
  • Domain-specific terms relevant to the package's function
  • Explicit trigger phrase list in the description frontmatter
  • Coverage of both imperative ("review this") and interrogative ("is this good?") forms

Scoring constraints: Fewer than 3 distinct trigger phrases caps at 2/5. Zero trigger phrases in the description caps at 1/5.

D3: Structural Completeness (20%)

Evaluates whether the package contains the sections needed to function reliably.

Signals:

  • Prerequisites or setup instructions (if applicable)
  • Multi-phase workflow or step-by-step procedure
  • Error handling guidance or edge case documentation
  • Output format specification (template, example, or schema)
  • Limitations or scope boundaries stated
  • Reference file table (if references/ directory exists)
  • Calibration rules or quality gates

Scoring constraints: A package with no workflow section caps at 2/5. A package with a workflow but no error handling or output format caps at 3/5.

D4: Content Depth (22%)

Evaluates the substantive quality of the package's guidance — whether it provides enough detail for an agent to execute well without human intervention.

Signals:

  • Multi-step workflows with decision points, not bare command lists
  • Error cases documented with recovery actions
  • Decision frameworks (when to do X vs Y, mode selection tables)
  • Verbatim output examples or templates
  • Severity classifications or scoring rubrics (where applicable)
  • Cross-cutting analysis or synthesis steps beyond simple checklists

Scoring constraints: A package consisting only of bare commands with no explanatory context caps at 2/5. Reference files count toward this dimension only if they contain substantive guidance (checklists, rubrics, criteria), not just link collections.

D5: Consistency and Integrity (12%)

Evaluates internal consistency and structural integrity.

Signals:

  • Directory name matches the name field in frontmatter exactly
  • All files referenced in the definition file exist on disk (reference files, scripts, assets)
  • Description content aligns with body content (description does not promise features the body does not deliver)
  • Consistent terminology throughout (same concept uses same term)
  • No broken internal links or dangling references
  • Self-containment: No cross-package references (../other-package/) in the definition file or reference files. Packages must be standalone — all referenced files must live within the package's own directory. Shared content should use the _templates/ sync system to maintain local copies.

Scoring constraints: A name mismatch between directory and frontmatter is a CRITICAL finding and caps at 1/5. Cross-package ../ references are a CRITICAL finding and cap at 1/5 — they break standalone packaging. Missing referenced files cap at 2/5.

D6: CONTRIBUTING.md Compliance (8%)

Evaluates adherence to the repository's contribution guidelines.

Signals:

  • Package name is kebab-case
  • Package name is 64 characters or fewer
  • Description is 1024 characters or fewer
  • No angle brackets in description
  • No pushy trigger language in description ("always use", "you must", "never do")
  • Valid YAML frontmatter syntax

Scoring constraints: Any single violation caps at 3/5. Multiple violations cap at 2/5. Invalid YAML that prevents parsing caps at 1/5.


Type-Specific Dimensions

In addition to the 6 shared dimensions, each package type has type-specific quality signals that influence D3 (Structural Completeness) and D4 (Content Depth) scoring.

Agent-Specific Signals

When evaluating an AGENT.md:

D3 additions:

  • model field specified (opus/sonnet/haiku)
  • color field specified
  • metadata.category specified
  • metadata.execution_phase specified (pre-write/post-write/pre-commit)
  • metadata.language_targets specified

D4 additions:

  • Decision tree or algorithm section present with clear phases
  • START/END or phase markers for structured execution flow
  • Severity classification for findings (CRITICAL/HIGH/MEDIUM/LOW)
  • Language-specific patterns documented (not just generic advice)
  • VIOLATION/BLOCK/PASS outcome paths defined

Scoring impact: An agent without a decision tree caps D4 at 2/5. An agent without model/category/phase metadata caps D3 at 3/5.

Hook-Specific Signals

When evaluating a HOOK.md:

D3 additions:

  • hook.events list specified (PreToolUse, PostToolUse, Stop, etc.)
  • hook.handler.type specified (command/python-module)
  • hook.handler.command specified
  • handler.sh or equivalent handler script exists in directory

D4 additions:

  • Handler script is functional (not placeholder/stub)
  • Handler reads stdin JSON correctly
  • Exit code semantics documented (non-zero blocks for PreToolUse)
  • Edge cases documented (what happens on timeout, malformed input)

Scoring impact: A hook without a handler script caps D4 at 1/5. A hook without documented events caps D3 at 2/5.

Rule-Specific Signals

When evaluating a RULE.md:

D3 additions:

  • metadata.scope specified (global/project)
  • metadata.applies_to.languages specified if scope is language-specific

D4 additions:

  • Concrete, actionable requirements (not vague principles)
  • Code examples for key rules
  • Anti-patterns shown alongside correct patterns
  • Thresholds/limits specified with numbers (not "reasonable" or "appropriate")

Scoring impact: A rule with only vague principles and no concrete requirements caps D4 at 2/5.

Command-Specific Signals

When evaluating a COMMAND.md:

D3 additions:

  • command.syntax specified with argument documentation
  • command.handler specified (inline/command)

D4 additions:

  • Numbered step-by-step workflow
  • Decision points clearly marked
  • Output format specified
  • Termination criteria defined (when to stop the workflow)

Scoring impact: A command without a step-by-step workflow caps D4 at 2/5.

Utility-Specific Signals

When evaluating a UTILITY.md:

D3 additions:

  • utility.runtime specified (python/node/shell)
  • utility.entry_point specified and the file exists
  • utility.executable specified

D4 additions:

  • Entry point script uses argparse or equivalent for CLI
  • Script has proper error handling (not bare except)
  • Usage examples in the body
  • No external dependencies beyond stdlib (or dependencies documented)

Scoring impact: A utility without a working entry_point script caps D4 at 1/5.

Preset-Specific Signals

When evaluating a PRESET.md:

D3 additions:

  • preset.packages specified with at least one type section
  • Each referenced package exists in the manifest
  • preset.compatibility.platforms specified

D4 additions:

  • Body explains why these packages work together
  • Describes the target workflow or use case
  • References are to existing packages (not aspirational)

Scoring impact: A preset referencing non-existent packages is a CRITICAL finding.


Show full SKILL.md (799 more words)Show less

Severity Classification

SeverityCriteriaScore Impact
CRITICALPackage cannot activate or breaks on load — missing frontmatter, name mismatch, invalid YAML, non-existent preset referencesCaps overall score at 40%
HIGHSignificant trigger gap or missing core section — no workflow, no error handling, zero trigger phrases, missing handler scriptCaps affected dimension at 3/5
MEDIUMWeak coverage, shallow content, few trigger synonyms, missing type-specific metadataDimension needs improvement but functions
LOWMinor polish — formatting inconsistencies, slightly short description, missing calibration rulesFix when convenient

Workflow

Phase 1: Input
  1. Determine audit mode from user input (see Mode Selection table above).
  2. For Quick Audit: validate the package directory exists and contains a recognized definition file. Detect package type from the definition file name: SKILL.md = skill, AGENT.md = agent, HOOK.md = hook, RULE.md = rule, COMMAND.md = command, UTILITY.md = utility, PRESET.md = preset. If the path points to a definition file directly, use its parent directory.
  3. For Full Audit: enumerate all directories under skills/, agents/, hooks/, rules/, commands/, utilities/, and presets/ that contain a recognized definition file. If a --type filter is specified, restrict to that type's directory.
  4. For each package to evaluate, note the directory name for D5 consistency checks and the package type for type-specific signal evaluation.
Phase 2: Analysis

For each package under evaluation:

  1. Read the definition file in full.
  2. Parse YAML frontmatter — extract name and description fields. If YAML parsing fails, record a CRITICAL finding and score D1 and D6 as 1/5.
  3. Check the references/ directory for existence and contents. Verify every file referenced in the definition file body exists on disk.
  4. Scan the definition file and all reference files for cross-package path references (../ patterns pointing outside the package directory). Flag any as CRITICAL D5 findings.
  5. Identify the package type and evaluate type-specific signals (see Type-Specific Dimensions above). Apply type-specific scoring caps to D3 and D4 as documented.
  6. Evaluate each of the 6 dimensions using the shared criteria above, the type-specific signals, and the detailed rubric in references/evaluation-rubric.md.
  7. Record findings with severity, dimension tag, description, and recommendation.
Phase 3: Scoring
  1. Score each dimension 1-5 using references/evaluation-rubric.md.
  2. Apply severity caps: if any CRITICAL finding exists, cap overall at 40% regardless of dimension scores.
  3. Compute weighted score: Overall% = (sum of dimension_score x weight) / 5 x 100.
  4. Determine verdict from the scale below.
RangeVerdict
90-100%Exemplary
80-89%Strong
70-79%Adequate
60-69%Needs Work
Below 60%Deficient
Phase 4: Report

Generate the structured output using the appropriate template below.


Output Format

Quick Audit Template
text
## Package Audit: {package-name} ({type})

| Dimension | Score | Weight | Weighted | Key Finding |
|-----------|-------|--------|----------|-------------|
| D1: Frontmatter Quality | X/5 | 20% | X.XXX | ... |
| D2: Trigger Coverage | X/5 | 18% | X.XXX | ... |
| D3: Structural Completeness | X/5 | 20% | X.XXX | ... |
| D4: Content Depth | X/5 | 22% | X.XXX | ... |
| D5: Consistency & Integrity | X/5 | 12% | X.XXX | ... |
| D6: CONTRIBUTING Compliance | X/5 | 8% | X.XXX | ... |

**Overall: XX% — {Verdict}**

### Type-Specific Findings

[Findings from type-specific signal evaluation, if any.]

### Findings

[Severity-sorted list. Each entry includes dimension tag, severity, description,
evidence, and recommendation.]

- **[CRITICAL] D5:** ...
- **[HIGH] D2:** ...
- **[MEDIUM] D4:** ...
- **[LOW] D3:** ...

### Score Calculation

D1: {score} x 0.20 = {result}
D2: {score} x 0.18 = {result}
D3: {score} x 0.20 = {result}
D4: {score} x 0.22 = {result}
D5: {score} x 0.12 = {result}
D6: {score} x 0.08 = {result}
Sum = {weighted_sum}
Overall = {weighted_sum} / 5 x 100 = {percentage}% — {Verdict}
Full Audit Template
text
## Package Repository Audit

| Package | Type | Overall | Verdict | Worst Dimension | Top Issue |
|---------|------|---------|---------|-----------------|-----------|
| {name} | {type} | XX% | {verdict} | {dimension} | {issue} |
| ... | ... | ... | ... | ... | ... |

### Per-Package Summaries

[Condensed Quick Audit for each package: scorecard table, overall score, top 3 findings.
Omit the full Score Calculation section in condensed mode.]

Error Handling

ProblemCauseFix
Definition file not found in directoryPath incorrect or file missingReport as CRITICAL; do not attempt evaluation; surface the path and stop
Unknown definition fileDirectory has no recognized definition file (SKILL.md, AGENT.md, HOOK.md, RULE.md, COMMAND.md, UTILITY.md, PRESET.md)Skip directory; report as warning in Full Audit; report as CRITICAL in Quick Audit
YAML frontmatter parse failureInvalid YAML syntax (unclosed quotes, bad indentation)Report as CRITICAL finding; score D1 and D6 as 1/5; continue evaluating the body content where parseable
references/ directory missingPackage has no reference filesNot an error — score D5 normally; check only that any files referenced in the definition file body actually exist on disk
references/ exists but referenced file is absentFile path in definition file body doesn't resolveRecord as a CRITICAL D5 finding; missing referenced files cap D5 at 2/5
Empty definition file (zero bytes or whitespace only)File created but never populatedTreat as CRITICAL; score all dimensions 1/5; overall verdict: Deficient
references/evaluation-rubric.md not foundEvaluator's own reference file missingNote the irony; evaluate using the criteria inline in this SKILL.md; flag D5 as a CRITICAL finding
Handler script missing for hookHOOK.md references a handler that doesn't existRecord as CRITICAL D4 finding; cap D4 at 1/5
Preset references non-existent packagePRESET.md lists a package not in the manifestRecord as CRITICAL finding; caps overall at 40%

Calibration Rules

  1. Score what exists, not what could exist — evaluate the package as-is, not its potential.
  2. Weight trigger coverage heavily for packages targeting broad domains (e.g., a GitHub skill covers issues, PRs, CI, releases, and API — it needs proportionally more trigger synonyms).
  3. A package with strong triggers but shallow content scores higher than deep content with poor triggers — activation is prerequisite to utility.
  4. Reference files count toward Content Depth only if they contain substantive guidance (checklists, rubrics, criteria), not link lists or stub files.
  5. When evaluating the package-evaluator itself, apply identical standards — no self-inflation.
  6. Frontmatter description quality is the single highest-leverage improvement for any package.
  7. Score type-specific signals proportionally — a hook's handler quality matters more than a rule's handler quality (rules have no handlers). Apply type-specific caps only when the signal is relevant to that package type.

© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/package-evaluator of Mathews-Tom/armory.

  • SKILL.md
  • evals/cases.yaml
  • references/evaluation-rubric.md

Open the folder on GitHubat commit 4594fb7

Compare with similar skills

Package Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Package Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Package Evaluator this skillMathews-Tom/armory328—~5.1kAutomated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
Dimensionsthedaviddias/Front-End-Checklist74k—~926Automated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Dimensions

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing image assets, markup, and CDN or build transforms related to Set explicit width and height on images.

    74k GitHub stars~926 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed

More from Mathews-Tom/armory

All 80 skills in this repo
  • Architecture Reviewer

    Mathews-Tom/armory

    Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.

    328 GitHub stars~4.6k tokensUpdated 2 days ago
    Auto-check passed
  • Concept To Image

    Mathews-Tom/armory

    Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.

    328 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check passed
  • Watch

    Mathews-Tom/armory

    A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…

    328 GitHub stars~2.8k tokensUpdated 2 days ago
    Auto-check passed
  • Code Refiner

    Mathews-Tom/armory

    Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.

    328 GitHub stars~3.1k tokensUpdated 2 days ago
    Auto-check passed
  • Concept To Video

    Mathews-Tom/armory

    Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.

    328 GitHub stars~4.9k tokensUpdated 2 days ago
    Auto-check passed
  • Decision Map

    Mathews-Tom/armory

    Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…

    328 GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed

Questions about Package Evaluator

What does Package Evaluator do?

Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports. Package Evaluator is an agent skill from Mathews-Tom/armory. Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports.

When should I use Package Evaluator?

Package Evaluator fits situations like: : evaluate package; audit agent quality; score this hook; skill quality check.

How do I install Package Evaluator in Claude Code?

Run `npx skills add Mathews-Tom/armory --skill package-evaluator -a claude-code`. Or copy the skill folder (skills/package-evaluator in Mathews-Tom/armory) into .claude/skills/package-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Package Evaluator in Codex?

Run `npx skills add Mathews-Tom/armory --skill package-evaluator -a codex`. Or copy the skill folder (skills/package-evaluator in Mathews-Tom/armory) into .agents/skills/package-evaluator in your project. Codex loads it when a task matches its description.

Can I use Package Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill package-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/package-evaluator, .gemini/skills/package-evaluator, .github/skills/package-evaluator and .opencode/skills/package-evaluator in your project.

What does Package Evaluator need to run?

SKILL.md names no scripts, command-line tools or credentials: Package Evaluator is instructions for the agent only.

Does Package Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Package Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Package Evaluator use?

Package Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Package Evaluator use?

About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.

What are the alternatives to Package Evaluator?

Skills that share tags, products or a category with Package Evaluator: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Dimensions (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Package Evaluator?

Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.

Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.