Arize Evaluator
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports.
$ npx skills add Mathews-Tom/armory --skill package-evaluator -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Mathews-Tom/armory package-evaluator --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/package-evaluator .claude/skills/package-evaluator && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "package-evaluator" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/package-evaluator into .claude/skills/package-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "package-evaluator", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Mathews-Tom/armory/tree/main/skills/package-evaluatorType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Mathews-Tom/armory --skill package-evaluator -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Mathews-Tom/armory package-evaluator --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/package-evaluator .agents/skills/package-evaluator && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "package-evaluator" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/package-evaluator into .agents/skills/package-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "package-evaluator", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mathews-Tom/armory --skill package-evaluator -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Mathews-Tom/armory package-evaluator --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/package-evaluator .cursor/skills/package-evaluator && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "package-evaluator" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/package-evaluator into .cursor/skills/package-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "package-evaluator", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Mathews-Tom/armory.git --path skills/package-evaluator--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Mathews-Tom/armory --skill package-evaluator -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Mathews-Tom/armory package-evaluator --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/package-evaluator .gemini/skills/package-evaluator && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "package-evaluator" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/package-evaluator into .gemini/skills/package-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "package-evaluator", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Mathews-Tom/armory package-evaluatorInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Mathews-Tom/armory --skill package-evaluator -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/package-evaluator .github/skills/package-evaluator && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "package-evaluator" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/package-evaluator into .github/skills/package-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "package-evaluator", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Mathews-Tom/armory --skill package-evaluator -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Mathews-Tom/armory package-evaluator --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Mathews-Tom/armory.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/package-evaluator .opencode/skills/package-evaluator && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "package-evaluator" agent skill from https://github.com/Mathews-Tom/armory/tree/main/skills/package-evaluator into .opencode/skills/package-evaluator/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "package-evaluator", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
package-evaluatorEvaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports.
Package Evaluator is an agent skill from Mathews-Tom/armory. Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports. Triggers on: "evaluate package", "audit agent quality", "score this hook", "package audit", "skill quality check". NOT for LLM prompts, use prompt-lab.
Its SKILL.md is about 5.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `evals/cases.yaml` and `references/evaluation-rubric.md`).
The repository describes itself as: Curated, production-grade skills for AI coding agents. Battle-tested workflows for developers who use AI seriously. The licence is MIT.
4 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 4594fb7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Package Evaluator loads about 5.1k tokens when it runs, and up to ~10k if it reads all its reference files. Until then it costs about 71 tokens; SKILL.md has 2,006 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Mathews-Tom/armory at commit 4594fb7, republished under its MIT licence (© Mathews-Tom). 2,006 words, ~5,097 tokens.
.claude/skills/package-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.Packages that do not activate on relevant queries waste the entire investment in writing them. A skill can have deep, well-structured content and still deliver zero value if its frontmatter description lacks the trigger phrases users actually type. An agent without a decision tree produces inconsistent results. A hook without a handler script is inert. Quality evaluation catches trigger gaps, missing sections, structural deficiencies, and shallow content before deployment — turning a package from a static document into a reliable tool.
| File | Contents |
|---|---|
references/evaluation-rubric.md | Detailed 1-5 scoring criteria per dimension, weight justifications, type-specific criteria, worked examples for calibration |
Two modes, selected by input:
| Input | Mode |
|---|---|
| Path to a specific package directory or definition file | Quick Audit |
| "all", "every package", no path specified | Full Audit |
| "--type agents" or type filter | Full Audit filtered to one type |
| Multiple specific paths | Quick Audit for each, then comparative summary |
Six dimensions, each scored 1-5. Weighted sum determines overall percentage.
Evaluates the YAML frontmatter block for completeness and discoverability.
Signals:
name field present and non-emptydescription field present and non-emptyScoring constraints: A description under 100 characters caps this dimension at 2/5.
A missing name or description field caps at 1/5.
Evaluates whether the package activates on the queries users actually type.
Signals:
Scoring constraints: Fewer than 3 distinct trigger phrases caps at 2/5. Zero trigger phrases in the description caps at 1/5.
Evaluates whether the package contains the sections needed to function reliably.
Signals:
Scoring constraints: A package with no workflow section caps at 2/5. A package with a workflow but no error handling or output format caps at 3/5.
Evaluates the substantive quality of the package's guidance — whether it provides enough detail for an agent to execute well without human intervention.
Signals:
Scoring constraints: A package consisting only of bare commands with no explanatory context caps at 2/5. Reference files count toward this dimension only if they contain substantive guidance (checklists, rubrics, criteria), not just link collections.
Evaluates internal consistency and structural integrity.
Signals:
name field in frontmatter exactly../other-package/) in the definition
file or reference files. Packages must be standalone — all referenced files must live
within the package's own directory. Shared content should use the _templates/ sync
system to maintain local copies.Scoring constraints: A name mismatch between directory and frontmatter is a CRITICAL
finding and caps at 1/5. Cross-package ../ references are a CRITICAL finding and cap
at 1/5 — they break standalone packaging. Missing referenced files cap at 2/5.
Evaluates adherence to the repository's contribution guidelines.
Signals:
Scoring constraints: Any single violation caps at 3/5. Multiple violations cap at 2/5. Invalid YAML that prevents parsing caps at 1/5.
In addition to the 6 shared dimensions, each package type has type-specific quality signals that influence D3 (Structural Completeness) and D4 (Content Depth) scoring.
When evaluating an AGENT.md:
D3 additions:
model field specified (opus/sonnet/haiku)color field specifiedmetadata.category specifiedmetadata.execution_phase specified (pre-write/post-write/pre-commit)metadata.language_targets specifiedD4 additions:
Scoring impact: An agent without a decision tree caps D4 at 2/5. An agent without model/category/phase metadata caps D3 at 3/5.
When evaluating a HOOK.md:
D3 additions:
hook.events list specified (PreToolUse, PostToolUse, Stop, etc.)hook.handler.type specified (command/python-module)hook.handler.command specifiedhandler.sh or equivalent handler script exists in directoryD4 additions:
Scoring impact: A hook without a handler script caps D4 at 1/5. A hook without documented events caps D3 at 2/5.
When evaluating a RULE.md:
D3 additions:
metadata.scope specified (global/project)metadata.applies_to.languages specified if scope is language-specificD4 additions:
Scoring impact: A rule with only vague principles and no concrete requirements caps D4 at 2/5.
When evaluating a COMMAND.md:
D3 additions:
command.syntax specified with argument documentationcommand.handler specified (inline/command)D4 additions:
Scoring impact: A command without a step-by-step workflow caps D4 at 2/5.
When evaluating a UTILITY.md:
D3 additions:
utility.runtime specified (python/node/shell)utility.entry_point specified and the file existsutility.executable specifiedD4 additions:
Scoring impact: A utility without a working entry_point script caps D4 at 1/5.
When evaluating a PRESET.md:
D3 additions:
preset.packages specified with at least one type sectionpreset.compatibility.platforms specifiedD4 additions:
Scoring impact: A preset referencing non-existent packages is a CRITICAL finding.
| Severity | Criteria | Score Impact |
|---|---|---|
| CRITICAL | Package cannot activate or breaks on load — missing frontmatter, name mismatch, invalid YAML, non-existent preset references | Caps overall score at 40% |
| HIGH | Significant trigger gap or missing core section — no workflow, no error handling, zero trigger phrases, missing handler script | Caps affected dimension at 3/5 |
| MEDIUM | Weak coverage, shallow content, few trigger synonyms, missing type-specific metadata | Dimension needs improvement but functions |
| LOW | Minor polish — formatting inconsistencies, slightly short description, missing calibration rules | Fix when convenient |
SKILL.md = skill, AGENT.md = agent, HOOK.md = hook, RULE.md = rule,
COMMAND.md = command, UTILITY.md = utility, PRESET.md = preset.
If the path points to a definition file directly, use its parent directory.skills/, agents/, hooks/,
rules/, commands/, utilities/, and presets/ that contain a recognized
definition file. If a --type filter is specified, restrict to that type's directory.For each package under evaluation:
name and description fields. If YAML parsing
fails, record a CRITICAL finding and score D1 and D6 as 1/5.references/ directory for existence and contents. Verify every file
referenced in the definition file body exists on disk.../ patterns pointing outside the package directory). Flag any as CRITICAL D5
findings.references/evaluation-rubric.md.references/evaluation-rubric.md.Overall% = (sum of dimension_score x weight) / 5 x 100.| Range | Verdict |
|---|---|
| 90-100% | Exemplary |
| 80-89% | Strong |
| 70-79% | Adequate |
| 60-69% | Needs Work |
| Below 60% | Deficient |
Generate the structured output using the appropriate template below.
## Package Audit: {package-name} ({type})
| Dimension | Score | Weight | Weighted | Key Finding |
|-----------|-------|--------|----------|-------------|
| D1: Frontmatter Quality | X/5 | 20% | X.XXX | ... |
| D2: Trigger Coverage | X/5 | 18% | X.XXX | ... |
| D3: Structural Completeness | X/5 | 20% | X.XXX | ... |
| D4: Content Depth | X/5 | 22% | X.XXX | ... |
| D5: Consistency & Integrity | X/5 | 12% | X.XXX | ... |
| D6: CONTRIBUTING Compliance | X/5 | 8% | X.XXX | ... |
**Overall: XX% — {Verdict}**
### Type-Specific Findings
[Findings from type-specific signal evaluation, if any.]
### Findings
[Severity-sorted list. Each entry includes dimension tag, severity, description,
evidence, and recommendation.]
- **[CRITICAL] D5:** ...
- **[HIGH] D2:** ...
- **[MEDIUM] D4:** ...
- **[LOW] D3:** ...
### Score Calculation
D1: {score} x 0.20 = {result}
D2: {score} x 0.18 = {result}
D3: {score} x 0.20 = {result}
D4: {score} x 0.22 = {result}
D5: {score} x 0.12 = {result}
D6: {score} x 0.08 = {result}
Sum = {weighted_sum}
Overall = {weighted_sum} / 5 x 100 = {percentage}% — {Verdict}## Package Repository Audit
| Package | Type | Overall | Verdict | Worst Dimension | Top Issue |
|---------|------|---------|---------|-----------------|-----------|
| {name} | {type} | XX% | {verdict} | {dimension} | {issue} |
| ... | ... | ... | ... | ... | ... |
### Per-Package Summaries
[Condensed Quick Audit for each package: scorecard table, overall score, top 3 findings.
Omit the full Score Calculation section in condensed mode.]| Problem | Cause | Fix |
|---|---|---|
| Definition file not found in directory | Path incorrect or file missing | Report as CRITICAL; do not attempt evaluation; surface the path and stop |
| Unknown definition file | Directory has no recognized definition file (SKILL.md, AGENT.md, HOOK.md, RULE.md, COMMAND.md, UTILITY.md, PRESET.md) | Skip directory; report as warning in Full Audit; report as CRITICAL in Quick Audit |
| YAML frontmatter parse failure | Invalid YAML syntax (unclosed quotes, bad indentation) | Report as CRITICAL finding; score D1 and D6 as 1/5; continue evaluating the body content where parseable |
references/ directory missing | Package has no reference files | Not an error — score D5 normally; check only that any files referenced in the definition file body actually exist on disk |
references/ exists but referenced file is absent | File path in definition file body doesn't resolve | Record as a CRITICAL D5 finding; missing referenced files cap D5 at 2/5 |
| Empty definition file (zero bytes or whitespace only) | File created but never populated | Treat as CRITICAL; score all dimensions 1/5; overall verdict: Deficient |
references/evaluation-rubric.md not found | Evaluator's own reference file missing | Note the irony; evaluate using the criteria inline in this SKILL.md; flag D5 as a CRITICAL finding |
| Handler script missing for hook | HOOK.md references a handler that doesn't exist | Record as CRITICAL D4 finding; cap D4 at 1/5 |
| Preset references non-existent package | PRESET.md lists a package not in the manifest | Record as CRITICAL finding; caps overall at 40% |
© Mathews-Tom, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 2 other files (references) in skills/package-evaluator of Mathews-Tom/armory.
Open the folder on GitHubat commit 4594fb7
Package Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Package Evaluator this skillMathews-Tom/armory | 328 | — | ~5.1k | Automated safety check: Pass | MIT | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 2 repos | ~8.1k | Automated safety check: Notes | MIT | |
| LLM Evaluationdavila7/claude-code-templates | 32k | 13 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Agent Evaluationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Dimensionsthedaviddias/Front-End-Checklist | 74k | — | ~926 | Automated safety check: Pass | MIT | |
| EvaluatorsArize-ai/phoenix | 12k | — | ~1.7k | Automated safety check: Pass | Custom licence |
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing image assets, markup, and CDN or build transforms related to Set explicit width and height on images.
Arize-ai/phoenix
Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.
sickn33/agentic-awesome-skills
A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.
Mathews-Tom/armory
Architecture reviews across 7 dimensions (structural, scalability, enterprise readiness, performance, security, ops, data) with scored reports.
Mathews-Tom/armory
Turn concepts into static HTML visuals exported as PNG or SVG files via HTML/CSS/SVG.
Mathews-Tom/armory
A skill your agent uses when analyzing an existing video URL or local recording: "watch this video", "analyze youtube video", "summarize this video", "youtube transcript", "find this moment", "what…
Mathews-Tom/armory
Deep code simplification and refactoring preserving behavior across Python, Go, TypeScript, Rust.
Mathews-Tom/armory
Turn concepts into animated explainer videos using Manim (Python) with MP4/GIF output, audio overlay, multi-scene composition.
Mathews-Tom/armory
Maps the unresolved architecture, policy, and scope decisions that must be answered before planning can start: one durable decision ticket per question on the issue tracker, typed and blocker-linked…
Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports. Package Evaluator is an agent skill from Mathews-Tom/armory. Evaluates Claude Code package quality across 6 dimensions for all 7 package types, producing scored audit reports.
Package Evaluator fits situations like: : evaluate package; audit agent quality; score this hook; skill quality check.
Run `npx skills add Mathews-Tom/armory --skill package-evaluator -a claude-code`. Or copy the skill folder (skills/package-evaluator in Mathews-Tom/armory) into .claude/skills/package-evaluator in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Mathews-Tom/armory --skill package-evaluator -a codex`. Or copy the skill folder (skills/package-evaluator in Mathews-Tom/armory) into .agents/skills/package-evaluator in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Mathews-Tom/armory --skill package-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/package-evaluator, .gemini/skills/package-evaluator, .github/skills/package-evaluator and .opencode/skills/package-evaluator in your project.
SKILL.md names no scripts, command-line tools or credentials: Package Evaluator is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Package Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.1k tokens (SKILL.md is roughly 20k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.4k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Package Evaluator: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars) and Dimensions (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Mathews-Tom (a GitHub user) maintains it in Mathews-Tom/armory, which has 328 GitHub stars. The repository holds 80 skills in this directory. The repository was last updated on October 6, 2026.
Source: Mathews-Tom/armory on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.