Agent skill

Agent Plugin Eval

by fabricioctelles in fabricioctelles/skills

Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification.

Apache-2.0Auto-check passedAgent Workflows

Install Agent Plugin Eval

skills CLI
$ npx skills add fabricioctelles/skills --skill agent-plugin-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fabricioctelles/skills agent-plugin-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/agent-plugin-eval .claude/skills/agent-plugin-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-plugin-eval
GitHub stars
106
Token cost
~2.1k tokens
SKILL.md length
959 words
Files
9 (incl. scripts, references)
Skills in repo
16
Repo updated
First seen
Licence
Apache-2.0

At a glance

Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification.

  • Works in 8 steps: Resolve the plugin root. Use a local… → Load the governing rules. Read… → Inventory every package path. Include… → …
  • Asked to review a plugin repo
  • SKILL.md covers Parameters, Safety boundary, Workflow and Gates and scoring, plus 3 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Agent Plugin Eval is an agent skill from fabricioctelles/skills. Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification. Use when asked to review a plugin repo, check plugin.json or mcp.json conformance, assess bundled skills and MCP servers, produce an evidence-cited 0–100 plugin scorecard, identify release blockers, or compare two agent plugins side by side.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 11 other files, including scripts and reference files (for example `agents/default.yaml`, `references/jev-integration.md` and `references/output-template.md`).

It sits in Agent Workflows, covering Hooks and plugins. It works with Model Context Protocol and Git. The repository describes itself as: A collection of skills for AI agents (Kiro, Cursor, Windsurf, Claude Code, and others). Each skill is a reusable module that teaches the agent to perform complex tasks with… The licence is Apache-2.0.

When your agent uses it

  • Asked to review a plugin repo
  • Check plugin.json
  • Mcp.json conformance
  • Assess bundled skills and MCP servers

Example prompts

  • “/agent-plugin-eval”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Resolve the plugin root. Use a local target in place. For a Git URL,
  2. Load the governing rules. Read references/spec-checklist.md and
  3. Inventory every package path. Include dotfiles, symlinks, immediate
  4. Run the deterministic scan. Execute
  5. Review components completely. Inspect every immediate
  6. Classify conformance before scoring. Use the exact failure boundaries in
  7. Score with cite-or-cut. Score all applicable rubric criteria from
  8. Answer in the requested language. Read

What it can do on your machine

Read from SKILL.md and the folder at commit f1de632. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 3 files in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Plugin Eval loads about 2.1k tokens when it runs, and up to ~7.4k if it reads all its reference files. Until then it costs about 96 tokens; SKILL.md has 959 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~7.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from fabricioctelles/skills at commit f1de632, republished under its Apache-2.0 licence (© fabricioctelles). 959 words, ~2,109 tokens.

Download SKILL.mdSave it as .claude/skills/agent-plugin-eval/SKILL.md (or your agent's skills folder). This skill also uses 8 other files; get the full folder from GitHub.
name
agent-plugin-eval
description
Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification. Use when asked to review a plugin repo, check plugin.json or mcp.json conformance, assess bundled skills and MCP servers, produce an evidence-cited 0–100 plugin scorecard, identify release blockers, or compare two agent plugins side by side.

Agent Plugin Evaluation

Treat the portable Agent Plugins specification as the authority. A client-native manifest (e.g., .codex-plugin/plugin.json, .claude/settings.json, .cursor/mcp.json) does not replace the required root plugin.json.

Parameters

ParameterDescriptionDefault
targetLocal repository/plugin path or Git URLAsk if missing
compareOptional second path or Git URLNone
outputScorecard destinationReply only; write only when requested
spec_versionAgent Plugins version to evaluateVersion declared by plugin.json, or 1.0.0

Safety boundary

Audit untrusted repositories statically. Do not run bundled executables, hooks, install scripts, package managers, MCP servers, or networked tests unless the user explicitly authorizes execution. Redact suspected secret values; report only their location and kind. A secret-like key or value is a suspicion, not confirmation: do not assign the FAIL gate without corroborating evidence such as a recognized live credential format, a trusted secret scanner, repository history/provenance, or user confirmation. Never test a credential against a service merely to confirm it.

Workflow

  1. Resolve the plugin root. Use a local target in place. For a Git URL, shallow-clone into a mktemp -d directory. A plugin root contains root plugin.json; if a repo has zero or multiple candidates, report the ambiguity instead of guessing. Done when every target maps to one explicit plugin root.
  2. Load the governing rules. Read references/spec-checklist.md and references/rubric.md. For Agent Plugins 1.0.0, use the bundled snapshot. For another declared version, or when the user asks for the latest spec, browse the canonical specification and schemas at agent-plugins.org and record the evaluated version and retrieval date. The normative text wins if it conflicts with JSON Schema.
  3. Inventory every package path. Include dotfiles, symlinks, immediate skill children, extension namespaces, executable files, and files ignored by Git. Resolve every symlink and package-relative path against the plugin root. Done when every discovered path is accounted for as portable core, client extension, supporting file, or containment violation.
  4. Run the deterministic scan. Execute python3 scripts/inspect_plugin.py <plugin-root> --json. Treat its output as evidence leads, not the final judgment. Confirm each reported issue in the source and add file:line or JSON-pointer evidence. Never weaken a normative finding merely because a client happens to accept it.
  5. Review components completely. Inspect every immediate skills/*/SKILL.md and every mcpServers entry. Validate Agent Skills against their own specification. Assess instructions, resources, scripts, MCP configuration, extension isolation, cohesion, and practical utility. If skill-evaluation is available, it may deepen individual skill-quality analysis, but it never replaces this plugin-level rubric.
  6. Classify conformance before scoring. Use the exact failure boundaries in references/spec-checklist.md: PASS, PARTIAL, or FAIL. Keep client compatibility separate from portable conformance. A client-specific feature may be excellent for that client and still add zero portable coverage.
  7. Score with cite-or-cut. Score all applicable rubric criteria from references/rubric.md. Every score needs specific evidence; every N/A needs a reason. Run scripts/score.py for the weighted result and gate cap; do not calculate it by hand. Done when all criteria and all findings are reconciled with the conformance status.
  8. Answer in the requested language. Read references/output-template.md and emit that structure. Lead with the verdict, distinguish blockers from recommendations, and provide concrete fixes. When compare is set, evaluate both independently before computing deltas; never force the same N/A set on both plugins.

Gates and scoring

  • PASS: no normative violation found; no score cap.
  • PARTIAL: non-fatal manifest deviation or invalid/skipped component; final score capped at 59.
  • FAIL: fatal manifest/package-root failure, root-manifest escape, or confirmed embedded credential; final score capped at 39.
  • Keep the uncapped score visible so authors can distinguish design quality from release-blocking conformance.

Invoke the calculator with one criterion:score:weight triple per criterion:

bash
python3 scripts/score.py --gate partial 1:90:3 2:80:3 3:NA:2
Show full SKILL.md (366 more words)Show less

Evaluation with Jev (Optional)

When TypeSafe Jev is available, use it for subjective quality criteria. Jev provides calibrated probability judgments that augment the deterministic checks.

When to use Jev
Evaluation TypeUse Jev?Method
Axes 1-3 conformanceNoDeterministic (inspect_plugin.py)
Axis 4 product qualityYesScore (UX, docs, errors)
Axis 2 quality criteriaYesScore (schema design, naming)
Gate classificationYesNoul (pass/fail categories)
Secret detectionYesNoul (suspected/not_suspected)
Quality checklistYesNoul (present/missing)
Discovery protocol
python
from typesafe import jev_available

if jev_available():
    from typesafe import Score, Noul
    # Use Jev for subjective criteria
else:
    # Fall back to heuristic scoring
Questions and integration

Questions are defined in scripts/jev_questions.json:

  • 7 Score questions: Axis 4 (UX coherence, documentation clarity, error handling) and Axis 2 (validation, schema design, naming, API elegance)
  • 20 Noul questions: Gates (G1-G4), secrets (4), quality checklist (7), component validity (4)

Score results (0.0-1.0) are averaged per axis and scaled to the rubric (0-25). Noul results provide categorical classifications for gates and checklists.

See references/jev-integration.md for full integration patterns and code examples.

Output format

When Jev is used, the scorecard includes a jev section:

json
{
  "jev": {
    "available": true,
    "quality_scores": { "ux_coherence": 0.72, ... },
    "gate_classifications": { "G1": {"label": "conformant", "passed": true} },
    "secret_findings": { "requires_review": false }
  }
}

Gotchas

  • The v1 portable core contains exactly Agent Skills and MCP servers. Hooks, commands, agents, apps, marketplaces, and distribution policy are client-specific unless placed in a valid extension namespace.
  • Missing optional skills/ or mcp.json is not an error. A present path of the wrong filesystem kind is an invalid component type.
  • Unknown root manifest fields are schema violations but have the spec's narrow non-fatal handling; most other manifest schema violations reject the whole plugin.
  • One invalid skill or MCP server must not be reported as if every independent component were invalid.
  • ${PLUGIN_ROOT} and ${PLUGIN_DATA} expand only in MCP args, env values, and cwd; never in command, URLs, or headers.
  • A high-quality client-native plugin can still fail the portable standard when root plugin.json is absent. Report both facts without averaging them away.
  • Keep possible credentials labeled “suspected” and redacted. A heuristic hit alone lowers the security score and demands remediation review, but does not become a confirmed-credential FAIL gate.

Final quality gate

  • Every target resolved to exactly one root
  • Every file, symlink, skill, MCP server, and extension inspected
  • Every normative violation mapped to its correct failure boundary
  • Every score cited and every N/A justified
  • Suspected secrets redacted
  • Score produced by scripts/score.py
  • Comparison deltas use independently computed scores

© fabricioctelles, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 8 other files (scripts, references) in skills/agent-plugin-eval of fabricioctelles/skills.

  • SKILL.md
  • agents/default.yaml
  • references/jev-integration.md
  • references/output-template.md
  • references/rubric.md
  • references/spec-checklist.md
  • scripts/inspect_plugin.py
  • scripts/jev_questions.json
  • scripts/score.py

Open the folder on GitHubat commit f1de632

Compare with similar skills

Agent Plugin Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Plugin Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Plugin Eval this skillfabricioctelles/skills106—~2.1kAutomated safety check: PassApache-2.0
MCP Integration for Pluginsanthropics/claude-plugins-official38k11 repos~3.1kAutomated safety check: PassApache-2.0
Crush Configurationcharmbracelet/crush29k—~3.7kAutomated safety check: PassCustom licence
Agent Repo Initstudy8677/repobrain1.3k1 repos~404Automated safety check: NotesMIT
OpenpetsOpenPetsHQ/openpets1.3k—~2.1kAutomated safety check: PassMIT
Claude Automation Recommenderanthropics/claude-plugins-official38k3 repos~2.7kAutomated safety check: NotesApache-2.0

Similar skills

  • MCP Integration for Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.

    38k GitHub starsUsed in 11 repos~3.1k tokens
    Agent WorkflowsAuto-check passed
  • Crush Configuration

    charmbracelet/crush

    Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.

    29k GitHub stars~3.7k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Agent Repo Init

    study8677/repobrain

    One-click initialization of a multi-agent repository from the RepoBrain template.

    1.3k GitHub starsUsed in 1 repo~404 tokens
    Agent WorkflowsAuto-check: notes
  • Openpets

    OpenPetsHQ/openpets

    A skill your agent uses whenever the user wants to build, extend, debug, test, validate, locally load, package, or publish an OpenPets plugin; work with the OpenPets Plugin SDK v3, plugin manifest…

    1.3k GitHub stars~2.1k tokensUpdated 11 days ago
    Agent WorkflowsAuto-check passed
  • Claude Automation Recommender

    anthropics/claude-plugins-official

    Official

    Scans a codebase and suggests which Claude Code hooks, subagents, skills, plugins and MCP servers fit its stack, without changing any files.

    38k GitHub starsUsed in 3 repos~2.7k tokens
    Agent WorkflowsAuto-check: notes
  • Compound Engineering Setup

    EveryInc/compound-engineering-plugin

    Checks Compound Engineering plugin health and repo-local config, or scaffolds a Compound Pack when you ask for one by id.

    25k GitHub stars~2k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from fabricioctelles/skills

All 16 skills in this repo
  • Motion Ad

    fabricioctelles/skills

    Produce a short motion-graphics video ad — a 15s Facebook/Instagram/TikTok spot — as a rendered MP4.

    106 GitHub stars~4.1k tokensUpdated 4 days ago
    Auto-check passed
  • Loop Architect

    fabricioctelles/skills

    Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them.

    106 GitHub stars~2.1k tokensUpdated 4 days ago
    Auto-check: notes
  • Ralph Loop Kiro Specs

    fabricioctelles/skills

    Automated iterative agent runner for spec-based development in Kiro.

    106 GitHub stars~2.6k tokensUpdated 4 days ago
    Auto-check passed
  • Security Specialist

    fabricioctelles/skills

    Runs security audits on codebases — full scans, diff reviews, threat models, vulnerability triage, remediation guidance, and finding tracking.

    106 GitHub stars~2.8k tokensUpdated 4 days ago
    Auto-check passed
  • Skill Evaluation

    fabricioctelles/skills

    Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…

    106 GitHub stars~3.8k tokensUpdated 4 days ago
    Auto-check passed
  • Okf Open Knowledge Format

    fabricioctelles/skills

    Create, validate, and enrich Open Knowledge Format (OKF) bundles — the open spec for representing organizational knowledge as markdown files with YAML frontmatter.

    106 GitHub stars~5.7k tokensUpdated 4 days ago
    Auto-check passed

Categories

Questions about Agent Plugin Eval

What does Agent Plugin Eval do?

Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification. Agent Plugin Eval is an agent skill from fabricioctelles/skills. Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification.

When should I use Agent Plugin Eval?

Agent Plugin Eval fits situations like: asked to review a plugin repo; check plugin.json; mcp.json conformance; assess bundled skills and MCP servers.

How do I install Agent Plugin Eval in Claude Code?

Run `npx skills add fabricioctelles/skills --skill agent-plugin-eval -a claude-code`. Or copy the skill folder (skills/agent-plugin-eval in fabricioctelles/skills) into .claude/skills/agent-plugin-eval in your project. Claude Code loads it when a task matches its description.

How do I install Agent Plugin Eval in Codex?

Run `npx skills add fabricioctelles/skills --skill agent-plugin-eval -a codex`. Or copy the skill folder (skills/agent-plugin-eval in fabricioctelles/skills) into .agents/skills/agent-plugin-eval in your project. Codex loads it when a task matches its description.

Can I use Agent Plugin Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fabricioctelles/skills --skill agent-plugin-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-plugin-eval, .gemini/skills/agent-plugin-eval, .github/skills/agent-plugin-eval and .opencode/skills/agent-plugin-eval in your project.

What does Agent Plugin Eval need to run?

Going by SKILL.md and its folder, Agent Plugin Eval needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Agent Plugin Eval access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Agent Plugin Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Agent Plugin Eval use?

Agent Plugin Eval is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Plugin Eval use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 5.3k tokens, read only when the agent opens those files.

What are the alternatives to Agent Plugin Eval?

Skills that share tags, products or a category with Agent Plugin Eval: MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars), Crush Configuration (charmbracelet/crush, 29k stars), Agent Repo Init (study8677/repobrain, 1.3k stars) and Openpets (OpenPetsHQ/openpets, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Plugin Eval?

fabricioctelles (a GitHub user) maintains it in fabricioctelles/skills, which has 106 GitHub stars. The repository holds 16 skills in this directory. The repository was last updated on October 4, 2026.

Source: fabricioctelles/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.