Agent skill

Skill Evaluation

by fabricioctelles in fabricioctelles/skills

Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…

Apache-2.0Auto-check passedAgent Workflows

Install Skill Evaluation

skills CLI
$ npx skills add fabricioctelles/skills --skill skill-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fabricioctelles/skills skill-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-evaluation .claude/skills/skill-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-evaluation
GitHub stars
106
Token cost
~3.8k tokens
SKILL.md length
1,876 words
Files
5 (incl. scripts, references)
Skills in repo
15
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…

  • Works in 9 steps: Read the target skill — SKILL.md, its… → Read references/mechanics.md — the… → Classify — use references/categories.md… → …
  • The user asks to evaluate
  • SKILL.md covers Source, Parameters, Criteria and Scoring Guide, plus 6 more sections
  • Runs Python scripts from its folder

What it does

Skill Evaluation is an agent skill from fabricioctelles/skills. Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering, Pruning). Produces an evidence-cited scorecard (0–100), a weighted overall score, and diagnosed failure modes with prioritized fixes. Use when the user asks to evaluate, rate, or audit a skill ("evaluate this skill", "skill scorecard", "review SKILL.md"), or to compare two skills.

Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/categories.md`, `references/mechanics.md` and `references/output-template.md`).

It sits in Agent Workflows, covering Agent evaluation and testing and Accessibility. The repository describes itself as: A collection of skills for AI agents (Kiro, Cursor, Windsurf, Claude Code, and others). Each skill is a reusable module that teaches the agent to perform complex tasks with… The licence is Apache-2.0.

When your agent uses it

  • The user asks to evaluate
  • Audit a skill (evaluate this skill
  • Skill scorecard
  • Review SKILL.md)

Example prompts

  • “s Claude Code best practices plus Matt Pocock”
  • “evaluate this skill”
  • “skill scorecard”
  • “/skill-evaluation”

Requirements

  • Python 3

Workflow steps

9 steps, taken from the first numbered list in SKILL.md.

  1. Read the target skill — SKILL.md, its frontmatter (check for
  2. Read references/mechanics.md — the vocabulary and tests Axes 1, 3,
  3. Classify — use references/categories.md and its decision tree to
  4. Score all applicable criteria — cite-or-cut: a criterion is only
  5. Trigger eval — empirical test of whether the skill's description
  6. Diagnose failure modes — done when every mode in the table below has
  7. Assess bonus patterns — the 4 carried over from v1, plus a fifth
  8. Compute the weighted score — run scripts/score.py with one
  9. Write the scorecard to the output path — read

What it can do on your machine

Read from SKILL.md and the folder at commit f1de632. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • claude.com
    • youtube.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Evaluation loads about 3.8k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 123 tokens; SKILL.md has 1,876 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~123
When it runs · the whole SKILL.md, loaded when a task matches
~3.8k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from fabricioctelles/skills at commit f1de632, republished under its Apache-2.0 licence (© fabricioctelles). 1,876 words, ~3,806 tokens.

Download SKILL.mdSave it as .claude/skills/skill-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
skill-evaluation
description
Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering, Pruning). Produces an evidence-cited scorecard (0–100), a weighted overall score, and diagnosed failure modes with prioritized fixes. Use when the user asks to evaluate, rate, or audit a skill ("evaluate this skill", "skill scorecard", "review SKILL.md"), or to compare two skills.
metadata.author
ft.ia.br
metadata.version
2.1.0
metadata.date
2026-07-03
metadata.repository
https://github.com/fabricioctelles/skills
metadata.license
Apache-2.0
metadata.category
code-quality-and-review

Skill Evaluation

If you need the vocabulary and tests behind Axes 1, 3, and 4 (leading words, completion criteria, context pointers, the deletion test, failure-mode definitions), read references/mechanics.md before scoring those axes.

Source

Parameters

ParameterDescriptionDefault
targetPath to skill directory or SKILL.md to evaluateAsk user
outputPath to write the scorecard<target>/EVALUATION.md
compareOptional second skill to compare side-by-sideNone

Also runs unattended: in CI, point target at skills changed in a PR and gate with scripts/score.py --fail-below 60 ... — non-zero exit below the threshold fails the check.

Criteria

18 criteria: 14 core, scored on every skill, plus 4 conditional criteria scored only when the skill's category makes them apply — otherwise mark N/A and exclude the criterion from both the numerator and denominator of the weighted average. Every score is 0–100 with evidence citing file, section, or line.

Axis 1 — Trigger (invocation)
#CriterionWeightKey question
1Invocation design2xIs model-invoked vs. user-invoked deliberate and fitting? Model-invoked pays context load (the description loads every turn); user-invoked pays cognitive load (the human is the index). A skill that only ever fires by hand should be user-invoked.
2Description quality2xModel-invoked: leading word up front, one trigger per branch (synonyms renaming the same branch are duplication), no identity that's redundant with the body. User-invoked (disable-model-invocation: true): a human-facing one-liner, no trigger list. Score against the mode the skill actually uses — never penalize a user-invoked skill for lacking trigger phrases.
Axis 2 — Structure
#CriterionWeightKey question
3Steps vs. reference clarity1xDoes the skill distinguish ordered steps from on-demand reference? All-reference and all-steps skills are both valid — score clarity, not the mix. Is related material co-located (definition, rules, caveats under one heading)?
4Branch-aware disclosure & pointers2xIs material every branch needs inline, and material only some branches need behind a context pointer? Does each pointer's wording say when to follow it ("if you need X, read Y")? A weakly worded pointer to must-have material is a variance bug.
5Conciseness (no sprawl)2xIs SKILL.md lean — under 500 lines as a ceiling, smaller is better — with every line earning its context cost?
6Coherent scope1xDoes the skill do one thing and compose with others, rather than covering too much?
Axis 3 — Steering
#CriterionWeightKey question
7Leading words2xDoes the skill use compact, high-prior terms ("vertical slice", "tight", "red") to anchor behavior, repeated consistently? Could any verbose passage collapse into one?
8Completion criteria & legwork2xSkills with steps: does each step end on a checkable, exhaustive completion criterion? A vague one invites premature completion. Skills that are pure reference: is there an exhaustiveness bar over the reference itself ("every rule applied")? If neither applies, mark N/A.
9Gotchas section2xIs there explicit capture of failure points, edge cases, footguns?
10Grounded in expertise2xDoes content come from observed failures and real project facts, or generic "best practices"?
11Avoids railroading1xDoes the skill leave room to adapt — procedures over declarations, defaults over menus — without over-prescribing?
Axis 4 — Pruning
#CriterionWeightKey question
12No-ops (deletion test)2xRunning the deletion test sentence by sentence: if removing a sentence leaves behavior unchanged, it's a no-op — including restatements of what the model already does by default. Cite line numbers for candidates.
13Single source of truth1xDoes each meaning live in exactly one place? Duplication between SKILL.md and references/ counts too.
14Relevance & sediment1xAre there stale lines, accumulated layers, or material that no longer influences what the skill does?
Conditional criteria

Score only when the skill's category (from references/categories.md) makes the criterion apply; otherwise mark N/A and drop it from the weighted average entirely.

#CriterionWeightApplies to category
15Setup flow1xlibrary-and-api-reference, data-fetching-and-analysis, ci-cd-and-deployment, infrastructure-operations
16Memory mechanism1xbusiness-process-automation, data-fetching-and-analysis, runbooks
17Scripts & libraries1xproduct-verification, code-scaffolding-and-templates, code-quality-and-review, data-fetching-and-analysis, infrastructure-operations
18On-demand hooks1xcode-quality-and-review, ci-cd-and-deployment

Override this table with judgment, in either direction: score a criterion for a skill outside these categories when it would clearly benefit (e.g., a non-product-verification skill that obviously needs a helper script), and mark it N/A even within an applicable category when the pattern doesn't fit the skill's shape (e.g., a pure-reference vocabulary skill filed under code-quality-and-review has nothing for a hook to enforce). Explain the override in the scorecard either way.

Overall score
overall = sum(score × weight) / sum(weight)

N/A criteria are excluded from both sums — never scored as 0, never counted as weight.

Scoring Guide

ScoreMeaning
0Not present at all
1–25Minimal/token effort, barely addresses the criterion
26–50Partially addressed but with significant gaps
51–75Solid implementation with room for improvement
76–90Strong implementation, minor gaps only
91–100Exemplary — would use as a reference for others

Grade Scale

GradeRangeMeaning
A80–100Production-quality, reference skill
B60–79Good skill, minor improvements needed
C40–59Functional but significant gaps
D20–39Needs substantial rework
F0–19Skeleton only, not production-ready

Workflow

  1. Read the target skill — SKILL.md, its frontmatter (check for disable-model-invocation), and every file in the skill directory.

  2. Read references/mechanics.md — the vocabulary and tests Axes 1, 3, and 4 depend on, including what makes a context pointer's wording effective.

  3. Classify — use references/categories.md and its decision tree to assign a category. The category determines which conditional criteria apply.

  4. Score all applicable criteria — cite-or-cut: a criterion is only scored once its justification cites specific evidence (file, section, or line); no citation, no score. Mark N/A wherever the conditional table, or your own judgment, says a criterion doesn't apply. Done when every applicable criterion carries a score and a citation, and every N/A a reason.

  5. Trigger eval — empirical test of whether the skill's description actually causes invocation. See the Trigger Eval section below for the full mechanic. Skip this step for user-invoked skills (disable-model-invocation: true) — they have no description to test.

  6. Diagnose failure modes — done when every mode in the table below has been checked against the skill and either cited (file:line) or dismissed.

  7. Assess bonus patterns — the 4 carried over from v1, plus a fifth:

    BonusApplies whenWhat to look for
    Validation loopsSkill produces output or modifies stateInstructs the agent to self-check before finalizing
    Output templatesSkill generates structured outputIncludes a concrete template/example of expected format
    Procedures over declarationsSkill teaches a methodTeaches how to approach problems, not what to produce for one case
    Defaults over menusSkill offers tool/approach choicesPicks a clear default, mentions alternatives briefly
    Trace-checkable steeringSkill uses leading wordsThe leading words are distinctive enough that a user could grep the agent's reasoning traces to confirm the skill actually fired

    Report each as Present / Absent / N/A.

  8. Compute the weighted score — run scripts/score.py with one criterion:score:weight triple per criterion (score NA to exclude); it prints both sums, the overall, and the grade. Don't do this arithmetic by hand.

  9. Write the scorecard to the output path — read references/output-template.md first (it also holds the comparison-mode template used when compare is set) and emit exactly that structure.

Show full SKILL.md (674 more words)Show less

Trigger Eval

Empirical test of whether the skill's description causes a model to invoke it when it should — and ignore it when it shouldn't. This is not a pass/fail gate; it produces observational data that feeds the scorecard and informs the failure-mode diagnosis.

When to run
  • Model-invoked skills only. User-invoked skills (disable-model-invocation: true) have no description to test — skip and mark the section N/A.
Prompt generation

Generate 10 prompts from the skill's description, scope, and gotchas:

  • 5 should-trigger — realistic user requests that fall squarely within the skill's stated scope. Vary phrasing: some use the skill's vocabulary, others describe the same need in naive/indirect language.
  • 5 should-not-trigger — requests that are adjacent but clearly outside scope (e.g., a sibling skill's territory, a task the description explicitly excludes, or a generic request a model handles without any skill).

Each prompt should read like something a real user would type — no meta-language about skills, no hints.

Sub-agent execution

Run each prompt in an independent sub-agent session with the target skill available. The sub-agent receives a single additional instruction appended to its system context:

At the end of your response, output exactly one line in this format:
SKILLS_USED: <comma-separated list of skill names you loaded during this task, or "none">

This instruction is generic — it does not name the skill under test or hint at what should be triggered. The sub-agent operates normally; it either loads the skill or doesn't based on the prompt alone.

Detection

Parse the SKILLS_USED: line from each sub-agent's response. Record per prompt:

FieldValue
PromptThe test prompt text
Expectedshould-trigger / should-not-trigger
Triggeredyes / no (was the target skill name in the list?)
Other skillsAny other skills that fired
What to report

Report raw counts — no pass/fail judgment:

  • Should-trigger hit rate — X/5 triggered
  • Should-not-trigger leak rate — X/5 triggered (lower is better)
  • Other skills observed — which siblings fired on the same prompts

These numbers feed criterion #1 (invocation design) and #2 (description quality) with empirical evidence, and may reveal failure modes like over-triggering or description weakness.

Practical notes
  • If the evaluation environment cannot spawn sub-agents (e.g., CI without agent access), skip the trigger eval and note "trigger eval: skipped (no agent access)" in the scorecard.
  • A single trial per prompt is acceptable given the observational (non-gating) nature. Run multiple trials only if results are ambiguous.
  • Keep prompts in the scorecard output so the skill author can reuse them as a regression set.

Failure-mode diagnosis

Name the failure mode, cite evidence, prescribe the defense. Each mode's defense is defined once in references/mechanics.md §5 — prescribe from there. This replaces a generic "top improvements" list.

ModeEvidence to look for
Premature completionVague completion criteria with future steps still visible
Weak steeringInstruction present but the agent doesn't reliably follow it
DuplicationSame meaning in 2+ places, including SKILL.md vs. references/
SedimentStale layers, outdated references, dead instructions
SprawlLong even with no duplication or sediment
No-opsLines that don't change behavior versus the model's default
Buried stepsInline reference so heavy it soaks the steps

After the table, write a Prioritized Actions section: 3–5 highest-impact actions derived directly from the detected failure modes, each citing its evidence.

Note: context overload — too many model-invoked skills competing for attention in one environment — is a portfolio-level problem, out of scope for evaluating a single skill. Record the description's context-load cost when it's notable; don't score the portfolio.

Gotchas

  • Tiny skills (under ~50 lines) flood the scorecard with N/A — score what's there; a small, sharp skill can reach grade A on few criteria.
  • Self-evaluation bias: when the skill under review is one you (or this session) wrote, apply the deletion test with extra skepticism — you will want your own lines to matter.
  • Fresh rewrites still carry duplication: sediment needs time to settle, but duplication can ship on day one. Run the pruning axis even on brand-new skills.

Quality Checklist

Final gate before delivering — each item names the step whose completion it re-checks, nothing new:

  • cite-or-cut held everywhere (step 4)
  • every N/A justified (step 4)
  • trigger eval run or skipped with reason (step 5)
  • every failure mode cited or dismissed (step 6)
  • 5 bonus patterns assessed (step 7)
  • score computed by scripts/score.py, not by hand (step 8)

© fabricioctelles, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (scripts, references) in skills/skill-evaluation of fabricioctelles/skills.

  • SKILL.md
  • references/categories.md
  • references/mechanics.md
  • references/output-template.md
  • scripts/score.py

Open the folder on GitHubat commit f1de632

Compare with similar skills

Skill Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Evaluation this skillfabricioctelles/skills106—~3.8kAutomated safety check: PassApache-2.0
Idea Signal MapperWILLOSCAR/research-units-pipeline-skills513—~336Automated safety check: PassNone
Vibe Sunsang Growthfivetaku/gptaku-plugins-codex128—~2kAutomated safety check: PassMIT
DocsPrefectHQ/fastmcp28k—~1kAutomated safety check: PassApache-2.0
Jarvis Setupethanplusai/jarvis843—~2.5kAutomated safety check: NotesCustom licence
CLI Anything BrowserHKUDS/CLI-Anything52k—~1.5kAutomated safety check: WarnApache-2.0

Similar skills

  • Idea Signal Mapper

    WILLOSCAR/research-units-pipeline-skills

    Map paper notes + taxonomy into a signal table of tensions, missing pieces, and promising academic axes for brainstorm discussion.

    513 GitHub stars~336 tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • Vibe Sunsang Growth

    fivetaku/gptaku-plugins-codex

    Growth report generator — analyzes converted Codex conversations and produces a progression report using the v2 level system (6 axes × 7 levels, 0.5 increments), leading with one level headline and…

    128 GitHub stars~2k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Docs

    PrefectHQ/fastmcp

    Write or revise a page under docs/ for gofastmcp.com. An agent skill from PrefectHQ/fastmcp.

    28k GitHub stars~1k tokensUpdated today
    Frontend & DesignAuto-check passed
  • Jarvis Setup

    ethanplusai/jarvis

    A skill your agent uses when helping someone install, configure, or debug a fresh clone of JARVIS (this repo) — especially "the mic doesn't work", "JARVIS says his language systems are down", any…

    843 GitHub stars~2.5k tokensUpdated 29 days ago
    Frontend & DesignAuto-check: notes
  • CLI Anything Browser

    HKUDS/CLI-Anything

    Browser automation CLI using DOMShell MCP server. An agent skill from HKUDS/CLI-Anything.

    52k GitHub stars~1.5k tokensUpdated 18 days ago
    Productivity & AutomationAuto-check: warnings
  • Web Interface Guidelines Reviewer

    vercel-labs/openreview

    Official

    Review UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "check accessibility", "audit design", "review UX", or "check my…

    1.7k GitHub starsUsed in 97 repos~308 tokens
    Frontend & DesignAuto-check passed

More from fabricioctelles/skills

All 15 skills in this repo
  • Motion Ad

    fabricioctelles/skills

    Produce a short motion-graphics video ad — a 15s Facebook/Instagram/TikTok spot — as a rendered MP4.

    106 GitHub stars~4.1k tokensUpdated 5 days ago
    Auto-check passed
  • Agent Plugin Eval

    fabricioctelles/skills

    Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification.

    106 GitHub stars~2.1k tokensUpdated 5 days ago
    Auto-check passed
  • Loop Architect

    fabricioctelles/skills

    Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them.

    106 GitHub stars~2.1k tokensUpdated 5 days ago
    Auto-check: notes
  • Pier Cloud

    fabricioctelles/skills

    This skill should be used when the user needs to consume the Pier Cloud (Lighthouse) API for cloud cost management — including JWT authentication, listing contexts, workspaces, workspace groups, and…

    106 GitHub stars~1.1k tokensUpdated 5 days ago
    Auto-check: notes
  • Ralph Loop Kiro Specs

    fabricioctelles/skills

    Automated iterative agent runner for spec-based development in Kiro.

    106 GitHub stars~2.6k tokensUpdated 5 days ago
    Auto-check passed
  • Security Specialist

    fabricioctelles/skills

    Runs security audits on codebases — full scans, diff reviews, threat models, vulnerability triage, remediation guidance, and finding tracking.

    106 GitHub stars~2.8k tokensUpdated 5 days ago
    Auto-check passed

Questions about Skill Evaluation

What does Skill Evaluation do?

Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…. Skill Evaluation is an agent skill from fabricioctelles/skills. Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering, Pruning).

When should I use Skill Evaluation?

Skill Evaluation fits situations like: the user asks to evaluate; audit a skill (evaluate this skill; skill scorecard; review SKILL.md).

How do I install Skill Evaluation in Claude Code?

Run `npx skills add fabricioctelles/skills --skill skill-evaluation -a claude-code`. Or copy the skill folder (skills/skill-evaluation in fabricioctelles/skills) into .claude/skills/skill-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Skill Evaluation in Codex?

Run `npx skills add fabricioctelles/skills --skill skill-evaluation -a codex`. Or copy the skill folder (skills/skill-evaluation in fabricioctelles/skills) into .agents/skills/skill-evaluation in your project. Codex loads it when a task matches its description.

Can I use Skill Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fabricioctelles/skills --skill skill-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-evaluation, .gemini/skills/skill-evaluation, .github/skills/skill-evaluation and .opencode/skills/skill-evaluation in your project.

What does Skill Evaluation need to run?

Going by SKILL.md and its folder, Skill Evaluation needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Skill Evaluation access the network?

SKILL.md names 2 domains. As links in the text: claude.com and youtube.com. This is read from the text; nothing was executed.

Is Skill Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skill Evaluation use?

Skill Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Evaluation use?

About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.

What are the alternatives to Skill Evaluation?

Skills that share tags, products or a category with Skill Evaluation: Idea Signal Mapper (WILLOSCAR/research-units-pipeline-skills, 513 stars), Vibe Sunsang Growth (fivetaku/gptaku-plugins-codex, 128 stars), Docs (PrefectHQ/fastmcp, 28k stars) and Jarvis Setup (ethanplusai/jarvis, 843 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Evaluation?

fabricioctelles (a GitHub user) maintains it in fabricioctelles/skills, which has 106 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 4, 2026.

Source: fabricioctelles/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.