Idea Signal Mapper
WILLOSCAR/research-units-pipeline-skills
Map paper notes + taxonomy into a signal table of tensions, missing pieces, and promising academic axes for brainstorm discussion.
Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…
$ npx skills add fabricioctelles/skills --skill skill-evaluation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install fabricioctelles/skills skill-evaluation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-evaluation .claude/skills/skill-evaluation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-evaluation" agent skill from https://github.com/fabricioctelles/skills/tree/main/skills/skill-evaluation into .claude/skills/skill-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/fabricioctelles/skills/tree/main/skills/skill-evaluationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add fabricioctelles/skills --skill skill-evaluation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install fabricioctelles/skills skill-evaluation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/skill-evaluation .agents/skills/skill-evaluation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-evaluation" agent skill from https://github.com/fabricioctelles/skills/tree/main/skills/skill-evaluation into .agents/skills/skill-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add fabricioctelles/skills --skill skill-evaluation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install fabricioctelles/skills skill-evaluation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/skill-evaluation .cursor/skills/skill-evaluation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-evaluation" agent skill from https://github.com/fabricioctelles/skills/tree/main/skills/skill-evaluation into .cursor/skills/skill-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/fabricioctelles/skills.git --path skills/skill-evaluation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add fabricioctelles/skills --skill skill-evaluation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install fabricioctelles/skills skill-evaluation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/skill-evaluation .gemini/skills/skill-evaluation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-evaluation" agent skill from https://github.com/fabricioctelles/skills/tree/main/skills/skill-evaluation into .gemini/skills/skill-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install fabricioctelles/skills skill-evaluationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add fabricioctelles/skills --skill skill-evaluation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/skill-evaluation .github/skills/skill-evaluation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-evaluation" agent skill from https://github.com/fabricioctelles/skills/tree/main/skills/skill-evaluation into .github/skills/skill-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add fabricioctelles/skills --skill skill-evaluation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install fabricioctelles/skills skill-evaluation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/fabricioctelles/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/skill-evaluation .opencode/skills/skill-evaluation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-evaluation" agent skill from https://github.com/fabricioctelles/skills/tree/main/skills/skill-evaluation into .opencode/skills/skill-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-evaluation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-evaluationEvaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…
Skill Evaluation is an agent skill from fabricioctelles/skills. Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering, Pruning). Produces an evidence-cited scorecard (0–100), a weighted overall score, and diagnosed failure modes with prioritized fixes. Use when the user asks to evaluate, rate, or audit a skill ("evaluate this skill", "skill scorecard", "review SKILL.md"), or to compare two skills.
Its SKILL.md is about 3.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including scripts and reference files (for example `references/categories.md`, `references/mechanics.md` and `references/output-template.md`).
It sits in Agent Workflows, covering Agent evaluation and testing and Accessibility. The repository describes itself as: A collection of skills for AI agents (Kiro, Cursor, Windsurf, Claude Code, and others). Each skill is a reusable module that teaches the agent to perform complex tasks with… The licence is Apache-2.0.
9 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit f1de632. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 1 file in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
claude.comyoutube.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Evaluation loads about 3.8k tokens when it runs, and up to ~8k if it reads all its reference files. Until then it costs about 123 tokens; SKILL.md has 1,876 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from fabricioctelles/skills at commit f1de632, republished under its Apache-2.0 licence (© fabricioctelles). 1,876 words, ~3,806 tokens.
.claude/skills/skill-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.If you need the vocabulary and tests behind Axes 1, 3, and 4 (leading words,
completion criteria, context pointers, the deletion test, failure-mode
definitions), read references/mechanics.md before scoring those axes.
writing-great-skills skill| Parameter | Description | Default |
|---|---|---|
target | Path to skill directory or SKILL.md to evaluate | Ask user |
output | Path to write the scorecard | <target>/EVALUATION.md |
compare | Optional second skill to compare side-by-side | None |
Also runs unattended: in CI, point target at skills changed in a PR and
gate with scripts/score.py --fail-below 60 ... — non-zero exit below the
threshold fails the check.
18 criteria: 14 core, scored on every skill, plus 4 conditional criteria scored only when the skill's category makes them apply — otherwise mark N/A and exclude the criterion from both the numerator and denominator of the weighted average. Every score is 0–100 with evidence citing file, section, or line.
| # | Criterion | Weight | Key question |
|---|---|---|---|
| 1 | Invocation design | 2x | Is model-invoked vs. user-invoked deliberate and fitting? Model-invoked pays context load (the description loads every turn); user-invoked pays cognitive load (the human is the index). A skill that only ever fires by hand should be user-invoked. |
| 2 | Description quality | 2x | Model-invoked: leading word up front, one trigger per branch (synonyms renaming the same branch are duplication), no identity that's redundant with the body. User-invoked (disable-model-invocation: true): a human-facing one-liner, no trigger list. Score against the mode the skill actually uses — never penalize a user-invoked skill for lacking trigger phrases. |
| # | Criterion | Weight | Key question |
|---|---|---|---|
| 3 | Steps vs. reference clarity | 1x | Does the skill distinguish ordered steps from on-demand reference? All-reference and all-steps skills are both valid — score clarity, not the mix. Is related material co-located (definition, rules, caveats under one heading)? |
| 4 | Branch-aware disclosure & pointers | 2x | Is material every branch needs inline, and material only some branches need behind a context pointer? Does each pointer's wording say when to follow it ("if you need X, read Y")? A weakly worded pointer to must-have material is a variance bug. |
| 5 | Conciseness (no sprawl) | 2x | Is SKILL.md lean — under 500 lines as a ceiling, smaller is better — with every line earning its context cost? |
| 6 | Coherent scope | 1x | Does the skill do one thing and compose with others, rather than covering too much? |
| # | Criterion | Weight | Key question |
|---|---|---|---|
| 7 | Leading words | 2x | Does the skill use compact, high-prior terms ("vertical slice", "tight", "red") to anchor behavior, repeated consistently? Could any verbose passage collapse into one? |
| 8 | Completion criteria & legwork | 2x | Skills with steps: does each step end on a checkable, exhaustive completion criterion? A vague one invites premature completion. Skills that are pure reference: is there an exhaustiveness bar over the reference itself ("every rule applied")? If neither applies, mark N/A. |
| 9 | Gotchas section | 2x | Is there explicit capture of failure points, edge cases, footguns? |
| 10 | Grounded in expertise | 2x | Does content come from observed failures and real project facts, or generic "best practices"? |
| 11 | Avoids railroading | 1x | Does the skill leave room to adapt — procedures over declarations, defaults over menus — without over-prescribing? |
| # | Criterion | Weight | Key question |
|---|---|---|---|
| 12 | No-ops (deletion test) | 2x | Running the deletion test sentence by sentence: if removing a sentence leaves behavior unchanged, it's a no-op — including restatements of what the model already does by default. Cite line numbers for candidates. |
| 13 | Single source of truth | 1x | Does each meaning live in exactly one place? Duplication between SKILL.md and references/ counts too. |
| 14 | Relevance & sediment | 1x | Are there stale lines, accumulated layers, or material that no longer influences what the skill does? |
Score only when the skill's category (from references/categories.md) makes
the criterion apply; otherwise mark N/A and drop it from the weighted
average entirely.
| # | Criterion | Weight | Applies to category |
|---|---|---|---|
| 15 | Setup flow | 1x | library-and-api-reference, data-fetching-and-analysis, ci-cd-and-deployment, infrastructure-operations |
| 16 | Memory mechanism | 1x | business-process-automation, data-fetching-and-analysis, runbooks |
| 17 | Scripts & libraries | 1x | product-verification, code-scaffolding-and-templates, code-quality-and-review, data-fetching-and-analysis, infrastructure-operations |
| 18 | On-demand hooks | 1x | code-quality-and-review, ci-cd-and-deployment |
Override this table with judgment, in either direction: score a criterion
for a skill outside these categories when it would clearly benefit (e.g., a
non-product-verification skill that obviously needs a helper script), and
mark it N/A even within an applicable category when the pattern doesn't fit
the skill's shape (e.g., a pure-reference vocabulary skill filed under
code-quality-and-review has nothing for a hook to enforce). Explain the
override in the scorecard either way.
overall = sum(score × weight) / sum(weight)N/A criteria are excluded from both sums — never scored as 0, never counted as weight.
| Score | Meaning |
|---|---|
| 0 | Not present at all |
| 1–25 | Minimal/token effort, barely addresses the criterion |
| 26–50 | Partially addressed but with significant gaps |
| 51–75 | Solid implementation with room for improvement |
| 76–90 | Strong implementation, minor gaps only |
| 91–100 | Exemplary — would use as a reference for others |
| Grade | Range | Meaning |
|---|---|---|
| A | 80–100 | Production-quality, reference skill |
| B | 60–79 | Good skill, minor improvements needed |
| C | 40–59 | Functional but significant gaps |
| D | 20–39 | Needs substantial rework |
| F | 0–19 | Skeleton only, not production-ready |
Read the target skill — SKILL.md, its frontmatter (check for
disable-model-invocation), and every file in the skill directory.
Read references/mechanics.md — the vocabulary and tests Axes 1, 3,
and 4 depend on, including what makes a context pointer's wording
effective.
Classify — use references/categories.md and its decision tree to
assign a category. The category determines which conditional criteria
apply.
Score all applicable criteria — cite-or-cut: a criterion is only scored once its justification cites specific evidence (file, section, or line); no citation, no score. Mark N/A wherever the conditional table, or your own judgment, says a criterion doesn't apply. Done when every applicable criterion carries a score and a citation, and every N/A a reason.
Trigger eval — empirical test of whether the skill's description
actually causes invocation. See the Trigger Eval section below for
the full mechanic. Skip this step for user-invoked skills
(disable-model-invocation: true) — they have no description to test.
Diagnose failure modes — done when every mode in the table below has been checked against the skill and either cited (file:line) or dismissed.
Assess bonus patterns — the 4 carried over from v1, plus a fifth:
| Bonus | Applies when | What to look for |
|---|---|---|
| Validation loops | Skill produces output or modifies state | Instructs the agent to self-check before finalizing |
| Output templates | Skill generates structured output | Includes a concrete template/example of expected format |
| Procedures over declarations | Skill teaches a method | Teaches how to approach problems, not what to produce for one case |
| Defaults over menus | Skill offers tool/approach choices | Picks a clear default, mentions alternatives briefly |
| Trace-checkable steering | Skill uses leading words | The leading words are distinctive enough that a user could grep the agent's reasoning traces to confirm the skill actually fired |
Report each as Present / Absent / N/A.
Compute the weighted score — run scripts/score.py with one
criterion:score:weight triple per criterion (score NA to exclude); it
prints both sums, the overall, and the grade. Don't do this arithmetic by
hand.
Write the scorecard to the output path — read
references/output-template.md first (it also holds the comparison-mode
template used when compare is set) and emit exactly that structure.
Empirical test of whether the skill's description causes a model to invoke it when it should — and ignore it when it shouldn't. This is not a pass/fail gate; it produces observational data that feeds the scorecard and informs the failure-mode diagnosis.
disable-model-invocation: true) have no description to test — skip and mark the section N/A.Generate 10 prompts from the skill's description, scope, and gotchas:
Each prompt should read like something a real user would type — no meta-language about skills, no hints.
Run each prompt in an independent sub-agent session with the target skill available. The sub-agent receives a single additional instruction appended to its system context:
At the end of your response, output exactly one line in this format:
SKILLS_USED: <comma-separated list of skill names you loaded during this task, or "none">This instruction is generic — it does not name the skill under test or hint at what should be triggered. The sub-agent operates normally; it either loads the skill or doesn't based on the prompt alone.
Parse the SKILLS_USED: line from each sub-agent's response. Record per
prompt:
| Field | Value |
|---|---|
| Prompt | The test prompt text |
| Expected | should-trigger / should-not-trigger |
| Triggered | yes / no (was the target skill name in the list?) |
| Other skills | Any other skills that fired |
Report raw counts — no pass/fail judgment:
These numbers feed criterion #1 (invocation design) and #2 (description quality) with empirical evidence, and may reveal failure modes like over-triggering or description weakness.
Name the failure mode, cite evidence, prescribe the defense. Each mode's
defense is defined once in references/mechanics.md §5 — prescribe from
there. This replaces a generic "top improvements" list.
| Mode | Evidence to look for |
|---|---|
| Premature completion | Vague completion criteria with future steps still visible |
| Weak steering | Instruction present but the agent doesn't reliably follow it |
| Duplication | Same meaning in 2+ places, including SKILL.md vs. references/ |
| Sediment | Stale layers, outdated references, dead instructions |
| Sprawl | Long even with no duplication or sediment |
| No-ops | Lines that don't change behavior versus the model's default |
| Buried steps | Inline reference so heavy it soaks the steps |
After the table, write a Prioritized Actions section: 3–5 highest-impact actions derived directly from the detected failure modes, each citing its evidence.
Note: context overload — too many model-invoked skills competing for attention in one environment — is a portfolio-level problem, out of scope for evaluating a single skill. Record the description's context-load cost when it's notable; don't score the portfolio.
Final gate before delivering — each item names the step whose completion it re-checks, nothing new:
scripts/score.py, not by hand (step 8)© fabricioctelles, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 4 other files (scripts, references) in skills/skill-evaluation of fabricioctelles/skills.
Open the folder on GitHubat commit f1de632
Skill Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Evaluation this skillfabricioctelles/skills | 106 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | |
| Idea Signal MapperWILLOSCAR/research-units-pipeline-skills | 513 | — | ~336 | Automated safety check: Pass | None | |
| Vibe Sunsang Growthfivetaku/gptaku-plugins-codex | 128 | — | ~2k | Automated safety check: Pass | MIT | |
| DocsPrefectHQ/fastmcp | 28k | — | ~1k | Automated safety check: Pass | Apache-2.0 | |
| Jarvis Setupethanplusai/jarvis | 843 | — | ~2.5k | Automated safety check: Notes | Custom licence | |
| CLI Anything BrowserHKUDS/CLI-Anything | 52k | — | ~1.5k | Automated safety check: Warn | Apache-2.0 |
WILLOSCAR/research-units-pipeline-skills
Map paper notes + taxonomy into a signal table of tensions, missing pieces, and promising academic axes for brainstorm discussion.
fivetaku/gptaku-plugins-codex
Growth report generator — analyzes converted Codex conversations and produces a progression report using the v2 level system (6 axes × 7 levels, 0.5 increments), leading with one level headline and…
PrefectHQ/fastmcp
Write or revise a page under docs/ for gofastmcp.com. An agent skill from PrefectHQ/fastmcp.
ethanplusai/jarvis
A skill your agent uses when helping someone install, configure, or debug a fresh clone of JARVIS (this repo) — especially "the mic doesn't work", "JARVIS says his language systems are down", any…
HKUDS/CLI-Anything
Browser automation CLI using DOMShell MCP server. An agent skill from HKUDS/CLI-Anything.
vercel-labs/openreview
Review UI code for Web Interface Guidelines compliance. Use when asked to "review my UI", "check accessibility", "audit design", "review UX", or "check my…
fabricioctelles/skills
Produce a short motion-graphics video ad — a 15s Facebook/Instagram/TikTok spot — as a rendered MP4.
fabricioctelles/skills
Audit, score, and compare repositories containing portable Agent Plugins against the official Agent Plugins specification.
fabricioctelles/skills
Design well-structured agent loops with best-practice coaching and cross-model review gates before you run them.
fabricioctelles/skills
This skill should be used when the user needs to consume the Pier Cloud (Lighthouse) API for cloud cost management — including JWT authentication, listing contexts, workspaces, workspace groups, and…
fabricioctelles/skills
Automated iterative agent runner for spec-based development in Kiro.
fabricioctelles/skills
Runs security audits on codebases — full scans, diff reviews, threat models, vulnerability triage, remediation guidance, and finding tracking.
Categories
Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering…. Skill Evaluation is an agent skill from fabricioctelles/skills. Evaluate any agent skill against a merged framework — Anthropic's Claude Code best practices plus Matt Pocock's writing-great-skills methodology — across 4 axes (Trigger, Structure, Steering, Pruning).
Skill Evaluation fits situations like: the user asks to evaluate; audit a skill (evaluate this skill; skill scorecard; review SKILL.md).
Run `npx skills add fabricioctelles/skills --skill skill-evaluation -a claude-code`. Or copy the skill folder (skills/skill-evaluation in fabricioctelles/skills) into .claude/skills/skill-evaluation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add fabricioctelles/skills --skill skill-evaluation -a codex`. Or copy the skill folder (skills/skill-evaluation in fabricioctelles/skills) into .agents/skills/skill-evaluation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fabricioctelles/skills --skill skill-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-evaluation, .gemini/skills/skill-evaluation, .github/skills/skill-evaluation and .opencode/skills/skill-evaluation in your project.
Going by SKILL.md and its folder, Skill Evaluation needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md names 2 domains. As links in the text: claude.com and youtube.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Skill Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.8k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.2k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Skill Evaluation: Idea Signal Mapper (WILLOSCAR/research-units-pipeline-skills, 513 stars), Vibe Sunsang Growth (fivetaku/gptaku-plugins-codex, 128 stars), Docs (PrefectHQ/fastmcp, 28k stars) and Jarvis Setup (ethanplusai/jarvis, 843 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
fabricioctelles (a GitHub user) maintains it in fabricioctelles/skills, which has 106 GitHub stars. The repository holds 15 skills in this directory. The repository was last updated on October 4, 2026.
Source: fabricioctelles/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.