Agent skill

Skill Evaluator

by gotalab in gotalab/skillport

Evaluates agent skills against Anthropic's best practices. An agent skill from gotalab/skillport.

MITAuto-check passedAgent Workflows

Install Skill Evaluator

skills CLI
$ npx skills add gotalab/skillport --skill skill-evaluator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install gotalab/skillport skill-evaluator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/gotalab/skillport.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.skills/experimental/skill-evaluator .claude/skills/skill-evaluator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-evaluator
GitHub stars
414
Token cost
~1.6k tokens
SKILL.md length
551 words
Files
7 (incl. scripts, references)
Skills in repo
3
Repo updated
First seen
Licence
MIT

At a glance

Evaluates agent skills against Anthropic's best practices. An agent skill from gotalab/skillport.

  • Works in 3 steps: Automated Validation → Manual Evaluation → Generate Report
  • Asked to review
  • SKILL.md covers Quick Start, Evaluation Workflow, Score Interpretation and References, plus 1 more section
  • Runs Python scripts from its folder

What it does

Skill Evaluator is an agent skill from gotalab/skillport. Evaluates agent skills against Anthropic's best practices. Use when asked to review, evaluate, assess, or audit a skill for quality. Analyzes SKILL.md structure, naming conventions, description quality, content organization, and identifies anti-patterns. Produces actionable improvement recommendations.

Its SKILL.md is about 1.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 9 other files, including scripts and reference files (for example `evaluations/README.md`, `evaluations/basic-skill-evaluation.json` and `evaluations/problematic-skill-evaluation.json`).

It sits in Agent Workflows. It works with Model Context Protocol. The repository describes itself as: Bring Agent Skills to Any AI Agent and Coding Agent — via CLI or MCP. Manage once, serve anywhere. The licence is MIT.

When your agent uses it

  • Asked to review
  • Audit a skill for quality

Example prompts

  • “Use the skill-evaluator skill to evaluate agent skills against Anthropic's best practices. An agent skill from gotalab/skillport”
  • “/skill-evaluator”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Automated Validation
  2. Manual Evaluation
  3. Generate Report

What it can do on your machine

Read from SKILL.md and the folder at commit 51334ae. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Evaluator loads about 1.6k tokens when it runs, and up to ~6.5k if it reads all its reference files. Until then it costs about 80 tokens; SKILL.md has 551 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~80
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from gotalab/skillport at commit 51334ae, republished under its MIT licence (© gotalab). 551 words, ~1,602 tokens.

Download SKILL.mdSave it as .claude/skills/skill-evaluator/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
skill-evaluator
description
Evaluates agent skills against Anthropic's best practices. Use when asked to review, evaluate, assess, or audit a skill for quality. Analyzes SKILL.md structure, naming conventions, description quality, content organization, and identifies anti-patterns. Produces actionable improvement recommendations.

Skill Evaluator (WIP)

Evaluates skills against Anthropic's official best practices for agent skill authoring. Produces structured evaluation reports with scores and actionable recommendations.

Quick Start

  1. Read the skill's SKILL.md and understand its purpose
  2. Run automated validation: scripts/validate_skill.py <skill-path>
  3. Perform manual evaluation against criteria below
  4. Generate evaluation report with scores and recommendations

Evaluation Workflow

Step 1: Automated Validation

Run the validation script first:

bash
scripts/validate_skill.py <path/to/skill>

This checks:

  • SKILL.md exists with valid YAML frontmatter
  • Name follows conventions (lowercase, hyphens, max 64 chars)
  • Description is present and under 1024 chars
  • Body is under 500 lines
  • File references are one-level deep
Step 2: Manual Evaluation

Evaluate each dimension and assign a score (1-5):

A. Naming (Weight: 10%)
ScoreCriteria
5Gerund form (-ing), clear purpose, memorable
4Descriptive, follows conventions
3Acceptable but could be clearer
2Vague or misleading
1Violates naming rules

Rules: Max 64 chars, lowercase + numbers + hyphens only, no reserved words (anthropic, claude), no XML tags.

Good: processing-pdfs, analyzing-spreadsheets, building-dashboards Bad: pdf, my-skill, ClaudeHelper, anthropic-tools

B. Description (Weight: 20%)
ScoreCriteria
5Clear functionality + specific activation triggers + third person
4Good description with some triggers
3Adequate but missing triggers or vague
2Too brief or unclear purpose
1Missing or unhelpful

Must include: What the skill does AND when to use it. Good: "Extracts text from PDFs. Use when working with PDF documents for text extraction, form parsing, or content analysis." Bad: "A skill for PDFs." or "Helps with documents."

C. Content Quality (Weight: 30%)
ScoreCriteria
5Concise, assumes Claude intelligence, actionable instructions
4Generally good, minor verbosity
3Some unnecessary explanations or redundancy
2Overly verbose or confusing
1Bloated, explains obvious concepts

Ask: "Does Claude really need this explanation?" Remove anything Claude already knows.

D. Structure & Organization (Weight: 25%)
ScoreCriteria
5Excellent progressive disclosure, clear navigation, optimal length
4Good organization, appropriate file splits
3Acceptable but could be better organized
2Poor organization, missing references, or bloated SKILL.md
1No structure, everything dumped in SKILL.md

Check:

  • SKILL.md under 500 lines
  • References are one-level deep (no nested chains)
  • Long reference files (>100 lines) have table of contents
  • Uses forward slashes in all paths
Show full SKILL.md (191 more words)Show less
E. Degrees of Freedom (Weight: 10%)
ScoreCriteria
5Perfect match: high freedom for flexible tasks, low for fragile operations
4Generally appropriate freedom levels
3Acceptable but could be better calibrated
2Mismatched: too rigid or too loose
1Completely wrong freedom level for the task type

Guideline:

  • High freedom (text): Multiple valid approaches, context-dependent
  • Medium freedom (parameterized): Preferred pattern exists, some variation OK
  • Low freedom (specific scripts): Fragile operations, exact sequence required
F. Anti-Pattern Check (Weight: 5%)

Deduct points for each anti-pattern found:

  • Too many options without clear recommendation (-1)
  • Time-sensitive information with date conditionals (-1)
  • Inconsistent terminology (-1)
  • Windows-style paths (backslashes) (-1)
  • Deeply nested references (more than one level) (-2)
  • Scripts that punt error handling to Claude (-1)
  • Magic numbers without justification (-1)
Step 3: Generate Report

Use this template:

markdown
# Skill Evaluation Report: [skill-name]

## Summary
- **Overall Score**: X.X/5.0
- **Recommendation**: [Ready for publication / Needs minor improvements / Needs major revision]

## Dimension Scores
| Dimension | Score | Weight | Weighted |
|-----------|-------|--------|----------|
| Naming | X/5 | 10% | X.XX |
| Description | X/5 | 20% | X.XX |
| Content Quality | X/5 | 30% | X.XX |
| Structure | X/5 | 25% | X.XX |
| Degrees of Freedom | X/5 | 10% | X.XX |
| Anti-Patterns | X/5 | 5% | X.XX |
| **Total** | | 100% | **X.XX** |

## Strengths
- [List 2-3 things done well]

## Areas for Improvement
- [List specific issues with actionable fixes]

## Anti-Patterns Found
- [List any anti-patterns detected]

## Recommendations
1. [Priority 1 fix]
2. [Priority 2 fix]
3. [Priority 3 fix]

## Pre-Publication Checklist
- [ ] Description is specific with activation triggers
- [ ] SKILL.md under 500 lines
- [ ] One-level-deep file references
- [ ] Forward slashes in all paths
- [ ] No time-sensitive information
- [ ] Consistent terminology
- [ ] Concrete examples provided
- [ ] Scripts handle errors explicitly
- [ ] All configuration values justified
- [ ] Required packages listed
- [ ] Tested with Haiku, Sonnet, Opus

Score Interpretation

Score RangeRatingAction
4.5 - 5.0ExcellentReady for publication
4.0 - 4.4GoodMinor improvements recommended
3.0 - 3.9AcceptableSeveral improvements needed
2.0 - 2.9Needs WorkMajor revision required
1.0 - 1.9PoorFundamental redesign needed

References

Examples

See evaluations/ for example evaluation scenarios.

© gotalab, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts, references) in .skills/experimental/skill-evaluator of gotalab/skillport.

  • SKILL.md
  • evaluations/README.md
  • evaluations/basic-skill-evaluation.json
  • evaluations/problematic-skill-evaluation.json
  • references/evaluation-criteria.md
  • references/scoring-rubric.md
  • scripts/validate_skill.py

Open the folder on GitHubat commit 51334ae

Compare with similar skills

Skill Evaluator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Evaluator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Evaluator this skillgotalab/skillport414—~1.6kAutomated safety check: PassMIT
AutoRAG Lite SetupMarker-Inc-Korea/AutoRAG5.1k—~3.3kAutomated safety check: PassMIT
Setting Up Papergraphlotchuazzz-crypto/papergraph-mcp285—~3.3kAutomated safety check: PassMIT
Control Chromewxtsky/byob132—~1.5kAutomated safety check: WarnMIT
Tapd Iteration AnalysisTencentBlueKing/bk-bcs840—~1.2kAutomated safety check: PassCustom licence
Investigate To Confluencetestdouble/han279—~2.3kAutomated safety check: PassMIT

Similar skills

  • AutoRAG Lite Setup

    Marker-Inc-Korea/AutoRAG

    Bootstraps and repairs the model-free AutoRAG Lite MCP server: installing it, initializing a config with approved search roots, building indexes and verifying discovery.

    5.1k GitHub stars~3.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Setting Up Papergraph

    lotchuazzz-crypto/papergraph-mcp

    A skill your agent uses when a user has cloned PaperGraph MCP and asks to install, initialize, configure, set up, or start using it with an agent or MCP client.

    285 GitHub stars~3.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Control Chrome

    wxtsky/byob

    Control and inspect the user's real Google Chrome through byob's local MCP tools.

    132 GitHub stars~1.5k tokensUpdated 10 days ago
    Agent WorkflowsAuto-check: warnings
  • Tapd Iteration Analysis

    TencentBlueKing/bk-bcs

    迭代执行流水线需求分析阶段。通过 TAPD MCP storiesget 拉取迭代中所有需求详情, 将 description 字段(完整 Markdown 内容)直接保存为需求文档,基于 parentid 字段自动识别父子关系和独立需求,从需求文档的"依赖关系"章节提取显式依赖进行 拓扑排序,确定实现顺序,更新迭代状态。

    840 GitHub stars~1.2k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Runs an evidence-based investigation of a bug, failure, or unexpected behavior with investigate and publishes the resulting investigation report to a user-specified Confluence location.

    279 GitHub stars~2.3k tokensUpdated 7 days ago
    Agent WorkflowsAuto-check passed
  • Siyuan MCP Markup Guide

    yangtaihong59/siyuan-plugins-mcp-sisyphus

    MCP SiYuan markup guide for rich Markdown written through block and document actions.

    114 GitHub stars~405 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from gotalab/skillport

  • Opus 4 5 Migration

    gotalab/skillport

    Migrate prompts and code from Claude Sonnet 4.0, Sonnet 4.5, or Opus 4.1 to Opus 4.5.

    414 GitHub starsUsed in 5 repos~1.2k tokens
    Auto-check passed
  • Git Branch Cleanup

    gotalab/skillport

    Analyzes and safely cleans up local Git branches. An agent skill from gotalab/skillport.

    414 GitHub stars~2k tokensUpdated 3 mo ago
    Auto-check passed

Questions about Skill Evaluator

What does Skill Evaluator do?

Evaluates agent skills against Anthropic's best practices. An agent skill from gotalab/skillport. Skill Evaluator is an agent skill from gotalab/skillport. Evaluates agent skills against Anthropic's best practices.

When should I use Skill Evaluator?

Skill Evaluator fits situations like: asked to review; audit a skill for quality.

How do I install Skill Evaluator in Claude Code?

Run `npx skills add gotalab/skillport --skill skill-evaluator -a claude-code`. Or copy the skill folder (.skills/experimental/skill-evaluator in gotalab/skillport) into .claude/skills/skill-evaluator in your project. Claude Code loads it when a task matches its description.

How do I install Skill Evaluator in Codex?

Run `npx skills add gotalab/skillport --skill skill-evaluator -a codex`. Or copy the skill folder (.skills/experimental/skill-evaluator in gotalab/skillport) into .agents/skills/skill-evaluator in your project. Codex loads it when a task matches its description.

Can I use Skill Evaluator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add gotalab/skillport --skill skill-evaluator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-evaluator, .gemini/skills/skill-evaluator, .github/skills/skill-evaluator and .opencode/skills/skill-evaluator in your project.

What does Skill Evaluator need to run?

Going by SKILL.md and its folder, Skill Evaluator needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Skill Evaluator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Evaluator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skill Evaluator use?

Skill Evaluator is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Evaluator use?

About 1.6k tokens (SKILL.md is roughly 6.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 4.9k tokens, read only when the agent opens those files.

What are the alternatives to Skill Evaluator?

Skills that share tags, products or a category with Skill Evaluator: AutoRAG Lite Setup (Marker-Inc-Korea/AutoRAG, 5.1k stars), Setting Up Papergraph (lotchuazzz-crypto/papergraph-mcp, 285 stars), Control Chrome (wxtsky/byob, 132 stars) and Tapd Iteration Analysis (TencentBlueKing/bk-bcs, 840 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Evaluator?

gotalab (a GitHub user) maintains it in gotalab/skillport, which has 414 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on June 24, 2026.

Source: gotalab/skillport on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.