Agent skill

Skill Judge

by shareAI-lab in shareAI-lab/lab-skills

Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

Apache-2.0Auto-check passedEducation

Install Skill Judge

skills CLI
$ npx skills add shareAI-lab/lab-skills --skill skill-judge -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install shareAI-lab/lab-skills skill-judge --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/shareAI-lab/lab-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/agent-development/skill-judge .claude/skills/skill-judge && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-judge
GitHub stars
315
Token cost
~1.9k tokens
SKILL.md length
876 words
Files
2 (incl. references)
Skills in repo
10
Repo updated
First seen
Licence
Apache-2.0

At a glance

Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

  • Works in 5 steps: Read SKILL.md completely. → List every bundled reference, script,… → Follow every resource link and loading… → …
  • Improving SKILL.md files and skill packages
  • SKILL.md covers Applicability, Choose the review depth, Read the complete package and Apply the hard gates first, plus 9 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Skill Judge is an agent skill from shareAI-lab/lab-skills. Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples. Use when reviewing, auditing, comparing, or improving SKILL.md files and skill packages. Produces evidence-backed findings, optional multi-dimensional scoring, and prioritized changes; treat the result as diagnostic guidance, not official certification or a universal standard.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/rubric.md`).

It sits in Education, covering Quizzes and assessments, Skill authoring and Agent evaluation and testing. The repository describes itself as: Skills distilled from the Lab's real work and collaboration practices. The licence is Apache-2.0.

When your agent uses it

  • Improving SKILL.md files and skill packages
  • Tasks that involve Quizzes and assessments
  • Tasks that involve Skill authoring

Example prompts

  • “/skill-judge”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Read SKILL.md completely.
  2. List every bundled reference, script, asset, and agent metadata file.
  3. Follow every resource link and loading instruction that affects behavior.
  4. Check whether referenced files exist and whether shipped scripts or examples are plausible and tested.
  5. Distinguish observed facts from assumptions about how an agent might behave.

What it can do on your machine

Read from SKILL.md and the folder at commit becee99. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Judge loads about 1.9k tokens when it runs, and up to ~3.8k if it reads all its reference files. Until then it costs about 105 tokens; SKILL.md has 876 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~105
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from shareAI-lab/lab-skills at commit becee99, republished under its Apache-2.0 licence (© shareAI-lab). 876 words, ~1,852 tokens.

Download SKILL.mdSave it as .claude/skills/skill-judge/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
skill-judge
description
Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples. Use when reviewing, auditing, comparing, or improving SKILL.md files and skill packages. Produces evidence-backed findings, optional multi-dimensional scoring, and prioritized changes; treat the result as diagnostic guidance, not official certification or a universal standard.

Skill Judge

Review a Skill as an executable knowledge package, not as an essay. Judge whether the package triggers at the right time, adds knowledge the AI model needs, routes to the right resources, and helps an agent produce better work.

Applicability

Use this as an opinionated diagnostic framework. It is informed by public specifications, examples, and practical use, but it is not an official compliance suite. Do not turn a numeric score into a claim of universal quality.

Choose the review depth

  • Quick review: identify blockers, the highest-value weakness, and the next change. Do not score unless a score helps the decision.
  • Full audit: inspect the entire package, score all dimensions, and produce prioritized recommendations. Read rubric.md completely before scoring.
  • Comparison: evaluate each Skill independently first, then compare evidence and trade-offs. Do not force a winner when the Skills serve different use cases.
  • Revision review: compare the current package with the previous version and judge whether the change improves real behavior, not only presentation.

Read the complete package

  1. Read SKILL.md completely.
  2. List every bundled reference, script, asset, and agent metadata file.
  3. Follow every resource link and loading instruction that affects behavior.
  4. Check whether referenced files exist and whether shipped scripts or examples are plausible and tested.
  5. Distinguish observed facts from assumptions about how an agent might behave.

Do not score a package from its frontmatter or first screen alone.

Apply the hard gates first

Report these before subjective scoring:

  • missing or invalid SKILL.md frontmatter;
  • folder name and name mismatch;
  • vague description that does not communicate function and trigger context;
  • broken resource paths or routing instructions;
  • advertised capabilities that the package does not implement;
  • unsafe scripts or instructions without proportionate guardrails;
  • contradictory requirements that make correct execution impossible.

A polished package with a hard-gate failure is not ready.

Evaluate actual knowledge value

Classify substantive sections using three labels:

  • Expert: non-obvious decisions, trade-offs, failure modes, domain procedures, or local knowledge the AI model is unlikely to supply reliably.
  • Activation: knowledge the AI model may know but benefits from seeing at the right moment.
  • Redundant: generic explanation, obvious advice, or repeated content that consumes context without changing behavior.

Use this classification diagnostically. Do not invent precise paragraph percentages when the boundary is ambiguous.

Ask:

  • What would the agent do better after loading this Skill?
  • Which content changes a decision rather than merely describing the domain?
  • Which instructions encode experience that is difficult to reconstruct on demand?
  • Which sections could disappear without affecting the result?

Review trigger quality

The description is the primary activation surface. It should make clear:

  • what the Skill does;
  • when it should be used;
  • which concrete tasks, artifacts, file types, tools, or user phrases should trigger it;
  • when a neighboring Skill is a better choice, if overlap is likely.

Do not reward keyword stuffing. A long description that triggers everywhere creates routing conflicts instead of discoverability.

Show full SKILL.md (394 more words)Show less

Review package architecture

Check whether the package uses progressive disclosure intentionally:

  • Keep core decisions and routing in SKILL.md.
  • Put detailed variants, schemas, examples, and domain references in bundled resources.
  • State when to load each resource and when not to load it.
  • Avoid orphan references that are shipped but never routed.
  • Avoid a large SKILL.md that loads an entire manual for every request.
  • Avoid fragmentation that forces the agent to open many tiny files for one simple task.

The target is the smallest context that preserves correct behavior, not a fixed line count.

Review behavioral quality

Look for:

  • explicit decision criteria where several paths are possible;
  • appropriate freedom for the task's risk and variability;
  • specific anti-patterns with reasons, not vague warnings;
  • failure handling, fallback paths, and realistic edge cases;
  • evidence that scripts and examples work as described;
  • boundaries that prevent the Skill from competing with unrelated or more specific Skills;
  • instructions that remain practical in the target environment.

Score only with evidence

For a full audit, use the 120-point rubric in rubric.md. For every dimension:

  1. cite concrete package evidence;
  2. name the behavioral consequence;
  3. assign a score using the published anchors;
  4. state what would materially raise the score.

Do not award points for length, polish, confident wording, or the number of bundled files.

Prioritize findings

Use four severity levels:

  • Blocker: prevents reliable activation, execution, safety, or installation.
  • High: materially degrades common outcomes or creates misleading behavior.
  • Medium: causes avoidable friction, context waste, or incomplete coverage.
  • Low: worthwhile refinement with limited behavioral effect.

Prioritize by expected improvement to real use, not by ease of editing.

Output format

Lead with the verdict.

markdown
# Skill Review: [name]

## Verdict
[Ready / usable with changes / not ready, plus one sentence explaining why]

## Applicability
[What context the judgment assumes and what it does not establish]

## Findings
1. **[Severity] [Issue]**
   - Evidence: [file and relevant location]
   - Effect: [how behavior or usability suffers]
   - Change: [specific improvement]

## Score
[Include only for a full audit: dimension table, total, and grade]

## Highest-value next changes
1. [First change]
2. [Second change]
3. [Third change]

## What is already strong
[Only evidence-backed strengths worth preserving]

Keep praise short. Make weaknesses specific enough that another agent can implement the improvement without reconstructing the audit.

Bad review patterns

  • Giving a high score because the package looks professional.
  • Treating every long procedure as expert knowledge.
  • Calling a practice-derived rubric an official standard.
  • Penalizing all repetition without considering activation value.
  • Recommending more files without a loading strategy.
  • Recommending fewer files when separation protects context.
  • Reporting stylistic preferences as behavioral defects.
  • Evaluating an example link inside a fenced code block as a broken live link.
  • Ignoring mismatches between advertised and implemented capabilities.

Final question

Ask: Does this package reliably transfer useful judgment or operational knowledge at the moment an agent needs it?

If the answer is unclear, the review is not finished.

© shareAI-lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in agent-development/skill-judge of shareAI-lab/lab-skills.

  • SKILL.md
  • references/rubric.md

Open the folder on GitHubat commit becee99

Compare with similar skills

Skill Judge next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Judge compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Judge this skillshareAI-lab/lab-skills315—~1.9kAutomated safety check: PassApache-2.0
Agent Evaluationseb1n/awesome-ai-agent-skills206—~1.4kAutomated safety check: PassMIT
Agent Launcher Orchestratoralirezarezvani/claude-skills28k—~1.3kAutomated safety check: PassMIT
Skill Reviewerdaymade/claude-code-skills1.4k—~2.1kAutomated safety check: PassMIT
Validate Skillmdjeremylongshore/tons-of-skills-marketplace2.8k—~6.2kAutomated safety check: PassMIT
Design EvaluationSeanJ1ang/design-judge-skills712—~3.1kAutomated safety check: PassApache-2.0

Similar skills

  • Agent Evaluation

    seb1n/awesome-ai-agent-skills

    Design reproducible evaluations for AI agents with representative task sets, explicit rubrics, appropriate graders, baselines, regression gates, and failure analysis.

    206 GitHub stars~1.4k tokensUpdated 2 mo ago
    EducationAuto-check passed
  • Agent Launcher Orchestrator

    alirezarezvani/claude-skills

    A skill your agent uses when a user wants to build, launch, grade, or schedule a Claude Managed Agent (CMA) in their own Anthropic account — "build me an agent", "launch this as a managed agent"…

    28k GitHub stars~1.3k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Skill Reviewer

    daymade/claude-code-skills

    Reviews skill quality with evidence-based design rubrics and read-only batch inventories.

    1.4k GitHub stars~2.1k tokensUpdated yesterday
    EducationAuto-check passed
  • Validate Skillmd

    jeremylongshore/tons-of-skills-marketplace

    Validate a SKILL.md file against the four-tier validation system: Tier 0 (locate), Tier 1 (standard or marketplace grading per the IS 100-point rubric), Tier 2 (static production gate —…

    2.8k GitHub stars~6.2k tokensUpdated yesterday
    EducationAuto-check passed
  • Design Evaluation

    SeanJ1ang/design-judge-skills

    Evaluate one design or a user-approved maturity-mapped batch through a transparent evidence-based rubric.

    712 GitHub stars~3.1k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Create Custom Grader

    NVIDIA/SkillEvaluator

    Official

    A skill your agent uses when converting an existing benchmark, rubric, verifier, task YAML/JSON, or domain check into SkillEvaluator BYOG/BYOT custom evaluation.

    554 GitHub stars~2.1k tokensUpdated yesterday
    EducationAuto-check passed

More from shareAI-lab/lab-skills

All 10 skills in this repo
  • Agent Builder

    shareAI-lab/lab-skills

    Helps design and build AI agents for any domain around a minimal loop of capabilities, knowledge and context, adding planning or subagents only when needed.

    315 GitHub stars~1.4k tokensUpdated 25 days ago
    Auto-check passed
  • Deep Architecture Research

    shareAI-lab/lab-skills

    Deeply research technical architecture, source code, mechanisms, SDKs, frameworks, project comparisons, and system-design options across repositories, history, official docs, issues, discussions…

    315 GitHub stars~1.3k tokensUpdated 25 days ago
    Auto-check passed
  • Neural Mechanism Research

    shareAI-lab/lab-skills

    Research why neural architectures and training methods work through forward computation, geometry, gradients, optimization dynamics, historical experiments, and competing explanations.

    315 GitHub stars~2.3k tokensUpdated 25 days ago
    Auto-check passed
  • Review AI Conversations

    shareAI-lab/lab-skills

    Recover and review local human-AI conversations from Claude Code, Codex, opencode, Grok Build, and Cursor.

    315 GitHub stars~1.1k tokensUpdated 25 days ago
    Auto-check passed
  • Understanding First Report

    shareAI-lab/lab-skills

    Reconstruct and report long-running or multi-turn research, architecture questions, reviews, decisions, completion results, and status as a clear, self-contained brief.

    315 GitHub stars~4.3k tokensUpdated 25 days ago
    Auto-check passed
  • Vibe Coding

    shareAI-lab/lab-skills

    Transform an AI agent into a disciplined software development partner with strong judgment, transparent decisions, proportionate verification, and craftsmanship.

    315 GitHub stars~2k tokensUpdated 25 days ago
    Auto-check passed

Questions about Skill Judge

What does Skill Judge do?

Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples. Skill Judge is an agent skill from shareAI-lab/lab-skills. Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

When should I use Skill Judge?

Skill Judge fits situations like: improving SKILL.md files and skill packages; tasks that involve Quizzes and assessments; tasks that involve Skill authoring.

How do I install Skill Judge in Claude Code?

Run `npx skills add shareAI-lab/lab-skills --skill skill-judge -a claude-code`. Or copy the skill folder (agent-development/skill-judge in shareAI-lab/lab-skills) into .claude/skills/skill-judge in your project. Claude Code loads it when a task matches its description.

How do I install Skill Judge in Codex?

Run `npx skills add shareAI-lab/lab-skills --skill skill-judge -a codex`. Or copy the skill folder (agent-development/skill-judge in shareAI-lab/lab-skills) into .agents/skills/skill-judge in your project. Codex loads it when a task matches its description.

Can I use Skill Judge in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add shareAI-lab/lab-skills --skill skill-judge -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-judge, .gemini/skills/skill-judge, .github/skills/skill-judge and .opencode/skills/skill-judge in your project.

What does Skill Judge need to run?

SKILL.md names no scripts, command-line tools or credentials: Skill Judge is instructions for the agent only.

Does Skill Judge access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Judge safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Judge use?

Skill Judge is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Judge use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 2k tokens, read only when the agent opens those files.

What are the alternatives to Skill Judge?

Skills that share tags, products or a category with Skill Judge: Agent Evaluation (seb1n/awesome-ai-agent-skills, 206 stars), Agent Launcher Orchestrator (alirezarezvani/claude-skills, 28k stars), Skill Reviewer (daymade/claude-code-skills, 1.4k stars) and Validate Skillmd (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Judge?

shareAI-lab (a GitHub organization) maintains it in shareAI-lab/lab-skills, which has 315 GitHub stars. The repository holds 10 skills in this directory. The repository was last updated on September 16, 2026.

Source: shareAI-lab/lab-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.