Agent skill

Skill Grader

by curiositech in curiositech/some_claude_skills

Evaluates Claude Agent Skills on 10 quality axes with letter grades (A+ through F) and specific improvement recommendations.

MITAuto-check passedDevelopment

Install Skill Grader

skills CLI
$ npx skills add curiositech/some_claude_skills --skill skill-grader -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install curiositech/some_claude_skills skill-grader --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/skill-grader .claude/skills/skill-grader && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-grader
GitHub stars
243
Token cost
~2.8k tokens
SKILL.md length
1,067 words
Files
3
Skills in repo
109
Repo updated
First seen
Licence
MIT

At a glance

Evaluates Claude Agent Skills on 10 quality axes with letter grades (A+ through F) and specific improvement recommendations.

  • Works in 6 steps: Read the entire skill folder — SKILL.md,… → Score each axis — Use the rubric below… → Convert to letter grade — See grade scale → …
  • Auditing a skill
  • SKILL.md covers When to Use, Grading Process, The 10 Evaluation Axes and Grade Scale, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Skill Grader is an agent skill from curiositech/some_claude_skills. Evaluates Claude Agent Skills on 10 quality axes with letter grades (A+ through F) and specific improvement recommendations. Use when auditing a skill, comparing skills, prioritizing improvements, or performing quality control on a skill library. Activate on "grade skill", "evaluate skill", "skill quality", "skill audit", "skill review", "rate skill". NOT for creating skills (use skill-architect), grading code quality, or evaluating non-skill documents.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files (for example `.claude-plugin/plugin.json` and `CHANGELOG.md`).

It sits in Development, covering Accessibility and Code quality. The repository describes itself as: Claude skills that make my life easier. The licence is MIT.

When your agent uses it

  • Auditing a skill
  • Comparing skills
  • Prioritizing improvements
  • Performing quality control on a skill library

Example prompts

  • “grade skill”
  • “evaluate skill”
  • “skill quality”
  • “/skill-grader”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Read the entire skill folder — SKILL.md, all references, scripts, CHANGELOG, README
  2. Score each axis — Use the rubric below (0-100 per axis)
  3. Convert to letter grade — See grade scale
  4. Compute overall grade — Weighted average (Description and Scope are 2x weight)
  5. Write 1-3 specific improvements per axis scoring below B+
  6. Produce the grading report in the output format below

What it can do on your machine

Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown and mermaid).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Grader loads about 2.8k tokens when it runs. Until then it costs about 118 tokens; SKILL.md has 1,067 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~118
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its MIT licence (© curiositech). 1,067 words, ~2,781 tokens.

Download SKILL.mdSave it as .claude/skills/skill-grader/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
skill-grader
description
Evaluates Claude Agent Skills on 10 quality axes with letter grades (A+ through F) and specific improvement recommendations. Use when auditing a skill, comparing skills, prioritizing improvements, or performing quality control on a skill library. Activate on "grade skill", "evaluate skill", "skill quality", "skill audit", "skill review", "rate skill". NOT for creating skills (use skill-architect), grading code quality, or evaluating non-skill documents.
allowed-tools
Read, Grep, Glob
argument-hint
[skill-path]
metadata.category
Productivity & Meta
metadata.tags
grader, grade-skill, evaluate-skill, skill-quality

Skill Grader

Structured evaluation rubric for Claude Agent Skills. Produces letter grades (A+ through F) on 10 axes plus an overall grade, with specific improvement recommendations for each axis.

Designed for sub-agents and non-expert reviewers who need a mechanical, repeatable process for assessing skill quality without deep domain expertise.


When to Use

✅ Use for:

  • Auditing a single skill's quality
  • Comparing skills against each other
  • Prioritizing which skills to improve first
  • Quality control sweeps across a skill library
  • Generating improvement roadmaps

❌ NOT for:

  • Creating new skills (use skill-architect)
  • Grading code quality or non-skill documents
  • Evaluating agent performance (different from skill quality)

Grading Process

mermaid
flowchart TD
  A[Read SKILL.md + all files] --> B[Score each of 10 axes]
  B --> C[Assign letter grade per axis]
  C --> D[Compute overall grade]
  D --> E[Write improvement recommendations]
  E --> F[Produce grading report]
Step-by-Step
  1. Read the entire skill folder — SKILL.md, all references, scripts, CHANGELOG, README
  2. Score each axis — Use the rubric below (0-100 per axis)
  3. Convert to letter grade — See grade scale
  4. Compute overall grade — Weighted average (Description and Scope are 2x weight)
  5. Write 1-3 specific improvements per axis scoring below B+
  6. Produce the grading report in the output format below

The 10 Evaluation Axes

Axis 1: Description Quality (Weight: 2x)

Does the description follow [What] [When] [Keywords]. NOT for [Exclusions]?

GradeCriteria
ASpecific verb+noun, domain keywords users would type, 2-5 explicit NOT exclusions, 25-50 words
BHas keywords and NOT clause, but slightly vague or missing synonym coverage
CToo generic, missing NOT clause, or >100 words of process detail
DSingle vague sentence ("helps with X") or name/description mismatch
FMissing or empty description
Axis 2: Scope Discipline (Weight: 2x)

Is the skill narrowly focused on one expertise type, or a catch-all?

GradeCriteria
AOne clear expertise domain, "When to Use" and "NOT for" sections both present and specific
BMostly focused, minor boundary ambiguity
CCovers 2-3 related but distinct domains, should probably be split
DCatch-all skill ("helps with anything related to X")
FNo scope boundaries defined at all
Axis 3: Progressive Disclosure

Does the skill follow the three-layer architecture (metadata → SKILL.md → references)?

GradeCriteria
ASKILL.md <300 lines, heavy content in references, reference index in SKILL.md with 1-line descriptions
BSKILL.md <500 lines, some references used, index present
CSKILL.md >500 lines, or all content inlined with no references
DSKILL.md >800 lines, or references exist but aren't indexed in SKILL.md
FSingle massive file with no structure
Axis 4: Anti-Pattern Coverage

Does the skill encode expert knowledge that prevents common mistakes?

GradeCriteria
A3+ anti-patterns with Novice/Expert/Timeline template, LLM-mistake notes
B1-2 anti-patterns with clear explanation
CAnti-patterns mentioned but no structured template
DNo anti-patterns, just positive instructions
FContains advice that IS an anti-pattern (outdated, harmful)
Axis 5: Self-Contained Tools

Does the skill include working tools (scripts, MCPs, subagents)?

GradeCriteria
AWorking scripts with CLI interface, error handling, dependency docs; OR valid "no tools needed" justification
BScripts exist and work but lack error handling or docs
CScripts referenced but are templates/pseudocode
DPhantom tools (SKILL.md references files that don't exist)
FReferences non-existent tools AND no acknowledgment

Note: Not every skill needs tools. A pure decision-tree skill can score A if tools aren't applicable.

Axis 6: Activation Precision

Would the skill activate correctly on relevant queries and stay silent on irrelevant ones?

GradeCriteria
ADescription has specific keywords matching user language, clear NOT clause, no obvious false-positive vectors
BGood keywords, minor false-positive risk
CGeneric keywords that overlap with other skills
DNo specific keywords, or NOT clause contradicts intended use
FDescription would cause constant false activation
Axis 7: Visual Artifacts

Does the skill use Mermaid diagrams, code examples, and tables effectively?

GradeCriteria
ADecision trees as Mermaid flowcharts, tables for comparisons, code examples for concrete patterns
BSome diagrams or tables, but key decision trees still in prose
CTables used but no Mermaid diagrams for processes
DProse-only, no visual structure
FWall of text with no formatting aids
Show full SKILL.md (422 more words)Show less
Axis 8: Output Contracts

Does the skill define what it produces in a format consumable by other agents?

GradeCriteria
AExplicit output format (JSON schema, markdown template, or structured sections), subagent-consumable
BOutput format implied but not explicitly documented
CNo output format, but content is structured enough to infer
DUnstructured prose output expected
FN/A (pure reference skill) — exempt from this axis
Axis 9: Temporal Awareness

Does the skill track when knowledge was current and what has changed?

GradeCriteria
ATimelines in anti-patterns, "as of [date]" markers, CHANGELOG with dates
BSome temporal context, CHANGELOG exists
CNo dates on knowledge, but CHANGELOG exists
DNo temporal context anywhere, knowledge could be stale
FContains demonstrably outdated advice without warning
Axis 10: Documentation Quality

README, CHANGELOG, and reference organization.

GradeCriteria
AREADME with quick start, CHANGELOG with dated versions, references well-organized with clear filenames
BREADME and CHANGELOG exist, references present
CSKILL.md is the only file, but it's well-structured
DNo README, no CHANGELOG, disorganized references
FSKILL.md is the only file and it's poorly structured

Grade Scale

LetterScore RangeMeaning
A+97-100Exemplary — sets the standard
A93-96Excellent — minor improvements possible
A-90-92Very good — a few small gaps
B+87-89Good — notable room for improvement
B83-86Solid — several areas need work
B-80-82Above average — meaningful gaps
C+77-79Average — significant improvements needed
C73-76Below average — major gaps
C-70-72Barely adequate
D+67-69Poor — fundamental issues
D63-66Very poor — needs major rework
D-60-62Near-failing quality
F<60Failing — start over

Overall Grade Computation

Axes 1 (Description) and 2 (Scope) carry 2x weight. All others carry 1x weight. If Axis 8 (Output Contracts) is marked exempt, remove it from the calculation.

Overall = (2×Axis1 + 2×Axis2 + Axis3 + Axis4 + Axis5 + Axis6 + Axis7 + Axis8 + Axis9 + Axis10) / 12

Convert the numeric average to a letter grade using the scale above.


Output Format

Produce this exact structure:

markdown
# Skill Grading Report: [skill-name]

**Graded**: [date]
**Overall Grade**: [letter] ([score]/100)

## Axis Grades

| # | Axis | Grade | Score | Key Finding |
|---|------|-------|-------|-------------|
| 1 | Description Quality | [grade] | [score] | [1-line finding] |
| 2 | Scope Discipline | [grade] | [score] | [1-line finding] |
| 3 | Progressive Disclosure | [grade] | [score] | [1-line finding] |
| 4 | Anti-Pattern Coverage | [grade] | [score] | [1-line finding] |
| 5 | Self-Contained Tools | [grade] | [score] | [1-line finding] |
| 6 | Activation Precision | [grade] | [score] | [1-line finding] |
| 7 | Visual Artifacts | [grade] | [score] | [1-line finding] |
| 8 | Output Contracts | [grade] | [score] | [1-line finding] |
| 9 | Temporal Awareness | [grade] | [score] | [1-line finding] |
| 10 | Documentation Quality | [grade] | [score] | [1-line finding] |

## Top 3 Improvements (Highest Impact)

1. **[Axis]: [Specific action]** — [Why this matters, expected grade improvement]
2. **[Axis]: [Specific action]** — [Why this matters, expected grade improvement]
3. **[Axis]: [Specific action]** — [Why this matters, expected grade improvement]

## Detailed Notes

### [Axis name] ([grade])
[2-3 sentences of specific feedback with examples from the skill]

[Repeat for each axis scoring below B+]

Quick Grading (Abbreviated)

For rapid triage across many skills, produce only:

markdown
| Skill | Overall | Desc | Scope | Disc | Anti | Tools | Activ | Visual | Output | Temp | Docs |
|-------|---------|------|-------|------|------|-------|-------|--------|--------|------|------|
| [name] | [grade] | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |

Anti-Patterns in Grading

Grade Inflation

Wrong: Giving B+ because "it's pretty good" without checking criteria. Right: Match observations to the rubric table literally. If the description lacks a NOT clause, it cannot score above C on Axis 1.

Missing Context

Wrong: Grading a pure decision-tree skill poorly on Axis 5 (tools) because it has no scripts. Right: Mark Axis 5 as "A — tools not applicable for this skill type."

Ignoring Phantoms

Wrong: Scoring Axis 5 as B because scripts are "referenced." Right: Actually check if every referenced file exists. If scripts/validate.py is mentioned but doesn't exist, that's D.

© curiositech, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files in .claude/skills/skill-grader of curiositech/some_claude_skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • CHANGELOG.md

Open the folder on GitHubat commit 6713fc7

Compare with similar skills

Skill Grader next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Grader compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Grader this skillcuriositech/some_claude_skills243—~2.8kAutomated safety check: PassMIT
Reviewquran/quran.com-frontend-next1.9k—~1.1kAutomated safety check: PassNone
Code ReviewOrange-OpenSource/Orange-Boosted-Bootstrap222—~846Automated safety check: PassMIT
Accessibility Cleanupnwjs/chromium.src160—~1.1kAutomated safety check: PassBSD-3-Clause
React Doctormakeplane/plane61k12 repos~657Automated safety check: PassAGPL-3.0
Codebase Review SwarmZaxbyHub/opencode-swarm490—~2.8kAutomated safety check: PassMIT

Similar skills

  • Review

    quran/quran.com-frontend-next

    Reviews PR(s) using comprehensive review guidelines including security, correctness, clean code, TypeScript, React patterns, i18n/RTL, performance, and accessibility.

    1.9k GitHub stars~1.1k tokensUpdated 5 mo ago
    DevelopmentAuto-check passed
  • Code Review

    Orange-OpenSource/Orange-Boosted-Bootstrap

    OUDS Web compliance code review. An agent skill from Orange-OpenSource/Orange-Boosted-Bootstrap.

    222 GitHub stars~846 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Accessibility Cleanup

    nwjs/chromium.src

    Finds common violations of the Android accessibility API in the Clank App Java code and attempts to address them (e.g.

    160 GitHub stars~1.1k tokensUpdated 5 days ago
    Frontend & DesignAuto-check passed
  • React Doctor

    makeplane/plane

    Scans React code for lint, accessibility, bundle size and architecture issues, reports a health score and checks that changes do not lower it.

    61k GitHub starsUsed in 12 repos~657 tokens
    Frontend & DesignAuto-check passed
  • Codebase Review Swarm

    ZaxbyHub/opencode-swarm

    Runs an evidence-gated, quote-grounded audit of a codebase for security, QA, accessibility, performance and more, and writes a verified report without changing source files.

    490 GitHub stars~2.8k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Best Practices

    tech-leads-club/agent-skills

    Apply modern web development best practices for security, compatibility, and code quality.

    7k GitHub stars~3.2k tokensUpdated 18 days ago
    Frontend & DesignAuto-check passed

More from curiositech/some_claude_skills

All 109 skills in this repo
  • Crisis Detection Intervention AI

    curiositech/some_claude_skills

    Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.

    243 GitHub starsUsed in 3 repos~3.8k tokens
    Auto-check passed
  • Form Validation Architect

    curiositech/some_claude_skills

    End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.

    243 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Competitive Cartographer

    curiositech/some_claude_skills

    Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.

    243 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • GitHub Actions Pipeline Builder

    curiositech/some_claude_skills

    Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.

    243 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: notes
  • Computer Vision Pipeline

    curiositech/some_claude_skills

    Build production computer vision pipelines for object detection, tracking, and video analysis.

    243 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed
  • Design Archivist

    curiositech/some_claude_skills

    Long-running design anthropologist that builds comprehensive visual databases from 500-1000 real-world examples, extracting color palettes, typography patterns, layout systems, and interaction…

    243 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about Skill Grader

What does Skill Grader do?

Evaluates Claude Agent Skills on 10 quality axes with letter grades (A+ through F) and specific improvement recommendations. Skill Grader is an agent skill from curiositech/some_claude_skills. Evaluates Claude Agent Skills on 10 quality axes with letter grades (A+ through F) and specific improvement recommendations.

When should I use Skill Grader?

Skill Grader fits situations like: auditing a skill; comparing skills; prioritizing improvements; performing quality control on a skill library.

How do I install Skill Grader in Claude Code?

Run `npx skills add curiositech/some_claude_skills --skill skill-grader -a claude-code`. Or copy the skill folder (.claude/skills/skill-grader in curiositech/some_claude_skills) into .claude/skills/skill-grader in your project. Claude Code loads it when a task matches its description.

How do I install Skill Grader in Codex?

Run `npx skills add curiositech/some_claude_skills --skill skill-grader -a codex`. Or copy the skill folder (.claude/skills/skill-grader in curiositech/some_claude_skills) into .agents/skills/skill-grader in your project. Codex loads it when a task matches its description.

Can I use Skill Grader in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill skill-grader -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-grader, .gemini/skills/skill-grader, .github/skills/skill-grader and .opencode/skills/skill-grader in your project.

What does Skill Grader need to run?

SKILL.md names no scripts, command-line tools or credentials: Skill Grader is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Grep, Glob.

Does Skill Grader access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Grader safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Grader use?

Skill Grader is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Grader use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Grader?

Skills that share tags, products or a category with Skill Grader: Review (quran/quran.com-frontend-next, 1.9k stars), Code Review (Orange-OpenSource/Orange-Boosted-Bootstrap, 222 stars), Accessibility Cleanup (nwjs/chromium.src, 160 stars) and React Doctor (makeplane/plane, 61k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Grader?

curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 243 GitHub stars. The repository holds 109 skills in this directory. The repository was last updated on September 6, 2026.

Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.