Agent skill

Skill Score

by davekilleen in davekilleen/Dex

Grade a Dex skill against the shape-aware quality rubric and report a ship/revise/no verdict with the exact fixes.

MITAuto-check passedEducation

Install Skill Score

skills CLI
$ npx skills add davekilleen/Dex --skill skill-score -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davekilleen/Dex skill-score --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davekilleen/Dex.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/skill-score .claude/skills/skill-score && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-score
GitHub stars
493
Token cost
~2.6k tokens
SKILL.md length
1,360 words
Files
3 (incl. scripts)
Skills in repo
61
Repo updated
First seen
Licence
MIT

At a glance

Grade a Dex skill against the shape-aware quality rubric and report a ship/revise/no verdict with the exact fixes.

  • Works in 6 steps: Prefer the script → Load and classify → Run the checks → …
  • You finish writing
  • SKILL.md covers Arguments, Step 0: Prefer the script, The hard gates (any one fails… and Tier 1 — universal must-pass…, plus 9 more sections
  • Runs Python scripts from its folder; calls python3

What it does

Skill Score is an agent skill from davekilleen/Dex. Grade a Dex skill against the shape-aware quality rubric and report a ship/revise/no verdict with the exact fixes. Use when you finish writing or editing a skill, when create-skill hands off a new package, before shipping a first-party skill, or when the user asks "is this skill any good / will it fire / score my skill". Also use proactively right after any SKILL.md is created or its description changes. Not for authoring a new skill from scratch (use create-skill) or fixing broken YAML frontmatter alone…

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including scripts (for example `evals/trigger-cases.yaml` and `scripts/score_skill.py`).

It sits in Education, covering Quizzes and assessments. The repository describes itself as: Your AI Chief of Staff — a personal operating system starter kit that adapts to your role. No coding required. The licence is MIT.

When your agent uses it

  • You finish writing
  • Editing a skill
  • Create-skill hands off a new package
  • Before shipping a first-party skill

Example prompts

  • “is this skill any good / will it fire / score my skill”
  • “/skill-score”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the step headings in SKILL.md.

  1. Prefer the script
  2. Load and classify
  3. Run the checks
  4. Model judgment (the parts a script can't)
  5. Run the trigger evals
  6. Verdict

What it can do on your machine

Read from SKILL.md and the folder at commit 227f78e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Score loads about 2.6k tokens when it runs. Until then it costs about 155 tokens; SKILL.md has 1,360 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~155
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from davekilleen/Dex at commit 227f78e, republished under its MIT licence (© davekilleen). 1,360 words, ~2,581 tokens.

Download SKILL.mdSave it as .claude/skills/skill-score/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
skill-score
description
Grade a Dex skill against the shape-aware quality rubric and report a ship/revise/no verdict with the exact fixes. Use when you finish writing or editing a skill, when create-skill hands off a new package, before shipping a first-party skill, or when the user asks "is this skill any good / will it fire / score my skill". Also use proactively right after any SKILL.md is created or its description changes. Not for authoring a new skill from scratch (use create-skill) or fixing broken YAML frontmatter alone (create-skill's validator does that); skill-score judges architecture and routing, not just format.
<!-- Generated from `.claude/skills/skill-score/SKILL.md` by `scripts/generate-agents-skills.py`. Do not edit. -->

/skill-score

Grade a skill the way the router and a real user will experience it — then say plainly whether it ships, and if not, exactly what to fix.

Two things make a skill good: it fires when it should (the description is the router) and it does the job safely and well when it fires (the body is a contract, not a how-to essay). This skill scores both, applies the hard safety gates that override the number, and returns a verdict.

Governing principle: hard on Core, gentle on the user's own creations. A Core / first-party skill that scores below the bar does not ship — that is a hard gate we hold ourselves to. A skill the user wrote for themselves is coached, never blocked: show the score, name the one change that would make it fire, offer to make it — but always create/keep their skill if they want it.


Arguments

$TARGET: Optional.

  • A skill name or path (triage, .agents/skills/triage/) → score that one skill.
  • --all → portfolio pass over every shipped skill: routing collisions, stale references, bad frontmatter, and per-skill grades in one table.
  • Empty → score the skill most recently created or edited this session; if none, ask which.

$ORIGIN: Optional. core (first-party, hard gate) or user (coach-don't-block). If omitted, infer: a -custom suffix or a path under .claude/skills-custom/ ⇒ user; anything else shipped in the repo ⇒ core.


Step 0: Prefer the script

The scoring math (Tier-1/Tier-2 point tally, length checks, when-trigger detection, reference-path existence, collision proximity) is deterministic. Run the script, don't recompute by hand:

python3 .agents/skills/skill-score/scripts/score_skill.py <path-to-skill-dir-or-SKILL.md> [--all] [--origin core|user] [--json]

The script returns the mechanical sub-scores, every hard-gate check it can decide from files alone, and a provisional grade. Then you (the model) do the judgment-only parts the script flags as NEEDS_MODEL: is the description actually distinguishable from its nearest neighbor in meaning (not just string distance)? Does the body inspect its own output before claiming success? Are the anti-triggers pointing at the right neighbor? Merge the script's tally with your judgment calls into the final verdict.

If the script cannot run (no Python), fall back to scoring by hand against the rubric below — say so in the output.


The hard gates (any one fails ⇒ verdict is NO, whatever the number)

These come straight from the ratified synthesis. A skill cannot ship if:

  1. Indistinguishable from a neighbor. Its description cannot be told apart from its nearest existing skill — the router would coin-flip between them. (Fix: sharper outcome + anti-trigger naming that neighbor.)
  2. Destructive / external / publish action without authority. It can delete, overwrite, send, post, or publish outside the user's vault without an explicit confirmation gate in the body.
  3. PII into a shared artifact. Secrets or personal content can flow into anything that leaves the machine (a published DexDiff profile, an uploaded page, an external message) without a redaction or confirmation step.
  4. Claims success without inspecting output. It tells the user "done / created / fixed" without reading back the thing it just produced or the tool result that proves it.

For a user-origin skill, gates 2–4 still warn loudly and gate 1 becomes advice ("this won't fire on its own unless we distinguish it from X — want me to?") — but they do not refuse to create the user's own skill. For a core-origin skill, any gate failure blocks the ship.


Tier 1 — universal must-pass (~60 pts). Every skill, every shape.

#CriterionPtsWhat "pass" looks like
T1.1Description carries a WHEN trigger12Frontmatter names concrete situations AND user phrases ("when the user says…"). The word when/whenever is necessary but not sufficient — it must describe a real firing situation.
T1.2Description has an anti-trigger10"Not for X; use Y" naming the real nearest neighbor, so the router can disambiguate.
T1.3Description states the outcome, not the mechanism8Leads with what the user gets, in plain language — no internal function/file names.
T1.4Thin body; detail externalized8Body is a router + contract. Long procedural/reference detail lives in references/. Soft cap ~200 lines; over ~350 is a fail unless justified.
T1.5Named quality bar + anti-patterns8The body says what "good output" is and names at least one failure mode to avoid.
T1.6Truthful degradation8When a prerequisite/tool is missing, it says so honestly (or skips silently by design) — never fakes success or invents a result.
T1.7Legible + composes6Refers to people/skills/artifacts by name not id; points to sibling skills instead of re-implementing them.

Tier-1 floor: a core skill must clear ≥50/60 on Tier 1 regardless of Tier 2.

Tier 2 — situational, scored only if the shape applies

Classify the skill's shape first, then score only the matching block. Do not cargo-cult these onto a plain conversational workflow — an inapplicable criterion is scored N/A, not zero.

ShapeCriterionPts
setup / integrationIdempotent writes + a re-run/repair path; states the setup contract10
dependent (needs a tool/feature)Doctor/degradation ladder: detect → explain → fix-path10
generative (produces user-facing content)Named bar + anti-slop rules + inspect-real-output-before-claiming12
multi-session / statefulState externalized + sized to a session; consider disable-model-invocation only if destructive/dev8
script-bearingAgent-native output contract: --diagnose/--dry-run, exit codes, machine-readable result10

Tier-2 available points vary by shape; normalize the final grade to 100 (earned / applicable * 100). The script does this.


Show full SKILL.md (500 more words)Show less

Grade bands

  • ≥85 — SHIP. Clears the bar. (Core: may merge. User: great, ship it.)
  • 70–84 — REVISE. Close. Return the specific criteria that lost points; usually 1–2 description fixes.
  • <70 — NO. Not ready. For core: does not ship. For user: coach, offer to sharpen, still create if they insist.
  • Any hard-gate failure — NO, printed above the number with the gate named.

Step 1: Load and classify

Read the target SKILL.md (+ its dir). Determine origin (core/user) and shape (workflow / setup / generative / multi-session / research / script-bearing — a skill can be more than one). Note which Tier-2 blocks apply.

Step 2: Run the checks

Run .agents/skills/skill-score/scripts/score_skill.py. It returns: Tier-1 sub-scores, applicable Tier-2 sub-scores, hard-gate results it can decide, referenced-path existence, nearest-neighbor by description proximity, and a provisional grade.

Step 3: Model judgment (the parts a script can't)

For each item the script marks NEEDS_MODEL, decide it and record one line of evidence:

  • Distinguishable-in-meaning? Read the nearest-neighbor's description; would you pick the right one from a real user utterance? If not → gate 1.
  • Inspects its own output? Find the place the skill declares done. Does it read back what it made? If not → gate 4.
  • Anti-trigger points at the true collision? The named neighbor must be the one users would actually confuse it with.
  • Degradation honest? Trace the missing-prereq path.

Step 4: Run the trigger evals

Load evals/trigger-cases.yaml for the skill (see shape below). For each case, decide whether this description would fire. Positives must fire; negatives/collisions must NOT (they should route to the named neighbor); the ambiguous case should ask; the missing-prereq case should degrade honestly; the failure-recovery case should not claim false success. Report pass/fail per case. A core skill that fails a positive or a collision case cannot score ≥85.

Step 5: Verdict

Print, in this order:

  1. VERDICT: SHIP / REVISE / NO + the numeric grade.
  2. Any hard-gate failures, each named, with the one change that clears it.
  3. Lost points, grouped Tier-1 / Tier-2, each with the concrete fix (ideally the rewritten line).
  4. Trigger-eval results (X/8 passed; list failures).
  5. For user-origin: the coaching line — never a refusal. For core-origin: the ship decision.

Keep it short and actionable. The output is a fix list, not an essay.


--all portfolio mode

Score every shipped skill and emit one report:

  • Grade table (skill | origin | shape | grade | verdict).
  • Routing collisions: pairs whose descriptions are too close to disambiguate.
  • Stale references: references//script paths named but absent.
  • Bad frontmatter: missing name/description, malformed YAML.
  • Portfolio summary: how many core skills are below the bar (the ship-blocking list).

This is the health check behind the description-rewrite / consolidation program — run it before and after the overhaul to measure the win.


Anti-patterns (for this skill itself)

  • Do not pass a skill just because its YAML is valid — that is create-skill's floor, not the bar.
  • Do not invent a Tier-2 penalty for a shape that doesn't apply.
  • Do not block a user's own skill. Coach it.
  • Do not claim a grade without running the checks (or saying you fell back to hand-scoring).

© davekilleen, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (scripts) in .agents/skills/skill-score of davekilleen/Dex.

  • SKILL.md
  • evals/trigger-cases.yaml
  • scripts/score_skill.py

Open the folder on GitHubat commit 227f78e

Compare with similar skills

Skill Score next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Score compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Score this skilldavekilleen/Dex493—~2.6kAutomated safety check: PassMIT
DeepTutor CLIHKUDS/DeepTutor41k—~2.8kAutomated safety check: PassApache-2.0
AI Engineering Placement Quizrohitg00/ai-engineering-from-scratch66k—~2kAutomated safety check: PassMIT
Codebase to Coursezarazhangrui/codebase-to-course5.7k—~4.4kAutomated safety check: PassNone
AI Engineering Phase Quizrohitg00/ai-engineering-from-scratch66k—~2.1kAutomated safety check: PassMIT
Scholar EvaluationK-Dense-AI/claude-scientific-writer2.4k2 repos~2.9kAutomated safety check: NotesMIT

Similar skills

  • DeepTutor CLI

    HKUDS/DeepTutor

    Teaches the agent to set up and run DeepTutor from the command line: chat and capabilities, knowledge bases, partners, memory, sessions, notebooks and the server or Web app.

    41k GitHub stars~2.8k tokensUpdated yesterday
    EducationAuto-check passed
  • AI Engineering Placement Quiz

    rohitg00/ai-engineering-from-scratch

    Runs a 10-question quiz across five areas to place a learner in the AI Engineering from Scratch curriculum, so they skip what they already know.

    66k GitHub stars~2k tokensUpdated today
    EducationAuto-check passed
  • Codebase to Course

    zarazhangrui/codebase-to-course

    Turns a codebase into an interactive single-page HTML course for non-technical learners, with scroll modules, animated diagrams, quizzes and plain-English code translations.

    5.7k GitHub stars~4.4k tokensUpdated 6 mo ago
    EducationAuto-check passed
  • AI Engineering Phase Quiz

    rohitg00/ai-engineering-from-scratch

    Quizzes you on a completed phase of the AI Engineering from Scratch course, taking a phase number or name and mapping it to that phase's directory.

    66k GitHub stars~2.1k tokensUpdated today
    EducationAuto-check passed
  • Scholar Evaluation

    K-Dense-AI/claude-scientific-writer

    Provide qualitative-first, evidence-traceable developmental review of scholarly works and audit low-stakes research-assessment rubrics with optional local quality controls.

    2.4k GitHub starsUsed in 2 repos~2.9k tokens
    EducationAuto-check: notes
  • Evaluation

    guanyang/open-agent-hub

    This skill should be used when building agent evaluation systems: deterministic checks, regression suites, multi-dimensional rubrics, quality gates, production monitoring, baseline comparison, and…

    977 GitHub starsUsed in 2 repos~4.2k tokens
    EducationAuto-check passed

More from davekilleen/Dex

All 61 skills in this repo
  • Dspy Ruby

    davekilleen/Dex

    This skill should be used when working with DSPy.rb, a Ruby framework for building type-safe, composable LLM applications.

    493 GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Diff Adopt Profile

    davekilleen/Dex

    Adopt a full published Heydex profile by handle ('set me up like @davekilleen').

    493 GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Diff Generate

    davekilleen/Dex

    Package one workflow — how you use Dex for a specific job — into a shareable DexDiff methodology doc.

    493 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Creating Agent Skills

    davekilleen/Dex

    Expert guidance for creating, writing, and refining Claude Code Skills.

    493 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed
  • Feedback

    davekilleen/Dex

    Report a Dex bug to the Dex team with zero homework — Dex investigates locally, builds a privacy-safe report, shows it to you (or auto-sends if you've chosen that), and tracks the ticket until it's…

    493 GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Dhh Rails Style

    davekilleen/Dex

    This skill should be used when writing Ruby and Rails code in DHH's distinctive 37signals style.

    493 GitHub starsUsed in 1 repo~1.7k tokens
    Auto-check passed

Categories

Questions about Skill Score

What does Skill Score do?

Grade a Dex skill against the shape-aware quality rubric and report a ship/revise/no verdict with the exact fixes. Skill Score is an agent skill from davekilleen/Dex. Grade a Dex skill against the shape-aware quality rubric and report a ship/revise/no verdict with the exact fixes.

When should I use Skill Score?

Skill Score fits situations like: you finish writing; editing a skill; create-skill hands off a new package; before shipping a first-party skill.

How do I install Skill Score in Claude Code?

Run `npx skills add davekilleen/Dex --skill skill-score -a claude-code`. Or copy the skill folder (.agents/skills/skill-score in davekilleen/Dex) into .claude/skills/skill-score in your project. Claude Code loads it when a task matches its description.

How do I install Skill Score in Codex?

Run `npx skills add davekilleen/Dex --skill skill-score -a codex`. Or copy the skill folder (.agents/skills/skill-score in davekilleen/Dex) into .agents/skills/skill-score in your project. Codex loads it when a task matches its description.

Can I use Skill Score in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davekilleen/Dex --skill skill-score -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-score, .gemini/skills/skill-score, .github/skills/skill-score and .opencode/skills/skill-score in your project.

What does Skill Score need to run?

Going by SKILL.md and its folder, Skill Score needs Python for the scripts in its folder and the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Skill Score access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Score safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skill Score use?

Skill Score is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Score use?

About 2.6k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Score?

Skills that share tags, products or a category with Skill Score: DeepTutor CLI (HKUDS/DeepTutor, 41k stars), AI Engineering Placement Quiz (rohitg00/ai-engineering-from-scratch, 66k stars), Codebase to Course (zarazhangrui/codebase-to-course, 5.7k stars) and AI Engineering Phase Quiz (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Score?

davekilleen (a GitHub user) maintains it in davekilleen/Dex, which has 493 GitHub stars. The repository holds 61 skills in this directory. The repository was last updated on October 8, 2026.

Source: davekilleen/Dex on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.