Agent skill

Spec Conformance Check

by mrmps in mrmps/classifier-dev

Check an artifact against a written spec one requirement at a time.

MITAuto-check passedDevelopment

Install Spec Conformance Check

skills CLI
$ npx skills add mrmps/classifier-dev --skill spec-conformance-check -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mrmps/classifier-dev spec-conformance-check --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mrmps/classifier-dev.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/spec-conformance-check .claude/skills/spec-conformance-check && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
spec-conformance-check
GitHub stars
424
Token cost
~1.5k tokens
SKILL.md length
522 words
Files
1
Skills in repo
21
Repo updated
First seen
Licence
MIT

At a glance

Check an artifact against a written spec one requirement at a time.

  • Works in 4 steps: Split the spec into requirements → One input per requirement → A real run → …
  • Tasks that involve Pull requests
  • SKILL.md covers When not to use it, 1. Split the spec into…, 2. One input per requirement and 3. A real run, plus 3 more sections
  • Calls python3; reaches classifier.dev

What it does

Spec Conformance Check is an agent skill from mrmps/classifier-dev. Check an artifact against a written spec one requirement at a time. Splits the spec into testable requirements, scores each met, partly met, not met or not applicable with a calibrated confidence, blocks on a confident not-met and sends the unsure rows to a person. Works on a PR description, an agent's answer, a generated test or an RFC. Use on "does this meet the spec", "did it do everything I asked", or a definition-of-done gate in CI.

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Pull requests. The repository describes itself as: Zero-shot text classification over plain HTTP — no API key, no account. One Cloudflare Worker, a CLI, and an MCP server. https://classifier.dev. The licence is MIT.

When your agent uses it

  • Tasks that involve Pull requests

Example prompts

  • “s answer, a generated test or an RFC. Use on”
  • “did it do everything I asked”
  • “/spec-conformance-check”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Split the spec into requirements
  2. One input per requirement
  3. A real run
  4. The gate

What it can do on your machine

Read from SKILL.md and the folder at commit 629df75. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • classifier.dev

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Spec Conformance Check loads about 1.5k tokens when it runs. Until then it costs about 116 tokens; SKILL.md has 522 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~116
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mrmps/classifier-dev at commit 629df75, republished under its MIT licence (© mrmps). 522 words, ~1,498 tokens.

Download SKILL.mdSave it as .claude/skills/spec-conformance-check/SKILL.md (or your agent's skills folder).
name
spec-conformance-check
description
Check an artifact against a written spec one requirement at a time. Splits the spec into testable requirements, scores each met, partly met, not met or not applicable with a calibrated confidence, blocks on a confident not-met and sends the unsure rows to a person. Works on a PR description, an agent's answer, a generated test or an RFC. Use on "does this meet the spec", "did it do everything I asked", or a definition-of-done gate in CI.
license
MIT

Check an artifact against a spec, one line at a time

Asking a model "does this meet the spec?" gets a paragraph that is generous, unfalsifiable and different next time. Asking once per requirement gets a table you can block a merge on, in one call: each input is one requirement plus the artifact.

When not to use it

  • When the spec is executable: a test suite beats this, so run the tests.
  • For behaviour the artifact does not describe. It reads the text in front of it, so a PR description that lies passes; point it at the diff too.
  • For requirements nobody wrote down. Write them down first.

1. Split the spec into requirements

One testable claim a line, in the spec's own words:

The endpoint responds at GET /v1/health.
The response body is JSON and includes the running version string.
The endpoint checks the database before reporting healthy.
An unhealthy dependency makes the endpoint answer 503.
Responses are not cached by intermediaries.
The endpoint is rate limited per IP.
The mobile app shows a maintenance banner when health fails.
  • Split on "and". "Returns JSON and is rate limited" scores as one blurred answer; two lines say which half is missing.
  • Keep the modal verb. "should" and "must" split a note from a block.
  • Leave in requirements this artifact cannot satisfy. Line 7 is about the mobile app, and not applicable is a real answer.

2. One input per requirement

The artifact goes into every input, after the requirement. The repetition is the point: each answer judges one thing. conform.py:

python
import json, sys, urllib.request

spec = [l.strip() for l in open("spec.txt") if l.strip()]
artifact = open("artifact.txt").read()
body = json.dumps({
    "labels": ["met", "partly met", "not met", "not applicable"],
    "instructions": "Each input is one requirement followed by the artifact under review. "
        "Decide whether the artifact, as written, meets that one requirement. Judge only "
        "what the artifact states; silence is not met. Use 'not applicable' when the "
        "requirement is outside what this artifact covers.",
    "inputs": [f"REQUIREMENT: {r}\n\nARTIFACT:\n{artifact}" for r in spec],
}).encode()
req = urllib.request.Request("https://classifier.dev/v1/classify", data=body,
    headers={"content-type": "application/json", "user-agent": "conformance/1.0"})

blocks = 0
for r, q in zip(json.load(urllib.request.urlopen(req))["results"], spec):
    c, label = r["confidence"], r["label"]
    block = label == "not met" and c >= 0.9
    gate = "BLOCK" if block else "ask" if c < 0.5 else "review" if c < 0.9 else "ok"
    blocks += block
    print(f"{label:<15}{c:.2f}  {gate:<7}{q}")
sys.exit(1 if blocks else 0)

Set the user-agent: Python's urllib default is blocked at the edge and returns 403 before the call is classified.

3. A real run

artifact.txt, a PR description:

PR: add GET /v1/health
Returns 200 with {"ok": true, "version": "1.4.0"} and runs SELECT 1 against the
database before answering. Response is JSON with a cache-control: no-store
header. Three unit tests added. I did not get to the per-IP rate limit; the
endpoint is currently unlimited.

python3 conform.py:

met            1.00  ok     The endpoint responds at GET /v1/health.
met            1.00  ok     The response body is JSON and includes the running version string.
met            1.00  ok     The endpoint checks the database before reporting healthy.
not met        0.68  review An unhealthy dependency makes the endpoint answer 503.
met            0.99  ok     Responses are not cached by intermediaries.
not met        0.99  BLOCK  The endpoint is rate limited per IP.
not applicable 0.93  ok     The mobile app shows a maintenance banner when health fails.

Exit status 1, so it gates the merge. Requirement 4 is the row that matters: the PR never mentions 503, so a not met in the review band means the model cannot tell whether that is missing from the description or from the code — a question for the author, not a block. It came back 0.68, 0.77 and 0.79 over three runs: read the band, never the second decimal.

Show full SKILL.md (218 more words)Show less

4. The gate

  • not met at 0.9 and above — block. One row is enough; print it, not a summary.
  • 0.5 to 0.9, any label — a person reads it. Omissions and vague requirements both land here.
  • Below 0.5 — the requirement is the problem. A line the model cannot place is compound or vague; rewrite it before arguing about the artifact.
  • partly met at any confidence — a person reads it. It means the artifact started on the requirement and stopped short.

"tier": "smart" re-asks answers under 0.7 of a reasoning model at seconds per row. Use it when several rows land there.

Pitfalls

  • Silence reads as met unless you say otherwise. "Judge only what the artifact states; silence is not met" in instructions stops a thin PR description passing a long spec.
  • A long artifact drowns a short requirement. Past a few thousand characters, pass only the relevant section.
  • Confidence is about the label, not the requirement's weight. A sure not applicable on a load-bearing line still needs a look.

What done looks like

Every requirement has a label, a confidence and a gate, in spec order. The command exits non-zero when a not met clears 0.9, the review band is a short list of rows with names against them, and the spec file goes into the next run unchanged.

© mrmps, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/spec-conformance-check of mrmps/classifier-dev.

Open the folder on GitHubat commit 629df75

Compare with similar skills

Spec Conformance Check next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Spec Conformance Check compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Spec Conformance Check this skillmrmps/classifier-dev424—~1.5kAutomated safety check: PassMIT
Finishing a Development Branchobra/superpowers296k5 repos~1.9kAutomated safety check: PassMIT
PR Babysitteropeninterpreter/openinterpreter69k3 repos~4.2kAutomated safety check: PassApache-2.0
Check PRonyx-dot-app/onyx32k2 repos~2.3kAutomated safety check: PassMIT
Understand Diff AnalysisEgonex-AI/Understand-Anything86k1 repos~1.4kAutomated safety check: PassMIT
PR Design DocOpenHands/OpenHands90k—~2.4kAutomated safety check: PassMIT

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    296k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • PR Babysitter

    openinterpreter/openinterpreter

    Watches an open GitHub pull request until it merges, handling review comments, diagnosing CI failures and retrying flaky checks along the way.

    69k GitHub starsUsed in 3 repos~4.2k tokens
    DevelopmentAuto-check passed
  • Check PR

    onyx-dot-app/onyx

    Checks a GitHub, GitLab, or Perforce (p4) pull request (or merge request, or shelved changelist) for unresolved review comments, failing status checks, and incomplete PR descriptions.

    32k GitHub starsUsed in 2 repos~2.3k tokens
    DevelopmentAuto-check passed
  • Understand Diff Analysis

    Egonex-AI/Understand-Anything

    Reads your git changes or a pull request against a prebuilt knowledge graph of the project to explain what changed, which components are affected and what is risky.

    86k GitHub starsUsed in 1 repo~1.4k tokens
    DevelopmentAuto-check passed
  • PR Design Doc

    OpenHands/OpenHands

    For a non-trivial pull request, write a self-contained HTML design doc under the temporary .pr/ directory and link a visibility-appropriate preview in the PR description, so maintainers grasp the…

    90k GitHub stars~2.4k tokensUpdated today
    DevelopmentAuto-check passed
  • WooCommerce Code Review

    woocommerce/woocommerce

    Reviews WooCommerce code changes against the project's standards, flagging backend PHP architecture, naming, documentation, data integrity and testing violations.

    11k GitHub starsUsed in 3 repos~1.1k tokens
    DevelopmentAuto-check passed

More from mrmps/classifier-dev

All 21 skills in this repo
  • Bulk Classify

    mrmps/classifier-dev

    Sort many texts into your own categories without reading them, using a keyless HTTP API that returns a calibrated confidence per answer.

    424 GitHub stars~3.1k tokensUpdated yesterday
    Auto-check passed
  • Computer Use Action Picker

    mrmps/classifier-dev

    Pick a browser or desktop agent's next action by choosing among the actions actually on screen instead of inventing one.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Content Moderation Gate

    mrmps/classifier-dev

    Check user-generated text against a written policy before it is published.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Label each context chunk keep, drop or replace-with-a-pointer and pass the survivors through byte for byte instead of summarising, with key-shaped chunks decided locally and never sent, and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Document Intake Routing

    mrmps/classifier-dev

    Label each page of an intake packet with a document type and a page role before extraction runs, so only confident pages reach an extractor and the rest reach a person.

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed
  • Headline Filter Map Reduce

    mrmps/classifier-dev

    Filter hundreds or thousands of headlines, search results or feed items against a written brief before opening any of them, using a two-stage cascade that spends a fast model on everything and a…

    424 GitHub stars~1.5k tokensUpdated yesterday
    Auto-check passed

Categories

Questions about Spec Conformance Check

What does Spec Conformance Check do?

Check an artifact against a written spec one requirement at a time. Spec Conformance Check is an agent skill from mrmps/classifier-dev. Check an artifact against a written spec one requirement at a time.

When should I use Spec Conformance Check?

Spec Conformance Check fits situations like: tasks that involve Pull requests.

How do I install Spec Conformance Check in Claude Code?

Run `npx skills add mrmps/classifier-dev --skill spec-conformance-check -a claude-code`. Or copy the skill folder (skills/spec-conformance-check in mrmps/classifier-dev) into .claude/skills/spec-conformance-check in your project. Claude Code loads it when a task matches its description.

How do I install Spec Conformance Check in Codex?

Run `npx skills add mrmps/classifier-dev --skill spec-conformance-check -a codex`. Or copy the skill folder (skills/spec-conformance-check in mrmps/classifier-dev) into .agents/skills/spec-conformance-check in your project. Codex loads it when a task matches its description.

Can I use Spec Conformance Check in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mrmps/classifier-dev --skill spec-conformance-check -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/spec-conformance-check, .gemini/skills/spec-conformance-check, .github/skills/spec-conformance-check and .opencode/skills/spec-conformance-check in your project.

What does Spec Conformance Check need to run?

Going by SKILL.md and its folder, Spec Conformance Check needs the command-line tools its instructions call (python3). Our summary lists: Python 3.

Does Spec Conformance Check access the network?

SKILL.md names 1 domain. In commands or code: classifier.dev; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Spec Conformance Check safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Spec Conformance Check use?

Spec Conformance Check is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Spec Conformance Check use?

About 1.5k tokens (SKILL.md is roughly 6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Spec Conformance Check?

Skills that share tags, products or a category with Spec Conformance Check: Finishing a Development Branch (obra/superpowers, 296k stars), PR Babysitter (openinterpreter/openinterpreter, 69k stars), Check PR (onyx-dot-app/onyx, 32k stars) and Understand Diff Analysis (Egonex-AI/Understand-Anything, 86k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Spec Conformance Check?

mrmps (a GitHub user) maintains it in mrmps/classifier-dev, which has 424 GitHub stars. The repository holds 21 skills in this directory. The repository was last updated on October 7, 2026.

Source: mrmps/classifier-dev on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.