Agent skill

Skill Compliance Checker

by affaan-m in affaan-m/ECC

Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

MITAuto-check passedAgent Workflows

Install Skill Compliance Checker

skills CLI
$ npx skills add affaan-m/ECC --skill skill-comply -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install affaan-m/ECC skill-comply --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/affaan-m/ECC.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-comply .claude/skills/skill-comply && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-comply
GitHub stars
277k
Used in
1 other repo
Token cost
~623 tokens
SKILL.md length
219 words
Files
22 (incl. scripts)
Skills in repo
683
Repo updated
First seen
Licence
MIT

At a glance

Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

  • Works in 6 steps: Auto-generating expected behavioral… → Auto-generating scenarios with… → Running claude -p and capturing tool… → …
  • Checking that a new rule is really followed by agents
  • SKILL.md covers Supported Targets, When to Activate, Usage and Key Concept: Prompt Independence, plus 1 more section
  • Runs Python scripts from its folder; calls uv and claude

What it does

Rather than assuming instructions are followed, it measures it. From any markdown file it generates a spec of expected behavior, creates scenarios whose prompts become less supportive (supportive, neutral, competing), runs claude -p while capturing tool calls through stream-json, classifies the calls against spec steps with an LLM, and checks temporal ordering deterministically.

The output is a self-contained report with the expected sequence, the scenario prompts, compliance scores per scenario and tool call timelines labeled by the classifier. Targets can be skills, rules such as testing.md or security.md, or agent definitions, though internal workflow checks for agents are not yet supported. Reports may add hook promotion suggestions for low-compliance steps. A set of Python scripts, prompt files and trace fixtures ship with it, and it runs with uv.

When your agent uses it

  • Checking that a new rule is really followed by agents
  • Testing whether a skill still triggers under unsupportive prompts
  • Comparing compliance across prompt strictness levels
  • Running periodic quality checks on an agent setup

Example prompts

  • “Run skill-comply on my testing.md rule and show the report.”
  • “Is this TDD skill actually being followed? Measure it.”
  • “Generate scenarios for my security rules file and score compliance.”

Requirements

  • uv with Python
  • The claude CLI, which the runner calls with -p

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Auto-generating expected behavioral sequences (specs) from any .md file
  2. Auto-generating scenarios with decreasing prompt strictness (supportive → neutral → competing)
  3. Running claude -p and capturing tool call traces via stream-json
  4. Classifying tool calls against spec steps using LLM (not regex)
  5. Checking temporal ordering deterministically
  6. Generating self-contained reports with spec, prompts, and timelines

What it can do on your machine

Read from SKILL.md and the folder at commit 2d515e4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 10 files in scripts/ (Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • claude

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Compliance Checker loads about 623 tokens when it runs. Until then it costs about 96 tokens; SKILL.md has 219 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~96
When it runs · the whole SKILL.md, loaded when a task matches
~623

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from affaan-m/ECC at commit 2d515e4, republished under its MIT licence (© affaan-m). 219 words, ~623 tokens.

Download SKILL.mdSave it as .claude/skills/skill-comply/SKILL.md (or your agent's skills folder). This skill also uses 21 other files; get the full folder from GitHub.
name
skill-comply
description
Visualize whether skills, rules, and agent definitions are actually followed — auto-generates scenarios at 3 prompt strictness levels, runs agents, classifies behavioral sequences, and reports compliance rates with full tool call timelines. Use when checking whether agents actually follow the skills, rules, and definitions they were given, rather than assuming they do.
metadata.origin
ECC
tools
Read, Bash

skill-comply: Automated Compliance Measurement

Measures whether coding agents actually follow skills, rules, or agent definitions by:

  1. Auto-generating expected behavioral sequences (specs) from any .md file
  2. Auto-generating scenarios with decreasing prompt strictness (supportive → neutral → competing)
  3. Running claude -p and capturing tool call traces via stream-json
  4. Classifying tool calls against spec steps using LLM (not regex)
  5. Checking temporal ordering deterministically
  6. Generating self-contained reports with spec, prompts, and timelines

Supported Targets

  • Skills (skills/*/SKILL.md): Workflow skills like search-first, TDD guides
  • Rules (rules/common/*.md): Mandatory rules like testing.md, security.md, git-workflow.md
  • Agent definitions (agents/*.md): Whether an agent gets invoked when expected (internal workflow verification not yet supported)

When to Activate

  • User runs /skill-comply <path>
  • User asks "is this rule actually being followed?"
  • After adding new rules/skills, to verify agent compliance
  • Periodically as part of quality maintenance

Usage

bash
# Full run
uv run python -m scripts.run ~/.claude/rules/common/testing.md

# Dry run (no cost, spec + scenarios only)
uv run python -m scripts.run --dry-run ~/.claude/skills/search-first/SKILL.md

# Custom models
uv run python -m scripts.run --gen-model haiku --model sonnet <path>

Key Concept: Prompt Independence

Measures whether a skill/rule is followed even when the prompt doesn't explicitly support it.

Report Contents

Reports are self-contained and include:

  1. Expected behavioral sequence (auto-generated spec)
  2. Scenario prompts (what was asked at each strictness level)
  3. Compliance scores per scenario
  4. Tool call timelines with LLM classification labels
Advanced (optional)

For users familiar with hooks, reports also include hook promotion recommendations for steps with low compliance. This is informational — the main value is the compliance visibility itself.

© affaan-m, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 21 other files (scripts) in skills/skill-comply of affaan-m/ECC.

  • SKILL.md
  • fixtures/compliant_trace.jsonl
  • fixtures/noncompliant_trace.jsonl
  • fixtures/tdd_spec.yaml
  • prompts/classifier.md
  • prompts/scenario_generator.md
  • prompts/spec_generator.md
  • pyproject.toml
  • scripts/__init__.py
  • scripts/classifier.py
  • scripts/grader.py
  • scripts/parser.py
  • scripts/report.py
  • scripts/run.py
  • scripts/runner.py
  • scripts/scenario_generator.py
  • scripts/spec_generator.py
  • scripts/utils.py
  • … and 4 more

Open the folder on GitHubat commit 2d515e4

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in affaan-m/ECC, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Skill Compliance Checker next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Compliance Checker compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Compliance Checker this skillaffaan-m/ECC277k1 repos~623Automated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Open-Science Skill Creatoraipoch/open-science5.5k—~1.7kAutomated safety check: PassApache-2.0
Skill Quality ReviewerGalaxy-Dawn/claude-scholar5.7k1 repos~3kAutomated safety check: PassMIT
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0
OpenCode Skill Creatorantongulin/opencode-skill-creator171—~8.1kAutomated safety check: PassApache-2.0

Similar skills

  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Open-Science Skill Creator

    aipoch/open-science

    Creates, revises, evaluates and publishes skills in the Open-Science app through its native host.skills composer, with optional test prompts and benchmarks.

    5.5k GitHub stars~1.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Skill Quality Reviewer

    Galaxy-Dawn/claude-scholar

    Scores a skill across description, content organization, writing style and structure, then produces letter grades and a prioritized improvement plan.

    5.7k GitHub starsUsed in 1 repo~3k tokens
    Agent WorkflowsAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated 4 days ago
    Agent WorkflowsAuto-check passed
  • OpenCode Skill Creator

    antongulin/opencode-skill-creator

    Walks you through drafting, testing, evaluating and tuning a skill for OpenCode, from an intake interview to description optimization.

    171 GitHub stars~8.1k tokensUpdated 9 days ago
    Agent WorkflowsAuto-check passed
  • Collects evidence from a failed or confusing skill-driven task and triggers skvm jit-optimize to propose concrete fixes to that skill's files.

    561 GitHub stars~2.9k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check: warnings

More from affaan-m/ECC

All 682 skills in this repo
  • Skill Stocktake

    affaan-m/ECC

    Audits your installed Claude skills and commands for quality, with a quick mode for recently changed skills and a full mode that evaluates all of them through subagents.

    277k GitHub starsUsed in 5 repos~3.1k tokens
    Auto-check passed
  • Ingests, indexes, searches, edits and monitors video, audio and live streams through the VideoDB Python SDK, returning stream links, clips and timestamps.

    277k GitHub starsUsed in 3 repos~3.5k tokens
    Auto-check: notes
  • Docs Governance

    affaan-m/ECC

    Route broad documentation-governance requests to existing ECC skills and run an opt-in, read-only audit of mapped documentation roles, links, ADR indexes, and evidence references.

    277k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Rules Distillation

    affaan-m/ECC

    Scans installed skills for principles that recur across them and proposes rule-file changes: append, revise, add a section, create a file or leave as covered.

    277k GitHub starsUsed in 2 repos~2.3k tokens
    Auto-check passed
  • Builds DRAFT counterparty agreements from one markdown template and a small JSON spec per party, with clauses picked by the party's role.

    277k GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Set an ECC-specific frontend design direction for production UI work.

    277k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed

Questions about Skill Compliance Checker

What does Skill Compliance Checker do?

Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces. Rather than assuming instructions are followed, it measures it. From any markdown file it generates a spec of expected behavior, creates scenarios whose prompts become less supportive (supportive, neutral, competing), runs claude -p while capturing tool calls through stream-json, classifies the calls against spec steps with an LLM, and checks temporal ordering deterministically.

When should I use Skill Compliance Checker?

Skill Compliance Checker fits situations like: checking that a new rule is really followed by agents; testing whether a skill still triggers under unsupportive prompts; comparing compliance across prompt strictness levels; running periodic quality checks on an agent setup.

How do I install Skill Compliance Checker in Claude Code?

Run `npx skills add affaan-m/ECC --skill skill-comply -a claude-code`. Or copy the skill folder (skills/skill-comply in affaan-m/ECC) into .claude/skills/skill-comply in your project. Claude Code loads it when a task matches its description.

How do I install Skill Compliance Checker in Codex?

Run `npx skills add affaan-m/ECC --skill skill-comply -a codex`. Or copy the skill folder (skills/skill-comply in affaan-m/ECC) into .agents/skills/skill-comply in your project. Codex loads it when a task matches its description.

Can I use Skill Compliance Checker in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add affaan-m/ECC --skill skill-comply -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-comply, .gemini/skills/skill-comply, .github/skills/skill-comply and .opencode/skills/skill-comply in your project.

What does Skill Compliance Checker need to run?

Going by SKILL.md and its folder, Skill Compliance Checker needs Python for the scripts in its folder and the command-line tools its instructions call (uv and claude). Our summary lists: uv with Python; The claude CLI, which the runner calls with -p.

Does Skill Compliance Checker access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Skill Compliance Checker safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skill Compliance Checker use?

Skill Compliance Checker is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Compliance Checker use?

About 623 tokens (SKILL.md is roughly 2.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Compliance Checker?

Skills that share tags, products or a category with Skill Compliance Checker: Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars), Open-Science Skill Creator (aipoch/open-science, 5.5k stars), Skill Quality Reviewer (Galaxy-Dawn/claude-scholar, 5.7k stars) and Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Compliance Checker?

affaan-m (a GitHub user) maintains it in affaan-m/ECC, which has 276,673 GitHub stars. The repository holds 683 skills in this directory. The repository was last updated on October 11, 2026.

Source: affaan-m/ECC on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.