Validate a SKILL.md file against the four-tier validation system: Tier 0 (locate), Tier 1 (standard or marketplace grading per the IS 100-point rubric), Tier 2 (static production gate —…

MITAuto-check passedEducation

Install Validate Skillmd

skills CLI
$ npx skills add jeremylongshore/tons-of-skills-marketplace --skill validate-skillmd -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jeremylongshore/tons-of-skills-marketplace validate-skillmd --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jeremylongshore/tons-of-skills-marketplace.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.curated/validate-skillmd .claude/skills/validate-skillmd && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
validate-skillmd
GitHub stars
2.8k
Token cost
~6.2k tokens
SKILL.md length
2,167 words
Files
3 (incl. references)
Skills in repo
3,342
Repo updated
First seen
Licence
MIT

At a glance

Validate a SKILL.md file against the four-tier validation system: Tier 0 (locate), Tier 1 (standard or marketplace grading per the IS 100-point rubric), Tier 2 (static production gate —…

  • Works in 8 steps: Locate the Skill → Run Validator → 5: Tier 2 — Static Production Gate → …
  • Creating a new skill
  • SKILL.md covers Overview, Prerequisites, Schema Reminder — Frontmatter… and Instructions, plus 4 more sections
  • Calls python3 and pnpm; needs GROQ_API_KEY and API_KEY

What it does

Validate Skillmd is an agent skill from jeremylongshore/tons-of-skills-marketplace. Validate a SKILL.md file against the four-tier validation system: Tier 0 (locate), Tier 1 (standard or marketplace grading per the IS 100-point rubric), Tier 2 (static production gate — allowed-tools accuracy, auth protocol, dead code, tool safety, orchestration bounds), and Tier 3 (JRig 7-layer behavioral eval, opt-in via --thorough). Use when creating a new skill, checking skill quality, preparing for marketplace submission, running deep quality analysis, or gating a skill for production. Trigger with "validate…

Its SKILL.md is about 6.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 3 other files, including reference files (for example `eval-spec.yaml` and `references/schema-reminder-frontmatter.md`). Compatibility notes: Designed for Claude Code; requires Python 3, optionally JRig CLI for Tier 3

It sits in Education, covering Quizzes and assessments and Skill authoring. The repository describes itself as: Model-agnostic agent-skills platform with a harness-free canonical layer, verified adapters, and the ccpi package manager. Explore at tonsofskills.com. The licence is MIT.

When your agent uses it

  • Creating a new skill
  • Checking skill quality
  • Preparing for marketplace submission
  • Running deep quality analysis

Example prompts

  • “validate this skill”
  • “grade my skill”
  • “deep eval”
  • “/validate-skillmd”

Requirements

  • Python 3
  • A credential in GROQ_API_KEY
  • A credential in API_KEY
  • Compatibility (from SKILL.md): Designed for Claude Code; requires Python 3, optionally JRig CLI for Tier 3
  • Pre-approved tools (allowed-tools): Read, Edit, Write, Bash(python3:*), Bash(j-rig:*), Bash(node:*), Bash(scripts/run-jrig-eval.sh:*), Glob, Grep, AskUserQuestion

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Locate the Skill
  2. Run Validator
  3. 5: Tier 2 — Static Production Gate
  4. 7: Tier 3 — JRig 7-Layer Behavioral Eval (opt-in via --thorough)
  5. Present Unified Report (all tiers)
  6. Prioritized Fix Recommendations
  7. Review Structural Advisors
  8. Auto-Fix (if requested or grade below B)

What it can do on your machine

Read from SKILL.md and the folder at commit cfae287. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Edit
    • Write
    • Bash(python3:*)
    • Bash(j-rig:*)
    • Bash(node:*)
    • Bash(scripts/run-jrig-eval.sh:*)
    • Glob
    • Grep
    • AskUserQuestion

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • pnpm

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • platform.claude.com
    • code.claude.com
    • agentskills.io

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • GROQ_API_KEY
    • API_KEY

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

  • Compatibility

    Designed for Claude Code; requires Python 3, optionally JRig CLI for Tier 3

    From compatibility in the SKILL.md frontmatter.

Context cost

Validate Skillmd loads about 6.2k tokens when it runs, and up to ~6.8k if it reads all its reference files. Until then it costs about 160 tokens; SKILL.md has 2,167 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~160
When it runs · the whole SKILL.md, loaded when a task matches
~6.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~6.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jeremylongshore/tons-of-skills-marketplace at commit cfae287, republished under its MIT licence (© jeremylongshore). 2,167 words, ~6,226 tokens.

Download SKILL.mdSave it as .claude/skills/validate-skillmd/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
validate-skillmd
description
Validate a SKILL.md file against the four-tier validation system: Tier 0 (locate), Tier 1 (standard or marketplace grading per the IS 100-point rubric), Tier 2 (static production gate — allowed-tools accuracy, auth protocol, dead code, tool safety, orchestration bounds), and Tier 3 (JRig 7-layer behavioral eval, opt-in via --thorough). Use when creating a new skill, checking skill quality, preparing for marketplace submission, running deep quality analysis, or gating a skill for production. Trigger with "validate this skill", "grade my skill", "deep eval", "check SKILL.md", "validate thorough", "/validate-skillmd".
allowed-tools
Read, Edit, Write, Bash(python3:*), Bash(j-rig:*), Bash(node:*), Bash(scripts/run-jrig-eval.sh:*), Glob, Grep, AskUserQuestion
compatibility
Designed for Claude Code; requires Python 3, optionally JRig CLI for Tier 3
version
7.2.0
author
Jeremy Longshore <jeremy@intentsolutions.io>
license
MIT
tags
validation, skill-quality, marketplace, grading, deep-eval, jrig, behavioral-eval
user-invocable
true
argument-hint
[path-to-SKILL.md] [--marketplace|--deep|--thorough]

Validate SKILL.md

Grade any SKILL.md file against the Intent Solutions rubric (validator v7.0 / schema 3.3.1). Four-tier validation: Tier 0 (locate), Tier 1 (standard or marketplace grading), Tier 2 (static production gate), Tier 3 (JRig behavioral eval — opt-in).

Source of truth: /skill-creator validation workflow + claude-code-plugins-plus-skills/000-docs/SCHEMA_CHANGELOG.md + j-rig-skill-binary-eval/ (Tier 3).

Overview

Schema 3.3.1 enforces the 8-field IS enterprise required-field set at marketplace tier (name, description, allowed-tools, version, author, license, compatibility, tags). Anthropic's spec floor (name + description only) sits underneath; the IS rubric sits on top. Modes:

  • Standard (default): Mirrors platform.claude.com/docs/en/agents-and-tools/agent-skills/overview exactly. Required: name, description. Everything else is silent unless invalid type/value. Fast (~10 sec).
  • Marketplace (--marketplace): 8-field enterprise required set + 100-point IS rubric. Missing required fields = ERROR, not warning. The --enterprise flag is a deprecated alias. Fast (~10 sec).
  • Deep (--deep): Intent Solutions Deep Evaluation Engine — 10 weighted dimensions, trust badges, Elo competitive ranking, optional LLM quality assessment via Groq. Fast (~30 sec).
  • Thorough (--thorough): Adds Tier 3 JRig behavioral eval on top of Tier 1+2. Runs 7-layer eval across Haiku/Sonnet/Opus. Slow (~10–30 min) and costs ~$2–5 per skill in API spend — opt-in only. Right for production-gating, not iterative authoring.

Performance + cost note: Tiers 0–2 run in seconds and are free. Tier 3 (JRig) is opt-in because behavioral eval across the model matrix is genuinely expensive. Default invocations stay fast; --thorough is for the moment a skill is being promoted to production or marketplace-verified.

Prerequisites

  • Python 3 with pyyaml installed
  • Validator script: claude-code-plugins-plus-skills/scripts/validate-skills-schema.py (v7.0+)
  • For --thorough (Tier 3): JRig CLI on PATH (jrig --version returns ≥ v0.14.0). Install: cd ~/000-projects/j-rig-binary-eval && pnpm install && pnpm build && pnpm link --global. Tier 3 is opt-in; the rest of the skill works without JRig installed.
Kernel schema first (canonical machine spec)

Validate frontmatter structure against references/kernel-schemas/v1/skill-frontmatter.schema.json FIRST — the kernel-pinned canonical machine spec (@intentsolutions/core@0.5.0, vendored; the STRICT v2 sibling under kernel-schemas/v2/ is not yet promoted to canonical). The prose spec references are supporting documentation only; on disagreement the kernel schema wins. $refs are absolute $id URIs — register every file under kernel-schemas/ to resolve them. Provenance + refresh: references/kernel-schemas/PROVENANCE.md.

Schema Reminder — Frontmatter Fields

Full field table: schema-reminder-frontmatter.md.

Instructions

Step 1: Locate the Skill

If path provided via $ARGUMENTS, use it directly. Otherwise:

  • Check current directory for SKILL.md
  • Ask user with AskUserQuestion

Common locations:

  • ~/.claude/skills/{name}/SKILL.md (global)
  • .claude/skills/{name}/SKILL.md (project)
Step 2: Run Validator
bash
# Standard tier (default — Anthropic spec exactly)
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py SKILL.md

# Marketplace tier (full 100-point rubric, polish recommendations as warnings)
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --marketplace SKILL.md

# Deep evaluation (10 dimensions, badges, Elo ranking)
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --deep SKILL.md

# Deep + LLM quality assessment via Groq (requires GROQ_API_KEY)
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --deep --thorough SKILL.md

# Deep eval with JSON/markdown/HTML report output
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --deep --report-format json SKILL.md
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --deep --report-format html SKILL.md

# Marketplace + write to compliance DB
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --marketplace --populate-db claude-code-plugins-plus-skills/freshie/inventory.sqlite SKILL.md

# Show D/F grade skills in full scan
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --marketplace --show-low-grades

# Minimum grade gate (exits 1 if any skill below threshold)
python3 claude-code-plugins-plus-skills/scripts/validate-skills-schema.py --marketplace --min-grade B SKILL.md

Default to marketplace tier when the user is preparing for marketplace submission. Use --deep for functional quality assessment beyond structural compliance. The --enterprise flag still works as a deprecated alias for --marketplace.

v6.0 Deep Evaluation Engine: 10 weighted dimensions (triggering accuracy 0.25, orchestration fitness 0.20, output quality 0.15, scope calibration 0.12, progressive disclosure 0.10, token efficiency 0.06, robustness 0.05, structural completeness 0.03, code template quality 0.02, ecosystem coherence 0.02). Anti-pattern detection with 5% penalty each. Elo competitive ranking. Trust badges (Flagship/Established/Emerging/Early). Optional LLM-as-judge via Groq free tier.

Step 2.5: Tier 2 — Static Production Gate

After Tier 1 grading and before any behavioral eval, run five inline static checks. These catch obvious production blockers in seconds without needing JRig. Always run regardless of mode (standard/marketplace/deep/thorough); each is binary pass/fail.

2.5.1 allowed-tools accuracy

Every tool declared in allowed-tools should actually be referenced somewhere in the skill body or its references//scripts/. Conversely, every tool the skill calls should be declared.

bash
# Extract declared tools (handles CSV string, space-separated, and YAML list forms — schema 3.3.1)
declared=$(python3 -c "import yaml,sys; fm=yaml.safe_load(open('SKILL.md').read().split('---')[1]); t=fm.get('allowed-tools',''); print(t if isinstance(t,str) else ' '.join(t))")

# Each declared base tool (Read, Write, Bash, etc.) must appear in the body
for tool in $(echo "$declared" | grep -oE '[A-Z][a-zA-Z]+' | sort -u); do
  if ! grep -q "$tool" "SKILL.md"; then
    echo "FAIL: tool '$tool' declared but not referenced in body"
  fi
done

Fail when: tool declared but never used (over-permissive — attack surface) OR tool used but never declared (will prompt user every invocation, defeating allowed-tools).

2.5.2 Auth protocol documented (when applicable)

If the skill mentions an external API (any URL, curl, fetch, MCP server, OAuth flow, API key reference), an authentication method must be documented in the body or in references/auth.md / references/api-surface.md.

bash
# Heuristic: look for API indicators
if grep -qE "(curl |fetch\(|mcp__|API_KEY|TOKEN|OAuth|Bearer )" "SKILL.md"; then
  if ! grep -qiE "(authentication|auth method|api key|bearer token|oauth flow|credentials)" "SKILL.md"; then
    echo "FAIL: external API referenced but no auth protocol documented"
  fi
fi

Fail when: API surface is referenced but a future engineer reading the skill couldn't tell how authentication happens.

2.5.3 Dead-code / unreachable-branch sanity

Conditional structures that can never fire (e.g., if false, mutually exclusive guards, sections that contradict an earlier hard-fail).

bash
# Conservative checks — flag for human review, don't auto-fail
grep -nE "^(if false|if \[ false \]|elif false)" "SKILL.md" && echo "WARN: literal-false branch found"
grep -cE "^### " "SKILL.md" # If section count grossly exceeds the table-of-contents count → drift

Warn when: a literal-false branch is found OR the body contains sections not present in the table of contents (silent drift).

2.5.4 Tool-safety combo flagging

Dangerous combinations require explicit justification:

ComboWhy dangerous
Bash (unscoped) + WebFetchCan fetch arbitrary content + execute it
Bash (unscoped) + WriteCan write executable scripts to arbitrary locations
Bash(curl:*) + Bash(sh:*)Curl-pipe-shell pattern
Bash(rm:*) not paired with explicit safe-pathsUnbounded delete authority
bash
# Check for unscoped Bash + dangerous companion
if grep -qE "^allowed-tools:.*\bBash\b" "SKILL.md" && \
   ! grep -qE "^allowed-tools:.*Bash\(" "SKILL.md" && \
   grep -qE "^allowed-tools:.*\b(Write|WebFetch)\b" "SKILL.md"; then
  if ! grep -qiE "(safety justification|why unscoped Bash|why Bash + )" "SKILL.md"; then
    echo "FAIL: unscoped Bash + Write/WebFetch without safety justification"
  fi
fi

Fail when: a dangerous combo is declared and the body has no ## Safety Justification section explaining why the wide scope is necessary.

2.5.5 Orchestration bounds

Skills are NOT plugins. A skill should not spawn other skills, delegate to other agents as a primary control flow, or self-coordinate across sessions. That's /skill-creator --forge territory and plugin-level orchestration. Skills do one job.

bash
# Look for orchestration smells in skills
if grep -qE "(spawn another skill|delegate to /|invoke .* skill|orchestrate across|self-coordinate)" "SKILL.md"; then
  echo "FAIL: skill appears to orchestrate other skills/agents — that belongs at the plugin layer"
fi

Fail when: the skill body claims it spawns/orchestrates other skills or agents as the primary control flow. Multi-agent synthesis WITHIN one skill invocation (calling subagents to specialize) is fine and expected; cross-skill orchestration is not.

Tier 2 verdict
  • All 5 checks pass → Tier 2 GREEN; proceed.
  • Any FAIL → Tier 2 RED; the skill is blocked from production promotion until resolved. Tier 3 (JRig) does not run if Tier 2 is RED — fail fast.
  • Only WARN → Tier 2 YELLOW; log warnings to the unified report; Tier 3 still runs.
Step 2.7: Tier 3 — JRig 7-Layer Behavioral Eval (opt-in via --thorough)

Default skipped. Tier 3 runs only when the user passes --thorough. Behavioral eval across the model matrix (Haiku / Sonnet / Opus) takes 10–30 minutes per skill and costs ~$2–$5 in API spend. Right for production-gate moments, not iterative authoring.

Prerequisites
  • JRig CLI on PATH: j-rig --version (note: bin name is j-rig with hyphen, not jrig)
  • Install: cd j-rig-skill-binary-eval/packages/cli && pnpm build && ln -sf $PWD/dist/index.js ~/.local/bin/j-rig
  • For Tier 3B (7-layer eval): API keys for Haiku, Sonnet, Opus configured per JRig docs

If JRig isn't on PATH, the skill emits a placeholder verdict and a one-line install hint, then continues to Step 3 (grade report) without blocking. Tier 3 absence does not mean a skill fails — only Tier 1+2 are mandatory.

Tier 3A: JRig package-integrity check (deterministic, fast, free)

Run JRig's check command on the skill directory (not the SKILL.md path). Returns deterministic pass/warn/error verdicts on package structure: SKILL.md exists + parses, name present, description length, deprecated patterns, time-sensitive content, etc.

bash
# JSON output for parsing into the unified report
j-rig check "$(dirname "SKILL.md")" --json

This is a separate concern from the IS spec-compliance check (Tier 1) — JRig's check is structural, not rubric-based. The Anthropic + AgentSkills.io spec snapshots in 000-docs/ are read by the IS validator (scripts/validate-skills-schema.py), not by JRig directly. The two are complementary: IS validator scores against the spec rubric; JRig verifies the package shape and surfaces structural anti-patterns.

Verdict mapping:

  • All severity: "pass" → Tier 3A GREEN
  • Any severity: "warning" → Tier 3A YELLOW (non-blocking; surfaced in unified report)
  • Any severity: "error" → Tier 3A RED (blocks production promotion)
Tier 3B: 7-Layer Behavioral Eval (execution-based, slow, $)
bash
# Ad hoc invocation — Sonnet only, with no write to the Freshie inventory
j-rig eval "$(dirname "SKILL.md")" --json

# Full model matrix with a durable Freshie evidence row. Run from the repository root.
# j-rig receives only a scratch DB; the local recorder owns the Freshie write.
scripts/run-jrig-eval.sh \
  --skill-dir "$(dirname "SKILL.md")" \
  --plugin "<catalog-plugin-name>" \
  --models haiku,sonnet,opus \
  --inventory-db freshie/inventory.sqlite

# Skip specific layers (when iterating)
j-rig eval "$(dirname "SKILL.md")" --no-trigger --no-functional --json

Eval spec source: JRig reads <skill-dir>/eval-spec.yaml if present, or use --spec <path> to point at one elsewhere. A spec is currently required — j-rig eval errors if neither is found. (Auto-generating a baseline spec from the SKILL.md frontmatter — should_trigger/should_not_trigger cases derived from the description trigger phrases — is the planned j-rig scaffold-spec <skill-dir> command; until it ships, author the spec by hand or copy skill/eval.yaml as a template.)

Layers:

  1. Trigger quality: precision/recall on user prompts that should and should not activate the skill
  2. Functional quality: task completion + output format match against gold cases
  3. Regression protection: sacred-case suite cannot break (skill is locked-down on these)
  4. Baseline value: skill output beats naked Claude on the same prompt; if not, flag for obsolescence
  5. Model variance: independent pass/fail per Haiku / Sonnet / Opus; flags any model where the skill collapses
  6. Rollout safety: prompt-leakage detection, unsafe-pattern surfacing, jailbreak resistance
  7. Cost / latency: per-invocation token + wall-clock — must beat declared budget
Tier 3 verdict
  • All 7 layers pass on all 3 models → Tier 3 GREEN; skill is JRig-Verified.
  • Any layer fails on any model → Tier 3 RED; report which layer + which model + recommended remediation. Block production promotion.
  • JRig unavailable → Tier 3 SKIPPED; emit a one-line note in the unified report; do not block.
Show full SKILL.md (851 more words)Show less
Persist results to Freshie

JRig writes runtime tables into whatever --db it receives, so it must never receive freshie/inventory.sqlite. The supported wrapper gives JRig a temporary database under /dev/shm, captures its JSON, then invokes the repository-owned recorder to upsert only the governed forge_proofs row:

bash
scripts/run-jrig-eval.sh \
  --skill-dir "$(dirname "SKILL.md")" \
  --plugin "<catalog-plugin-name>" \
  --models haiku,sonnet,opus \
  --inventory-db freshie/inventory.sqlite

JRig's runtime tables remain in the temporary database and never enter the Freshie/Dolt export. The durable integration surface is forge_proofs, written by scripts/record-jrig-proofs.mjs:

sql
SELECT plugin_name, passed, layers_passed, baseline_delta, verified_at
FROM forge_proofs
WHERE verification_type = 'tier3-jrig';

Freshie remains the sole local writer of inventory and grade history. A future schema change may project selected JRig evidence into skill_compliance, but JRig itself still must not write that database directly:

  • jrig_passed (boolean)
  • jrig_tier_blocked (1–7 if any)
  • jrig_baseline_delta (numeric — skill output vs. naked Claude on same prompt)

Until then, the join above is the integration surface.

Snapshot refresh workflow (quarterly)

Tier 3A reads versioned snapshots, NOT live Anthropic / AgentSkills.io docs. Live-fetching from CI is a rate-limit + flakiness risk. The snapshot refresh is a separate PR cadence:

  1. Quarterly cron (or manual trigger) fetches latest specs from code.claude.com/docs/en/skills and agentskills.io/specification.
  2. Writes to 000-docs/anthropic-skills-spec-snapshot.md + 000-docs/agentskills-spec-snapshot.md.
  3. Opens a PR with the diff.
  4. PR review = the human gate that catches breaking spec changes before they reach validation.

This isolates "the spec changed" events from per-skill validation runs.

Step 3: Present Unified Report (all tiers)

Parse the combined output of Tiers 1–3 (or 1+2 when Tier 3 is skipped) and present:

Production verdict: PASS / FAIL / VERIFIED (when Tier 3 ran and all 7 layers green)

┌─ TIER 1: Marketplace grade ─────────────────────────────┐
│ Grade: [LETTER] ([SCORE]/100)                           │
│                                                          │
│ Pillar              | Score | Notes                      │
│ Progressive Disc.   | X/30  | Token economy, structure   │
│ Ease of Use         | X/25  | Metadata, discoverability  │
│ Utility             | X/20  | Problem solving, examples  │
│ Spec Compliance     | X/15  | Frontmatter, naming        │
│ Writing Style       | X/10  | Voice, objectivity         │
│ Modifiers           | +/-X  | Bonuses/penalties          │
└──────────────────────────────────────────────────────────┘

┌─ TIER 2: Static production gate ────────────────────────┐
│ allowed-tools accuracy:    PASS / FAIL                   │
│ Auth protocol documented:  PASS / FAIL / N/A             │
│ Dead code / drift:         PASS / WARN                   │
│ Tool-safety combo:         PASS / FAIL                   │
│ Orchestration bounds:      PASS / FAIL                   │
│ Verdict:                   GREEN / YELLOW / RED          │
└──────────────────────────────────────────────────────────┘

┌─ TIER 3: JRig behavioral eval (only if --thorough) ─────┐
│ 3A spec compliance:        PASS / FAIL                   │
│ 3B Layer 1 trigger:        Haiku|Sonnet|Opus → P/F      │
│    Layer 2 functional:     Haiku|Sonnet|Opus → P/F      │
│    Layer 3 regression:     Haiku|Sonnet|Opus → P/F      │
│    Layer 4 baseline:       Haiku|Sonnet|Opus → P/F      │
│    Layer 5 model variance: Haiku|Sonnet|Opus → P/F      │
│    Layer 6 rollout safety: Haiku|Sonnet|Opus → P/F      │
│    Layer 7 cost/latency:   Haiku|Sonnet|Opus → P/F      │
│ Baseline delta:            +N% vs. naked Claude          │
│ Verdict:                   VERIFIED / BLOCKED / SKIPPED  │
└──────────────────────────────────────────────────────────┘

Final verdict logic:

Tier 1Tier 2Tier 3Final
≥BGREENVERIFIEDPRODUCTION READY (JRig-Verified)
≥BGREENSKIPPEDPRODUCTION READY (unverified)
≥BYELLOW*PRODUCTION READY with warnings
anyRED*BLOCKED — fix Tier 2 fails before promoting
<B**BLOCKED — bring grade to B+ before promoting
≥BGREENBLOCKEDBLOCKED — JRig found behavioral regression; fix before promoting

Grade scale: A (90+), B (80-89), C (70-79), D (60-69), F (<60)

Step 4: Prioritized Fix Recommendations

List fixes sorted by point value (highest first):

Top improvements:

  1. {fix description} (+N pts)
  2. {fix description} (+N pts)
  3. {fix description} (+N pts)

Common high-value fixes:

  • Add "Use when" to description (+3 pts marketplace tier)
  • Add "Trigger with" to description (+3 pts marketplace tier)
  • Extract long content to references/ (up to +10 pts on token_economy)
  • Add missing sections: Overview (+4), Prerequisites (+2), Output (+2), Error Handling (+2)
  • Migrate compatible-with → compatibility (deprecation warning fix; run batch-remediate.py --migrate-compatible-with)
  • Add compatibility: field with one of the AgentSkills.io examples (+1 pt on metadata quality)
  • Add external resource links (+1 pt modifier)
  • Add DCI directives for discovery (file existence, git status, tool versions) (+1 pt modifier)
  • Add TOC to reference files >100 lines (+1 pt modifier)
  • Add feedback loops for quality-critical workflows (+2 pts utility)
  • Remove time-sensitive information (+1 pt modifier)
  • Ensure consistent terminology throughout (+1 pt writing style)
Step 5: Review Structural Advisors

The validator emits INFO-level structural suggestions (marketplace tier):

  • Split to commands: 3+ kebab-case ## operation-name sections detected without commands/ directory → suggest splitting into individual commands/*.md files
  • Offload to references: Body sections >20 lines (Output, Error Handling, Examples, etc.) without references/ directory → suggest moving to references/ with relative markdown links
  • DCI opportunities: Skill performs file existence checks, git operations, or tool version detection without DCI → suggest !command`` directives
Step 6: Auto-Fix (if requested or grade below B)

If grade < B (80), ask user: "Fix issues automatically?"

If approved, apply in order:

  1. Add missing sections (Overview, Prerequisites, Output, Error Handling, Examples)
  2. Add "Use when" / "Trigger with" to description if missing
  3. Move author/version/license from nested metadata to top-level (or vice versa per AgentSkills.io spec — both valid)
  4. Migrate compatible-with → compatibility via batch-remediate.py --migrate-compatible-with
  5. Fix text references to use relative markdown links: [file](references/file.md); keep ${CLAUDE_SKILL_DIR}/ for DCI/bash only
  6. Split long SKILL.md (>500 lines) into references/ with relative links
  7. Scope unscoped Bash tools: Bash → Bash(command:*)
  8. Add DCI directives for common discovery patterns (file checks, git status)
  9. If 3+ operation sections found, offer to split into commands/*.md files
  10. Add TOC to reference files >100 lines
  11. Flag time-sensitive content for review (dates, version numbers, URLs that may go stale)

After fixes, re-run validator and show before/after comparison.

Output

Final report includes:

  • Production verdict (PASS / FAIL / VERIFIED)
  • Tier 1 letter grade with numeric score + pillar breakdown table
  • Tier 2 static-gate pass/fail per check (5 checks)
  • Tier 3 JRig results per layer × per model (when --thorough); SKIPPED otherwise
  • Warning/error list from validator
  • Prioritized fix recommendations with point values
  • Before/after comparison if auto-fix was applied
  • Source citation for each spec-grounded claim
  • One-line install hint for JRig if Tier 3 was requested but JRig is missing

Error Handling

ErrorRecovery
File not foundSuggest Glob to find SKILL.md files nearby
Python not availableRead SKILL.md manually, check frontmatter by hand
Validator script missingFall back to manual checks against rubric pillars
YAML parse errorReport the parse error line, suggest fix
compatible-with deprecation warningRun batch-remediate.py --migrate-compatible-with

Examples

Validate a specific skill:

/validate-skillmd ~/.claude/skills/repo-sweep/SKILL.md

Marketplace grading (recommended for marketplace submissions):

/validate-skillmd --marketplace path/to/SKILL.md

Natural language:

grade my skill
check skill quality

Resources

© jeremylongshore, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in skills/.curated/validate-skillmd of jeremylongshore/tons-of-skills-marketplace.

  • SKILL.md
  • eval-spec.yaml
  • references/schema-reminder-frontmatter.md

Open the folder on GitHubat commit cfae287

Compare with similar skills

Validate Skillmd next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Validate Skillmd compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Validate Skillmd this skilljeremylongshore/tons-of-skills-marketplace2.8k—~6.2kAutomated safety check: PassMIT
Skill JudgeshareAI-lab/lab-skills315—~1.9kAutomated safety check: PassApache-2.0
Agent Launcher Orchestratoralirezarezvani/claude-skills28k—~1.3kAutomated safety check: PassMIT
Darwin SkillHHU3637kr/skills1451 repos~2.2kAutomated safety check: PassNone
Skill Authoringgrafana/skills282—~1.8kAutomated safety check: PassApache-2.0
Prompt LabMathews-Tom/armory329—~2.1kAutomated safety check: PassMIT

Similar skills

  • Skill Judge

    shareAI-lab/lab-skills

    Evaluate Agent Skill design quality with an opinionated, practice-derived rubric informed by public specifications and examples.

    315 GitHub stars~1.9k tokensUpdated 24 days ago
    EducationAuto-check passed
  • Agent Launcher Orchestrator

    alirezarezvani/claude-skills

    A skill your agent uses when a user wants to build, launch, grade, or schedule a Claude Managed Agent (CMA) in their own Anthropic account — "build me an agent", "launch this as a managed agent"…

    28k GitHub stars~1.3k tokensUpdated 1 mo ago
    EducationAuto-check passed
  • Darwin Skill

    HHU3637kr/skills

    Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch.

    145 GitHub starsUsed in 1 repo~2.2k tokens
    Agent WorkflowsAuto-check passed
  • Skill Authoring

    grafana/skills

    Official

    Author, audit, and improve Grafana SKILL.md files against Anthropic's published Agent Skills guidance and the four-dimension rubric the grafana/skills CI gate uses (conciseness, actionability…

    282 GitHub stars~1.8k tokensUpdated 2 days ago
    DevOps & CloudAuto-check passed
  • Prompt Lab

    Mathews-Tom/armory

    LLM prompt engineering: analyzes failure modes, generates variants (direct, few-shot, CoT), designs rubrics, produces test suites.

    329 GitHub stars~2.1k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Author Skill

    ericrisco/rsc-harness

    A skill your agent uses when authoring a NEW rsc skill or editing an existing one — scoping it to one job, writing the description that decides whether it ever loads, splitting the body into…

    180 GitHub stars~4.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed

More from jeremylongshore/tons-of-skills-marketplace

All 3,342 skills in this repo
  • Performing Security Code Review

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to conduct a security-focused code review using the security-agent plugin.

    2.8k GitHub starsUsed in 2 repos~1.3k tokens
    Auto-check: notes
  • Adapting Transfer Learning Models

    jeremylongshore/tons-of-skills-marketplace

    Build this skill automates the adaptation of pre-trained machine learning models using transfer learning techniques.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Agent Context Loader

    jeremylongshore/tons-of-skills-marketplace

    Execute proactive auto-loading: automatically detects and loads agents.md files.

    2.8k GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Aggregating Performance Metrics

    jeremylongshore/tons-of-skills-marketplace

    Aggregate and centralize performance metrics from applications, systems, databases, caches, and services.

    2.8k GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Analyzing Capacity Planning

    jeremylongshore/tons-of-skills-marketplace

    Execute this skill enables AI assistant to analyze capacity requirements and plan for future growth.

    2.8k GitHub stars~947 tokensUpdated today
    Auto-check passed
  • Analyzing Database Indexes

    jeremylongshore/tons-of-skills-marketplace

    Process use when you need to work with database indexing. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2k tokensUpdated today
    Auto-check passed

Questions about Validate Skillmd

What does Validate Skillmd do?

Validate a SKILL.md file against the four-tier validation system: Tier 0 (locate), Tier 1 (standard or marketplace grading per the IS 100-point rubric), Tier 2 (static production gate —…. Validate Skillmd is an agent skill from jeremylongshore/tons-of-skills-marketplace.md file against the four-tier validation system: Tier 0 (locate), Tier 1 (standard or marketplace grading per the IS 100-point rubric), Tier 2 (static production gate — allowed-tools accuracy, auth protocol, dead code, tool safety, orchestration bounds), and Tier 3 (JRig 7-layer behavioral eval, opt-in via --thorough).

When should I use Validate Skillmd?

Validate Skillmd fits situations like: creating a new skill; checking skill quality; preparing for marketplace submission; running deep quality analysis.

How do I install Validate Skillmd in Claude Code?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill validate-skillmd -a claude-code`. Or copy the skill folder (skills/.curated/validate-skillmd in jeremylongshore/tons-of-skills-marketplace) into .claude/skills/validate-skillmd in your project. Claude Code loads it when a task matches its description.

How do I install Validate Skillmd in Codex?

Run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill validate-skillmd -a codex`. Or copy the skill folder (skills/.curated/validate-skillmd in jeremylongshore/tons-of-skills-marketplace) into .agents/skills/validate-skillmd in your project. Codex loads it when a task matches its description.

Can I use Validate Skillmd in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jeremylongshore/tons-of-skills-marketplace --skill validate-skillmd -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/validate-skillmd, .gemini/skills/validate-skillmd, .github/skills/validate-skillmd and .opencode/skills/validate-skillmd in your project.

What does Validate Skillmd need to run?

Going by SKILL.md and its folder, Validate Skillmd needs the command-line tools its instructions call (python3 and pnpm) and credentials named GROQ_API_KEY and API_KEY. Our summary lists: Python 3; A credential in GROQ_API_KEY; A credential in API_KEY. Its frontmatter pre-approves these tools: Read, Edit, Write, Bash(python3:*), Bash(j-rig:*), Bash(node:*), Bash(scripts/run-jrig-eval.sh:*), Glob, Grep, AskUserQuestion. Compatibility (from SKILL.md): Designed for Claude Code; requires Python 3, optionally JRig CLI for Tier 3.

Does Validate Skillmd access the network?

SKILL.md names 3 domains. As links in the text: platform.claude.com, code.claude.com and agentskills.io. This is read from the text; nothing was executed.

Is Validate Skillmd safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Validate Skillmd use?

Validate Skillmd is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Validate Skillmd use?

About 6.2k tokens (SKILL.md is roughly 25k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 529 tokens, read only when the agent opens those files.

What are the alternatives to Validate Skillmd?

Skills that share tags, products or a category with Validate Skillmd: Skill Judge (shareAI-lab/lab-skills, 315 stars), Agent Launcher Orchestrator (alirezarezvani/claude-skills, 28k stars), Darwin Skill (HHU3637kr/skills, 145 stars) and Skill Authoring (grafana/skills, 282 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Validate Skillmd?

jeremylongshore (a GitHub user) maintains it in jeremylongshore/tons-of-skills-marketplace, which has 2,827 GitHub stars. The repository holds 3,342 skills in this directory. The repository was last updated on October 10, 2026.

Source: jeremylongshore/tons-of-skills-marketplace on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.