Agent skill

Metric Validation Harness

by pproenca in pproenca/dot-skills

Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments…

MITAuto-check passed

Install Metric Validation Harness

skills CLI
$ npx skills add pproenca/dot-skills --skill metric-validation-harness -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install pproenca/dot-skills metric-validation-harness --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/metric-validation-harness .claude/skills/metric-validation-harness && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
metric-validation-harness
GitHub stars
215
Token cost
~1.7k tokens
SKILL.md length
627 words
Files
28 (incl. scripts, references)
Skills in repo
41
Repo updated
First seen
Licence
MIT

At a glance

Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments…

  • Ever someone proposes
  • SKILL.md covers When to Apply, Workflow Overview, The Adapter Contract and How to Run, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls bash and python3
  • Asks is this metric any good

What it does

Metric Validation Harness is an agent skill from pproenca/dot-skills. Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments that try to falsify each property a good metric must have. Checks determinism (same input, same number across runs and hash seeds), invariance to cosmetic edits (also an anti-gaming probe), monotonicity under construct-increasing edits, discrimination, robustness on edge inputs, near-linear tractability, and construct…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 32 other files, including scripts and reference files (for example `config.json`, `gotchas.md` and `metadata.json`).

The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.

When your agent uses it

  • Ever someone proposes
  • Asks is this metric any good
  • Suspects a score tracks LOC
  • Jumps between runs

Example prompts

  • “is this metric any good”
  • “/metric-validation-harness”

Requirements

  • Python 3
  • A Bash shell

What it can do on your machine

Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 10 files in scripts/ (Python and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • bash
    • python3

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Metric Validation Harness loads about 1.7k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 233 tokens; SKILL.md has 627 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~233
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 627 words, ~1,735 tokens.

Download SKILL.mdSave it as .claude/skills/metric-validation-harness/SKILL.md (or your agent's skills folder). This skill also uses 27 other files; get the full folder from GitHub.
name
metric-validation-harness
description
Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments that try to falsify each property a good metric must have. Checks determinism (same input, same number across runs and hash seeds), invariance to cosmetic edits (also an anti-gaming probe), monotonicity under construct-increasing edits, discrimination, robustness on edge inputs, near-linear tractability, and construct validity (convergent, discriminant vs LOC, predictive AUC, lift over a baseline). Trigger whenever someone proposes, reviews, tunes, or ships a metric, score, or index, asks "is this metric any good", suspects a score tracks LOC or jumps between runs, or builds a deterministic optimization target. It is the empirical companion to the deterministic-metric-design skill and is read-only.

Metric Validation Harness

Point this harness at a candidate metric and a corpus, and it runs experiments that try to falsify each property a trustworthy, optimizable metric must have. It is the empirical companion to deterministic-metric-design: that skill tells you to prove monotonicity, invariance, determinism, and construct validity; this skill runs the experiment and reports PASS/FAIL, each result mapped to the design-skill category it checks.

Read-only. It computes and reports; it never modifies your metric, the corpus, or any external state. Safe to run unsupervised.

When to Apply

  • Someone proposes, reviews, tunes, or ships a metric / score / index and you need evidence it is sound
  • A score "feels off" — you suspect it tracks LOC, jumps between runs, or saturates
  • You are about to let an agent optimize a metric and need to know it can't be gamed by cosmetic edits
  • You built a candidate per deterministic-metric-design and want to empirically confirm the properties you argued for
  • You are choosing between two metrics and need to know which actually predicts the outcome (and beats a trivial baseline)

Workflow Overview

config.json / env  →  resolve metric_cmd, corpus, thresholds (env > config > bundled default)
        │
        ▼
   verify.sh ──► determinism ─ invariance ─ monotonicity ─ robustness ─ tractability ─ validity
        │            (each property check maps to a deterministic-metric-design category)
        ▼
   PASS / FAIL per property  →  exit 0 (all pass) or 1 (any group failed)

The Adapter Contract

Your metric is any command that takes a path as its last argument and prints exactly one number to stdout:

bash
$ python3 mymetric.py path/to/file.py
42

Language-agnostic — Python, a shell one-liner, a compiled binary, anything. Diagnostics go to stderr; stdout is the number only. A bundled example metric (scripts/examples/metric_ast_nodes.py, AST-node count) ships so the harness runs out of the box.

How to Run

bash
# 1. Validate the bundled example metric (works with zero setup):
bash scripts/verify.sh

# 2. Validate YOUR metric — set metric_cmd in config.json, or override per-run:
METRIC_CMD="python3 /abs/path/mymetric.py" bash scripts/verify.sh

# 3. Prove the harness itself works (positive + negative cases):
bash scripts/selftest.sh

# 4. Sanity-check your adapter prints one number:
bash scripts/run-metric.sh path/to/file.py

verify.sh runs every check and prints a final PASS/FAIL. Each check is also runnable on its own (e.g. bash scripts/check-determinism.sh).

What It Checks

CheckMaps to (design skill)What it doesPASS condition
check-determinism.shdet-Runs the metric twice + under PYTHONHASHSEED 0/1identical number every time
check-invariance.shprop- / game-Adds comments/blank lines/whitespace (cosmetic)score unchanged (else it's gameable)
check-monotonicity.shprop-Appends a code block (construct-increasing) + checks spreadscore non-decreasing; not saturated
check-robustness.shprop-Empty + single-statement edge inputsfinite, in declared range, no crash
check-tractability.pycomp-Times the metric on growing inputswithin budget, sub-quadratic growth
check-validity.pyvalid-Spearman vs accepted, vs LOC; AUC vs outcomeconvergent high, discriminant not ~LOC, predictive beats baseline

Statistics (Spearman, AUC/Mann–Whitney) are pure Python stdlib — no numpy/scipy.

Show full SKILL.md (267 more words)Show less

Setup & Configuration

The harness runs with zero config against the bundled example. To validate your own metric, set fields in config.json (or override any of them with the matching UPPER_CASE environment variable per run):

config.jsonEnv overrideMeaning
metric_cmdMETRIC_CMDyour metric command (path-printing → number)
baseline_cmdBASELINE_CMDtrivial baseline (default: bundled LOC)
corpus_dirCORPUS_DIRartifacts the property checks iterate over
labels_csvLABELS_CSVpath[,outcome][,accepted] for validity
declared_min / declared_maxDECLARED_MIN / DECLARED_MAXrange the robustness check enforces

Validity thresholds are env-tunable: CONVERGENT_MIN, DISCRIMINANT_MAX, PREDICTIVE_MIN (defaults are lenient — tighten for a real run; see gotchas.md).

Empty config fields fall back to the bundled demo, so the skill never crashes on missing setup — it runs the example instead.

Tool Requirements

  • python3 (3.8+) — runs the metric, the transforms, and the stats
  • bash and awk — the orchestrator and numeric comparisons (scripts are macOS bash 3.2-safe)

No network, no external packages.

Interpreting Results

A FAIL names the property and the design-skill rule to consult. Examples:

  • cosmetic noise moved the score → the metric reads surface text; see prop-prove-invariance-under-irrelevant-transforms and game-make-cheapest-improvement-the-right-one.
  • score DROPPED after adding code → non-monotonic; optimizing it can reward worse code (prop-prove-monotonicity).
  • |Spearman(metric, LOC)| too high → it's LOC relabeled (valid-discriminant-not-just-loc).
  • deterministic-metric-design — the design half. Use it to construct the metric (define the construct, choose a computable proxy, pick the scale, argue the properties); use this harness to empirically verify what you argued.
  • same-results-less-code, complexity-optimizer, knip-deadcode — prescriptive code-reduction skills; validate any reduction metric you build to drive them with this harness before letting an agent optimize against it.

See references/workflow.md for per-check details, how to wire up your own metric and corpus, and troubleshooting.

© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 27 other files (scripts, references) in skills/.experimental/metric-validation-harness of pproenca/dot-skills.

  • SKILL.md
  • .gitignore
  • config.json
  • gotchas.md
  • metadata.json
  • references/workflow.md
  • scripts/check-determinism.sh
  • scripts/check-invariance.sh
  • scripts/check-monotonicity.sh
  • scripts/check-robustness.sh
  • scripts/check-tractability.py
  • scripts/check-validity.py
  • scripts/examples/metric_ast_nodes.py
  • scripts/examples/metric_loc.py
  • scripts/fixtures/corpus.csv
  • scripts/fixtures/corpus/f1_dense.py
  • … and 12 more

Open the folder on GitHubat commit cf93c57

Compare with similar skills

Metric Validation Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Metric Validation Harness compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Metric Validation Harness this skillpproenca/dot-skills215—~1.7kAutomated safety check: PassMIT
GPU Live Metric ValidationDataDog/datadog-agent3.8k—~1.1kAutomated safety check: PassApache-2.0
SQL Optimizationgithub/awesome-copilot40k2 repos~2.3kAutomated safety check: PassMIT
Agent Performance Optimizerruvnet/ruflo74k2 repos~3.6kAutomated safety check: PassMIT
Form Validationthedaviddias/Front-End-Checklist74k—~633Automated safety check: PassMIT
Database Optimizerdavila7/claude-code-templates32k8 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • GPU Live Metric Validation

    DataDog/datadog-agent

    Official

    Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures.

    3.8k GitHub stars~1.1k tokensUpdated today
    DevOps & CloudAuto-check passed
  • SQL Optimization

    github/awesome-copilot

    Official

    Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…

    40k GitHub starsUsed in 2 repos~2.3k tokens
    DatabasesAuto-check passed
  • Agent skill for performance-optimizer - invoke with $agent-performance-optimizer

    74k GitHub starsUsed in 2 repos~3.6k tokens
    Auto-check passed
  • Form Validation

    thedaviddias/Front-End-Checklist

    A skill your agent uses when reviewing templates, rendered HTML, or shared components related to Validate forms accessibly.

    74k GitHub stars~633 tokensUpdated 2 days ago
    Frontend & DesignAuto-check passed
  • Database Optimizer

    davila7/claude-code-templates

    Expert database optimizer specializing in modern performance tuning, query optimization, and scalable architectures.

    32k GitHub starsUsed in 8 repos~2.5k tokens
    DatabasesAuto-check passed
  • Prompt Optimizer

    affaan-m/ECC

    分析原始提示,识别意图和差距,匹配ECC组件(技能/命令/代理/钩子),并输出一个可直接粘贴的优化提示。仅提供咨询角色——绝不自行执行任务。触发时机:当用户说“优化提示”、“改进我的提示”、“如何编写提示”、“帮我优化这个指令”或明确要求提高提示质量时。中文等效表达同样触发:“优化prompt”、“改进prompt”、“怎么写prompt”、“帮我优化这个指令”。不触发时机:当用户希望直接执行任…

    275k GitHub starsUsed in 2 repos~2.4k tokens
    DevelopmentAuto-check passed

More from pproenca/dot-skills

All 41 skills in this repo
  • Audio Voice Recovery

    pproenca/dot-skills

    Audio forensics and voice recovery guidelines for CSI-level audio analysis.

    215 GitHub stars~3.3k tokensUpdated 1 mo ago
    Auto-check passed
  • Codemod React Pipeline

    pproenca/dot-skills

    Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.

    215 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed
  • Dev Rfc

    pproenca/dot-skills

    Create well-structured RFCs and technical proposals for software projects.

    215 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Dx Harness

    pproenca/dot-skills

    Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.

    215 GitHub stars~1.5k tokensUpdated 1 mo ago
    Auto-check passed
  • Language Spec Author

    pproenca/dot-skills

    Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…

    215 GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Python Pep Author

    pproenca/dot-skills

    Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…

    215 GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Metric Validation Harness

What does Metric Validation Harness do?

Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments…. Metric Validation Harness is an agent skill from pproenca/dot-skills. Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments that try to falsify each property a good metric must have.

When should I use Metric Validation Harness?

Metric Validation Harness fits situations like: ever someone proposes; asks is this metric any good; suspects a score tracks LOC; jumps between runs.

How do I install Metric Validation Harness in Claude Code?

Run `npx skills add pproenca/dot-skills --skill metric-validation-harness -a claude-code`. Or copy the skill folder (skills/.experimental/metric-validation-harness in pproenca/dot-skills) into .claude/skills/metric-validation-harness in your project. Claude Code loads it when a task matches its description.

How do I install Metric Validation Harness in Codex?

Run `npx skills add pproenca/dot-skills --skill metric-validation-harness -a codex`. Or copy the skill folder (skills/.experimental/metric-validation-harness in pproenca/dot-skills) into .agents/skills/metric-validation-harness in your project. Codex loads it when a task matches its description.

Can I use Metric Validation Harness in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill metric-validation-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/metric-validation-harness, .gemini/skills/metric-validation-harness, .github/skills/metric-validation-harness and .opencode/skills/metric-validation-harness in your project.

What does Metric Validation Harness need to run?

Going by SKILL.md and its folder, Metric Validation Harness needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (bash and python3). Our summary lists: Python 3; A Bash shell.

Does Metric Validation Harness access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Metric Validation Harness safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Metric Validation Harness use?

Metric Validation Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Metric Validation Harness use?

About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Metric Validation Harness?

Skills that share tags, products or a category with Metric Validation Harness: GPU Live Metric Validation (DataDog/datadog-agent, 3.8k stars), SQL Optimization (github/awesome-copilot, 40k stars), Agent Performance Optimizer (ruvnet/ruflo, 74k stars) and Form Validation (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Metric Validation Harness?

pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 215 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on August 15, 2026.

Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.