GPU Live Metric Validation
DataDog/datadog-agent
Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures.
Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments…
$ npx skills add pproenca/dot-skills --skill metric-validation-harness -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install pproenca/dot-skills metric-validation-harness --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/.experimental/metric-validation-harness .claude/skills/metric-validation-harness && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "metric-validation-harness" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/metric-validation-harness into .claude/skills/metric-validation-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "metric-validation-harness", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/metric-validation-harnessType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add pproenca/dot-skills --skill metric-validation-harness -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install pproenca/dot-skills metric-validation-harness --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/.experimental/metric-validation-harness .agents/skills/metric-validation-harness && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "metric-validation-harness" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/metric-validation-harness into .agents/skills/metric-validation-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "metric-validation-harness", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill metric-validation-harness -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install pproenca/dot-skills metric-validation-harness --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/.experimental/metric-validation-harness .cursor/skills/metric-validation-harness && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "metric-validation-harness" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/metric-validation-harness into .cursor/skills/metric-validation-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "metric-validation-harness", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/pproenca/dot-skills.git --path skills/.experimental/metric-validation-harness--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add pproenca/dot-skills --skill metric-validation-harness -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install pproenca/dot-skills metric-validation-harness --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/.experimental/metric-validation-harness .gemini/skills/metric-validation-harness && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "metric-validation-harness" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/metric-validation-harness into .gemini/skills/metric-validation-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "metric-validation-harness", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install pproenca/dot-skills metric-validation-harnessInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add pproenca/dot-skills --skill metric-validation-harness -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/.experimental/metric-validation-harness .github/skills/metric-validation-harness && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "metric-validation-harness" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/metric-validation-harness into .github/skills/metric-validation-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "metric-validation-harness", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add pproenca/dot-skills --skill metric-validation-harness -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install pproenca/dot-skills metric-validation-harness --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/pproenca/dot-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/.experimental/metric-validation-harness .opencode/skills/metric-validation-harness && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "metric-validation-harness" agent skill from https://github.com/pproenca/dot-skills/tree/master/skills/.experimental/metric-validation-harness into .opencode/skills/metric-validation-harness/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "metric-validation-harness", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
metric-validation-harnessEmpirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments…
Metric Validation Harness is an agent skill from pproenca/dot-skills. Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments that try to falsify each property a good metric must have. Checks determinism (same input, same number across runs and hash seeds), invariance to cosmetic edits (also an anti-gaming probe), monotonicity under construct-increasing edits, discrimination, robustness on edge inputs, near-linear tractability, and construct…
Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 32 other files, including scripts and reference files (for example `config.json`, `gotchas.md` and `metadata.json`).
The repository describes itself as: A collection of AI agent skills following the Agent Skills open format. The licence is MIT.
Read from SKILL.md and the folder at commit cf93c57. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 10 files in scripts/ (Python and Shell, from the files we listed), which the agent can run.
Shell commands in SKILL.md call:
bashpython3From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Metric Validation Harness loads about 1.7k tokens when it runs, and up to ~3.4k if it reads all its reference files. Until then it costs about 233 tokens; SKILL.md has 627 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from pproenca/dot-skills at commit cf93c57, republished under its MIT licence (© pproenca). 627 words, ~1,735 tokens.
.claude/skills/metric-validation-harness/SKILL.md (or your agent's skills folder). This skill also uses 27 other files; get the full folder from GitHub.Point this harness at a candidate metric and a corpus, and it runs experiments that try to falsify each property a trustworthy, optimizable metric must have. It is the empirical companion to deterministic-metric-design: that skill tells you to prove monotonicity, invariance, determinism, and construct validity; this skill runs the experiment and reports PASS/FAIL, each result mapped to the design-skill category it checks.
Read-only. It computes and reports; it never modifies your metric, the corpus, or any external state. Safe to run unsupervised.
deterministic-metric-design and want to empirically confirm the properties you argued forconfig.json / env → resolve metric_cmd, corpus, thresholds (env > config > bundled default)
│
▼
verify.sh ──► determinism ─ invariance ─ monotonicity ─ robustness ─ tractability ─ validity
│ (each property check maps to a deterministic-metric-design category)
▼
PASS / FAIL per property → exit 0 (all pass) or 1 (any group failed)Your metric is any command that takes a path as its last argument and prints exactly one number to stdout:
$ python3 mymetric.py path/to/file.py
42Language-agnostic — Python, a shell one-liner, a compiled binary, anything. Diagnostics go to stderr; stdout is the number only. A bundled example metric (scripts/examples/metric_ast_nodes.py, AST-node count) ships so the harness runs out of the box.
# 1. Validate the bundled example metric (works with zero setup):
bash scripts/verify.sh
# 2. Validate YOUR metric — set metric_cmd in config.json, or override per-run:
METRIC_CMD="python3 /abs/path/mymetric.py" bash scripts/verify.sh
# 3. Prove the harness itself works (positive + negative cases):
bash scripts/selftest.sh
# 4. Sanity-check your adapter prints one number:
bash scripts/run-metric.sh path/to/file.pyverify.sh runs every check and prints a final PASS/FAIL. Each check is also runnable on its own (e.g. bash scripts/check-determinism.sh).
| Check | Maps to (design skill) | What it does | PASS condition |
|---|---|---|---|
check-determinism.sh | det- | Runs the metric twice + under PYTHONHASHSEED 0/1 | identical number every time |
check-invariance.sh | prop- / game- | Adds comments/blank lines/whitespace (cosmetic) | score unchanged (else it's gameable) |
check-monotonicity.sh | prop- | Appends a code block (construct-increasing) + checks spread | score non-decreasing; not saturated |
check-robustness.sh | prop- | Empty + single-statement edge inputs | finite, in declared range, no crash |
check-tractability.py | comp- | Times the metric on growing inputs | within budget, sub-quadratic growth |
check-validity.py | valid- | Spearman vs accepted, vs LOC; AUC vs outcome | convergent high, discriminant not ~LOC, predictive beats baseline |
Statistics (Spearman, AUC/Mann–Whitney) are pure Python stdlib — no numpy/scipy.
The harness runs with zero config against the bundled example. To validate your own metric, set fields in config.json (or override any of them with the matching UPPER_CASE environment variable per run):
| config.json | Env override | Meaning |
|---|---|---|
metric_cmd | METRIC_CMD | your metric command (path-printing → number) |
baseline_cmd | BASELINE_CMD | trivial baseline (default: bundled LOC) |
corpus_dir | CORPUS_DIR | artifacts the property checks iterate over |
labels_csv | LABELS_CSV | path[,outcome][,accepted] for validity |
declared_min / declared_max | DECLARED_MIN / DECLARED_MAX | range the robustness check enforces |
Validity thresholds are env-tunable: CONVERGENT_MIN, DISCRIMINANT_MAX, PREDICTIVE_MIN (defaults are lenient — tighten for a real run; see gotchas.md).
Empty config fields fall back to the bundled demo, so the skill never crashes on missing setup — it runs the example instead.
python3 (3.8+) — runs the metric, the transforms, and the statsbash and awk — the orchestrator and numeric comparisons (scripts are macOS bash 3.2-safe)No network, no external packages.
A FAIL names the property and the design-skill rule to consult. Examples:
prop-prove-invariance-under-irrelevant-transforms and game-make-cheapest-improvement-the-right-one.prop-prove-monotonicity).valid-discriminant-not-just-loc).deterministic-metric-design — the design half. Use it to construct the metric (define the construct, choose a computable proxy, pick the scale, argue the properties); use this harness to empirically verify what you argued.same-results-less-code, complexity-optimizer, knip-deadcode — prescriptive code-reduction skills; validate any reduction metric you build to drive them with this harness before letting an agent optimize against it.See references/workflow.md for per-check details, how to wire up your own metric and corpus, and troubleshooting.
© pproenca, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 27 other files (scripts, references) in skills/.experimental/metric-validation-harness of pproenca/dot-skills.
Open the folder on GitHubat commit cf93c57
Metric Validation Harness next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Metric Validation Harness this skillpproenca/dot-skills | 215 | — | ~1.7k | Automated safety check: Pass | MIT | |
| GPU Live Metric ValidationDataDog/datadog-agent | 3.8k | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | |
| SQL Optimizationgithub/awesome-copilot | 40k | 2 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Agent Performance Optimizerruvnet/ruflo | 74k | 2 repos | ~3.6k | Automated safety check: Pass | MIT | |
| Form Validationthedaviddias/Front-End-Checklist | 74k | — | ~633 | Automated safety check: Pass | MIT | |
| Database Optimizerdavila7/claude-code-templates | 32k | 8 repos | ~2.5k | Automated safety check: Pass | MIT |
DataDog/datadog-agent
Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures.
github/awesome-copilot
Universal SQL performance optimization assistant for comprehensive query tuning, indexing strategies, and database performance analysis across all SQL databases (MySQL, PostgreSQL, SQL Server…
ruvnet/ruflo
Agent skill for performance-optimizer - invoke with $agent-performance-optimizer
thedaviddias/Front-End-Checklist
A skill your agent uses when reviewing templates, rendered HTML, or shared components related to Validate forms accessibly.
davila7/claude-code-templates
Expert database optimizer specializing in modern performance tuning, query optimization, and scalable architectures.
affaan-m/ECC
分析原始提示,识别意图和差距,匹配ECC组件(技能/命令/代理/钩子),并输出一个可直接粘贴的优化提示。仅提供咨询角色——绝不自行执行任务。触发时机:当用户说“优化提示”、“改进我的提示”、“如何编写提示”、“帮我优化这个指令”或明确要求提高提示质量时。中文等效表达同样触发:“优化prompt”、“改进prompt”、“怎么写prompt”、“帮我优化这个指令”。不触发时机:当用户希望直接执行任…
pproenca/dot-skills
Audio forensics and voice recovery guidelines for CSI-level audio analysis.
pproenca/dot-skills
Guided, scripted pipeline for running JSX/TSX/React codemods safely across large legacy codebases.
pproenca/dot-skills
Create well-structured RFCs and technical proposals for software projects.
pproenca/dot-skills
Developer-experience friction auditing and fixing — slow onboarding, repeated manual setup steps, missing bootstrap/reset/seed scripts, undiscoverable conventions.
pproenca/dot-skills
Turn a rough idea for a language into a complete, implementable specification — a DSL, query, config/data, template, or protocol language — by interviewing the author dimension by dimension until…
pproenca/dot-skills
Drafting Python Enhancement Proposals (PEPs) — proposing a Python language feature, a standard library change, an interoperability standard, or an informational/process document for the Python…
Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments…. Metric Validation Harness is an agent skill from pproenca/dot-skills. Empirically validates a software metric before trusting or optimizing it — point it at any candidate metric (a command that takes a path and prints one number) plus a corpus, and it runs experiments that try to falsify each property a good metric must have.
Metric Validation Harness fits situations like: ever someone proposes; asks is this metric any good; suspects a score tracks LOC; jumps between runs.
Run `npx skills add pproenca/dot-skills --skill metric-validation-harness -a claude-code`. Or copy the skill folder (skills/.experimental/metric-validation-harness in pproenca/dot-skills) into .claude/skills/metric-validation-harness in your project. Claude Code loads it when a task matches its description.
Run `npx skills add pproenca/dot-skills --skill metric-validation-harness -a codex`. Or copy the skill folder (skills/.experimental/metric-validation-harness in pproenca/dot-skills) into .agents/skills/metric-validation-harness in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add pproenca/dot-skills --skill metric-validation-harness -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/metric-validation-harness, .gemini/skills/metric-validation-harness, .github/skills/metric-validation-harness and .opencode/skills/metric-validation-harness in your project.
Going by SKILL.md and its folder, Metric Validation Harness needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (bash and python3). Our summary lists: Python 3; A Bash shell.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Metric Validation Harness is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.7k tokens (SKILL.md is roughly 6.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Metric Validation Harness: GPU Live Metric Validation (DataDog/datadog-agent, 3.8k stars), SQL Optimization (github/awesome-copilot, 40k stars), Agent Performance Optimizer (ruvnet/ruflo, 74k stars) and Form Validation (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
pproenca (a GitHub user) maintains it in pproenca/dot-skills, which has 215 GitHub stars. The repository holds 41 skills in this directory. The repository was last updated on August 15, 2026.
Source: pproenca/dot-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.