Agent skill

Benchmark Validator

by Intelligent-Internet in Intelligent-Internet/zenith

Benchmark validation procedure for one assigned benchmark-related target.

Apache-2.0Auto-check passedAgent Workflows

Install Benchmark Validator

skills CLI
$ npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Intelligent-Internet/zenith benchmark-validator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/benchmark-validator .claude/skills/benchmark-validator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-validator
GitHub stars
334
Token cost
~1.8k tokens
SKILL.md length
720 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Benchmark validation procedure for one assigned benchmark-related target.

  • Works in 6 steps: Identify artifacts → Audit evidence integrity → Define adversarial validation plan → …
  • Agent Workflows work in your project
  • SKILL.md covers Inputs, Procedure and Report
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmark Validator is an agent skill from Intelligent-Internet/zenith. Benchmark validation procedure for one assigned benchmark-related target. For optimization EXP- targets, independently classify candidate outcome. For engineering VAL- or legacy engineering targets, prove or disprove the required benchmark/performance assertion.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. It works with Model Context Protocol. The repository describes itself as: Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP. The licence is Apache-2.0.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “/benchmark-validator”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Identify artifacts
  2. Audit evidence integrity
  3. Define adversarial validation plan
  4. Remeasure independently
  5. Run correctness and guardrails
  6. Classify outcome

What it can do on your machine

Read from SKILL.md and the folder at commit a8d9b57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Validator loads about 1.8k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 720 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~71
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Intelligent-Internet/zenith at commit a8d9b57, republished under its Apache-2.0 licence (© Intelligent-Internet). 720 words, ~1,788 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-validator/SKILL.md (or your agent's skills folder).
name
benchmark-validator
description
Benchmark validation procedure for one assigned benchmark-related target. For optimization EXP-* targets, independently classify candidate outcome. For engineering VAL-* or legacy engineering targets, prove or disprove the required benchmark/performance assertion.

Benchmark Validator

Use this skill for a validation assignment targeting exactly one benchmark-related assertion.

For optimization EXP-* targets, you are not selecting the winner or optimizing further. You independently decide whether the experiment produced a trustworthy outcome under its contract.

For optimization EXP-* targets, act as an adversarial tester for the selected candidate. The contract is the minimum bar, not the whole test plan. Before promotion, define and run a compact but comprehensive validation plan that attacks the candidate's likely correctness and performance failure modes.

For engineering VAL-* or legacy engineering targets, do not use optimization outcome semantics. Pass only when the assigned benchmark, performance, correctness, and guardrail behavior required by the assignment and contract is proven with fresh evidence.

Inputs

Read:

  • Validation assignment.
  • The single assigned benchmark-related contract.
  • AGENTS.md.
  • Experiment ledger path cited by the assignment or contract.
  • Candidate artifact/ref/patch/checksum cited by the assignment, contract, or ledger.
  • Measurement protocol, correctness/guardrail commands, protected files, and source/baseline refs cited by the contract.
  • For optimization targets, also inspect visible runtime resources read-only when useful: changed files, benchmark scripts, fixtures, public outputs, baseline/reference artifacts, verifier-adjacent code, logs, and generated artifacts. Do not mutate these resources, copy protected artifacts into a submission, or make submitted code depend on evaluator-only locations.

If the assignment targets more than one benchmark-related assertion, fail the assignment as too broad and request attention.

Procedure

  1. Identify artifacts

    • Parent/baseline ref.
    • Candidate ref or patch artifact.
    • Benchmark command, run count, aggregation, variance policy.
    • Correctness and guardrail commands.
    • Protected benchmark/scoring/data/verifier files.
  2. Audit evidence integrity

    • Inspect candidate changes affecting benchmark scripts, scoring, data, fixtures, golden outputs, correctness verifiers, generated expected outputs, seeds, or measurement config.
    • If proof is weakened or unapproved metric surfaces changed, classify an optimization experiment invalid; for an engineering target, mark the item passed=false.
  3. Define adversarial validation plan

    • For optimization EXP-*, do this before remeasurement. Do not only run the contract's benchmark command.
    • From the contract, ledger, changed files, source, available fixtures, benchmark scripts, baseline/reference artifacts, public outputs, and visible runtime resources, identify the correctness and performance risk axes this candidate could break.
    • Build a compact test plan that includes the contract-required checks plus targeted/adversarial/proxy cases for changed surfaces. Cover important input, data, media, config, size, timing, cache, concurrency, edge, stress, and failure-mode axes when they are relevant to the candidate.
    • For each case, state the expected correctness evidence, the failure boundary, and whether timing evidence is needed. Record limitations for risk axes that cannot be tested in the validation environment.
    • If the contract or ledger does not provide enough oracle, candidate binding, or correctness information to build a meaningful adversarial plan, classify the optimization experiment invalid rather than promoting from aggregate speed alone.
  4. Remeasure independently

    • Reconstruct parent and candidate states using the method assigned by the assignment or contract.
    • Avoid destructive operations on the shared workspace.
    • Run parent and candidate under the declared protocol.
    • Capture raw outputs, aggregate values, spread/noise, environment, and deviations.
  5. Run correctness and guardrails

    • Required correctness/quality/compatibility/safety checks must run before objective metric comparison can promote a candidate.
    • For optimization EXP-*, run the adversarial validation plan as well as the contract-required checks.
    • Record per-case correctness and speed evidence before relying on aggregate metric output.
    • Failed checks make the candidate rejected, not promoted.
    • Missing, stale, skipped, or un-runnable checks make the experiment invalid.
  6. Classify outcome

    • First determine whether the assigned target is an optimization EXP-* target or an engineering/legacy target.
    • promoted: credible integrity, setup complete, correctness/guardrails pass, metric meets promotion rule.
    • rejected: credible integrity and setup, but correctness/guardrails fail or metric does not win.
    • budget_exhausted: declared budget prevented completion and the contract allows that honest outcome.
    • invalid: missing ledger/ref/checks, broken candidate binding, compromised evidence, or unverifiable setup.
Show full SKILL.md (106 more words)Show less

For optimization EXP-* targets, use passed=true for promoted, rejected, or legitimate budget_exhausted; use passed=false for invalid.

For engineering VAL-* or legacy engineering targets, use normal validation semantics: passed=true only when the assigned benchmark/performance assertion and required guardrails pass with fresh evidence. A rejected candidate, failed guardrail, missing evidence, or budget-exhausted result is passed=false unless the contract explicitly defines that outcome as the required behavior.

  1. Write regression ledger only for invalid
    • If passed=false, write <regressions_dir>/<item_id>.md with setup, command, expected credible measurement or required behavior, observed invalidating issue, and evidence artifact paths.

Report

markdown
## Experiment
- Item ID: <id>
- Parent artifact: <ref/path>
- Candidate artifact: <ref/path>
- Ledger: <path>

## Evidence Integrity
- Changed protected/evidence-sensitive paths: <list>
- Verdict: <credible | invalid>

## Measurements
- Parent: <runs, aggregate, spread, raw paths>
- Candidate: <runs, aggregate, spread, raw paths>
- Improvement: <calculation and threshold>

## Adversarial Validation Plan
- Risk axes: <candidate-specific axes attacked>
- Cases: <case -> expected evidence -> failure boundary>
- Limitations: <untested risk axes and why>

## Correctness And Guardrails
- `<command>` -> exit <code>, <result>
- Per-case adversarial results: <case -> passed/rejected/invalid evidence>

## Outcome
- <promoted | rejected | budget_exhausted | invalid>
- Reason: <contract-tied explanation>

## Limitations
- <setup or measurement caveats>

Call end_node with exactly one item for the assigned target id, then exit immediately.

© Intelligent-Internet, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in zenith/src/zenith_harness/bundled/skills/benchmark-validator of Intelligent-Internet/zenith.

Open the folder on GitHubat commit a8d9b57

Compare with similar skills

Benchmark Validator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Validator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Validator this skillIntelligent-Internet/zenith334—~1.8kAutomated safety check: PassApache-2.0
Google Antigravity SDKgoogle-antigravity/antigravity-sdk-python3.6k—~2.1kAutomated safety check: NotesApache-2.0
Chatgpt AppsHaohao-end/openagent8071 repos~4.9kAutomated safety check: PassApache-2.0
Autocontext for Hermesgreyhaven-ai/autocontext1.3k—~2.5kAutomated safety check: PassApache-2.0
Workflow Schema Tuningbreaking-brake/cc-wf-studio5.4k—~1.3kAutomated safety check: PassCustom licence
Documentation Serverandrea9293/mcp-documentation-server343—~2.3kAutomated safety check: PassMIT

Similar skills

  • Google Antigravity SDK

    google-antigravity/antigravity-sdk-python

    Design, implement, and debug autonomous AI agents and multi-agent systems using the Google Antigravity (AGY) SDK.

    3.6k GitHub stars~2.1k tokensUpdated 10 days ago
    Agent WorkflowsAuto-check: notes
  • Chatgpt Apps

    Haohao-end/openagent

    Build, scaffold, refactor, and troubleshoot ChatGPT Apps SDK applications that combine an MCP server and widget UI.

    807 GitHub starsUsed in 1 repo~4.9k tokens
    Agent WorkflowsAuto-check passed
  • Autocontext for Hermes

    greyhaven-ai/autocontext

    Lets a Hermes agent run Autocontext scenarios, inspect Hermes curator state, export reusable knowledge and prepare local MLX or CUDA training data through the autoctx CLI.

    1.3k GitHub stars~2.5k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Workflow Schema Tuning

    breaking-brake/cc-wf-studio

    Guides edits to cc-wf-studio's workflow schema so AI agents generate better workflows, treating schema text as prompt engineering rather than validation.

    5.4k GitHub stars~1.3k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Documentation Server

    andrea9293/mcp-documentation-server

    A skill your agent uses when you need to store, retrieve, search, or manage documents in a local knowledge base with semantic search and hybrid (vector + full-text) retrieval.

    343 GitHub stars~2.3k tokensUpdated 1 mo ago
    Agent WorkflowsAuto-check passed
  • Clawmem

    yoloshii/ClawMem

    ClawMem operational reference for agents at query time — the 3-rule escalation gate, MCP tool routing, the 4 query-optimization levers, pipeline behavior (query vs intentsearch), composite scoring…

    210 GitHub stars~7.5k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed

More from Intelligent-Internet/zenith

  • Engineering Mission Playbook

    Intelligent-Internet/zenith

    A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…

    334 GitHub stars~8.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Scrutiny Validator

    Intelligent-Internet/zenith

    Adversarial scrutiny procedure for engineering validation assignments.

    334 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • User Testing Validator

    Intelligent-Internet/zenith

    Real-surface validation coordinator for engineering validation assignments.

    334 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Browser

    Intelligent-Internet/zenith

    Automates browser and Electron app interactions for user-flow validation.

    334 GitHub stars~6.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Optimization Mission Playbook

    Intelligent-Internet/zenith

    Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…

    334 GitHub stars~11k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Benchmark Validator

What does Benchmark Validator do?

Benchmark validation procedure for one assigned benchmark-related target. Benchmark Validator is an agent skill from Intelligent-Internet/zenith. Benchmark validation procedure for one assigned benchmark-related target.

When should I use Benchmark Validator?

Benchmark Validator fits situations like: agent Workflows work in your project.

How do I install Benchmark Validator in Claude Code?

Run `npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a claude-code`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/benchmark-validator in Intelligent-Internet/zenith) into .claude/skills/benchmark-validator in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Validator in Codex?

Run `npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a codex`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/benchmark-validator in Intelligent-Internet/zenith) into .agents/skills/benchmark-validator in your project. Codex loads it when a task matches its description.

Can I use Benchmark Validator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Intelligent-Internet/zenith --skill benchmark-validator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-validator, .gemini/skills/benchmark-validator, .github/skills/benchmark-validator and .opencode/skills/benchmark-validator in your project.

What does Benchmark Validator need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark Validator is instructions for the agent only.

Does Benchmark Validator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Validator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Validator use?

Benchmark Validator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Validator use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Validator?

Skills that share tags, products or a category with Benchmark Validator: Google Antigravity SDK (google-antigravity/antigravity-sdk-python, 3.6k stars), Chatgpt Apps (Haohao-end/openagent, 807 stars), Autocontext for Hermes (greyhaven-ai/autocontext, 1.3k stars) and Workflow Schema Tuning (breaking-brake/cc-wf-studio, 5.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Validator?

Intelligent-Internet (a GitHub organization) maintains it in Intelligent-Internet/zenith, which has 334 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 6, 2026.

Source: Intelligent-Internet/zenith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.