Agent skill

Verifier Evaluations

by morluto in morluto/jacobian

Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring.

MITAuto-check passedResearch & Science

Install Verifier Evaluations

skills CLI
$ npx skills add morluto/jacobian --skill verifier-evaluations -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install morluto/jacobian verifier-evaluations --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/morluto/jacobian.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/verifier-evaluations .claude/skills/verifier-evaluations && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
verifier-evaluations
GitHub stars
205
Token cost
~661 tokens
SKILL.md length
314 words
Files
4 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring.

  • Tasks that involve Design review and critique
  • SKILL.md covers Establish the contract, Replay and score and Complete the repair
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Math and symbolic computation

What it does

Verifier Evaluations is an agent skill from morluto/jacobian. Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring.

Its SKILL.md is about 660 tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `agents/openai.yaml`, `references/replay-and-attacks.md` and `references/submission-contracts.md`).

It sits in Research & Science, covering Design review and critique and Math and symbolic computation. It works with Model Context Protocol. The repository describes itself as: Composable mathematics tools for agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Design review and critique
  • Tasks that involve Math and symbolic computation

Example prompts

  • “/verifier-evaluations”

What it can do on your machine

Read from SKILL.md and the folder at commit 9dc2aaf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Verifier Evaluations loads about 661 tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 29 tokens; SKILL.md has 314 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~661
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from morluto/jacobian at commit 9dc2aaf, republished under its MIT licence (© morluto). 314 words, ~661 tokens.

Download SKILL.mdSave it as .claude/skills/verifier-evaluations/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.
name
verifier-evaluations
description
Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring.

Verifier Evaluations

A verifier decides a benchmark's mathematical predicate from frozen input and a bounded submission. It does not grade prose, confidence, tool use, or equality with one preferred solution. Use harbor-benchmarks only when the task also needs Harbor packaging, environment changes, or execution guidance.

Establish the contract

Choose the smallest checkable submission: a typed result, a result with a necessary finite witness, or a supported formal proof. Accept mathematically equivalent representations unless canonicalization is an explicit task outcome. The visible instruction and schema must describe every enforced condition and must not leak the solution or derived conclusions.

For a new or changed submission shape, read submission contracts, including witness selection, schema reductions, and independent result fields. Keep instruction, schema, verifier, gold, public contract, and host tests consistent when changing that shape.

Replay and score

Bound input before parsing; enforce exact types and shapes before computation. Replay mathematics against the frozen verifier copy. Malformed submissions must produce a deterministic false predicate and reward artifact, not a host exception. Input binding, mathematical correctness, and declared witness validity remain separate diagnostics where independently observable; required gates combine in reward. Default reward is binary. Partial credit requires explicitly declared, independent mathematical subclaims.

For implementation, diagnostic binding exceptions, artifact/path handling, or regression tests, read replay and attacks. Preserve alternate valid witnesses and a discriminating wrong mathematical claim alongside malformed-input cases. Natural-language proofs need human review unless formalization or an executable certificate makes them checkable.

Complete the repair

Run focused behavioral attacks, the selected Oracle, and the repository's planned Harbor gate when the verifier or task contract changes. Refresh selected Dockerfile checksums after verifier edits through the task preparation workflow. Shared support changes require the affected task-local copies and Oracles; historical snapshots remain unchanged.

Report the actual command, task digest, Oracle result, and deferred validation. A verifier test, Oracle pass, repository gate, and causal benchmark comparison establish different evidence.

© morluto, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 3 other files (references) in .agents/skills/verifier-evaluations of morluto/jacobian.

  • SKILL.md
  • agents/openai.yaml
  • references/replay-and-attacks.md
  • references/submission-contracts.md

Open the folder on GitHubat commit 9dc2aaf

Compare with similar skills

Verifier Evaluations next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Verifier Evaluations compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Verifier Evaluations this skillmorluto/jacobian205—~661Automated safety check: PassMIT
Azsdk Common Pipeline TroubleshootingAzure/azure-sdk-for-android121—~496Automated safety check: PassMIT
Product Designqf-studio/navigator354—~5.1kAutomated safety check: NotesMIT
Research RefinezjYao36/Auto-Research-Refine1287 repos~6.9kAutomated safety check: NotesNone
Read GitHubAgentTeam-TaichuAI/ScienceClaw6702 repos~638Automated safety check: PassNone
Proof Run Orchestratorwanshuiyin/Auto-claude-code-research-in-sleep17k1 repos~4.7kAutomated safety check: PassMIT

Similar skills

  • Azsdk Common Pipeline Troubleshooting

    Azure/azure-sdk-for-android

    Official

    Diagnose and resolve failures in Azure SDK CI and generation pipelines.

    121 GitHub stars~496 tokensUpdated 4 mo ago
    Agent WorkflowsAuto-check passed
  • Product Design

    qf-studio/navigator

    Automates design review, token extraction, component mapping, and implementation planning.

    354 GitHub stars~5.1k tokensUpdated yesterday
    Agent WorkflowsAuto-check: notes
  • Research Refine

    zjYao36/Auto-Research-Refine

    Turns a vague research direction into a focused, problem-anchored method plan through up to five review rounds with a second model.

    128 GitHub starsUsed in 7 repos~6.9k tokens
    Research & ScienceAuto-check: notes
  • Read GitHub

    AgentTeam-TaichuAI/ScienceClaw

    Read and search GitHub repository documentation via gitmcp.io MCP service.

    670 GitHub starsUsed in 2 repos~638 tokens
    Research & ScienceAuto-check passed
  • Proof Run Orchestrator

    wanshuiyin/Auto-claude-code-research-in-sleep

    Runs a mathematical proof project as a stateful pipeline of run directories: a local attempt first, then a manual GPT Pro handoff package, with an optional DeepSeek audit.

    17k GitHub starsUsed in 1 repo~4.7k tokens
    Research & ScienceAuto-check passed
  • Firstdata

    MLT-OSS/FirstData

    Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.).

    183 GitHub stars~3.1k tokensUpdated 5 days ago
    Research & ScienceAuto-check passed

More from morluto/jacobian

All 11 skills in this repo
  • Evaluate Jacobian reliability using recently resolved conjectures as held-out probes.

    205 GitHub stars~816 tokensUpdated 3 days ago
    Auto-check passed
  • Harbor Benchmarks

    morluto/jacobian

    Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets.

    205 GitHub stars~690 tokensUpdated 3 days ago
    Auto-check passed
  • Design or audit a Jacobian operation’s mathematical contract, boundedness, exact results, and composition.

    205 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Extract reusable Jacobian capabilities from a mathematical solution corpus, rather than one agent trajectory.

    205 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Investigate MCP tool availability, discovery, and selection, including controlled adoption evaluations.

    205 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Review mathematical agent trajectories for evidence-backed Jacobian improvements; do not resume solving.

    205 GitHub stars~848 tokensUpdated 3 days ago
    Auto-check passed

Questions about Verifier Evaluations

What does Verifier Evaluations do?

Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring. Verifier Evaluations is an agent skill from morluto/jacobian. Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring.

When should I use Verifier Evaluations?

Verifier Evaluations fits situations like: tasks that involve Design review and critique; tasks that involve Math and symbolic computation.

How do I install Verifier Evaluations in Claude Code?

Run `npx skills add morluto/jacobian --skill verifier-evaluations -a claude-code`. Or copy the skill folder (.agents/skills/verifier-evaluations in morluto/jacobian) into .claude/skills/verifier-evaluations in your project. Claude Code loads it when a task matches its description.

How do I install Verifier Evaluations in Codex?

Run `npx skills add morluto/jacobian --skill verifier-evaluations -a codex`. Or copy the skill folder (.agents/skills/verifier-evaluations in morluto/jacobian) into .agents/skills/verifier-evaluations in your project. Codex loads it when a task matches its description.

Can I use Verifier Evaluations in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add morluto/jacobian --skill verifier-evaluations -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/verifier-evaluations, .gemini/skills/verifier-evaluations, .github/skills/verifier-evaluations and .opencode/skills/verifier-evaluations in your project.

What does Verifier Evaluations need to run?

SKILL.md names no scripts, command-line tools or credentials: Verifier Evaluations is instructions for the agent only.

Does Verifier Evaluations access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Verifier Evaluations safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Verifier Evaluations use?

Verifier Evaluations is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Verifier Evaluations use?

About 661 tokens (SKILL.md is roughly 2.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.7k tokens, read only when the agent opens those files.

What are the alternatives to Verifier Evaluations?

Skills that share tags, products or a category with Verifier Evaluations: Azsdk Common Pipeline Troubleshooting (Azure/azure-sdk-for-android, 121 stars), Product Design (qf-studio/navigator, 354 stars), Research Refine (zjYao36/Auto-Research-Refine, 128 stars) and Read GitHub (AgentTeam-TaichuAI/ScienceClaw, 670 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Verifier Evaluations?

morluto (a GitHub user) maintains it in morluto/jacobian, which has 205 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 5, 2026.

Source: morluto/jacobian on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.