Agent skill

Harbor Benchmarks

by morluto in morluto/jacobian

Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets.

MITAuto-check passedResearch & Science

Install Harbor Benchmarks

skills CLI
$ npx skills add morluto/jacobian --skill harbor-benchmarks -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install morluto/jacobian harbor-benchmarks --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/morluto/jacobian.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/harbor-benchmarks .claude/skills/harbor-benchmarks && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
harbor-benchmarks
GitHub stars
205
Token cost
~690 tokens
SKILL.md length
288 words
Files
3 (incl. references)
Skills in repo
11
Repo updated
First seen
Licence
MIT

At a glance

Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets.

  • Tasks that involve Math and symbolic computation
  • SKILL.md covers Choose the artifact and Package and validate
  • Calls make

What it does

Harbor Benchmarks is an agent skill from morluto/jacobian. Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets.

Its SKILL.md is about 690 tokens, which your agent loads only when the skill is triggered. The skill folder holds 4 other files, including reference files (for example `agents/openai.yaml` and `references/packaging.md`).

It sits in Research & Science, covering Math and symbolic computation. It works with Model Context Protocol and Python. The repository describes itself as: Composable mathematics tools for agents. The licence is MIT.

When your agent uses it

  • Tasks that involve Math and symbolic computation

Example prompts

  • “/harbor-benchmarks”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9dc2aaf. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • make

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Harbor Benchmarks loads about 690 tokens when it runs, and up to ~1k if it reads all its reference files. Until then it costs about 26 tokens; SKILL.md has 288 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~26
When it runs · the whole SKILL.md, loaded when a task matches
~690
With references · SKILL.md plus every file in references/, read only if the agent opens them
~1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from morluto/jacobian at commit 9dc2aaf, republished under its MIT licence (© morluto). 288 words, ~690 tokens.

Download SKILL.mdSave it as .claude/skills/harbor-benchmarks/SKILL.md (or your agent's skills folder). This skill also uses 2 other files; get the full folder from GitHub.
name
harbor-benchmarks
description
Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets.

Harbor Benchmarks

Build a bounded mathematical experiment with a hidden, replayable verifier. Jacobian availability is an experimental condition, not a prerequisite for task admission. Task correctness depends on the submitted mathematics, not the agent's tool choices or transcript.

Choose the artifact

  • Put meaningful agent-facing mathematical tasks with replayable outcomes in Harbor, including capability gaps and tool-free baselines.
  • Put provider/environment feasibility spikes under tests/fixtures/providers/.
  • Keep public known-answer regressions and human-reviewed proof corpora separate from held-out comparative scoring.

For submission design, mathematical predicates, scoring, or verifier repairs, use verifier-evaluations. Ordinary packaging and execution do not require that skill. A task's instruction and schema must expose its complete protocol while keeping solutions, Oracle fixtures, verifier code, and host caches hidden.

Package and validate

Each task is a direct dataset child with one member record:

text
benchmarks/datasets/<dataset>/<task-id>/
benchmarks/datasets/<dataset>/members/<task-id>.toml

Use frozen offline inputs, pinned images and dependencies, and an explicit execution profile. Verifier Dockerfiles build from tests/; avoid parent COPY, host paths, floating tags, and symlinks. For the affected contract, use benchmark contracts. For local commands and planner selection, use the Harbor section of CONTRIBUTING.md.

sh
make harbor-plan BASE=origin/main
make harbor-prepare-task DATASET=<dataset> TASKS="<task-id>"
make harbor-validate-task DATASET=<dataset> TASKS="<task-id>"

Preparation formats selected task Python and refreshes contracts/checksums; validation runs the selected static, host, and Oracle checks without a model agent. Run the selected verifier's behavioral attacks when its contract changes. Report any deferred affected Oracle coverage.

For task-local support migrations, older contracts, checksum updates, or an intentional snapshot, read packaging details. Do not create a snapshot merely because a task contract changed.

Complete the requested task preparation and validation, repairing failures caused by the change within the authorized scope. Report the command, task digest, selected scope, Oracle result, and remaining evidence gaps. Distinguish local validation from a causal model comparison; model calls and external publication require the applicable user authorization.

© morluto, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 2 other files (references) in .agents/skills/harbor-benchmarks of morluto/jacobian.

  • SKILL.md
  • agents/openai.yaml
  • references/packaging.md

Open the folder on GitHubat commit 9dc2aaf

Compare with similar skills

Harbor Benchmarks next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Harbor Benchmarks compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Harbor Benchmarks this skillmorluto/jacobian205—~690Automated safety check: PassMIT
Setup MedsciAperivue/medsci-skills329—~960Automated safety check: PassMIT
Deep Research MCP Guidepminervini/deep-research-mcp112—~5.8kAutomated safety check: PassMIT
Aer Statspaibrycewang-stanford/Auto-Empirical-Research-Skills4.5k—~3kAutomated safety check: PassCustom licence
SympyzLanqing/codex-claude-academic-skills4.6k16 repos~3.4kAutomated safety check: PassMIT
Edu Analytic Geometrywy51ai/edulab1.4k1 repos~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Setup Medsci

    Aperivue/medsci-skills

    A skill your agent uses when a skill fails for a missing tool or the environment needs checking.

    329 GitHub stars~960 tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Deep Research MCP Guide

    pminervini/deep-research-mcp

    Explains how to run, integrate and debug the deep-research-mcp project through its CLI, Python API or MCP server, with OpenAI, Gemini and DR-Tulu backends.

    112 GitHub stars~5.8k tokensUpdated 10 days ago
    Research & ScienceAuto-check passed
  • Aer Statspai

    brycewang-stanford/Auto-Empirical-Research-Skills

    A skill your agent uses when aer-identification has fixed the design, after methodology choice and before aer-robustness or aer-tables-figures, to run an AER-track analysis with StatsPAI — the…

    4.5k GitHub stars~3k tokensUpdated 3 days ago
    Research & ScienceAuto-check passed
  • Sympy

    zLanqing/codex-claude-academic-skills

    A skill your agent uses when working with symbolic mathematics in Python.

    4.6k GitHub starsUsed in 16 repos~3.4k tokens
    Research & ScienceAuto-check passed
  • 把一道解析几何题解成一个自包含的交互教学网页:左栏题面 + 动态控制台(一个 可变参数滑块驱动实时重算的几何量 + 理论范围/定值指示),中栏 KaTeX 分步解析,右栏 2D Canvas 动态几何画板(椭圆/双曲线/抛物线/圆 + 动直线/动点 + 向量 + 标注 + 画笔涂鸦)。

    1.4k GitHub starsUsed in 1 repo~1.6k tokens
    Research & ScienceAuto-check passed
  • NSFC Literature Review Writer

    HuiyuLi-2000/Chinese-Grant-Writer-Skills

    Writes the research-status literature review and critique section of an NSFC grant proposal, backed by a bundled multi-source literature search.

    432 GitHub starsUsed in 1 repo~1.4k tokens
    Research & ScienceAuto-check: notes

More from morluto/jacobian

All 11 skills in this repo
  • Evaluate Jacobian reliability using recently resolved conjectures as held-out probes.

    205 GitHub stars~816 tokensUpdated 3 days ago
    Auto-check passed
  • Design or audit a Jacobian operation’s mathematical contract, boundedness, exact results, and composition.

    205 GitHub stars~2k tokensUpdated 3 days ago
    Auto-check passed
  • Extract reusable Jacobian capabilities from a mathematical solution corpus, rather than one agent trajectory.

    205 GitHub stars~1.9k tokensUpdated 3 days ago
    Auto-check passed
  • Investigate MCP tool availability, discovery, and selection, including controlled adoption evaluations.

    205 GitHub stars~1k tokensUpdated 3 days ago
    Auto-check passed
  • Review mathematical agent trajectories for evidence-backed Jacobian improvements; do not resume solving.

    205 GitHub stars~848 tokensUpdated 3 days ago
    Auto-check passed
  • Verifier Evaluations

    morluto/jacobian

    Design, audit, or repair mathematical benchmark verifiers, submission contracts, and scoring.

    205 GitHub stars~661 tokensUpdated 3 days ago
    Auto-check passed

Questions about Harbor Benchmarks

What does Harbor Benchmarks do?

Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets. Harbor Benchmarks is an agent skill from morluto/jacobian. Author, package, validate, or run mathematical evaluations as Jacobian Harbor datasets.

When should I use Harbor Benchmarks?

Harbor Benchmarks fits situations like: tasks that involve Math and symbolic computation.

How do I install Harbor Benchmarks in Claude Code?

Run `npx skills add morluto/jacobian --skill harbor-benchmarks -a claude-code`. Or copy the skill folder (.agents/skills/harbor-benchmarks in morluto/jacobian) into .claude/skills/harbor-benchmarks in your project. Claude Code loads it when a task matches its description.

How do I install Harbor Benchmarks in Codex?

Run `npx skills add morluto/jacobian --skill harbor-benchmarks -a codex`. Or copy the skill folder (.agents/skills/harbor-benchmarks in morluto/jacobian) into .agents/skills/harbor-benchmarks in your project. Codex loads it when a task matches its description.

Can I use Harbor Benchmarks in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add morluto/jacobian --skill harbor-benchmarks -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/harbor-benchmarks, .gemini/skills/harbor-benchmarks, .github/skills/harbor-benchmarks and .opencode/skills/harbor-benchmarks in your project.

What does Harbor Benchmarks need to run?

Going by SKILL.md and its folder, Harbor Benchmarks needs the command-line tools its instructions call (make). Our summary lists: Python 3.

Does Harbor Benchmarks access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Harbor Benchmarks safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Harbor Benchmarks use?

Harbor Benchmarks is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Harbor Benchmarks use?

About 690 tokens (SKILL.md is roughly 2.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 343 tokens, read only when the agent opens those files.

What are the alternatives to Harbor Benchmarks?

Skills that share tags, products or a category with Harbor Benchmarks: Setup Medsci (Aperivue/medsci-skills, 329 stars), Deep Research MCP Guide (pminervini/deep-research-mcp, 112 stars), Aer Statspai (brycewang-stanford/Auto-Empirical-Research-Skills, 4.5k stars) and Sympy (zLanqing/codex-claude-academic-skills, 4.6k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Harbor Benchmarks?

morluto (a GitHub user) maintains it in morluto/jacobian, which has 205 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 5, 2026.

Source: morluto/jacobian on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.