Agent skill

Mh Eval

by HangYu8123 in HangYu8123/mini-harness

mini-harness · A/B benchmark of the harness on this repo — two disposable worktrees (harness on · off), the same questions run in each as headless Claude Code sessions, token/time measured with…

No licenceAuto-check passedDevelopment

Install Mh Eval

skills CLI
$ npx skills add HangYu8123/mini-harness --skill mh-eval -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install HangYu8123/mini-harness mh-eval --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/HangYu8123/mini-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/mh-eval .claude/skills/mh-eval && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
mh-eval
GitHub stars
199
Token cost
~1.3k tokens
SKILL.md length
693 words
Files
3
Skills in repo
8
Repo updated
First seen
Licence
None found

At a glance

mini-harness · A/B benchmark of the harness on this repo — two disposable worktrees (harness on · off), the same questions run in each as headless Claude Code sessions, token/time measured with…

  • Works in 6 steps: Questions. With no --questions, the… → setup — creates the arms from the… → run — one headless session per arm per… → …
  • Tasks that involve Git worktrees
  • SKILL.md covers Procedure and Rules
  • Runs Python scripts from its folder; calls claude and git

What it does

Mh Eval is an agent skill from HangYu8123/mini-harness. mini-harness · A/B benchmark of the harness on this repo — two disposable worktrees (harness on · off), the same questions run in each as headless Claude Code sessions, token/time measured with mh.sh usage, the release criterion usage --compare (ON ≤ 1.10 × OFF), the 2N reports graded head-to-head, and a report with improvement suggestions. Runs only when invoked (/mh-eval · /mini-harness eval).

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `mh_eval.py` and `questions.example.md`).

It sits in Development, covering Git worktrees. The repository describes itself as: A local minimized harness, with simplified output and logs for self-evolve.

When your agent uses it

  • Tasks that involve Git worktrees

Example prompts

  • “/mh-eval”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Questions. With no --questions, the helper copies questions.example.md; better: write /QUESTIONS.md first, ## Q — per question, 6–10…
  2. setup — creates the arms from the current working tree (uncommitted changes and untracked files included), repairs symlink stubs on…
  3. run — one headless session per arm per question (claude -p … --output-format json --dangerously-skip-permissions, cwd = the arm, the arms…
  4. measure — mh.sh usage --json --all per arm, then usage --compare (the pack's release criterion), METRICS.md, and copies of the ON arm's…
  5. report — writes REPORT.md with the metrics, every pair's ## Run report, and six empty headings. Fill them yourself from traj/*.md…
  6. Verdict. State the compare verdict, the quality score (pairs won by each arm), and whether the harness should ship as is, ship with the…

What it can do on your machine

Read from SKILL.md and the folder at commit baa821f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • claude
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Mh Eval loads about 1.3k tokens when it runs. Until then it costs about 103 tokens; SKILL.md has 693 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~103
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

Without a licence we can't republish the file, so here is its outline and opening line. It has 693 words (~1,325 tokens).

“Helper: py "${CLAUDE_SKILL_DIR}/mh_eval.py" [options] (Codex: python3; the skill folder is .agents/skills/mh-eval standalone). Every artifact lands under /analysis/eval//: QUESTIONS.md, RUN.json, traj/q_{on,off}.{json,md}, usage_{on,off}.json, COMPARE.txt, METRICS.md, harness_traj/ (the ON arm's sealed records), repo_info_on/, REPORT.md. The two arms are git worktrees beside the repo…”

— opening of SKILL.md by HangYu8123
name
mh-eval
argument-hint
[all | setup | run | measure | report | clean | status] [--questions <file>] [--model <alias>] [--effort <level>] [--chat continue|fresh|split] [--only Q1,Q2]…

Read the full SKILL.md on GitHub

Files

SKILL.md and 2 other files in skills/mh-eval of HangYu8123/mini-harness.

  • SKILL.md
  • mh_eval.py
  • questions.example.md

Open the folder on GitHubat commit baa821f

Compare with similar skills

Mh Eval next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Mh Eval compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Mh Eval this skillHangYu8123/mini-harness199—~1.3kAutomated safety check: PassNone
Finishing a Development Branchobra/superpowers297k5 repos~1.9kAutomated safety check: PassMIT
Migrate Core Code to Submodulestinyhumansai/openhuman42k—~2.6kAutomated safety check: PassGPL-3.0
Finishing A Development Branchfarm-fe/farm5.6k34 repos~1.8kAutomated safety check: PassMIT
Git Worktree Cleanuplobehub/lobehub83k—~2.8kAutomated safety check: PassCustom licence
Keep Codex Fastvibeforge1111/keep-codex-fast1.6k—~3.1kAutomated safety check: PassMIT

Similar skills

  • Walks the last step of a branch: confirm tests pass, detect the git environment, ask how to integrate, carry out your choice and clean up the worktree.

    297k GitHub starsUsed in 5 repos~1.9k tokens
    DevelopmentAuto-check passed
  • Migrate Core Code to Submodules

    tinyhumansai/openhuman

    Plans and carries out moving non-host-specific code and its tests from the OpenHuman core into vendored tiny submodule libraries, then releases the submodule and re-pins the host.

    42k GitHub stars~2.6k tokensUpdated today
    DevelopmentAuto-check passed
  • A skill your agent uses when implementation is complete, all tests pass, and you need to decide how to integrate the work - guides completion of development work by presenting structured options for…

    5.6k GitHub starsUsed in 34 repos~1.8k tokens
    DevelopmentAuto-check passed
  • Git Worktree Cleanup

    lobehub/lobehub

    Audits stale Git worktrees and branches with a bundled script, classifies each one, and deletes only after you approve the exact candidates.

    83k GitHub stars~2.8k tokensUpdated today
    DevelopmentAuto-check passed
  • Keep Codex Fast

    vibeforge1111/keep-codex-fast

    A skill your agent uses when Codex feels slow or bloated, when local sessions/logs/worktrees/config have grown over time, or when a user wants safe maintenance for Codex Desktop/CLI state.

    1.6k GitHub stars~3.1k tokensUpdated 5 mo ago
    DevelopmentAuto-check passed
  • Pre-Release PR Triage

    jamiepine/voicebox

    Sorts a backlog of open pull requests into must-merge, candidate, superseded and deferred, writes a triage doc and works the merge loop before a release.

    57k GitHub stars~3.1k tokensUpdated 3 days ago
    DevelopmentAuto-check passed

More from HangYu8123/mini-harness

All 8 skills in this repo
  • Breakdown PR

    HangYu8123/mini-harness

    Plan and optionally execute a large pull request split into small stacked PRs.

    199 GitHub stars~2.8k tokensUpdated 17 days ago
    Auto-check passed
  • Code Review And Quality

    HangYu8123/mini-harness

    Multi-axis, review-only code review of a diff across six axes — request achievement, correctness, readability/simplicity, architecture, security, and performance — with severity-labelled findings.

    199 GitHub stars~1.8k tokensUpdated 17 days ago
    Auto-check passed
  • Code Simplification

    HangYu8123/mini-harness

    Simplify code for clarity while preserving exact behavior — reduce nesting, split long functions, remove redundancy and dead code, and fix unclear names.

    199 GitHub stars~1.2k tokensUpdated 17 days ago
    Auto-check passed
  • Mh Init

    HangYu8123/mini-harness

    mini-harness · initialize or re-initialize a repo's memory under .harness/repoinfo/ with a multi-agent pass — three-perspective codebase overview, ranked scripts map, issue scan, git-seeded update…

    199 GitHub stars~2.6k tokensUpdated 17 days ago
    Auto-check passed
  • Mh Loop

    HangYu8123/mini-harness

    mini-harness · repeat a pass toward a verifiable goal until a check passes or a safety stop fires — spec → validate → confirm → loop (pass · observe · exit check · ledger) → review → record.

    199 GitHub stars~2.8k tokensUpdated 17 days ago
    Auto-check passed
  • Mh Wiki

    HangYu8123/mini-harness

    mini-harness · consolidate execution trajectories (.harness/exectraj/) into repoinfo/harnesseffect.md — which parts of the harness helped or hurt, root causes, and one proposed change at a time.

    199 GitHub stars~1.6k tokensUpdated 17 days ago
    Auto-check passed

Categories

Questions about Mh Eval

What does Mh Eval do?

mini-harness · A/B benchmark of the harness on this repo — two disposable worktrees (harness on · off), the same questions run in each as headless Claude Code sessions, token/time measured with…. Mh Eval is an agent skill from HangYu8123/mini-harness.10 × OFF), the 2N reports graded head-to-head, and a report with improvement suggestions.

When should I use Mh Eval?

Mh Eval fits situations like: tasks that involve Git worktrees.

How do I install Mh Eval in Claude Code?

Run `npx skills add HangYu8123/mini-harness --skill mh-eval -a claude-code`. Or copy the skill folder (skills/mh-eval in HangYu8123/mini-harness) into .claude/skills/mh-eval in your project. Claude Code loads it when a task matches its description.

How do I install Mh Eval in Codex?

Run `npx skills add HangYu8123/mini-harness --skill mh-eval -a codex`. Or copy the skill folder (skills/mh-eval in HangYu8123/mini-harness) into .agents/skills/mh-eval in your project. Codex loads it when a task matches its description.

Can I use Mh Eval in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add HangYu8123/mini-harness --skill mh-eval -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/mh-eval, .gemini/skills/mh-eval, .github/skills/mh-eval and .opencode/skills/mh-eval in your project.

What does Mh Eval need to run?

Going by SKILL.md and its folder, Mh Eval needs Python for the scripts in its folder and the command-line tools its instructions call (claude and git). Our summary lists: Python 3.

Does Mh Eval access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Mh Eval safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Mh Eval use?

No licence was found for Mh Eval or its repository. Without one, default copyright applies: ask the author before reusing or redistributing it.

How many tokens does Mh Eval use?

About 1.3k tokens (SKILL.md is roughly 5.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Mh Eval?

Skills that share tags, products or a category with Mh Eval: Finishing a Development Branch (obra/superpowers, 297k stars), Migrate Core Code to Submodules (tinyhumansai/openhuman, 42k stars), Finishing A Development Branch (farm-fe/farm, 5.6k stars) and Git Worktree Cleanup (lobehub/lobehub, 83k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Mh Eval?

HangYu8123 (a GitHub user) maintains it in HangYu8123/mini-harness, which has 199 GitHub stars. The repository holds 8 skills in this directory. The repository was last updated on September 22, 2026.

Source: HangYu8123/mini-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.