Agent skill

Omh Agent Evaluation

by rlaope in rlaope/oh-my-hermes

[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

MITAuto-check passedAgent Workflows

Install Omh Agent Evaluation

skills CLI
$ npx skills add rlaope/oh-my-hermes --skill omh-agent-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install rlaope/oh-my-hermes omh-agent-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/rlaope/oh-my-hermes.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/omh-agent-evaluation .claude/skills/omh-agent-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
omh-agent-evaluation
GitHub stars
3.2k
Token cost
~2.1k tokens
SKILL.md length
1,061 words
Files
2 (incl. references)
Skills in repo
143
Repo updated
First seen
Licence
MIT

At a glance

[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

  • The user says: agent-evaluation
  • SKILL.md covers Why This Exists, Do Not Use When, Examples and Completion Checklist, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Agent evaluation

What it does

Omh Agent Evaluation is an agent skill from rlaope/oh-my-hermes. [omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor benchmark, compare agents, compare codex claude.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including reference files (for example `references/self-evaluation-loops.md`).

It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: All in one plugin for Hermes Agent ⚚ the coding intelligence, a long-term memory system and model optimized workflow packages. The licence is MIT.

When your agent uses it

  • The user says: agent-evaluation
  • Agent evaluation
  • Agent benchmark
  • Executor evaluation

Example prompts

  • “/omh-agent-evaluation”

What it can do on your machine

Read from SKILL.md and the folder at commit 41de9dc. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Omh Agent Evaluation loads about 2.1k tokens when it runs, and up to ~3.7k if it reads all its reference files. Until then it costs about 87 tokens; SKILL.md has 1,061 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~87
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from rlaope/oh-my-hermes at commit 41de9dc, republished under its MIT licence (© rlaope). 1,061 words, ~2,117 tokens.

Download SKILL.mdSave it as .claude/skills/omh-agent-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
omh-agent-evaluation
description
[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Use when the user says: agent-evaluation, agent evaluation, agent eval, agent benchmark, executor evaluation, executor benchmark, compare agents, compare codex claude.

Agent Evaluation

This is a Hermes-native agent-evaluation workflow skill.

Why This Exists

agent-evaluation gives OMH a way to improve executor choice empirically, not by vibes, while preserving executor-neutral product language across Codex, Claude Code, Hermes, and generic runtimes.

Do Not Use When

  • The user needs current runtime readiness only; use executor-runtime-readiness.
  • The user already selected an executor and wants implementation; use the coding handoff or delivery workflow.
  • The user asks for workflow learning from a single failed route; use workflow-learning.
  • The ask is to find and fix runtime, memory, cost, or rendering hotspots rather than score executor or model output quality; use ultraperf.

Examples

Good example:

  • Prompt: agent-evaluation Codex와 Claude Code를 같은 버그 수정 태스크로 비교해서 어떤 런타임을 기본으로 둘지 판단해줘.
  • Expected behavior: Prepare paired_run_decision/v1 requirements and a scenario-specific recommendation.
  • Why: The request compares executor choices and needs fair evaluation boundaries.

Bad example:

  • Prompt: agent-evaluation 실행 증거 없이 Codex가 항상 최고라고 결론내줘.
  • Expected behavior: Reject universal ranking and require observed runs or mark the recommendation as ungrounded.
  • Why: Agent evaluation must be reproducible and evidence-backed.

Completion Checklist

  • Confirm the workflow target, evidence boundary, and stop condition are named.
  • Report which outputs are prepared, observed, blocked, or missing.
  • Name the smallest next verification or handoff instead of claiming completion from narration.

Recovery Notes

  • If required context is missing, ask one blocking question or route back to the narrower workflow.
  • If runtime or wrapper evidence is unavailable, keep the status as not_observed and expose the next observable action.

Workflow Lane

  • Current lane: Automation and status (achievements, workspace-audit, production-audit, live-incident-response, automation-blueprint, github-event-ops, github-issue-intake, buzz, +39 more) - schedules, status, health, and ops review.
  • If intent belongs to another lane, hand back to oh-my-hermes or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: omh-routing/references/skill-common-rail.md.

Use When

Use when Hermes should design or summarize a fair comparison of Codex, Claude Code, Hermes coding, or generic executors for a bounded task set.

Strong routing signals: `agent-evaluation`, `agent evaluation`, `agent eval`, `agent benchmark`, `executor evaluation`, `executor benchmark`, `compare agents`, `compare codex claude`, `agent tournament`, `which agent is better`, `에이전트 평가`, `에이전트 비교`, `실행자 평가`, `코덱스 클로드 비교`

Catalog Metadata

Category: operations Phase: agent-evaluation Hermes role: operator Quality tier: agent-eval-gated Reasoning demand: light

Quality bar:

  • Define tasks, rubric, isolation, budgets, and stop rules before comparing agents.
  • Use the same inputs and success criteria across candidates unless the difference is the variable under test.
  • Require receipt-authenticated observed_at provenance before public parse or validation can return pass or fail.
  • Report quality, correctness, time, cost, tool coverage, verification, and review gaps separately.
  • Score the trajectory as its own dimension beside the outcome and never fold the two into one total: it asks whether the run searched before it asserted, took approval before an irreversible side effect, used an available tool instead of guessing, and verified before claiming done. A run that reached the right answer with one of those skipped or out of order scores lower on trajectory than one that did not, and the report carries both dimensions.
  • When the question is an agent judging and improving its own output rather than comparing executors, load omh-agent-evaluation/references/self-evaluation-loops.md and pick the loop shape from it - reflection, evaluator-optimizer, or test-driven refinement - remembering that an executable check outranks a judge whenever one exists.
  • Declare all three stop rules before the loop runs - a maximum iteration count, a score threshold chosen in advance, and a no-improvement break - and report the iteration count, the final score, and which of the three ended the run. A loop whose only stop is that the output looks good now is a defect.
  • Write criteria before generation and score a rubric dimension by dimension beside its total: criteria derived from an output describe it instead of testing it, and a single number hides which dimension failed.
  • Recommend executor choice per scenario and confidence, not as a universal ranking.

Handoff policy:

Keep evaluation design and scoring in Hermes. Actual executor runs, costs, timings, tool calls, code edits, and review results must come from observed runtime or supplied artifacts.

Show full SKILL.md (392 more words)Show less

Required inputs:

  • candidate executors or agents
  • task set and fixtures
  • success criteria and scoring rubric
  • allowed tools, budget, timebox, and isolation policy
  • observed run artifacts when comparing completed attempts

Expected outputs:

  • paired_run_decision/v1
  • not-evidence boundary

Artifact expectations:

  • paired_run_decision/v1 with per-task input digests, explicit criteria, baseline and variant exposure, attempted-run and per-dispatch time budgets, signed observed_at receipt provenance, and a scoped Pareto outcome

Safety rules:

  • Do not claim an executor is better from anecdotes, brand names, or unobserved runs.
  • Do not send secrets, credentials, private data, or production tasks into evaluation without explicit authority.
  • Keep benchmark design, observed run evidence, scoring, and executor selection separate.
  • A judge score is never correctness: it licenses no claim that the output is right, tested, reviewed, or shippable, and a model scoring its own output is the weakest evidence class - labelled as such, never reported as verification.
  • A judge with no agreement measured against human labels is unqualified: report it as unqualified, never as a result, and take the qualification procedure and its agreement thresholds from omh-agent-evaluation/references/self-evaluation-loops.md.
  • A signed local Hermes-child receipt proves that OMH recorded a process-sealed confirmed local dispatch event; it does not prove executor internals or protect evidence from the owning OS user.

Runtime Evidence

Preferred harness for this skill: agent-evaluation.

sh
omh runtime record --skill agent-evaluation --harness agent-evaluation --status started

Record observed delegation results; otherwise return not_available or not_observed. Prepared OMH routing is not execution, review, CI, merge-readiness, or merge evidence.

  • Treat wrapper memory/context summaries as advisory local context, not proof of opaque Hermes memory reads or changes. Preserve workflow intent and stop conditions; verify before claiming completion. Reply in the user's own words and the host's own voice: its SOUL.md persona owns reply language, tone, speech level, and sentence endings, progress updates included (where it sets no language, use the one the user wrote in), and OMH shapes structure and content only; OMH's record terms (surface, lane, wrapper, handoff, evidence boundary, not_observed) stay in records and tool calls, never in the sentence the user reads unless they ask about one; and when a stop condition or a decision the user owns ends the turn, offer the next action as a question rather than declaring what will not be done.

Use Hermes-native subagent/delegation features when available: native subagents -> Hermes delegation when available, otherwise sequential lanes.

Shared product, compatibility, topology, memory, harness, and execution rules: omh-routing/references/skill-common-rail.md. Load it when applicable; otherwise name an unavailable capability.

© rlaope, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (references) in skills/omh-agent-evaluation of rlaope/oh-my-hermes.

  • SKILL.md
  • references/self-evaluation-loops.md

Open the folder on GitHubat commit 41de9dc

Compare with similar skills

Omh Agent Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Omh Agent Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Omh Agent Evaluation this skillrlaope/oh-my-hermes3.2k—~2.1kAutomated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers297k3 repos~1.7kAutomated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch66k—~1kAutomated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph74k—~950Automated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    297k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    74k GitHub stars~950 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from rlaope/oh-my-hermes

All 143 skills in this repo
  • Omh Accessibility Audit

    rlaope/oh-my-hermes

    [omh] Screen-reader or keyboard accessibility gaps: prepare WCAG, keyboard, focus, screen-reader, target-size, and reflow evidence gates for UI surfaces.

    3.2k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Omh Agent Instructions

    rlaope/oh-my-hermes

    [omh] Agent instruction file for a repo -- AGENTS.md, CLAUDE.md, a Cursor rule: write or update what an agent cannot derive from the code, inside a marked region, with every command verified or…

    3.2k GitHub stars~2.2k tokensUpdated today
    Auto-check passed
  • Omh Agent Ops Review

    rlaope/oh-my-hermes

    [omh] AI agent progress for managers: help managers inspect AI-agent progress, blockers, quality gates, and throughput levers.

    3.2k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Omh AI Slop Cleaner

    rlaope/oh-my-hermes

    [omh] Messy or AI-generated code to clean up: delete AI-generated slop, dead code, and duplication while observable behavior stays identical.

    3.2k GitHub stars~2.7k tokensUpdated today
    Auto-check passed
  • Omh App Debugging

    rlaope/oh-my-hermes

    [omh] Application code misbehaves -- a wrong value, a flaky test, a lost update: reproduce it first, form competing hypotheses, discriminate them with the cheapest observation, and only then fix the…

    3.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed
  • Omh Apple Design

    rlaope/oh-my-hermes

    [omh] Designing or reviewing an iOS, macOS, or Apple-style UI: prepare native Apple UI or Apple marketing product-visual direction, review, and improvement briefs with evidence-backed remediation…

    3.2k GitHub stars~2.3k tokensUpdated today
    Auto-check passed

Categories

Questions about Omh Agent Evaluation

What does Omh Agent Evaluation do?

[omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics. Omh Agent Evaluation is an agent skill from rlaope/oh-my-hermes. [omh] Choosing between coding agents on evidence: compare executor or agent choices on reproducible tasks using quality, cost, time, tool, and evidence metrics.

When should I use Omh Agent Evaluation?

Omh Agent Evaluation fits situations like: the user says: agent-evaluation; agent evaluation; agent benchmark; executor evaluation.

How do I install Omh Agent Evaluation in Claude Code?

Run `npx skills add rlaope/oh-my-hermes --skill omh-agent-evaluation -a claude-code`. Or copy the skill folder (skills/omh-agent-evaluation in rlaope/oh-my-hermes) into .claude/skills/omh-agent-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Omh Agent Evaluation in Codex?

Run `npx skills add rlaope/oh-my-hermes --skill omh-agent-evaluation -a codex`. Or copy the skill folder (skills/omh-agent-evaluation in rlaope/oh-my-hermes) into .agents/skills/omh-agent-evaluation in your project. Codex loads it when a task matches its description.

Can I use Omh Agent Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add rlaope/oh-my-hermes --skill omh-agent-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/omh-agent-evaluation, .gemini/skills/omh-agent-evaluation, .github/skills/omh-agent-evaluation and .opencode/skills/omh-agent-evaluation in your project.

What does Omh Agent Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Omh Agent Evaluation is instructions for the agent only.

Does Omh Agent Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Omh Agent Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Omh Agent Evaluation use?

Omh Agent Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Omh Agent Evaluation use?

About 2.1k tokens (SKILL.md is roughly 8.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.6k tokens, read only when the agent opens those files.

What are the alternatives to Omh Agent Evaluation?

Skills that share tags, products or a category with Omh Agent Evaluation: MCP Server Builder (anthropics/skills, 180k stars), Diagnosing Superpowers Sessions (obra/superpowers, 297k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Skill Release Gate (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Omh Agent Evaluation?

rlaope (a GitHub user) maintains it in rlaope/oh-my-hermes, which has 3,233 GitHub stars. The repository holds 143 skills in this directory. The repository was last updated on October 8, 2026.

Source: rlaope/oh-my-hermes on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.