Agent skill

Run Eval Harbor

by lobehub in lobehub/lobehub

Runs existing Harbor evaluation jobs against a local LobeHub build or LobeHub Cloud, with preflight checks, resume support, and failure triage.

Custom licenceAuto-check: notesAgent Workflows

Install Run Eval Harbor

skills CLI
$ npx skills add lobehub/lobehub --skill run-eval-harbor -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install lobehub/lobehub run-eval-harbor --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/lobehub/lobehub.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/run-eval-harbor .claude/skills/run-eval-harbor && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run-eval-harbor
GitHub stars
83k
Token cost
~1k tokens
SKILL.md length
510 words
Files
20 (incl. scripts, references)
Skills in repo
50
Repo updated
First seen
Licence
Custom licence

At a glance

Runs existing Harbor evaluation jobs against a local LobeHub build or LobeHub Cloud, with preflight checks, resume support, and failure triage.

  • Works in 6 steps: Target: local production server from… → CLI: checkout build from apps/cli, or… → Exact LH_AGENT_ID; the selected agent… → …
  • Running an existing Harbor evaluation job against local LobeHub
  • SKILL.md covers Ask First, Guardrails, Run and Diagnose, plus 1 more section
  • Runs Shell and Python scripts from its folder; calls bun

What it does

Before preparing or running anything, this skill asks for six independent choices and never infers them: local or cloud target, a checkout or published CLI build, the exact agent ID, the eval repository path, whether to start a new job or resume one, and the right credentials for the chosen target.

In local mode it starts LobeHub itself on a fixed port while leaving infrastructure to Compose, and requires explicit confirmation that local ports are unreachable from untrusted networks before it will bootstrap the stack; in cloud mode it never starts local infrastructure or rewrites server addresses. Preflight checks are read-only and specific to the chosen target, and the skill explicitly excludes authoring Harbor tasks or product acceptance, which belong to other tools.

When your agent uses it

  • Running an existing Harbor evaluation job against local LobeHub
  • Diagnosing a failed or stalled Harbor job
  • Resuming an interrupted evaluation run

Example prompts

  • “Run this Harbor eval job against my local LobeHub build.”
  • “Resume the eval job that failed partway through.”
  • “Check why the cloud Harbor run isn't connecting to the gateway.”

Requirements

  • A configured LH_AGENT_ID
  • Provider credentials for the selected agent
  • Docker Compose for local mode

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Target: local production server from this checkout, or cloud/remote.
  2. CLI: checkout build from apps/cli, or published npm release.
  3. Exact LH_AGENT_ID; the selected agent already owns its model.
  4. Eval repository path and whether the user wants a new job or a resume.
  5. Credentials: for cloud, require its CLI API key in the eval repository's
  6. For local, obtain explicit confirmation that port 3210 and every configured

What it can do on your machine

Read from SKILL.md and the folder at commit 35d442e. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 12 files in scripts/ (Shell and Python, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • bun

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run Eval Harbor loads about 1k tokens when it runs, and up to ~2.4k if it reads all its reference files. Until then it costs about 72 tokens; SKILL.md has 510 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~72
When it runs · the whole SKILL.md, loaded when a task matches
~1k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~2.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteMentions a .env fileSKILL.md:21
    ignored `.env`. For local, ask whether the selected agent's provider
  • NoteMentions a .env fileSKILL.md:38
    - Create `docker-compose/eval/.env` from `.env.example` only when absent; never

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

Its licence (Custom licence) doesn't allow us to republish the file, so here is its outline and opening line. It has 510 words (~1,029 tokens).

“Run existing Harbor evaluations against either this checkout's isolated local production harness or a remote LobeHub target. Use create-task to author or grade tasks and acceptance for product acceptance.”

— opening of SKILL.md by lobehub, Custom licence
name
run-eval-harbor

Read the full SKILL.md on GitHub

Files

SKILL.md and 19 other files (scripts, references) in .agents/skills/run-eval-harbor of lobehub/lobehub.

  • SKILL.md
  • agents/openai.yaml
  • references/cloud.md
  • references/local.md
  • scripts/.gitignore
  • scripts/bootstrap.sh
  • scripts/lh/agent.py
  • scripts/lh/template/check-lh.sh.j2
  • scripts/lh/template/connect-lh.sh.j2
  • scripts/lh/template/install-lh.sh.j2
  • scripts/lh/template/run-agent.sh.j2
  • scripts/lh/template/supervisord.conf.j2
  • scripts/preflight.sh
  • scripts/run-smoke.sh
  • scripts/server.sh
  • scripts/smoke
  • … and 4 more

Open the folder on GitHubat commit 35d442e

Compare with similar skills

Run Eval Harbor next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run Eval Harbor compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run Eval Harbor this skilllobehub/lobehub83k—~1kAutomated safety check: NotesCustom licence
Diagnosing Superpowers Sessionsobra/superpowers297k3 repos~1.7kAutomated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph74k—~950Automated safety check: PassMIT
Skill Compliance Checkeraffaan-m/ECC276k1 repos~623Automated safety check: PassMIT
Waza Skill Evaluatormicrosoft/waza1.4k—~2kAutomated safety check: PassMIT
Waza Interactivemicrosoft/waza1.4k—~1.3kAutomated safety check: PassMIT

Similar skills

  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    297k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    74k GitHub stars~950 tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Measures whether agents actually follow a skill, rule or agent definition by generating scenarios at three strictness levels and scoring tool-call traces.

    276k GitHub starsUsed in 1 repo~623 tokens
    Agent WorkflowsAuto-check passed
  • Waza Skill Evaluator

    microsoft/waza

    Official

    Evaluates agent skills with a Go CLI that runs YAML-defined benchmarks, compares runs and scores the quality of SKILL.md frontmatter.

    1.4k GitHub stars~2k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Waza Interactive

    microsoft/waza

    Official

    Walks you through creating, running and reading waza evals for an agent skill, then proposes concrete fixes when tasks fail or the score is low.

    1.4k GitHub stars~1.3k tokensUpdated 2 days ago
    Agent WorkflowsAuto-check passed
  • Deep Researcher Maintain CI

    NVIDIA-AI-Blueprints/deep-researcher-agent

    A skill your agent uses when changing Deep Researcher Agent continuous integration, pre-commit, or contributor governance — editing .github/workflows/ (ci, ui, skills-eval, request-nvskills-ci)…

    885 GitHub stars~1.5k tokensUpdated today
    Agent WorkflowsAuto-check: notes

More from lobehub/lobehub

All 50 skills in this repo
  • Builds single-file interactive HTML prototypes rendered with the real LobeHub UI components and written as production-style React, so they can later be split into files.

    83k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Auto-check passed
  • Git Worktree Cleanup

    lobehub/lobehub

    Audits stale Git worktrees and branches with a bundled script, classifies each one, and deletes only after you approve the exact candidates.

    83k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Maintains LobeHub's model-backed alint rule set: writing rules, removing false positives against real code, deciding warn versus error and tracking token cost.

    83k GitHub stars~1.9k tokensUpdated today
    Auto-check passed
  • Guides building LobeHub builtin agent tools, from the manifest and execution runtime to executors, chat UI renders and registry wiring.

    83k GitHub stars~2.4k tokensUpdated today
    Auto-check passed
  • Explains how LobeHub client code fetches data through services, SWR store hooks and cache keys, and when to avoid useEffect fetching or duplicated state.

    83k GitHub stars~1.7k tokensUpdated today
    Auto-check passed

Questions about Run Eval Harbor

What does Run Eval Harbor do?

Runs existing Harbor evaluation jobs against a local LobeHub build or LobeHub Cloud, with preflight checks, resume support, and failure triage. Before preparing or running anything, this skill asks for six independent choices and never infers them: local or cloud target, a checkout or published CLI build, the exact agent ID, the eval repository path, whether to start a new job or resume one, and the right credentials for the chosen target.

When should I use Run Eval Harbor?

Run Eval Harbor fits situations like: running an existing Harbor evaluation job against local LobeHub; diagnosing a failed or stalled Harbor job; resuming an interrupted evaluation run.

How do I install Run Eval Harbor in Claude Code?

Run `npx skills add lobehub/lobehub --skill run-eval-harbor -a claude-code`. Or copy the skill folder (.agents/skills/run-eval-harbor in lobehub/lobehub) into .claude/skills/run-eval-harbor in your project. Claude Code loads it when a task matches its description.

How do I install Run Eval Harbor in Codex?

Run `npx skills add lobehub/lobehub --skill run-eval-harbor -a codex`. Or copy the skill folder (.agents/skills/run-eval-harbor in lobehub/lobehub) into .agents/skills/run-eval-harbor in your project. Codex loads it when a task matches its description.

Can I use Run Eval Harbor in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add lobehub/lobehub --skill run-eval-harbor -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run-eval-harbor, .gemini/skills/run-eval-harbor, .github/skills/run-eval-harbor and .opencode/skills/run-eval-harbor in your project.

What does Run Eval Harbor need to run?

Going by SKILL.md and its folder, Run Eval Harbor needs a shell and Python for the scripts in its folder and the command-line tools its instructions call (bun). Our summary lists: A configured LH_AGENT_ID; Provider credentials for the selected agent; Docker Compose for local mode.

Does Run Eval Harbor access the network?

SKILL.md names 1 domain. As links in the text: github.com. This is read from the text; nothing was executed.

Is Run Eval Harbor safe to install?

Our automated static check of SKILL.md found notes only (mentions a .env file), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Run Eval Harbor use?

Run Eval Harbor has a licence file (the repository's licence) that doesn't match a standard licence. Read it on GitHub before reusing the skill.

How many tokens does Run Eval Harbor use?

About 1k tokens (SKILL.md is roughly 4.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 1.3k tokens, read only when the agent opens those files.

What are the alternatives to Run Eval Harbor?

Skills that share tags, products or a category with Run Eval Harbor: Diagnosing Superpowers Sessions (obra/superpowers, 297k stars), CodeGraph Agent Eval (colbymchenry/codegraph, 74k stars), Skill Compliance Checker (affaan-m/ECC, 276k stars) and Waza Skill Evaluator (microsoft/waza, 1.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run Eval Harbor?

lobehub (a GitHub organization) maintains it in lobehub/lobehub, which has 83,074 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: lobehub/lobehub on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.