OpenHarness End-to-End Evals
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
A skill your agent uses when a nontrivial change needs end-to-end verification before committing or shipping
$ npx skills add nyldn/claude-octopus --skill skill-verify -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install nyldn/claude-octopus skill-verify --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/nyldn/claude-octopus.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-verify .claude/skills/skill-verify && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "skill-verify" agent skill from https://github.com/nyldn/claude-octopus/tree/main/skills/skill-verify into .claude/skills/skill-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-verify", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/nyldn/claude-octopus/tree/main/skills/skill-verifyType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add nyldn/claude-octopus --skill skill-verify -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install nyldn/claude-octopus skill-verify --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nyldn/claude-octopus.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/skill-verify .agents/skills/skill-verify && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "skill-verify" agent skill from https://github.com/nyldn/claude-octopus/tree/main/skills/skill-verify into .agents/skills/skill-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-verify", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add nyldn/claude-octopus --skill skill-verify -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install nyldn/claude-octopus skill-verify --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nyldn/claude-octopus.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/skill-verify .cursor/skills/skill-verify && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "skill-verify" agent skill from https://github.com/nyldn/claude-octopus/tree/main/skills/skill-verify into .cursor/skills/skill-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-verify", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/nyldn/claude-octopus.git --path skills/skill-verify--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add nyldn/claude-octopus --skill skill-verify -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install nyldn/claude-octopus skill-verify --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nyldn/claude-octopus.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/skill-verify .gemini/skills/skill-verify && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "skill-verify" agent skill from https://github.com/nyldn/claude-octopus/tree/main/skills/skill-verify into .gemini/skills/skill-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-verify", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install nyldn/claude-octopus skill-verifyInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add nyldn/claude-octopus --skill skill-verify -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/nyldn/claude-octopus.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/skill-verify .github/skills/skill-verify && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "skill-verify" agent skill from https://github.com/nyldn/claude-octopus/tree/main/skills/skill-verify into .github/skills/skill-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-verify", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add nyldn/claude-octopus --skill skill-verify -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install nyldn/claude-octopus skill-verify --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/nyldn/claude-octopus.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/skill-verify .opencode/skills/skill-verify && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "skill-verify" agent skill from https://github.com/nyldn/claude-octopus/tree/main/skills/skill-verify into .opencode/skills/skill-verify/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "skill-verify", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skill-verifyA skill your agent uses when a nontrivial change needs end-to-end verification before committing or shipping
Skill Verify is an agent skill from nyldn/claude-octopus. Use when a nontrivial change needs end-to-end verification before committing or shipping
Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).
It sits in Testing & QA. The repository describes itself as: Run multiple AI models against the same research, design, or coding task. Surface disagreements before you ship. The licence is MIT.
5 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 4d152db. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
npmgitFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use npm and git, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Skill Verify loads about 1.4k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 632 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from nyldn/claude-octopus at commit 4d152db, republished under its MIT licence (© nyldn). 632 words, ~1,441 tokens.
.claude/skills/skill-verify/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.Host: Codex CLI — This skill was designed for Claude Code and adapted for Codex. Cross-reference commands use installed skill names in Codex rather than
/octo:*slash commands. Use the active Codex shell and subagent tools. Do not claim a provider, model, or host subagent is available until the current session exposes it. For host tool equivalents, seeskills/blocks/codex-host-adapter.md.
Load the selected feature's policy provenance and open-marker count through the existing boundary adapter. Cite the exact policy passage and conflicting plan action for any verified policy violation. Report unresolved decisions and which tasks depend on them. Stable task completion needs parent verification tied to the current task-contract digest; a checkbox, model claim or historical manifest alone cannot prove completion. Retain result-to-task mappings through correction, cancellation and resume.
<HARD-GATE>
NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
</HARD-GATE>
If you haven't run the verification command in this turn, you cannot claim it passes.
Before claiming any success or expressing satisfaction:
Skip any step = the claim is unverified.
| Excuse | Reality |
|---|---|
| "I ran the tests earlier this session" | Earlier is not fresh. Code changed since. Run again. |
| "The edit was trivial, it can't break anything" | Trivial edits break builds daily. The gate has no size exemption. |
| "The subagent reported success" | Agent reports are claims, not evidence. Verify independently. |
| "CI will catch it anyway" | CI is the safety net, not the verification. Verify before push. |
| "I'm confident this works" | Confidence is not evidence. Run the command. |
| "Running the full suite is slow" | Then run the targeted suite — but run something, fresh. |
| Claim | Requires | NOT Sufficient |
|---|---|---|
| Tests pass | Test command output showing 0 failures | Previous run, "should pass" |
| Build succeeds | Build command exit 0 | Linter passing |
| Bug fixed | Reproduce original symptom: now passes | "Code changed, should work" |
| Regression test works | Red (fail without fix) → Green (pass with fix) | Test passes once |
| Subagent completed task | git diff shows expected changes | Subagent says "done" |
| Requirements met | Line-by-line checklist against spec | Tests passing |
| Provider dispatch worked | Output contains expected content | No error ≠ success |
If you catch yourself thinking any of these, STOP:
| Thought | What to do instead |
|---|---|
| "Should work now" | Run the verification |
| "I'm confident" | Confidence ≠ evidence |
| "Just this once" | No exceptions |
| "The linter passed" | Linter ≠ tests ≠ build |
| "The agent said it worked" | Verify independently |
| "It's a small change" | Small changes cause big bugs |
In Claude Octopus workflows, verification is especially critical because:
After any multi-provider workflow:
# Verify synthesis file exists and is recent
ls -la ~/.claude-octopus/results/*-synthesis-*.md | tail -1
# Verify it has content (not just headers)
wc -l ~/.claude-octopus/results/*-synthesis-*.md | tail -1ALWAYS before:
In orchestrate.sh workflows:
probe (discover) — verify synthesis file existsgrasp (define) — verify consensus score meets thresholdtangle (develop) — verify tests pass, not just that code was writtenink (deliver) — verify review actually ran, not just that it was dispatched$ npm test
✓ user.create() saves to database (45ms)
✓ user.create() validates email (12ms)
Tests: 2 passed, 2 total
All 2 tests pass. ← Claim backed by output.I've implemented the feature. It should work now. The tests should pass.
← No test was run. "Should" is not evidence.1. Write test → run → FAIL (expected, proves test detects the bug)
2. Implement fix → run → PASS (proves fix works)
3. Revert fix → run → FAIL (proves test isn't false-positive)
4. Restore fix → run → PASS (final confirmation)This skill is referenced by:
flow-develop.md — verification gate after implementationflow-deliver.md — verification gate before deliveryskill-code-review.md — verify review findings before reportingskill-tdd.md — red-green cycle requires evidence at each stepskill-factory.md — autonomous pipeline must verify at every phase© nyldn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 1 other file in skills/skill-verify of nyldn/claude-octopus.
Open the folder on GitHubat commit 4d152db
We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in nyldn/claude-octopus, which our catalogue first saw on October 7, 2026.
Skill Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Skill Verify this skillnyldn/claude-octopus | 4.2k | 1 repos | ~1.4k | Automated safety check: Pass | MIT | |
| OpenHarness End-to-End EvalsHKUDS/OpenHarness | 16k | 1 repos | ~2.1k | Automated safety check: Notes | MIT | |
| Clawteam DevHKUDS/ClawTeam | 5.5k | 1 repos | ~1.1k | Automated safety check: Pass | MIT | |
| Acceptance Evidence for Deliverieslobehub/lobehub | 83k | — | ~9.7k | Automated safety check: Pass | Apache-2.0 | |
| Qwen Code E2E TestingQwenLM/qwen-code | 28k | — | ~2.1k | Automated safety check: Pass | Apache-2.0 | |
| Skill Testdatabricks-solutions/ai-dev-kit | 1.9k | — | ~1.9k | Automated safety check: Pass | Custom licence |
HKUDS/OpenHarness
Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.
HKUDS/ClawTeam
A skill your agent uses when working inside the ClawTeam repository itself: local development, debugging, reviewing, testing, validating multi-agent flows, or checking whether a code change actually…
lobehub/lobehub
Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.
QwenLM/qwen-code
Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.
databricks-solutions/ai-dev-kit
Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.
TriliumNext/Trilium
A skill your agent uses when adding, changing, or reviewing an LLM/MCP tool in Trilium (the defineTools definitions under packages/trilium-core/src/services/llm/tools/ —…
nyldn/claude-octopus
Quick execution for ad-hoc tasks without full workflow overhead — use for small, self-contained requests
nyldn/claude-octopus
Thorough research across multiple sources — use for complex topics needing broad synthesis
nyldn/claude-octopus
OWASP compliance, vulnerability scanning, and adversarial red team testing — use for security reviews
nyldn/claude-octopus
Audit codebases for quality, consistency, and broken patterns — use for pre-release or tech debt review
nyldn/claude-octopus
Extract patterns and anatomy from URLs — use to reverse-engineer content strategies from live pages
nyldn/claude-octopus
Auto-detect work context (Dev vs Knowledge) — use to tailor workflows based on current task type
Categories
A skill your agent uses when a nontrivial change needs end-to-end verification before committing or shipping. Skill Verify is an agent skill from nyldn/claude-octopus.
Skill Verify fits situations like: A nontrivial change needs end-to-end verification before committing.
Run `npx skills add nyldn/claude-octopus --skill skill-verify -a claude-code`. Or copy the skill folder (skills/skill-verify in nyldn/claude-octopus) into .claude/skills/skill-verify in your project. Claude Code loads it when a task matches its description.
Run `npx skills add nyldn/claude-octopus --skill skill-verify -a codex`. Or copy the skill folder (skills/skill-verify in nyldn/claude-octopus) into .agents/skills/skill-verify in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nyldn/claude-octopus --skill skill-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-verify, .gemini/skills/skill-verify, .github/skills/skill-verify and .opencode/skills/skill-verify in your project.
Going by SKILL.md and its folder, Skill Verify needs the command-line tools its instructions call (npm and git).
SKILL.md contains no URLs. Its commands use npm and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Skill Verify is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Skill Verify: OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars), Clawteam Dev (HKUDS/ClawTeam, 5.5k stars), Acceptance Evidence for Deliveries (lobehub/lobehub, 83k stars) and Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
nyldn (a GitHub user) maintains it in nyldn/claude-octopus, which has 4,182 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on October 7, 2026.
Source: nyldn/claude-octopus on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.