Agent skill

Skill Verify

by nyldn in nyldn/claude-octopus

A skill your agent uses when a nontrivial change needs end-to-end verification before committing or shipping

MITAuto-check passedTesting & QA

Install Skill Verify

skills CLI
$ npx skills add nyldn/claude-octopus --skill skill-verify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install nyldn/claude-octopus skill-verify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/nyldn/claude-octopus.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/skill-verify .claude/skills/skill-verify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-verify
GitHub stars
4.2k
Used in
1 other repo
Token cost
~1.4k tokens
SKILL.md length
632 words
Files
2
Skills in repo
62
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when a nontrivial change needs end-to-end verification before committing or shipping

  • Works in 5 steps: IDENTIFY — What command proves this claim? → RUN — Execute the full command (fresh,… → READ — Full output, check exit code,… → …
  • A nontrivial change needs end-to-end verification before committing
  • SKILL.md covers Portable feature evidence, The Iron Law, The Gate and Rationalization Table, plus 6 more sections
  • Calls npm and git

What it does

Skill Verify is an agent skill from nyldn/claude-octopus. Use when a nontrivial change needs end-to-end verification before committing or shipping

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `agents/openai.yaml`).

It sits in Testing & QA. The repository describes itself as: Run multiple AI models against the same research, design, or coding task. Surface disagreements before you ship. The licence is MIT.

When your agent uses it

  • A nontrivial change needs end-to-end verification before committing

Example prompts

  • “/skill-verify”

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. IDENTIFY — What command proves this claim?
  2. RUN — Execute the full command (fresh, not cached)
  3. READ — Full output, check exit code, count failures
  4. VERIFY — Does output actually confirm the claim?
  5. ONLY THEN — State the claim WITH evidence

What it can do on your machine

Read from SKILL.md and the folder at commit 4d152db. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • npm
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use npm and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Verify loads about 1.4k tokens when it runs. Until then it costs about 25 tokens; SKILL.md has 632 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~25
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from nyldn/claude-octopus at commit 4d152db, republished under its MIT licence (© nyldn). 632 words, ~1,441 tokens.

Download SKILL.mdSave it as .claude/skills/skill-verify/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
skill-verify
description
Use when a nontrivial change needs end-to-end verification before committing or shipping
disable-model-invocation
true

Host: Codex CLI — This skill was designed for Claude Code and adapted for Codex. Cross-reference commands use installed skill names in Codex rather than /octo:* slash commands. Use the active Codex shell and subagent tools. Do not claim a provider, model, or host subagent is available until the current session exposes it. For host tool equivalents, see skills/blocks/codex-host-adapter.md.

Verification Gate

Portable feature evidence

Load the selected feature's policy provenance and open-marker count through the existing boundary adapter. Cite the exact policy passage and conflicting plan action for any verified policy violation. Report unresolved decisions and which tasks depend on them. Stable task completion needs parent verification tied to the current task-contract digest; a checkbox, model claim or historical manifest alone cannot prove completion. Retain result-to-task mappings through correction, cancellation and resume.

The Iron Law

<HARD-GATE>
NO COMPLETION CLAIMS WITHOUT FRESH VERIFICATION EVIDENCE
</HARD-GATE>

If you haven't run the verification command in this turn, you cannot claim it passes.

The Gate

Before claiming any success or expressing satisfaction:

  1. IDENTIFY — What command proves this claim?
  2. RUN — Execute the full command (fresh, not cached)
  3. READ — Full output, check exit code, count failures
  4. VERIFY — Does output actually confirm the claim?
  5. ONLY THEN — State the claim WITH evidence

Skip any step = the claim is unverified.

Rationalization Table

ExcuseReality
"I ran the tests earlier this session"Earlier is not fresh. Code changed since. Run again.
"The edit was trivial, it can't break anything"Trivial edits break builds daily. The gate has no size exemption.
"The subagent reported success"Agent reports are claims, not evidence. Verify independently.
"CI will catch it anyway"CI is the safety net, not the verification. Verify before push.
"I'm confident this works"Confidence is not evidence. Run the command.
"Running the full suite is slow"Then run the targeted suite — but run something, fresh.

What Counts as Evidence

ClaimRequiresNOT Sufficient
Tests passTest command output showing 0 failuresPrevious run, "should pass"
Build succeedsBuild command exit 0Linter passing
Bug fixedReproduce original symptom: now passes"Code changed, should work"
Regression test worksRed (fail without fix) → Green (pass with fix)Test passes once
Subagent completed taskgit diff shows expected changesSubagent says "done"
Requirements metLine-by-line checklist against specTests passing
Provider dispatch workedOutput contains expected contentNo error ≠ success
Show full SKILL.md (242 more words)Show less

Red Flags — STOP and Verify

If you catch yourself thinking any of these, STOP:

ThoughtWhat to do instead
"Should work now"Run the verification
"I'm confident"Confidence ≠ evidence
"Just this once"No exceptions
"The linter passed"Linter ≠ tests ≠ build
"The agent said it worked"Verify independently
"It's a small change"Small changes cause big bugs

Multi-Provider Context

In Claude Octopus workflows, verification is especially critical because:

  • Provider outputs can be hallucinated — Codex, Antigravity, Copilot, and other providers may claim success without evidence
  • Consensus ≠ correctness — three models agreeing doesn't mean they're right
  • Synthesis files may be stale — check timestamps, don't assume freshness
  • orchestrate.sh exit code 0 ≠ quality — the script ran, but did it produce good output?

After any multi-provider workflow:

bash
# Verify synthesis file exists and is recent
ls -la ~/.claude-octopus/results/*-synthesis-*.md | tail -1

# Verify it has content (not just headers)
wc -l ~/.claude-octopus/results/*-synthesis-*.md | tail -1

When to Apply

ALWAYS before:

  • Committing code
  • Creating PRs
  • Marking tasks complete
  • Moving to next workflow phase
  • Reporting results to user
  • Claiming a bug is fixed

In orchestrate.sh workflows:

  • After probe (discover) — verify synthesis file exists
  • After grasp (define) — verify consensus score meets threshold
  • After tangle (develop) — verify tests pass, not just that code was written
  • After ink (deliver) — verify review actually ran, not just that it was dispatched

Examples

Correct: Evidence-Based Claim
$ npm test
  ✓ user.create() saves to database (45ms)
  ✓ user.create() validates email (12ms)
  Tests: 2 passed, 2 total

All 2 tests pass. ← Claim backed by output.
Incorrect: Claim Without Evidence
I've implemented the feature. It should work now. The tests should pass.
← No test was run. "Should" is not evidence.
Correct: Regression Test Red-Green
1. Write test → run → FAIL (expected, proves test detects the bug)
2. Implement fix → run → PASS (proves fix works)
3. Revert fix → run → FAIL (proves test isn't false-positive)
4. Restore fix → run → PASS (final confirmation)

Integration with Other Skills

This skill is referenced by:

  • flow-develop.md — verification gate after implementation
  • flow-deliver.md — verification gate before delivery
  • skill-code-review.md — verify review findings before reporting
  • skill-tdd.md — red-green cycle requires evidence at each step
  • skill-factory.md — autonomous pipeline must verify at every phase

© nyldn, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/skill-verify of nyldn/claude-octopus.

  • SKILL.md
  • agents/openai.yaml

Open the folder on GitHubat commit 4d152db

Used in 1 other repository

We found 1 copy of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in nyldn/claude-octopus, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Skill Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Verify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Verify this skillnyldn/claude-octopus4.2k1 repos~1.4kAutomated safety check: PassMIT
OpenHarness End-to-End EvalsHKUDS/OpenHarness16k1 repos~2.1kAutomated safety check: NotesMIT
Clawteam DevHKUDS/ClawTeam5.5k1 repos~1.1kAutomated safety check: PassMIT
Acceptance Evidence for Deliverieslobehub/lobehub83k—~9.7kAutomated safety check: PassApache-2.0
Qwen Code E2E TestingQwenLM/qwen-code28k—~2.1kAutomated safety check: PassApache-2.0
Skill Testdatabricks-solutions/ai-dev-kit1.9k—~1.9kAutomated safety check: PassCustom licence

Similar skills

  • Validates OpenHarness features by running real multi-turn agent loops with live LLM calls against an unfamiliar codebase, checking actual tool execution.

    16k GitHub starsUsed in 1 repo~2.1k tokens
    Testing & QAAuto-check: notes
  • Clawteam Dev

    HKUDS/ClawTeam

    A skill your agent uses when working inside the ClawTeam repository itself: local development, debugging, reviewing, testing, validating multi-agent flows, or checking whether a code change actually…

    5.5k GitHub starsUsed in 1 repo~1.1k tokens
    Testing & QAAuto-check passed
  • Verifies a delivery end to end by driving the real product on a CLI, web, desktop or iOS Simulator surface, capturing evidence and publishing a round with the lh CLI.

    83k GitHub stars~9.7k tokensUpdated today
    Testing & QAAuto-check passed
  • Qwen Code E2E Testing

    QwenLM/qwen-code

    Guides end-to-end testing of the Qwen Code CLI in headless mode with real model calls, MCP test servers and inspection of raw API traffic.

    28k GitHub stars~2.1k tokensUpdated today
    Testing & QAAuto-check passed
  • Skill Test

    databricks-solutions/ai-dev-kit

    Testing framework for evaluating Databricks skills. An agent skill from databricks-solutions/ai-dev-kit.

    1.9k GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Adding LLM MCP Tools

    TriliumNext/Trilium

    A skill your agent uses when adding, changing, or reviewing an LLM/MCP tool in Trilium (the defineTools definitions under packages/trilium-core/src/services/llm/tools/ —…

    38k GitHub stars~2.5k tokensUpdated today
    Testing & QAAuto-check passed

More from nyldn/claude-octopus

All 62 skills in this repo
  • Octopus Quick

    nyldn/claude-octopus

    Quick execution for ad-hoc tasks without full workflow overhead — use for small, self-contained requests

    4.2k GitHub starsUsed in 1 repo~2.2k tokens
    Auto-check passed
  • Octopus Research

    nyldn/claude-octopus

    Thorough research across multiple sources — use for complex topics needing broad synthesis

    4.2k GitHub starsUsed in 1 repo~1.9k tokens
    Auto-check passed
  • Octopus Security Audit

    nyldn/claude-octopus

    OWASP compliance, vulnerability scanning, and adversarial red team testing — use for security reviews

    4.2k GitHub starsUsed in 1 repo~2.3k tokens
    Auto-check passed
  • Skill Audit

    nyldn/claude-octopus

    Audit codebases for quality, consistency, and broken patterns — use for pre-release or tech debt review

    4.2k GitHub starsUsed in 1 repo~3.2k tokens
    Auto-check passed
  • Skill Content Pipeline

    nyldn/claude-octopus

    Extract patterns and anatomy from URLs — use to reverse-engineer content strategies from live pages

    4.2k GitHub starsUsed in 1 repo~3.9k tokens
    Auto-check passed
  • Skill Context Detection

    nyldn/claude-octopus

    Auto-detect work context (Dev vs Knowledge) — use to tailor workflows based on current task type

    4.2k GitHub starsUsed in 1 repo~2.6k tokens
    Auto-check passed

Questions about Skill Verify

What does Skill Verify do?

A skill your agent uses when a nontrivial change needs end-to-end verification before committing or shipping. Skill Verify is an agent skill from nyldn/claude-octopus.

When should I use Skill Verify?

Skill Verify fits situations like: A nontrivial change needs end-to-end verification before committing.

How do I install Skill Verify in Claude Code?

Run `npx skills add nyldn/claude-octopus --skill skill-verify -a claude-code`. Or copy the skill folder (skills/skill-verify in nyldn/claude-octopus) into .claude/skills/skill-verify in your project. Claude Code loads it when a task matches its description.

How do I install Skill Verify in Codex?

Run `npx skills add nyldn/claude-octopus --skill skill-verify -a codex`. Or copy the skill folder (skills/skill-verify in nyldn/claude-octopus) into .agents/skills/skill-verify in your project. Codex loads it when a task matches its description.

Can I use Skill Verify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add nyldn/claude-octopus --skill skill-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-verify, .gemini/skills/skill-verify, .github/skills/skill-verify and .opencode/skills/skill-verify in your project.

What does Skill Verify need to run?

Going by SKILL.md and its folder, Skill Verify needs the command-line tools its instructions call (npm and git).

Does Skill Verify access the network?

SKILL.md contains no URLs. Its commands use npm and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Skill Verify safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Skill Verify use?

Skill Verify is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Verify use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Verify?

Skills that share tags, products or a category with Skill Verify: OpenHarness End-to-End Evals (HKUDS/OpenHarness, 16k stars), Clawteam Dev (HKUDS/ClawTeam, 5.5k stars), Acceptance Evidence for Deliveries (lobehub/lobehub, 83k stars) and Qwen Code E2E Testing (QwenLM/qwen-code, 28k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Verify?

nyldn (a GitHub user) maintains it in nyldn/claude-octopus, which has 4,182 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on October 7, 2026.

Source: nyldn/claude-octopus on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.