Agent skill

Feature Verify

by sd0xdev in sd0xdev/sd0x-harness

Feature verification (READ-ONLY, P0-P5). An agent skill from sd0xdev/sd0x-harness.

MITAuto-check: notesTesting & QA

Install Feature Verify

skills CLI
$ npx skills add sd0xdev/sd0x-harness --skill feature-verify -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sd0xdev/sd0x-harness feature-verify --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sd0xdev/sd0x-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/feature-verify .claude/skills/feature-verify && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
feature-verify
GitHub stars
192
Token cost
~3.2k tokens
SKILL.md length
1,141 words
Files
5 (incl. references)
Skills in repo
89
Repo updated
First seen
Licence
MIT

At a glance

Feature verification (READ-ONLY, P0-P5). An agent skill from sd0xdev/sd0x-harness.

  • Works in 3 steps: Get diff: git diff main...HEAD… → Map changed files → affected endpoints →… → Identify L1 regression endpoints, L2…
  • : verifying feature behavior after deployment
  • SKILL.md covers Trigger, When NOT to Use, Core Principle and Degradation Matrix, plus 11 more sections
  • Calls claude, curl and git

What it does

Feature Verify is an agent skill from sd0xdev/sd0x-harness. Feature verification (READ-ONLY, P0-P5). Use when: verifying feature behavior after deployment, validating API responses, diagnosing production issues, post-deploy smoke test. Not for: modifying data (use feature-dev), code review (use codex-review-fast), writing tests (use codex-test-gen), security audit (use codex-security).

Its SKILL.md is about 3.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/blackbox-testing.md`, `references/environments.md` and `references/output-template.md`).

It sits in Testing & QA, covering QA and bug reports, Security review and Deployment. The repository describes itself as: The harness layer for Claude Code — a reference implementation of harness engineering with hook-enforced dual review, state-machine gates that survive context compaction, and… The licence is MIT.

When your agent uses it

  • : verifying feature behavior after deployment
  • Validating API responses
  • Diagnosing production issues
  • Post-deploy smoke test

Example prompts

  • “/feature-verify”

Requirements

  • Pre-approved tools (allowed-tools): Read, Grep, Glob, Bash, WebFetch, Task, Skill

Workflow steps

3 steps, taken from the first numbered list in SKILL.md.

  1. Get diff: git diff main...HEAD --name-only (or user-provided scope)
  2. Map changed files → affected endpoints → dependency chains
  3. Identify L1 regression endpoints, L2 trigger cases, L3 passive targets

What it can do on your machine

Read from SKILL.md and the folder at commit a4d4bc1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Grep
    • Glob
    • Bash
    • WebFetch
    • Task
    • Skill

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • claude
    • curl
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Feature Verify loads about 3.2k tokens when it runs, and up to ~9.4k if it reads all its reference files. Until then it costs about 86 tokens; SKILL.md has 1,141 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~86
When it runs · the whole SKILL.md, loaded when a task matches
~3.2k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NotePre-approves every shell command (allowed-tools: Bash)SKILL.md
    allowed-tools: Read, Grep, Glob, Bash, WebFetch, Task, Skill

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sd0xdev/sd0x-harness at commit a4d4bc1, republished under its MIT licence (© sd0xdev). 1,141 words, ~3,221 tokens.

Download SKILL.mdSave it as .claude/skills/feature-verify/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
feature-verify
description
Feature verification (READ-ONLY, P0-P5). Use when: verifying feature behavior after deployment, validating API responses, diagnosing production issues, post-deploy smoke test. Not for: modifying data (use feature-dev), code review (use codex-review-fast), writing tests (use codex-test-gen), security audit (use codex-security).
allowed-tools
Read, Grep, Glob, Bash, WebFetch, Task, Skill
context
fork

Feature Verify — Runtime-First API Verification

Trigger

  • Keywords: verify, investigate, diagnose, check if working, post-deploy, smoke test, validate
  • User wants to confirm deployed feature behavior
  • User provides environment access (API URL, log system, credentials)

When NOT to Use

NeedUse Instead
Modify data or state/feature-dev
Code quality review/codex-review-fast
Generate unit tests/codex-test-gen
Security audit/codex-security
Run local tests/verify
Review test coverage/codex-test-review

Core Principle

⚠️ ALL OPERATIONS MUST BE READ-ONLY ⚠️

Claude independent analysis → Codex third-perspective confirmation → Integrated verdict

Tool safety note: allowed-tools includes Bash for curl/log queries. Read-only enforcement is behavioral — all commands MUST be reviewed against references/safety-rules.md before execution. Codex independently verifies compliance at P5.

Degradation Matrix

Auto-detect from references/environments.md configuration:

LevelAvailable ResourcesP3 APIP4 ObservationConfidence Cap
L4API + Log + MetricsFullLog + MetricsHigh
L3API + LogFullLog onlyHigh
L2-APIAPI onlyFullResponse-onlyMedium
L2-OBSLog only (API unreachable)SkipTime-window scanMedium
L1No runtime accessSkip P3/P4Code review onlyLow

Auto-detection logic (see references/environments.md § Degradation Detection):

API StatusLog SystemMetricsLevel
ReachableYesYesL4
ReachableYesNoL3
ReachableNo—L2-API
UnreachableYes—L2-OBS
UnreachableNo—L1

Fail-closed: If Endpoint Allowlist section is missing, skip P3 (cannot call unverified endpoints). At L1, skip P3 and P4. Provide code-review-based analysis only with Low confidence. At L2-OBS, skip P3 (API unreachable); execute P4 time-window scan and background service observation only.

Workflow

mermaid
sequenceDiagram
    participant C as Claude
    participant U as User
    participant API as Target API
    participant Log as Log System
    participant Cx as Codex

    C->>C: P0: Scope & Safety
    C->>C: P1: Diff-Lite Scoping
    C->>U: P2: Test Charter (approve?)
    U->>C: Approved
    C->>API: P3: API Execute (read-only)
    C->>Log: P4: Observation Correlate
    C->>Cx: P5: Codex independent review
    Cx-->>C: Codex verdict
    C->>U: P5: Integrated Verdict Report

P0: Scope & Safety

Read safety-rules.md and environments.md.

CheckMethodFail Action
Environment select--env flag or ask user; load from references/environments.mdDefault to test
Read-only confirmedReview references/safety-rules.md and load the endpoint allowlist. This row runs before any request is made — see below—
API reachableDeterministic health-check (3x, 2s timeout — see references/environments.md)Unreachable + Log config → L2-OBS; Unreachable + no Log → L1
Deployment alignedCompare local HEAD with deployed versionMismatch → warn, lower confidence
Degradation levelCheck references/environments.md for log/metrics configSet level (L1-L4)

The health check is a request, so the allowlist gates it too. A reviewer found this skill calling the configured health endpoint before enforcing its own deny-all policy — which is the one request the policy could never have approved, because nothing had loaded the allowlist yet. Validate the health endpoint and its method against the allowlist first; if the allowlist is missing, or the health endpoint is not on it, make no request and degrade on that basis (unreachable-equivalent), recording why. "It is only a health check" is exactly the reasoning the deny-all policy exists to refuse.

P1: Diff-Lite Scoping

Read blackbox-testing.md § P1.

Scope only — no code quality judgment.

  1. Get diff: git diff main...HEAD --name-only (or user-provided scope)
  2. Map changed files → affected endpoints → dependency chains
  3. Identify L1 regression endpoints, L2 trigger cases, L3 passive targets

Fallback: If no git diff available, ask user for feature description and build scope manually.

--level override: If user passes --level L2-API, skip log/metrics cases even if configured. --level L2-OBS forces observation-only mode. --level L2 defaults to L2-API for backward compatibility.

P2: Test Charter

Read blackbox-testing.md § P2.

Generate test cases dynamically from P1 results:

TypeGoalWhen
L1 RegressionAffected API returns expected resultsL2-API+ (N/A for L2-OBS)
L2 Active TriggerNew code path exercised, verify responseL2-API+ (N/A for L2-OBS)
L3 Passive ObserveBackground service running, check logsL3+ only
M1 MetricsMetrics correctly emitted with right labelsL4 only

User approval gate: Present charter table to user for confirmation before proceeding to P3. User may add/remove/modify cases.

P3: API Execute

Prerequisites: P2 approved, degradation level is L2-API or higher (L2-API/L3/L4). L2-OBS skips P3 entirely (API unreachable).

For each test case:

  1. Load headers from references/environments.md (generate unique request ID per call)
  2. Send request — only allowlisted endpoints (references/safety-rules.md)
  3. Record: HTTP status, response code, key response fields, request ID, latency
  4. Single request at a time (no concurrent/load testing)
  5. Use fixed test parameters from references/environments.md (no real user data)
bash
# Example execution pattern
make_headers
REQ_ID=$(extract_request_id)
# Timing comes from curl, not from `date`: `date +%s%3N` is GNU-only and on macOS prints a literal
# `3N`, so the subtraction that used to live here produced garbage on the platform this repo runs on.
RESP=$(curl -s -w "\n%{http_code}\n%{time_total}" -X {{ METHOD }} "$HOST/{{ ENDPOINT }}" \
  "${HEADERS[@]}" -d '{{ PAYLOAD }}')
LATENCY=$(echo "$RESP" | tail -1)          # seconds, millisecond resolution
HTTP_CODE=$(echo "$RESP" | tail -2 | head -1)
BODY=$(echo "$RESP" | sed '$d' | sed '$d')

P4: Observation Correlate

Read blackbox-testing.md § P4.

Prerequisites: Degradation level L2-OBS or L3+.

L2-OBS mode: Skip subsection A (no P3 requests to correlate). Execute B (time-window scan) and C (background service observation). Observation window: deploy_time → now (fallback: user-specified or last 30min).

Show full SKILL.md (469 more words)Show less
A. Per-Request Log Correlation (L1/L2 test case types, requires L3+)

For each P3 request, query logs by request ID with fallback strategy:

  1. Primary: request ID exact match
  2. Fallback: alternate field names
  3. Fallback: endpoint + time window

Retry: 30s fast → 120s delayed → mark unreachable.

B. Time-Window Scan (all cases)

Scan test period for anomalies (error + warn levels).

C. L3 Background Service Observation (if applicable)

Query logs for schedule/cron tags with 120s delay.

D. Metrics Observation (L4 only, if applicable)

Query metrics system for affected metrics, verify labels and values.

E. Blind Spot Analysis

Record what cannot be observed through black-box testing. List in report for /codex-test-review follow-up.

P5: Verdict

Per-Endpoint Verdict
VerdictCondition
PassL1 passed + L2 has expected signal + L3 normal + M1 correct (N/A items don't block)
WarnL1 passed but L2 signal missing, or L3/M1 has non-blocking anomaly
BlockedL1 failed, or regression detected, or M1 shows incorrect labels
InconclusiveAPI/log/metrics unreachable, insufficient evidence
Confidence Level
LevelCondition
HighL3/L4 + Claude and Codex agree
MediumL2-API (API-only) or L2-OBS (observation-only) or partial agreement
LowL1 (no runtime) or Claude and Codex diverge
Dual Verification (Claude + Codex)
  1. Claude analysis: Form independent conclusion from P3 + P4 evidence
  2. Codex review: Use /codex-brainstorm with P1 scope + P3 results + P4 observations (see references/blackbox-testing.md § P5)
  3. Integrated verdict: Synthesize both perspectives

Codex must independently verify (see references/blackbox-testing.md § P5 prompt):

  • No write operations were performed during P3
  • Each endpoint called was on the Endpoint Allowlist (references/environments.md)
  • All HTTP methods match allowlist (GET or allowlisted POST)
  • Verdict is justified by evidence
Output

Generate report using output-template.md.

Verdict is independent: Report may recommend follow-up skills (/codex-review-fast, /verify, /codex-test-review) but does NOT auto-invoke them.

Production Guardrails

RuleDescription
Single requestOne request at a time (no load testing)
Fixed parametersUse test parameters from references/environments.md
Read-only onlyOnly allowlisted endpoints (references/safety-rules.md)
No PIINo real user credentials, keys, or sensitive data in payloads
Rate awareRespect API rate limits

Verification Checklist

  • P0: Environment selected, reachable, deployment aligned
  • P0: Degradation level determined
  • P1: Affected endpoints mapped from diff (or user input)
  • P2: Test charter approved by user
  • P3: All API calls are read-only and on allowlist (L2-API+)
  • P3: L2-OBS correctly skips API execution
  • P3: Each call recorded with HTTP status, request ID, latency
  • P4: Log correlation attempted for each request (L3+)
  • P4: Time-window scan completed (L2-OBS or L3+)
  • P4: L2-OBS time-window scan uses correct observation window
  • P4: Blind spots documented
  • P5: Claude analysis formed independently
  • P5: Codex review completed independently
  • P5: Integrated verdict with confidence level
  • Report follows references/output-template.md format

References

FileContentRead At
environments.mdAPI endpoints, auth headers, log/metrics config, test paramsP0, P3
safety-rules.mdRead-only rules, endpoint allowlist, forbidden opsP0, P3
blackbox-testing.mdDiff-lite scoping, test charter design, log verification, blind spotsP1, P2, P4, P5
output-template.mdVerdict report formatP5

Examples

Input: /feature-verify "User Auth API" --env test
Action: P0(reachable? → L3) → P1(diff → /api/auth/*) → P2(L1+L2 charter, user approves)
        → P3(curl read-only endpoints) → P4(log correlation) → P5(verdict: Pass, High)
Input: /feature-verify "Payment query" --env prod --level L2
Action: P0(prod, forced L2) → P1(diff → /api/payment/query) → P2(L1+L2, no L3)
        → P3(curl) → P4(response-only) → P5(verdict: Pass, Medium)
Input: /feature-verify "Background sync job" --env staging
Action: P0(staging, L3) → P1(diff → cron changes) → P2(L3 passive only)
        → P3(skip — no API endpoint) → P4(log observation for schedule tag) → P5(verdict)
Input: /feature-verify "Cache optimization" (no env configured)
Action: P0(no config → L1) → P1(diff → cache service) → P2(code review only)
        → P3(skip) → P4(skip) → P5(verdict: Inconclusive, Low — recommend configuring references/environments.md)
Input: /feature-verify "Order processing" --env prod
Action: P0(prod, API unreachable 3/3, Log config present → L2-OBS)
        → P1(diff → /api/order/*) → P2(L3 passive + time-window only, no L1/L2 active)
        → P3(skip — API unreachable) → P4(time-window scan: deploy→now, background observation)
        → P5(verdict: Pass/Warn/Inconclusive, Medium)

© sd0xdev, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in skills/feature-verify of sd0xdev/sd0x-harness.

  • SKILL.md
  • references/blackbox-testing.md
  • references/environments.md
  • references/output-template.md
  • references/safety-rules.md

Open the folder on GitHubat commit a4d4bc1

Compare with similar skills

Feature Verify next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Feature Verify compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Feature Verify this skillsd0xdev/sd0x-harness192—~3.2kAutomated safety check: NotesMIT
Money Qualityiamzifei/show-me-the-money1k—~5.7kAutomated safety check: PassCustom licence
Vss E2E Smokeopen-edge-platform/edge-ai-libraries169—~1.5kAutomated safety check: PassApache-2.0
Review Codetobihagemann/turbo407—~3.2kAutomated safety check: PassMIT
Write Fix Briefn1m21n/Infinite264—~2.9kAutomated safety check: PassCustom licence
Spec Dogfoodleo-kuang-ai/spec-first107—~6.1kAutomated safety check: NotesMIT

Similar skills

  • Money Quality

    iamzifei/show-me-the-money

    Code and product quality gates for shipping with confidence.

    1k GitHub stars~5.7k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Vss E2E Smoke

    open-edge-platform/edge-ai-libraries

    Run this skill whenever the user asks to verify my VSS install works, smoke test VSS, check whether the deployment succeeded, or run an end-to-end test of summary/search for the…

    169 GitHub stars~1.5k tokensUpdated today
    Testing & QAAuto-check passed
  • Review Code

    tobihagemann/turbo

    Review code for bugs, security vulnerabilities, API misuse, consistency issues, simplicity problems, or test coverage gaps and low-value tests by running internal reviews and a peer review in…

    407 GitHub stars~3.2k tokensUpdated today
    DevelopmentAuto-check passed
  • Write Fix Brief

    n1m21n/Infinite

    Turn a bug report, review findings, broken-UI screenshot or new-node idea into a verified, file/line-precise implementation prompt after checking the real code.

    264 GitHub stars~2.9k tokensUpdated today
    Testing & QAAuto-check passed
  • Spec Dogfood

    leo-kuang-ai/spec-first

    Hands-off, diff-scoped browser QA of the active branch or PR.

    107 GitHub stars~6.1k tokensUpdated 13 days ago
    Testing & QAAuto-check: notes
  • Deslop

    foryourhealth111-pixel/Vibe-Skills

    Remove AI-generated code slop from a branch: unnecessary comments, redundant defensive checks, boilerplate, style drift, and type casts.

    3.6k GitHub stars~331 tokensUpdated 1 mo ago
    DevelopmentAuto-check passed

More from sd0xdev/sd0x-harness

All 89 skills in this repo
  • Adr

    sd0xdev/sd0x-harness

    Write an Architecture Decision Record (ADR) for a feature — Context / Decision / Status / Consequences / Alternatives, filed as docs/features/<feature/adr-<NNN-<title.md with a 3-digit zero-padded…

    192 GitHub stars~4.8k tokensUpdated today
    Auto-check passed
  • Load PR Review

    sd0xdev/sd0x-harness

    Load GitHub PR review comments into AI session — analyze, triage, plan.

    192 GitHub stars~4.4k tokensUpdated today
    Auto-check passed
  • Next Step

    sd0xdev/sd0x-harness

    Change-aware next step advisor. An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Obsidian CLI

    sd0xdev/sd0x-harness

    Obsidian vault integration via official CLI. An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Orchestrate

    sd0xdev/sd0x-harness

    Agent-driven workflow orchestration (v1 report-only). An agent skill from sd0xdev/sd0x-harness.

    192 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • PR Comment

    sd0xdev/sd0x-harness

    Post friendly review comments to a GitHub PR — prepare locally, preview, then submit as atomic review.

    192 GitHub stars~1.5k tokensUpdated today
    Auto-check passed

Questions about Feature Verify

What does Feature Verify do?

Feature verification (READ-ONLY, P0-P5). An agent skill from sd0xdev/sd0x-harness. Feature Verify is an agent skill from sd0xdev/sd0x-harness. Feature verification (READ-ONLY, P0-P5).

When should I use Feature Verify?

Feature Verify fits situations like: : verifying feature behavior after deployment; validating API responses; diagnosing production issues; post-deploy smoke test.

How do I install Feature Verify in Claude Code?

Run `npx skills add sd0xdev/sd0x-harness --skill feature-verify -a claude-code`. Or copy the skill folder (skills/feature-verify in sd0xdev/sd0x-harness) into .claude/skills/feature-verify in your project. Claude Code loads it when a task matches its description.

How do I install Feature Verify in Codex?

Run `npx skills add sd0xdev/sd0x-harness --skill feature-verify -a codex`. Or copy the skill folder (skills/feature-verify in sd0xdev/sd0x-harness) into .agents/skills/feature-verify in your project. Codex loads it when a task matches its description.

Can I use Feature Verify in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sd0xdev/sd0x-harness --skill feature-verify -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/feature-verify, .gemini/skills/feature-verify, .github/skills/feature-verify and .opencode/skills/feature-verify in your project.

What does Feature Verify need to run?

Going by SKILL.md and its folder, Feature Verify needs the command-line tools its instructions call (claude, curl and git). Its frontmatter pre-approves these tools: Read, Grep, Glob, Bash, WebFetch, Task, Skill.

Does Feature Verify access the network?

SKILL.md contains no URLs. Its commands use curl and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Feature Verify safe to install?

Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Feature Verify use?

Feature Verify is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Feature Verify use?

About 3.2k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.2k tokens, read only when the agent opens those files.

What are the alternatives to Feature Verify?

Skills that share tags, products or a category with Feature Verify: Money Quality (iamzifei/show-me-the-money, 1k stars), Vss E2E Smoke (open-edge-platform/edge-ai-libraries, 169 stars), Review Code (tobihagemann/turbo, 407 stars) and Write Fix Brief (n1m21n/Infinite, 264 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Feature Verify?

sd0xdev (a GitHub user) maintains it in sd0xdev/sd0x-harness, which has 192 GitHub stars. The repository holds 89 skills in this directory. The repository was last updated on October 8, 2026.

Source: sd0xdev/sd0x-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.