Agent skill

User Testing Validator

by Intelligent-Internet in Intelligent-Internet/zenith

Real-surface validation coordinator for engineering validation assignments.

Apache-2.0Auto-check passedProduct & Project Management

Install User Testing Validator

skills CLI
$ npx skills add Intelligent-Internet/zenith --skill user-testing-validator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Intelligent-Internet/zenith user-testing-validator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/user-testing-validator .claude/skills/user-testing-validator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
user-testing-validator
GitHub stars
334
Token cost
~1.8k tokens
SKILL.md length
717 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Real-surface validation coordinator for engineering validation assignments.

  • Works in 6 steps: Prepare setup → Partition lanes when useful → Exercise each assertion → …
  • Tasks that involve User research
  • SKILL.md covers Inputs, Surface Selection, Procedure and Minimum Evidence Floors, plus 1 more section
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

User Testing Validator is an agent skill from Intelligent-Internet/zenith. Real-surface validation coordinator for engineering validation assignments. Exercises assigned assertions through browser, API, CLI, background, generated-artifact, migration/data, public-library, or parity surfaces and returns per-target verdicts with fresh evidence.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Product & Project Management, covering User research. It works with Model Context Protocol. The repository describes itself as: Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve User research

Example prompts

  • “/user-testing-validator”

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Prepare setup
  2. Partition lanes when useful
  3. Exercise each assertion
  4. Save evidence
  5. Write regression ledgers for failures
  6. Synthesize verdicts

What it can do on your machine

Read from SKILL.md and the folder at commit a8d9b57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

User Testing Validator loads about 1.8k tokens when it runs. Until then it costs about 73 tokens; SKILL.md has 717 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~73
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Intelligent-Internet/zenith at commit a8d9b57, republished under its Apache-2.0 licence (© Intelligent-Internet). 717 words, ~1,750 tokens.

Download SKILL.mdSave it as .claude/skills/user-testing-validator/SKILL.md (or your agent's skills folder).
name
user-testing-validator
description
Real-surface validation coordinator for engineering validation assignments. Exercises assigned assertions through browser, API, CLI, background, generated-artifact, migration/data, public-library, or parity surfaces and returns per-target verdicts with fresh evidence.

User Testing Validator

Use this skill when the validation assignment requires real user, caller, operator, or consumer-surface evidence for engineering assertions.

Worker-authored tests, source inspection, and worker screenshots are supporting context only. Fresh validator-collected evidence is the verdict source.

Inputs

Read:

  • Validation assignment, assigned target ids, requested surface/method, and any assignment-level setup or dependency notes.
  • Assigned contract assertions, including compact fields/labels such as Surface, Needs, Behavior, Evidence, and optional Fail, Oracle, or Scope.
  • AGENTS.md.
  • Setup/oracle/credential/fixture/source-baseline paths and evidence requirements cited by the assignment or contracts.
  • Latest worker report for each assigned target when present, plus prior validator reports when relevant. Treat reports as claims, not proof; do not let them anchor the verdict.

Surface Selection

Choose the surface that matches the contract:

  • Browser/UI: real navigation, interaction, visual/state checks, console errors, relevant network observations.
  • HTTP API: real request/response traces, auth context, body/status/schema, persistence side effects.
  • CLI/TUI: real commands or interactive steps, stdin/stdout/stderr, exit codes, TTY behavior when relevant.
  • Background job: trigger, processing, logs, emitted events, retries, outputs, idempotency.
  • Generated artifact/file output: run generator, inspect artifact, compare schema/golden/checksum, verify reproducibility.
  • Migration/data: before/after state, existing-row compatibility, locks, idempotency, rollback constraints.
  • Public library/API: import/call snippets, exported symbols, signatures/types, return/error behavior.
  • Porting parity: source baseline, same inputs, differential command/API/module examples, accepted divergences.

Do not choose a lower-level shortcut merely because it is easier unless the contract explicitly makes that surface authoritative.

Procedure

  1. Prepare setup

    • Run required setup assigned by the assignment, contract, skill, or AGENTS.md.
    • Create disposable probes when useful to expose bugs: temporary tests, scripts, sample repos, fuzz cases, fixtures, data prefixes, or input corpora. Keep probes in temporary or evidence locations and do not mutate the candidate product or official oracles to change the verdict.
    • Respect off-limits surfaces from the user request, assignment, contract, or skill. Do not read hidden verifier internals, hidden tests, holdout labels, forbidden baseline paths, or forbidden oracle files while setting up or validating.
    • Parse each target's Needs before exercising the surface. Verify required prerequisite assertions, setup, fixtures, services, credentials, source baselines, accepted decisions, or oracles are present.
    • Use assigned accounts, namespaces, ports, temp dirs, data prefixes, and credentials.
    • If setup or a required Needs entry fails, attempt one non-disruptive recovery. If still blocked, fail affected targets with setup evidence.
  2. Partition lanes when useful

    • Use flow-validator subagents for independent surface groups when resources can be isolated.
    • Give each lane target ids, contract paths or bodies, surface/method, relevant Needs, required Evidence, allowed resources, unique evidence subdirectory or artifact prefix, non-goals, and output schema.
    • Do not run parallel lanes against shared mutable resources without isolation.
    • Reject or fail lane results whose artifacts cannot be attributed to exact target ids. Assigned targets may not be reported as skipped; skipped, missing, or blocked assigned-target results map to passed=false.
  3. Exercise each assertion

    • Perform the actor workflow from setup to expected result.
    • Capture evidence named by the target's Evidence field and validation assignment.
    • Compare observed behavior to Behavior, Surface, Needs, Evidence, and any optional Fail, Oracle, or Scope constraints.
  4. Save evidence

    • Write screenshots, traces, logs, terminal captures, raw outputs, or other artifacts under <evidence_dir>.
    • Use descriptive filenames or subdirectories that include target id and lane id when applicable. Do not overwrite another target or lane's artifacts.
  5. Write regression ledgers for failures

    • For each failed item, write <regressions_dir>/<item_id>.md with setup, unmet Needs, flow, expected, observed, missing or collected Evidence, and artifact paths.
  6. Synthesize verdicts

    • passed=true only when required fresh evidence exists and observed behavior matches the contract and any assignment-added checks.
    • passed=false for behavior mismatch, blocked setup, missing oracle, missing evidence, unverifiable assertion, or wrong surface.
    • For subagent lanes, any fail, blocked, assigned-target skipped, missing assigned-target result, or missing required artifact maps to parent passed=false for that target.
Show full SKILL.md (99 more words)Show less

Minimum Evidence Floors

  • Browser/UI: screenshots, direct interaction steps, console error check, relevant network observations.
  • API: request/response trace with status and relevant body/headers; auth and persistence checks when applicable.
  • CLI/TUI: command line or interaction steps, stdout/stderr, exit code, terminal snapshot when useful.
  • Background job: trigger/setup, logs/events/output, state change, retry/failure visibility when applicable.
  • Generated artifact: artifact path, generation command, diff/checksum/schema/golden check when useful.
  • Migration/data: before/after state and compatibility/idempotency evidence.
  • Public library/API: import/call snippet and observed public behavior.
  • Porting parity: source baseline, differential/golden output, accepted-divergence evidence.

Missing required evidence means passed=false.

Report

markdown
## Setup
- <services, commands, accounts, namespaces, source baselines>

## Per-Target Verdicts
### <target-id>: passed | failed
- Surface: <browser/API/CLI/job/artifact/data/library/parity>
- Needs: <checked prerequisites, unmet dependencies, or None>
- Steps: <what was exercised>
- Evidence: <artifact paths and observations>
- Reason: <why contract passed or failed>

## Frictions And Blockers
- <deduped by root cause>

## Review Limitations
- <anything not tested and why>

## Guidance Suggestions
- <optional setup/skill/AGENTS.md suggestions for orchestrator>

Call end_node with one item per target, then exit immediately.

© Intelligent-Internet, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in zenith/src/zenith_harness/bundled/skills/user-testing-validator of Intelligent-Internet/zenith.

Open the folder on GitHubat commit a8d9b57

Compare with similar skills

User Testing Validator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

User Testing Validator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
User Testing Validator this skillIntelligent-Internet/zenith334—~1.8kAutomated safety check: PassApache-2.0
Produck Feedback To Buildtryproduck/produck-skills511—~1kAutomated safety check: PassApache-2.0
Rethink Surveysooiyeefei/ccc494—~3.5kAutomated safety check: PassMIT
Foundation Meeting Recapproduct-on-purpose/pm-skills713—~2.6kAutomated safety check: PassApache-2.0
Ouroboros PM InterviewQ00/ouroboros6.2k—~5.7kAutomated safety check: PassMIT
Jira Natural Language Interfacejjmartres/opencode1333 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Produck Feedback To Build

    tryproduck/produck-skills

    Pulls full in-context user feedback tickets through the Produck MCP server and turns them into an aligned product change instead of a guess.

    511 GitHub stars~1k tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • Rethink Surveys

    ooiyeefei/ccc

    Design, critique, or scaffold surveys grounded in Caroline Jarrett, Dillman, and Tourangeau methods.

    494 GitHub stars~3.5k tokensUpdated 2 mo ago
    Product & Project ManagementAuto-check passed
  • Foundation Meeting Recap

    product-on-purpose/pm-skills

    Produces a topic-segmented post-meeting summary for attendees with decisions highlighted and actions captured inline per topic (plus a consolidated action view at the end).

    713 GitHub stars~2.6k tokensUpdated 3 days ago
    Productivity & AutomationAuto-check passed
  • Runs a guided product-manager interview that classifies each question automatically and produces a Product Requirements Document.

    6.2k GitHub stars~5.7k tokensUpdated yesterday
    Product & Project ManagementAuto-check passed
  • Lets an agent view, create, update and transition Jira issues in natural language, automatically choosing between the jira CLI and Atlassian MCP tools.

    133 GitHub starsUsed in 3 repos~1.7k tokens
    Product & Project ManagementAuto-check passed
  • Load Context

    vishalmdi/ai-native-pm-os

    Loads all PM context files and prepares Claude for a productive session.

    106 GitHub stars~562 tokensUpdated 5 mo ago
    Product & Project ManagementAuto-check passed

More from Intelligent-Internet/zenith

  • Engineering Mission Playbook

    Intelligent-Internet/zenith

    A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…

    334 GitHub stars~8.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Benchmark Validator

    Intelligent-Internet/zenith

    Benchmark validation procedure for one assigned benchmark-related target.

    334 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Scrutiny Validator

    Intelligent-Internet/zenith

    Adversarial scrutiny procedure for engineering validation assignments.

    334 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Browser

    Intelligent-Internet/zenith

    Automates browser and Electron app interactions for user-flow validation.

    334 GitHub stars~6.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Optimization Mission Playbook

    Intelligent-Internet/zenith

    Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…

    334 GitHub stars~11k tokensUpdated 1 mo ago
    Auto-check passed

Questions about User Testing Validator

What does User Testing Validator do?

Real-surface validation coordinator for engineering validation assignments. User Testing Validator is an agent skill from Intelligent-Internet/zenith. Real-surface validation coordinator for engineering validation assignments.

When should I use User Testing Validator?

User Testing Validator fits situations like: tasks that involve User research.

How do I install User Testing Validator in Claude Code?

Run `npx skills add Intelligent-Internet/zenith --skill user-testing-validator -a claude-code`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/user-testing-validator in Intelligent-Internet/zenith) into .claude/skills/user-testing-validator in your project. Claude Code loads it when a task matches its description.

How do I install User Testing Validator in Codex?

Run `npx skills add Intelligent-Internet/zenith --skill user-testing-validator -a codex`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/user-testing-validator in Intelligent-Internet/zenith) into .agents/skills/user-testing-validator in your project. Codex loads it when a task matches its description.

Can I use User Testing Validator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Intelligent-Internet/zenith --skill user-testing-validator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/user-testing-validator, .gemini/skills/user-testing-validator, .github/skills/user-testing-validator and .opencode/skills/user-testing-validator in your project.

What does User Testing Validator need to run?

SKILL.md names no scripts, command-line tools or credentials: User Testing Validator is instructions for the agent only.

Does User Testing Validator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is User Testing Validator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does User Testing Validator use?

User Testing Validator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does User Testing Validator use?

About 1.8k tokens (SKILL.md is roughly 7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to User Testing Validator?

Skills that share tags, products or a category with User Testing Validator: Produck Feedback To Build (tryproduck/produck-skills, 511 stars), Rethink Surveys (ooiyeefei/ccc, 494 stars), Foundation Meeting Recap (product-on-purpose/pm-skills, 713 stars) and Ouroboros PM Interview (Q00/ouroboros, 6.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains User Testing Validator?

Intelligent-Internet (a GitHub organization) maintains it in Intelligent-Internet/zenith, which has 334 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 6, 2026.

Source: Intelligent-Internet/zenith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.