Agent skill

Scrutiny Validator

by Intelligent-Internet in Intelligent-Internet/zenith

Adversarial scrutiny procedure for engineering validation assignments.

Apache-2.0Auto-check passedAgent Workflows

Install Scrutiny Validator

skills CLI
$ npx skills add Intelligent-Internet/zenith --skill scrutiny-validator -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Intelligent-Internet/zenith scrutiny-validator --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Intelligent-Internet/zenith.git skills-src && mkdir -p .claude/skills && cp -r skills-src/zenith/src/zenith_harness/bundled/skills/scrutiny-validator .claude/skills/scrutiny-validator && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
scrutiny-validator
GitHub stars
337
Token cost
~1.8k tokens
SKILL.md length
770 words
Files
1
Skills in repo
6
Repo updated
First seen
Licence
Apache-2.0

At a glance

Adversarial scrutiny procedure for engineering validation assignments.

  • Works in 7 steps: Establish scope → Run hard-gate commands → Review implementation → …
  • Agent Workflows work in your project
  • SKILL.md covers Inputs, Procedure, Failure Severity and Report
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Scrutiny Validator is an agent skill from Intelligent-Internet/zenith. Adversarial scrutiny procedure for engineering validation assignments. Runs hard-gate commands, reviews the current implementation and evidence integrity against assigned contracts, can use feature-reviewer lanes, and returns per-target verdicts.

Its SKILL.md is about 1.8k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows. It works with Model Context Protocol. The repository describes itself as: Zenith: a continuous-improvement harness for long-running agent tasks. Turns Claude Code, Codex, or Hermes into a multi-agent mission orchestrator via MCP/ACP. The licence is Apache-2.0.

When your agent uses it

  • Agent Workflows work in your project

Example prompts

  • “/scrutiny-validator”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Establish scope
  2. Run hard-gate commands
  3. Review implementation
  4. Review evidence integrity
  5. Use feature-reviewer lanes when useful
  6. Write regression ledger entries
  7. Return verdicts

What it can do on your machine

Read from SKILL.md and the folder at commit a8d9b57. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are markdown).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Scrutiny Validator loads about 1.8k tokens when it runs. Until then it costs about 66 tokens; SKILL.md has 770 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~1.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Intelligent-Internet/zenith at commit a8d9b57, republished under its Apache-2.0 licence (© Intelligent-Internet). 770 words, ~1,798 tokens.

Download SKILL.mdSave it as .claude/skills/scrutiny-validator/SKILL.md (or your agent's skills folder).
name
scrutiny-validator
description
Adversarial scrutiny procedure for engineering validation assignments. Runs hard-gate commands, reviews the current implementation and evidence integrity against assigned contracts, can use feature-reviewer lanes, and returns per-target verdicts.

Scrutiny Validator

Use this skill when the validation assignment asks for implementation scrutiny, hard-gate command review, evidence-integrity audit, or source review for engineering targets.

Scrutiny is not a substitute for real user/caller/operator surface validation. Use user-testing-validator or another surface-specific validator when the contract requires behavior through a real interface.

Inputs

Read:

  • Validation assignment, including exact target ids, assigned validation method, and any assignment-level setup or dependency notes.
  • Assigned assertions. Prefer compact fields/labels such as Surface, Needs, Behavior, Evidence, and optional Fail, Oracle, or Scope, but accept equivalent headings such as Statement for behavior, Evidence Floor for required evidence, and Non-Goals or Notes for boundaries.
  • AGENTS.md.
  • Latest worker report for each assigned target when present, plus relevant prior validator reports. Treat reports as claims, not proof.
  • Current product checkout files related to the targets. Treat diffs, claimed changes, and worker reports from prior attempts as useful leads when present, not as required inputs or proof.
  • Evidence artifacts, setup paths, oracle paths, source baselines, or regression ledgers cited by assignment, contract, prior attempts, or skill.

Procedure

  1. Establish scope

    • Identify exact target ids and the current product behavior, files, and evidence surfaces being scrutinized.
    • Parse every assigned assertion into behavior, surface, required evidence, prerequisites, oracle, failure conditions, and scope boundaries. Accept equivalent headings such as Statement, Evidence Floor, Non-Goals, and Notes when they provide the same meaning.
    • If behavior, surface, or required evidence cannot be determined from the assertion and validation assignment, mark the target unverifiable.
    • Map each target's Needs and Evidence requirements before reviewing implementation or evidence surfaces. Do not assume prerequisites, setup, fixtures, services, source baselines, or accepted decisions that are not present.
    • Map each target to implementation files and evidence-producing files.
  2. Run hard-gate commands

    • Run tests, lint, typecheck, build, generated-output checks, migration checks, or contract-required commands relevant to the target.
    • Capture command, exit code, focused output, and saved artifact path when the output is too large or important for the report alone.
    • Save durable scrutiny artifacts under <evidence_dir> when useful: command logs, focused outputs, generated-output diffs, fixture/golden review notes, or source-baseline comparison notes.
    • Do not hide exit codes with output truncation pipelines.
  3. Review implementation

    • Confirm the current implementation satisfies the contract behavior, surface, required evidence, and any fail, oracle, scope, non-goal, or notes constraints, and accounts for every relevant prerequisite.
    • Check edge cases, error paths, data integrity, auth/authz, compatibility, migration safety, idempotency, and public API behavior when relevant.
    • Flag responsibility drift outside assignment or contract scope.
  4. Review evidence integrity

    • Inspect changed or evidence-sensitive tests, fixtures, benchmark scripts, verifier code, golden files, generated expected outputs, source baselines, data files, and cited evidence artifacts when they are allowed and relevant.
    • Restrict that inspection to allowed paths. Do not read hidden verifier internals, hidden tests, holdout labels, forbidden baseline paths, or other off-limits surfaces; use allowed public artifacts and report unverifiable proof gaps instead.
    • Compare the available proof against each target's Evidence field. Missing required evidence owned by this validation task means passed=false.
    • If the contract requires real user/caller/operator surface evidence, do not pass it through scrutiny alone unless the validation assignment explicitly scopes this run as the scrutiny lane and names a sibling validator that owns the real-surface evidence.
    • Fail targets when proof was weakened, mocked away, or made benchmark/test-only.
  5. Use feature-reviewer lanes when useful

    • For broad implementation areas or multiple feature areas, spawn feature-reviewer with one bounded review question, assigned target ids, contract paths or bodies, relevant current-checkout files or surfaces, relevant evidence artifacts or command output, claimed changes or prior reports as leads when useful, allowed commands/probes/evidence-write locations, off-limits surfaces, and expected output shape.
    • The subagent report is advisory. You own the verdict.
  6. Write regression ledger entries

    • For each failed or unverifiable item, write <regressions_dir>/<item_id>.md with the failing command, review finding, observation, unmet Needs, missing Evidence, artifact paths, and reproduction details.
  7. Return verdicts

    • One items[] entry per assigned target.
    • passed=true only when every check owned by this scrutiny assignment passes, relevant contract prerequisites are accounted for, scrutiny evidence artifacts are collected, and no integrity or responsibility-drift issue blocks the target.
    • If required real-surface evidence is outside this assignment's method, name the sibling validator that owns it. Do not claim scrutiny alone proves the full target.
Show full SKILL.md (68 more words)Show less

Failure Severity

Blocking findings include:

  • Missing implementation for a target.
  • Required command failure tied to the target.
  • Unmet prerequisite from Needs.
  • Missing required evidence owned by this validation task.
  • Evidence-sensitive changes that invalidate proof.
  • Security, data-integrity, migration, compatibility, or public API break.
  • Required oracle/source baseline missing.
  • Unverifiable assertion.

Prefer passed=false with concrete limitations over speculative passes.

Report

markdown
## Hard-Gate Commands
- `<command>` -> exit <code>. <focused output>

## Per-Target Verdicts
### <target-id>: passed | failed
- Needs: <checked prerequisites, unmet dependencies, or None>
- Evidence: <commands/source review/artifacts collected and any named sibling evidence owner>
- Review: <why contract is or is not satisfied>

## Evidence Integrity
- <changed proof surfaces, missing evidence, and credibility impact>

## Responsibility Drift
- <none or concrete drift>

## Review Limitations
- <unreviewed areas and why>

## Guidance Suggestions
- <optional skill/AGENTS.md/setup guidance for orchestrator>

Call end_node with one item per assigned target, then exit immediately.

© Intelligent-Internet, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in zenith/src/zenith_harness/bundled/skills/scrutiny-validator of Intelligent-Internet/zenith.

Open the folder on GitHubat commit a8d9b57

Compare with similar skills

Scrutiny Validator next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Scrutiny Validator compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Scrutiny Validator this skillIntelligent-Internet/zenith337—~1.8kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
MCP Server BuildershareAI-lab/learn-claude-code78k5 repos~1.2kAutomated safety check: PassMIT
MCP Integration for Pluginsanthropics/claude-plugins-official38k11 repos~3.1kAutomated safety check: PassApache-2.0
MemPalace Memory SearchMemPalace/mempalace59k—~1.4kAutomated safety check: PassMIT
Crush Configurationcharmbracelet/crush29k—~3.7kAutomated safety check: PassCustom licence

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • MCP Server Builder

    shareAI-lab/learn-claude-code

    Walks through building MCP servers in Python or TypeScript that expose tools, resources and prompts to Claude, with templates, registration and testing.

    78k GitHub starsUsed in 5 repos~1.2k tokens
    Agent WorkflowsAuto-check passed
  • MCP Integration for Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to bundle Model Context Protocol servers in a Claude Code plugin, covering config files, stdio, SSE, HTTP and WebSocket server types, and authentication.

    38k GitHub starsUsed in 11 repos~3.1k tokens
    Agent WorkflowsAuto-check passed
  • MemPalace Memory Search

    MemPalace/mempalace

    Mines project files and conversation exports into a local, searchable memory palace and recalls past work by semantic search through the mempalace CLI.

    59k GitHub stars~1.4k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • Crush Configuration

    charmbracelet/crush

    Explains how to configure the Crush coding agent with crushrc or crush.json, covering providers, models, LSPs, MCP servers, hooks, permissions and config precedence.

    29k GitHub stars~3.7k tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Context Mode Output Sandbox

    mksglu/context-mode

    Routes large command, file, API and browser output through context-mode tools so only the needed result enters the agent's context, instead of dumping it via Bash.

    26k GitHub stars~4.1k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from Intelligent-Internet/zenith

  • Engineering Mission Playbook

    Intelligent-Internet/zenith

    A skill your agent uses when planning or replanning engineering missions that create, change, port, migrate, integrate, or preserve durable codebase behavior across UI, API, CLI, background jobs…

    337 GitHub stars~8.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Benchmark Validator

    Intelligent-Internet/zenith

    Benchmark validation procedure for one assigned benchmark-related target.

    337 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • User Testing Validator

    Intelligent-Internet/zenith

    Real-surface validation coordinator for engineering validation assignments.

    337 GitHub stars~1.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Agent Browser

    Intelligent-Internet/zenith

    Automates browser and Electron app interactions for user-flow validation.

    337 GitHub stars~6.9k tokensUpdated 1 mo ago
    Auto-check passed
  • Optimization Mission Playbook

    Intelligent-Internet/zenith

    Domain playbook for optimization missions — any task whose goal is to move a metric: performance, latency, throughput, memory, cost, score, quality, compression, ranking, solver, model/eval, and…

    337 GitHub stars~11k tokensUpdated 1 mo ago
    Auto-check passed

Categories

Questions about Scrutiny Validator

What does Scrutiny Validator do?

Adversarial scrutiny procedure for engineering validation assignments. Scrutiny Validator is an agent skill from Intelligent-Internet/zenith. Adversarial scrutiny procedure for engineering validation assignments.

When should I use Scrutiny Validator?

Scrutiny Validator fits situations like: agent Workflows work in your project.

How do I install Scrutiny Validator in Claude Code?

Run `npx skills add Intelligent-Internet/zenith --skill scrutiny-validator -a claude-code`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/scrutiny-validator in Intelligent-Internet/zenith) into .claude/skills/scrutiny-validator in your project. Claude Code loads it when a task matches its description.

How do I install Scrutiny Validator in Codex?

Run `npx skills add Intelligent-Internet/zenith --skill scrutiny-validator -a codex`. Or copy the skill folder (zenith/src/zenith_harness/bundled/skills/scrutiny-validator in Intelligent-Internet/zenith) into .agents/skills/scrutiny-validator in your project. Codex loads it when a task matches its description.

Can I use Scrutiny Validator in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Intelligent-Internet/zenith --skill scrutiny-validator -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/scrutiny-validator, .gemini/skills/scrutiny-validator, .github/skills/scrutiny-validator and .opencode/skills/scrutiny-validator in your project.

What does Scrutiny Validator need to run?

SKILL.md names no scripts, command-line tools or credentials: Scrutiny Validator is instructions for the agent only.

Does Scrutiny Validator access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Scrutiny Validator safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Scrutiny Validator use?

Scrutiny Validator is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Scrutiny Validator use?

About 1.8k tokens (SKILL.md is roughly 7.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Scrutiny Validator?

Skills that share tags, products or a category with Scrutiny Validator: MCP Server Builder (anthropics/skills, 180k stars), MCP Server Builder (shareAI-lab/learn-claude-code, 78k stars), MCP Integration for Plugins (anthropics/claude-plugins-official, 38k stars) and MemPalace Memory Search (MemPalace/mempalace, 59k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Scrutiny Validator?

Intelligent-Internet (a GitHub organization) maintains it in Intelligent-Internet/zenith, which has 337 GitHub stars. The repository holds 6 skills in this directory. The repository was last updated on September 6, 2026.

Source: Intelligent-Internet/zenith on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.