Agent skill

Output Validation

by benchflow-ai in benchflow-ai/skillsbench

Local self-check of instructions and mask outputs (format/range/consistency) without using GT.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Output Validation

skills CLI
$ npx skills add benchflow-ai/skillsbench --skill output-validation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install benchflow-ai/skillsbench output-validation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/benchflow-ai/skillsbench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tasks/dynamic-object-aware-egomotion/environment/skills/output-validation .claude/skills/output-validation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
output-validation
GitHub stars
1.8k
Token cost
~462 tokens
SKILL.md length
103 words
Files
1
Skills in repo
180
Repo updated
First seen
Licence
Apache-2.0

At a glance

Local self-check of instructions and mask outputs (format/range/consistency) without using GT.

  • Tasks that involve LLM guardrails
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Tasks that involve Verification before completion

What it does

Output Validation is an agent skill from benchflow-ai/skillsbench. Local self-check of instructions and mask outputs (format/range/consistency) without using GT.

Its SKILL.md is about 460 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM guardrails and Verification before completion. The repository describes itself as: SkillsBench evaluates how well skills work and how effective agents are at using them. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve LLM guardrails
  • Tasks that involve Verification before completion

Example prompts

  • “/output-validation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 9a1f4dd. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Output Validation loads about 462 tokens when it runs. Until then it costs about 28 tokens; SKILL.md has 103 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~28
When it runs · the whole SKILL.md, loaded when a task matches
~462

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from benchflow-ai/skillsbench at commit 9a1f4dd, republished under its Apache-2.0 licence (© benchflow-ai). 103 words, ~462 tokens.

Download SKILL.mdSave it as .claude/skills/output-validation/SKILL.md (or your agent's skills folder).
name
output-validation
description
Local self-check of instructions and mask outputs (format/range/consistency) without using GT.

When to use

  • After generating your outputs (interval instructions, masks, etc.), before submission/hand-off.

Checks

  • Key format: every key is "{start}->{end}", integers only, start<=end.
  • Coverage: max frame index ≤ video total-1; consistent with your sampling policy.
  • Frame count: NPZ f_{i}_* count equals sampled frame count; no gaps or missing components.
  • CSR integrity: each frame has data/indices/indptr; len(indptr)==H+1; indptr[-1]==indices.size; indices within [0,W).
  • Value validity: JSON values are non-empty string lists; labels in the allowed set.

Reference snippet

python
import json, numpy as np, cv2
VIDEO_PATH = "<path/to/video>"
INSTRUCTIONS_PATH = "<path/to/interval_instructions.json>"
MASKS_PATH = "<path/to/masks.npz>"
cap=cv2.VideoCapture(VIDEO_PATH)
n=int(cap.get(cv2.CAP_PROP_FRAME_COUNT)); H=int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)); W=int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
j=json.load(open(INSTRUCTIONS_PATH))
npz=np.load(MASKS_PATH)
for k,v in j.items():
    s,e=k.split("->"); assert s.isdigit() and e.isdigit()
    s=int(s); e=int(e); assert 0<=s<=e<n
    for lbl in v: assert isinstance(lbl,str)
frames=0
while f"f_{frames}_data" in npz: frames+=1
assert frames>0
assert npz["shape"][0]==H and npz["shape"][1]==W
indptr=npz["f_0_indptr"]; indices=npz["f_0_indices"]
assert indptr.shape[0]==H+1 and indptr[-1]==indices.size
assert indices.size==0 or (indices.min()>=0 and indices.max()<W)

Self-check list

  • JSON keys/values pass format checks.
  • Max frame index within video range and near sampled max.
  • NPZ frame count matches sampling; keys consecutive.
  • CSR structure and shape validated.

© benchflow-ai, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in tasks/dynamic-object-aware-egomotion/environment/skills/output-validation of benchflow-ai/skillsbench.

Open the folder on GitHubat commit 9a1f4dd

Compare with similar skills

Output Validation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Output Validation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Output Validation this skillbenchflow-ai/skillsbench1.8k—~462Automated safety check: PassApache-2.0
Formax Skill Captureyusifeng/formax195—~530Automated safety check: PassMIT
Fable Modemrtooher/fable-mode870—~1kAutomated safety check: PassNone
Mcaf ML AI Deliverymanagedcode/Storage138—~1kAutomated safety check: PassMIT
Gate MCP Skillholon-run/uxc115—~765Automated safety check: PassMIT
Hive MCP Skillholon-run/uxc115—~971Automated safety check: PassMIT

Similar skills

  • Formax Skill Capture

    yusifeng/formax

    A skill your agent uses when we want to turn a just-finished Formax workflow (e.g.

    195 GitHub stars~530 tokensUpdated 2 mo ago
    AI & LLM EngineeringAuto-check passed
  • Fable Mode

    mrtooher/fable-mode

    Enforces staged execution discipline on large tasks: a written stage plan, delegation to named fable agents where the runtime supports it, a failable verification check at each stage, and a…

    870 GitHub stars~1k tokensUpdated 1 mo ago
    AI & LLM EngineeringAuto-check passed
  • Mcaf ML AI Delivery

    managedcode/Storage

    Apply ML/AI project delivery guidance for data exploration, feasibility, experimentation, testing, responsible AI, and operating ML systems.

    138 GitHub stars~1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Gate MCP Skill

    holon-run/uxc

    Use Gate MCP through UXC for public spot and futures market data workflows with a fixed streamable-http endpoint and read-first guardrails.

    115 GitHub stars~765 tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Hive MCP Skill

    holon-run/uxc

    Use Hive Intelligence MCP through UXC for broad crypto market, onchain, portfolio, and risk workflows with help-first discovery and convenience-layer guardrails.

    115 GitHub stars~971 tokensUpdated 22 days ago
    AI & LLM EngineeringAuto-check passed
  • Delegation Brief

    mohitagw15856/pm-claude-skills

    Delegate so the work comes back right the first time — the brief that transfers outcome, context, and constraints (not just the task), the autonomy level stated explicitly, and the check-in design…

    1.4k GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed

More from benchflow-ai/skillsbench

All 180 skills in this repo
  • Lean4 Memories

    benchflow-ai/skillsbench

    This skill should be used when working on Lean 4 formalization projects to maintain persistent memory of successful proof patterns, failed approaches, project conventions, and user preferences…

    1.8k GitHub stars~3.2k tokensUpdated 2 mo ago
    Auto-check passed
  • Senior Data Engineer

    benchflow-ai/skillsbench

    World-class data engineering skill for building scalable data pipelines, ETL/ELT systems, real-time streaming, and data infrastructure.

    1.8k GitHub stars~5.9k tokensUpdated 2 mo ago
    Auto-check passed
  • Ac Branch Pi Model

    benchflow-ai/skillsbench

    AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching acopf-math-model.md and MATPOWER branch fields.

    1.8k GitHub stars~1.1k tokensUpdated 2 mo ago
    Auto-check passed
  • Civ6lib

    benchflow-ai/skillsbench

    Civilization 6 district mechanics library. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~1.7k tokensUpdated 2 mo ago
    Auto-check passed
  • D3 Visualization

    benchflow-ai/skillsbench

    Build deterministic, verifiable data visualizations with D3.js (v6).

    1.8k GitHub stars~1.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Dc Power Flow

    benchflow-ai/skillsbench

    DC power flow analysis for power systems. An agent skill from benchflow-ai/skillsbench.

    1.8k GitHub stars~717 tokensUpdated 2 mo ago
    Auto-check passed

Questions about Output Validation

What does Output Validation do?

Local self-check of instructions and mask outputs (format/range/consistency) without using GT. Output Validation is an agent skill from benchflow-ai/skillsbench. Local self-check of instructions and mask outputs (format/range/consistency) without using GT.

When should I use Output Validation?

Output Validation fits situations like: tasks that involve LLM guardrails; tasks that involve Verification before completion.

How do I install Output Validation in Claude Code?

Run `npx skills add benchflow-ai/skillsbench --skill output-validation -a claude-code`. Or copy the skill folder (tasks/dynamic-object-aware-egomotion/environment/skills/output-validation in benchflow-ai/skillsbench) into .claude/skills/output-validation in your project. Claude Code loads it when a task matches its description.

How do I install Output Validation in Codex?

Run `npx skills add benchflow-ai/skillsbench --skill output-validation -a codex`. Or copy the skill folder (tasks/dynamic-object-aware-egomotion/environment/skills/output-validation in benchflow-ai/skillsbench) into .agents/skills/output-validation in your project. Codex loads it when a task matches its description.

Can I use Output Validation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add benchflow-ai/skillsbench --skill output-validation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/output-validation, .gemini/skills/output-validation, .github/skills/output-validation and .opencode/skills/output-validation in your project.

What does Output Validation need to run?

SKILL.md names no scripts, command-line tools or credentials: Output Validation is instructions for the agent only. Our summary lists: Python 3.

Does Output Validation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Output Validation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Output Validation use?

Output Validation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Output Validation use?

About 462 tokens (SKILL.md is roughly 1.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Output Validation?

Skills that share tags, products or a category with Output Validation: Formax Skill Capture (yusifeng/formax, 195 stars), Fable Mode (mrtooher/fable-mode, 870 stars), Mcaf ML AI Delivery (managedcode/Storage, 138 stars) and Gate MCP Skill (holon-run/uxc, 115 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Output Validation?

benchflow-ai (a GitHub organization) maintains it in benchflow-ai/skillsbench, which has 1,832 GitHub stars. The repository holds 180 skills in this directory. The repository was last updated on July 23, 2026.

Source: benchflow-ai/skillsbench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.