Agent skill

Skill Validation And Evaluation

by BingHanOfUESTC in BingHanOfUESTC/open_agent_team

A skill your agent uses when validating a newly created or updated agent skill with structure checks, routing positive and negative examples, contamination checks, progressive disclosure review, and…

MITAuto-check passedAgent Workflows

Install Skill Validation And Evaluation

skills CLI
$ npx skills add BingHanOfUESTC/open_agent_team --skill skill-validation-and-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install BingHanOfUESTC/open_agent_team skill-validation-and-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/BingHanOfUESTC/open_agent_team.git skills-src && mkdir -p .claude/skills && cp -r skills-src/teams/skill_creator_team/skills/skill-validation-and-evaluation .claude/skills/skill-validation-and-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
skill-validation-and-evaluation
GitHub stars
106
Token cost
~303 tokens
SKILL.md length
20 words
Files
2 (incl. scripts)
Skills in repo
30
Repo updated
First seen
Licence
MIT

At a glance

A skill your agent uses when validating a newly created or updated agent skill with structure checks, routing positive and negative examples, contamination checks, progressive disclosure review, and…

  • Validating a newly created
  • SKILL.md covers Required Validation, Validation Cases Format, Scoring and Script
  • Runs Python scripts from its folder
  • Updated agent skill with structure checks

What it does

Skill Validation And Evaluation is an agent skill from BingHanOfUESTC/open_agent_team. Use this skill when validating a newly created or updated agent skill with structure checks, routing positive and negative examples, contamination checks, progressive disclosure review, and quality scoring before delivery.

Its SKILL.md is about 300 tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files, including scripts (for example `scripts/validate_skill.py`).

It sits in Agent Workflows. The repository describes itself as: Build persistent multi-agent teams that collaborate like real organizations to deliver complex tasks. The licence is MIT.

When your agent uses it

  • Validating a newly created
  • Updated agent skill with structure checks
  • Routing positive and negative examples
  • Contamination checks

Example prompts

  • “/skill-validation-and-evaluation”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 7e28736. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 1 file in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Skill Validation And Evaluation loads about 303 tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 20 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~303

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from BingHanOfUESTC/open_agent_team at commit 7e28736, republished under its MIT licence (© BingHanOfUESTC). 20 words, ~303 tokens.

Download SKILL.mdSave it as .claude/skills/skill-validation-and-evaluation/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
skill-validation-and-evaluation
description
Use this skill when validating a newly created or updated agent skill with structure checks, routing positive and negative examples, contamination checks, progressive disclosure review, and quality scoring before delivery.

Skill Validation And Evaluation

Required Validation

text
Structure check: SKILL.md exists, has YAML frontmatter, name, description.
Routing check: 2+ positive examples should trigger the skill.
Negative routing check: 1+ example should not trigger the skill.
Content check: workflow, inputs, outputs, gotchas, resources.
Contamination check: no copied expected answer, private data, or one-off conclusion.
Progressive disclosure check: no unnecessary long content in SKILL.md.

Validation Cases Format

markdown
# Validation Cases

## Positive Case 1
User request:
Expected trigger reason:

## Positive Case 2
User request:
Expected trigger reason:

## Negative Case
User request:
Why it should not trigger:

Scoring

Use 0-10:

text
8.5-10: pass
7.0-8.4: conditional_pass with required fixes
<7.0: fail

Script

Use scripts/validate_skill.py <skill_dir> for basic structure validation.

© BingHanOfUESTC, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file (scripts) in teams/skill_creator_team/skills/skill-validation-and-evaluation of BingHanOfUESTC/open_agent_team.

  • SKILL.md
  • scripts/validate_skill.py

Open the folder on GitHubat commit 7e28736

Compare with similar skills

Skill Validation And Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Skill Validation And Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Skill Validation And Evaluation this skillBingHanOfUESTC/open_agent_team106—~303Automated safety check: PassMIT
MCP Server Builderanthropics/skills180k63 repos~2.3kAutomated safety check: PassApache-2.0
Hook Development for Claude Code Pluginsanthropics/claude-plugins-official38k10 repos~4.1kAutomated safety check: NotesApache-2.0
Using Superpowersfarm-fe/farm5.6k35 repos~1.4kAutomated safety check: PassMIT
Executing Plans Inlineobra/superpowers297k2 repos~5.1kAutomated safety check: PassMIT
Skill CreatorAzure/azqr79689 repos~8.2kAutomated safety check: PassApache-2.0

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 63 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Hook Development for Claude Code Plugins

    anthropics/claude-plugins-official

    Official

    Explains how to write Claude Code plugin hooks, both prompt-based checks and bash commands, for events such as PreToolUse, Stop and SessionStart.

    38k GitHub starsUsed in 10 repos~4.1k tokens
    Agent WorkflowsAuto-check: notes
  • Using Superpowers

    farm-fe/farm

    A skill your agent uses when starting any conversation - establishes how to find and use skills, requiring Skill tool invocation before ANY response including clarifying questions

    5.6k GitHub starsUsed in 35 repos~1.4k tokens
    Agent WorkflowsAuto-check passed
  • Executing Plans Inline

    obra/superpowers

    Has the agent carry out an implementation plan itself, task by task in the current session, keeping a ledger, proving each step with a test and ending with one whole-branch review.

    297k GitHub starsUsed in 2 repos~5.1k tokens
    Agent WorkflowsAuto-check passed
  • Skill Creator

    Azure/azqr

    Official

    Create new skills, modify and improve existing skills, and measure skill performance.

    796 GitHub starsUsed in 89 repos~8.2k tokens
    Agent WorkflowsAuto-check passed
  • Claude Code Agent Development

    anthropics/claude-plugins-official

    Official

    Explains how to write agents for Claude Code plugins: the markdown file with YAML frontmatter, trigger descriptions, model and color settings, and system prompt design.

    38k GitHub starsUsed in 7 repos~2.8k tokens
    Agent WorkflowsAuto-check passed

More from BingHanOfUESTC/open_agent_team

All 30 skills in this repo
  • Autoresearch Orchestration

    BingHanOfUESTC/open_agent_team

    A skill your agent uses when autoresearchteam needs to manage an end-to-end research lifecycle from Boss brief to literature, idea, implementation, experiments, synthesis, and LaTeX delivery.

    106 GitHub stars~691 tokensUpdated 3 mo ago
    Auto-check passed
  • Experiment Iteration Loop

    BingHanOfUESTC/open_agent_team

    A skill your agent uses when designing, running, logging, diagnosing, or iterating experiments for the selected research idea.

    106 GitHub stars~446 tokensUpdated 3 mo ago
    Auto-check passed
  • Latex Paper Artifact

    BingHanOfUESTC/open_agent_team

    A skill your agent uses when producing the final arXiv-style LaTeX paper, BibTeX, figures, tables, reproducibility statement, appendix, and artifact manifest.

    106 GitHub stars~635 tokensUpdated 3 mo ago
    Auto-check passed
  • Literature Evidence Mapping

    BingHanOfUESTC/open_agent_team

    A skill your agent uses when discovering papers, mapping related work, reading papers deeply, checking citations, or building evidence cards for an auto research task.

    106 GitHub stars~949 tokensUpdated 3 mo ago
    Auto-check passed
  • Media Asset Sourcing

    BingHanOfUESTC/open_agent_team

    A skill your agent uses when PPT slides need external images, screenshots, logos, product photos, stock-like visuals, diagrams, or videos that must be downloaded, inserted, or linked with source and…

    106 GitHub stars~819 tokensUpdated 3 mo ago
    Auto-check passed
  • Originality Safety Guard

    BingHanOfUESTC/open_agent_team

    A skill your agent uses when checking creative fiction for originality, name safety, public figure misuse, historical figure misuse, news-event overfitting, recognizable plot borrowing, IP-like…

    106 GitHub stars~653 tokensUpdated 3 mo ago
    Auto-check passed

Categories

Questions about Skill Validation And Evaluation

What does Skill Validation And Evaluation do?

A skill your agent uses when validating a newly created or updated agent skill with structure checks, routing positive and negative examples, contamination checks, progressive disclosure review, and…. Skill Validation And Evaluation is an agent skill from BingHanOfUESTC/open_agent_team. Use this skill when validating a newly created or updated agent skill with structure checks, routing positive and negative examples, contamination checks, progressive disclosure review, and quality scoring before delivery.

When should I use Skill Validation And Evaluation?

Skill Validation And Evaluation fits situations like: validating a newly created; updated agent skill with structure checks; routing positive and negative examples; contamination checks.

How do I install Skill Validation And Evaluation in Claude Code?

Run `npx skills add BingHanOfUESTC/open_agent_team --skill skill-validation-and-evaluation -a claude-code`. Or copy the skill folder (teams/skill_creator_team/skills/skill-validation-and-evaluation in BingHanOfUESTC/open_agent_team) into .claude/skills/skill-validation-and-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Skill Validation And Evaluation in Codex?

Run `npx skills add BingHanOfUESTC/open_agent_team --skill skill-validation-and-evaluation -a codex`. Or copy the skill folder (teams/skill_creator_team/skills/skill-validation-and-evaluation in BingHanOfUESTC/open_agent_team) into .agents/skills/skill-validation-and-evaluation in your project. Codex loads it when a task matches its description.

Can I use Skill Validation And Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add BingHanOfUESTC/open_agent_team --skill skill-validation-and-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/skill-validation-and-evaluation, .gemini/skills/skill-validation-and-evaluation, .github/skills/skill-validation-and-evaluation and .opencode/skills/skill-validation-and-evaluation in your project.

What does Skill Validation And Evaluation need to run?

Going by SKILL.md and its folder, Skill Validation And Evaluation needs Python for the scripts in its folder. Our summary lists: Python 3.

Does Skill Validation And Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Skill Validation And Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Skill Validation And Evaluation use?

Skill Validation And Evaluation is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Skill Validation And Evaluation use?

About 303 tokens (SKILL.md is roughly 1.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Skill Validation And Evaluation?

Skills that share tags, products or a category with Skill Validation And Evaluation: MCP Server Builder (anthropics/skills, 180k stars), Hook Development for Claude Code Plugins (anthropics/claude-plugins-official, 38k stars), Using Superpowers (farm-fe/farm, 5.6k stars) and Executing Plans Inline (obra/superpowers, 297k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Skill Validation And Evaluation?

BingHanOfUESTC (a GitHub user) maintains it in BingHanOfUESTC/open_agent_team, which has 106 GitHub stars. The repository holds 30 skills in this directory. The repository was last updated on June 23, 2026.

Source: BingHanOfUESTC/open_agent_team on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.