Agent skill

Agent Evaluation

by Prism-Shadow in Prism-Shadow/penguin-harness

Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

Apache-2.0Auto-check passedAgent Workflows

Install Agent Evaluation

skills CLI
$ npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Prism-Shadow/penguin-harness agent-evaluation --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-evaluation .claude/skills/agent-evaluation && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
agent-evaluation
GitHub stars
2.5k
Token cost
~2.5k tokens
SKILL.md length
1,219 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

  • Tasks that involve Agent evaluation and testing
  • SKILL.md covers Before you start, Contract, Prepare and Run and verify, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Agent Evaluation is an agent skill from Prism-Shadow/penguin-harness. Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.

When your agent uses it

  • Tasks that involve Agent evaluation and testing

Example prompts

  • “/agent-evaluation”

What it can do on your machine

Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Agent Evaluation loads about 2.5k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 1,219 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~39
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 1,219 words, ~2,525 tokens.

Download SKILL.mdSave it as .claude/skills/agent-evaluation/SKILL.md (or your agent's skills folder).
name
agent-evaluation
description
Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

Agent Evaluation

Handle one evaluation request from a run_subagent caller: run the specified Test Agent on one Benchmark Case once, score that execution privately, and return one protocol result.

The top-level Benchmark Designer or Optimizer owns all Case and Run loops, concurrency, and follow-up handling. This worker handles no other Case or Run, launches no evaluator or subagent, modifies no Agent or Benchmark, and never writes scoreboard.yaml. Use the Penguin CLI only to launch the specified Test Agent; do not use it to create another phase, designer, optimizer, or evaluator.

Operate silently. Call tools without progress messages. Across all streamed and final responses, the only worker-authored text must be the final plain protocol YAML. Emit no narration, headings, Markdown fences, summaries, private scoring details, or other text.

Before you start

Use this Skill only for a complete request from a run_subagent caller. If the request is incomplete or inconsistent, return invalid_request through the protocol instead of asking the user a question.

Contract

Require exactly one value for every field below:

text
protocol_version: 1
case_id: <case_id>
run: <1_based_run_index>
expected_version: <tested_agent_state_version>
test_agent_id: <test_agent_id>
benchmark_id: <benchmark_id>
provider: <provider>
model_id: <model_id>

One request represents one Test Agent execution. The run value identifies that execution; it is not a repeat count. provider and model_id must both be non-empty and select that exact configured model. If a required field is missing, duplicated, or conflicting, return invalid_request without creating a Workspace or launching the Test Agent.

Return a scored result when the Test Agent ran and the Rubric could be applied. Wrong, malformed, or missing Test Agent output is still a scored result. Return an evaluation failure when the request, Benchmark, launch, version check, Trace binding, or scoring process prevents a valid score.

Resolve the Project, Test Agent, Benchmark, and Case only from the explicit request and Environment App Data Dir. Reject traversal, symlink escape, or any path outside the requested Test Agent and Benchmark. Never read a Project configuration file, credential, or vault.

Prepare

Use the App Data Dir from the Environment:

text
TEST_AGENT_DIR = <app_data_dir>/agents/<test_agent_id>
BENCHMARK_DIR = <app_data_dir>/benchmarks/<benchmark_id>

The Benchmark is Project-level and is not owned by the Test Agent: it sits beside agents/ and may evaluate several Agents. test_agent_id names the Agent this request evaluates; return it as agent_id.

Reject path traversal, symlink escape, or any resolved path outside the requested Test Agent and Benchmark. Inspect only the requested Agent State, Benchmark config and Case, isolated Test Workspace, and Traces needed to verify this execution. Do not inspect another Agent, Project secrets, hidden configuration, or unrelated Workspaces or Traces.

Require agent_state/system_config.yaml, benchmark_config.toml, <case_id>/statement/README.md, and <case_id>/rubric/README.md. Return benchmark_invalid when benchmark_config.toml says status = "failed": a Benchmark whose calibration failed is not evaluated. Treat run only as the caller-owned label for this evaluation and return it unchanged; do not read or validate the total Run count. The top-level Agent State version, defaulting to 1, must equal expected_version; otherwise return version_changed. Read and snapshot model.thinking_level from this Target Agent config, using the normal Agent-config default medium only when the field is absent. This configured value is the evaluation thinking_level; do not require or read thinking metadata from a Trace.

Before launch, snapshot every file under the Case's statement/ and rubric/ directories. Require a usable Rubric whose scoring items total exactly 100 points. Create a unique Workspace under <test_agent_dir>/workspaces/, resolve it to an absolute canonical path, and verify that the resolved path remains under that directory. Copy only statement/ into it. The Test Agent may see the Statement and its own State, but never the Rubric, Gold answers, scoring rules, or Evaluator reasoning.

Show full SKILL.md (649 more words)Show less

Run and verify

Use an existing verified Penguin CLI or repository-local launcher. Do not install or probe a launcher. Snapshot the isolated Workspace and record the existing Trace files.

Resolve PROJECT_DIR, then derive and verify PROJECT_ID, then derive and verify PENGUIN_HOME. Perform these as separate shell statements in this order. Never compress the assignments onto one command line, derive a value before its input exists, or substitute another Penguin home. Before launch, confirm that PROJECT_ID equals the basename of PROJECT_DIR and PENGUIN_HOME equals its dirname.

Start one foreground execution with a fresh top-level Session. With an explicit pair, use:

bash
PROJECT_DIR="<app_data_dir>"   # the App Data Dir value from your Environment section is the project root
PROJECT_ID="$(basename "$PROJECT_DIR")"
PENGUIN_HOME="$(dirname "$PROJECT_DIR")"
export PENGUIN_HOME
penguin run \
  --message "Read README.md in the current Workspace and complete the task exactly as specified there." \
  --provider "<provider>" --model-id "<model_id>" --project-id "$PROJECT_ID" \
  --agent-id "<test_agent_id>" --workspace "<absolute_unique_workspace_path>" \
  --approve allow-all --source benchmark

--source benchmark files the Test Session under the Evaluations folder of the Web App's session list rather than the Test Agent's active conversations; never omit it.

Use the exact requested Agent, Project, absolute Workspace path, and model pair. Never omit either model flag and never fall back to a Project default. If a launch fails, retry only when unchanged Workspace and Trace evidence proves that the Test Agent did not start. Every retry must follow a new diagnosis and apply a specific correction; never repeat an unchanged launch. Do not impose a numeric retry limit while distinct safe repairs remain. Return evaluation_failed when no new repair remains, external configuration is required, or it is unclear whether the Test Agent started.

Verify after the run that the State version, configured model.thinking_level, and both directory snapshots are unchanged. Return version_changed when the State version or configured thinking level differs and benchmark_invalid when the Statement or Rubric differs.

Inspect only new or changed Traces. Bind exactly one root Test Trace whose Workspace, Agent State path, provider, and model match this request. Ignore unrelated parallel Traces and exclude the root Trace's directly referenced child Sessions. Return evaluation_failed if there is no unique match. Read the actual non-empty provider and model_id from the bound root Trace's session_meta; return evaluation_failed if either is unavailable. Use the unchanged Target Agent configuration snapshot—not Trace metadata—for thinking_level.

Score

Inspect only the isolated Workspace, the bound root Trace, its directly referenced child Traces, and the private Rubric. Apply every scoring item and allowed equivalent. Keep Rubric contents, Gold answers, per-item scoring, and scoring rationale private.

A wrong answer, missing artifact, malformed output, or task failure attributable to the Test Agent is scored behavior and returns status: ok. A launcher, Trace-binding, or Evaluator failure is not scored. Return benchmark_invalid when the Rubric cannot be applied and evaluation_failed when the score is non-finite or outside 0..100.

Set duration_ms from the root Test Session. Compute cost only from reliable final cumulative usage or cost already recorded in that Session and directly referenced child Traces found in the same bounded pass. Never browse, query a pricing service, or infer cost from external model prices. If the required data is unavailable, return cost: null. Missing cost data must not invalidate a score.

Round score to two decimal places. Preserve a non-null cost at the precision recorded in the Trace; do not round it. Write duration_ms as a non-negative integer rounded to the nearest millisecond.

Return

Return the required YAML as the only worker-authored text. Do not wrap it in backticks or a Markdown fence.

If the caller reports that your response formatting was invalid, use the scored or failed result already present in this Session and resend only the clean protocol YAML. Do not call tools, relaunch the Test Agent, rescore, or add an explanation.

For a scored result:

text
protocol_version: 1
status: ok
case_id: <case_id>
run: <run>
agent_id: <test_agent_id>
expected_version: <version>
provider: <actual_provider>
model_id: <actual_model_id>
thinking_level: <configured_thinking_level>
score: <0_to_100>
cost: <number_or_null>
duration_ms: <non_negative_integer>
session_id: <test_session_id>

For an evaluation failure, use null for an identity field that was missing or conflicting:

text
protocol_version: 1
status: failed
case_id: <case_id_or_null>
run: <run_or_null>
agent_id: <test_agent_id_or_null>
expected_version: <version_or_null>
provider: <provider_or_null>
model_id: <model_id_or_null>
thinking_level: <thinking_level_or_null>
failure_code: <stable_failure_code>

Use four failure codes:

  • invalid_request: the request is incomplete or inconsistent.
  • benchmark_invalid: the Statement, Rubric, or scoring contract is invalid.
  • version_changed: the Test Agent version does not match the request or changed during evaluation.
  • evaluation_failed: launch could not be safely repaired, or Trace binding or scoring failed.

Never include score, cost, duration, Session id, private data, or optimization advice on failure.

© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/agent-tuning/skills/agent-evaluation of Prism-Shadow/penguin-harness.

Open the folder on GitHubat commit d56d9ce

Compare with similar skills

Agent Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Agent Evaluation compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Agent Evaluation this skillPrism-Shadow/penguin-harness2.5k—~2.5kAutomated safety check: PassApache-2.0
MCP Server Builderanthropics/skills180k64 repos~2.3kAutomated safety check: PassApache-2.0
Diagnosing Superpowers Sessionsobra/superpowers296k3 repos~1.7kAutomated safety check: PassMIT
Darwin Skill Optimizeralchaincyf/darwin-skill6.2k1 repos~4.7kAutomated safety check: PassMIT
Skill Release Gaterohitg00/ai-engineering-from-scratch66k—~1kAutomated safety check: PassMIT
CodeGraph Agent Evalcolbymchenry/codegraph73k—~950Automated safety check: PassMIT

Similar skills

  • MCP Server Builder

    anthropics/skills

    Official

    Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.

    180k GitHub starsUsed in 64 repos~2.3k tokens
    Agent WorkflowsAuto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed
  • Darwin Skill Optimizer

    alchaincyf/darwin-skill

    Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.

    6.2k GitHub starsUsed in 1 repo~4.7k tokens
    Agent WorkflowsAuto-check passed
  • Skill Release Gate

    rohitg00/ai-engineering-from-scratch

    Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.

    66k GitHub stars~1k tokensUpdated yesterday
    Agent WorkflowsAuto-check passed
  • CodeGraph Agent Eval

    colbymchenry/codegraph

    Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.

    73k GitHub stars~950 tokensUpdated today
    Agent WorkflowsAuto-check passed
  • Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.

    23k GitHub stars~3.4k tokensUpdated today
    Agent WorkflowsAuto-check passed

More from Prism-Shadow/penguin-harness

All 31 skills in this repo
  • A2ui

    Prism-Shadow/penguin-harness

    Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…

    2.5k GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Penguin Harness Dev

    Prism-Shadow/penguin-harness

    A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…

    2.5k GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Bento Slides

    Prism-Shadow/penguin-harness

    Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.

    2.5k GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Penguin Harness Manual Test

    Prism-Shadow/penguin-harness

    A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…

    2.5k GitHub stars~1.3k tokensUpdated today
    Auto-check passed
  • Penguin Harness Frontend

    Prism-Shadow/penguin-harness

    A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…

    2.5k GitHub stars~6.4k tokensUpdated today
    Auto-check passed
  • Browser Automation

    Prism-Shadow/penguin-harness

    Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…

    2.5k GitHub stars~2.9k tokensUpdated today
    Auto-check: warnings

Categories

Questions about Agent Evaluation

What does Agent Evaluation do?

Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result. Agent Evaluation is an agent skill from Prism-Shadow/penguin-harness. Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.

When should I use Agent Evaluation?

Agent Evaluation fits situations like: tasks that involve Agent evaluation and testing.

How do I install Agent Evaluation in Claude Code?

Run `npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a claude-code`. Or copy the skill folder (plugins/agent-tuning/skills/agent-evaluation in Prism-Shadow/penguin-harness) into .claude/skills/agent-evaluation in your project. Claude Code loads it when a task matches its description.

How do I install Agent Evaluation in Codex?

Run `npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a codex`. Or copy the skill folder (plugins/agent-tuning/skills/agent-evaluation in Prism-Shadow/penguin-harness) into .agents/skills/agent-evaluation in your project. Codex loads it when a task matches its description.

Can I use Agent Evaluation in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-evaluation, .gemini/skills/agent-evaluation, .github/skills/agent-evaluation and .opencode/skills/agent-evaluation in your project.

What does Agent Evaluation need to run?

SKILL.md names no scripts, command-line tools or credentials: Agent Evaluation is instructions for the agent only.

Does Agent Evaluation access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Agent Evaluation safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Agent Evaluation use?

Agent Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Agent Evaluation use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Agent Evaluation?

Skills that share tags, products or a category with Agent Evaluation: MCP Server Builder (anthropics/skills, 180k stars), Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Skill Release Gate (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Agent Evaluation?

Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,455 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.