MCP Server Builder
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-evaluation --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-evaluation .claude/skills/agent-evaluation && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-evaluation" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluation into .claude/skills/agent-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-evaluation", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-evaluation --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-evaluation .agents/skills/agent-evaluation && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-evaluation" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluation into .agents/skills/agent-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-evaluation", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-evaluation --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-evaluation .cursor/skills/agent-evaluation && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-evaluation" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluation into .cursor/skills/agent-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-evaluation", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Prism-Shadow/penguin-harness.git --path plugins/agent-tuning/skills/agent-evaluation--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-evaluation --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-evaluation .gemini/skills/agent-evaluation && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-evaluation" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluation into .gemini/skills/agent-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-evaluation", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Prism-Shadow/penguin-harness agent-evaluationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-evaluation .github/skills/agent-evaluation && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-evaluation" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluation into .github/skills/agent-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-evaluation", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-evaluation --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-evaluation .opencode/skills/agent-evaluation && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-evaluation" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-evaluation into .opencode/skills/agent-evaluation/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-evaluation", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-evaluationRun one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.
Agent Evaluation is an agent skill from Prism-Shadow/penguin-harness. Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.
Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Agent Workflows, covering Agent evaluation and testing. The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Evaluation loads about 2.5k tokens when it runs. Until then it costs about 39 tokens; SKILL.md has 1,219 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 1,219 words, ~2,525 tokens.
.claude/skills/agent-evaluation/SKILL.md (or your agent's skills folder).Handle one evaluation request from a run_subagent caller: run the specified Test Agent on one Benchmark Case once, score that execution privately, and return one protocol result.
The top-level Benchmark Designer or Optimizer owns all Case and Run loops, concurrency, and follow-up handling. This worker handles no other Case or Run, launches no evaluator or subagent, modifies no Agent or Benchmark, and never writes scoreboard.yaml. Use the Penguin CLI only to launch the specified Test Agent; do not use it to create another phase, designer, optimizer, or evaluator.
Operate silently. Call tools without progress messages. Across all streamed and final responses, the only worker-authored text must be the final plain protocol YAML. Emit no narration, headings, Markdown fences, summaries, private scoring details, or other text.
Use this Skill only for a complete request from a run_subagent caller. If the request is incomplete or inconsistent, return invalid_request through the protocol instead of asking the user a question.
Require exactly one value for every field below:
protocol_version: 1
case_id: <case_id>
run: <1_based_run_index>
expected_version: <tested_agent_state_version>
test_agent_id: <test_agent_id>
benchmark_id: <benchmark_id>
provider: <provider>
model_id: <model_id>One request represents one Test Agent execution. The run value identifies that execution; it is not a repeat count. provider and model_id must both be non-empty and select that exact configured model. If a required field is missing, duplicated, or conflicting, return invalid_request without creating a Workspace or launching the Test Agent.
Return a scored result when the Test Agent ran and the Rubric could be applied. Wrong, malformed, or missing Test Agent output is still a scored result. Return an evaluation failure when the request, Benchmark, launch, version check, Trace binding, or scoring process prevents a valid score.
Resolve the Project, Test Agent, Benchmark, and Case only from the explicit request and Environment App Data Dir. Reject traversal, symlink escape, or any path outside the requested Test Agent and Benchmark. Never read a Project configuration file, credential, or vault.
Use the App Data Dir from the Environment:
TEST_AGENT_DIR = <app_data_dir>/agents/<test_agent_id>
BENCHMARK_DIR = <app_data_dir>/benchmarks/<benchmark_id>The Benchmark is Project-level and is not owned by the Test Agent: it sits beside agents/ and may evaluate several Agents. test_agent_id names the Agent this request evaluates; return it as agent_id.
Reject path traversal, symlink escape, or any resolved path outside the requested Test Agent and Benchmark. Inspect only the requested Agent State, Benchmark config and Case, isolated Test Workspace, and Traces needed to verify this execution. Do not inspect another Agent, Project secrets, hidden configuration, or unrelated Workspaces or Traces.
Require agent_state/system_config.yaml, benchmark_config.toml, <case_id>/statement/README.md, and <case_id>/rubric/README.md. Return benchmark_invalid when benchmark_config.toml says status = "failed": a Benchmark whose calibration failed is not evaluated. Treat run only as the caller-owned label for this evaluation and return it unchanged; do not read or validate the total Run count. The top-level Agent State version, defaulting to 1, must equal expected_version; otherwise return version_changed. Read and snapshot model.thinking_level from this Target Agent config, using the normal Agent-config default medium only when the field is absent. This configured value is the evaluation thinking_level; do not require or read thinking metadata from a Trace.
Before launch, snapshot every file under the Case's statement/ and rubric/ directories. Require a usable Rubric whose scoring items total exactly 100 points. Create a unique Workspace under <test_agent_dir>/workspaces/, resolve it to an absolute canonical path, and verify that the resolved path remains under that directory. Copy only statement/ into it. The Test Agent may see the Statement and its own State, but never the Rubric, Gold answers, scoring rules, or Evaluator reasoning.
Use an existing verified Penguin CLI or repository-local launcher. Do not install or probe a launcher. Snapshot the isolated Workspace and record the existing Trace files.
Resolve PROJECT_DIR, then derive and verify PROJECT_ID, then derive and verify PENGUIN_HOME. Perform these as separate shell statements in this order. Never compress the assignments onto one command line, derive a value before its input exists, or substitute another Penguin home. Before launch, confirm that PROJECT_ID equals the basename of PROJECT_DIR and PENGUIN_HOME equals its dirname.
Start one foreground execution with a fresh top-level Session. With an explicit pair, use:
PROJECT_DIR="<app_data_dir>" # the App Data Dir value from your Environment section is the project root
PROJECT_ID="$(basename "$PROJECT_DIR")"
PENGUIN_HOME="$(dirname "$PROJECT_DIR")"
export PENGUIN_HOME
penguin run \
--message "Read README.md in the current Workspace and complete the task exactly as specified there." \
--provider "<provider>" --model-id "<model_id>" --project-id "$PROJECT_ID" \
--agent-id "<test_agent_id>" --workspace "<absolute_unique_workspace_path>" \
--approve allow-all --source benchmark--source benchmark files the Test Session under the Evaluations folder of the Web App's session list rather than the Test Agent's active conversations; never omit it.
Use the exact requested Agent, Project, absolute Workspace path, and model pair. Never omit either model flag and never fall back to a Project default. If a launch fails, retry only when unchanged Workspace and Trace evidence proves that the Test Agent did not start. Every retry must follow a new diagnosis and apply a specific correction; never repeat an unchanged launch. Do not impose a numeric retry limit while distinct safe repairs remain. Return evaluation_failed when no new repair remains, external configuration is required, or it is unclear whether the Test Agent started.
Verify after the run that the State version, configured model.thinking_level, and both directory snapshots are unchanged. Return version_changed when the State version or configured thinking level differs and benchmark_invalid when the Statement or Rubric differs.
Inspect only new or changed Traces. Bind exactly one root Test Trace whose Workspace, Agent State path, provider, and model match this request. Ignore unrelated parallel Traces and exclude the root Trace's directly referenced child Sessions. Return evaluation_failed if there is no unique match. Read the actual non-empty provider and model_id from the bound root Trace's session_meta; return evaluation_failed if either is unavailable. Use the unchanged Target Agent configuration snapshot—not Trace metadata—for thinking_level.
Inspect only the isolated Workspace, the bound root Trace, its directly referenced child Traces, and the private Rubric. Apply every scoring item and allowed equivalent. Keep Rubric contents, Gold answers, per-item scoring, and scoring rationale private.
A wrong answer, missing artifact, malformed output, or task failure attributable to the Test Agent is scored behavior and returns status: ok. A launcher, Trace-binding, or Evaluator failure is not scored. Return benchmark_invalid when the Rubric cannot be applied and evaluation_failed when the score is non-finite or outside 0..100.
Set duration_ms from the root Test Session. Compute cost only from reliable final cumulative usage or cost already recorded in that Session and directly referenced child Traces found in the same bounded pass. Never browse, query a pricing service, or infer cost from external model prices. If the required data is unavailable, return cost: null. Missing cost data must not invalidate a score.
Round score to two decimal places. Preserve a non-null cost at the precision recorded in the Trace; do not round it. Write duration_ms as a non-negative integer rounded to the nearest millisecond.
Return the required YAML as the only worker-authored text. Do not wrap it in backticks or a Markdown fence.
If the caller reports that your response formatting was invalid, use the scored or failed result already present in this Session and resend only the clean protocol YAML. Do not call tools, relaunch the Test Agent, rescore, or add an explanation.
For a scored result:
protocol_version: 1
status: ok
case_id: <case_id>
run: <run>
agent_id: <test_agent_id>
expected_version: <version>
provider: <actual_provider>
model_id: <actual_model_id>
thinking_level: <configured_thinking_level>
score: <0_to_100>
cost: <number_or_null>
duration_ms: <non_negative_integer>
session_id: <test_session_id>For an evaluation failure, use null for an identity field that was missing or conflicting:
protocol_version: 1
status: failed
case_id: <case_id_or_null>
run: <run_or_null>
agent_id: <test_agent_id_or_null>
expected_version: <version_or_null>
provider: <provider_or_null>
model_id: <model_id_or_null>
thinking_level: <thinking_level_or_null>
failure_code: <stable_failure_code>Use four failure codes:
invalid_request: the request is incomplete or inconsistent.benchmark_invalid: the Statement, Rubric, or scoring contract is invalid.version_changed: the Test Agent version does not match the request or changed during evaluation.evaluation_failed: launch could not be safely repaired, or Trace binding or scoring failed.Never include score, cost, duration, Session id, private data, or optimization advice on failure.
© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/agent-tuning/skills/agent-evaluation of Prism-Shadow/penguin-harness.
Open the folder on GitHubat commit d56d9ce
Agent Evaluation next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Evaluation this skillPrism-Shadow/penguin-harness | 2.5k | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | |
| MCP Server Builderanthropics/skills | 180k | 64 repos | ~2.3k | Automated safety check: Pass | Apache-2.0 | |
| Diagnosing Superpowers Sessionsobra/superpowers | 296k | 3 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Darwin Skill Optimizeralchaincyf/darwin-skill | 6.2k | 1 repos | ~4.7k | Automated safety check: Pass | MIT | |
| Skill Release Gaterohitg00/ai-engineering-from-scratch | 66k | — | ~1k | Automated safety check: Pass | MIT | |
| CodeGraph Agent Evalcolbymchenry/codegraph | 73k | — | ~950 | Automated safety check: Pass | MIT |
anthropics/skills
Guides the design and implementation of Model Context Protocol servers in TypeScript or Python, from tool naming and error messages to evaluation.
obra/superpowers
Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.
alchaincyf/darwin-skill
Scores SKILL.md files on a nine-dimension rubric, then improves them in a keep-or-revert loop with independent judge agents, test prompts, git history and human checkpoints.
rohitg00/ai-engineering-from-scratch
Evaluates an Agent Skill bundle before release for structure, trigger quality, artifact improvement, script correctness, safety, installed-tree integrity and host portability.
colbymchenry/codegraph
Benchmarks how much CodeGraph helps a coding agent on a real repository, comparing runs with and without it for a chosen local or published version.
dotnet/maui
Mines local Copilot CLI session logs for dotnet/maui to rank costly or failing runs, tag recurring failure modes, propose repo edits and emit guard evals.
Prism-Shadow/penguin-harness
Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…
Prism-Shadow/penguin-harness
A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…
Prism-Shadow/penguin-harness
Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.
Prism-Shadow/penguin-harness
A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…
Prism-Shadow/penguin-harness
A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…
Prism-Shadow/penguin-harness
Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…
Categories
Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result. Agent Evaluation is an agent skill from Prism-Shadow/penguin-harness. Run one specified Test Agent on one specified Benchmark Case exactly once, privately score that execution, and return one protocol result.
Agent Evaluation fits situations like: tasks that involve Agent evaluation and testing.
Run `npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a claude-code`. Or copy the skill folder (plugins/agent-tuning/skills/agent-evaluation in Prism-Shadow/penguin-harness) into .claude/skills/agent-evaluation in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a codex`. Or copy the skill folder (plugins/agent-tuning/skills/agent-evaluation in Prism-Shadow/penguin-harness) into .agents/skills/agent-evaluation in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill agent-evaluation -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-evaluation, .gemini/skills/agent-evaluation, .github/skills/agent-evaluation and .opencode/skills/agent-evaluation in your project.
SKILL.md names no scripts, command-line tools or credentials: Agent Evaluation is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Evaluation is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Evaluation: MCP Server Builder (anthropics/skills, 180k stars), Diagnosing Superpowers Sessions (obra/superpowers, 296k stars), Darwin Skill Optimizer (alchaincyf/darwin-skill, 6.2k stars) and Skill Release Gate (rohitg00/ai-engineering-from-scratch, 66k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,455 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.
Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.