Agent skill

Benchmark Design

by Prism-Shadow in Prism-Shadow/penguin-harness

Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.

Apache-2.0Auto-check passed

Install Benchmark Design

skills CLI
$ npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install Prism-Shadow/penguin-harness benchmark-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agent-tuning/skills/benchmark-design .claude/skills/benchmark-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
benchmark-design
GitHub stars
2.5k
Token cost
~5.7k tokens
SKILL.md length
3,118 words
Files
1
Skills in repo
31
Repo updated
First seen
Licence
Apache-2.0

At a glance

Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.

  • Works in 7 steps: Validate the Test Agent, target… → Write a Capability Contract that defines… → Plan the complete initial Case set and… → …
  • SKILL.md covers Before you start, Workflow, Setup and access and Build the Benchmark, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Benchmark Design is an agent skill from Prism-Shadow/penguin-harness. Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.

Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.

Example prompts

  • “/benchmark-design”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Validate the Test Agent, target capability, resolved evaluation Runtime, and evaluation access.
  2. Write a Capability Contract that defines the observable process to measure, common weaker behavior, and the general Agent State…
  3. Plan the complete initial Case set and point allocation. For each Case, privately state the intended behavior, a plausible shortcut for a…
  4. Complete one valid evaluation for every planned Case. Together these results form Pilot iteration 1; finish this complete set before…
  5. For later Pilot iterations, use scores and Traces to reconstruct how the Test Agent solved each Case. A single iteration may refine…
  6. Freeze the first valid Pilot revision that meets the desired baseline score. If none does within the requested valid-iteration limit…
  7. Freeze the selected revision and record its complete one-Run-per-Case Pilot result as the Formal Baseline when every cell is valid, the…

What it can do on your machine

Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Benchmark Design loads about 5.7k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 3,118 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~29
When it runs · the whole SKILL.md, loaded when a task matches
~5.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 3,118 words, ~5,658 tokens.

Download SKILL.mdSave it as .claude/skills/benchmark-design/SKILL.md (or your agent's skills folder).
name
benchmark-design
description
Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.

Benchmark Design

Build a multi-Case Benchmark for one Test Agent, calibrate its difficulty with one Run per Case, and record the selected frozen Pilot as the Formal Baseline.

This Skill changes the Benchmark, never the Test Agent. It does not run or score the Test Agent. Delegate every evaluation with run_subagent, and tell each worker to use agent-evaluation. Stop after the Baseline; do not begin optimization.

Before you start

If the request does not identify a Test Agent, target capability, desired baseline score, and Pilot iteration limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them. Treat the current Agent as the Builder. A user-specified evaluation (provider, model_id) takes priority; otherwise inherit the current Builder Session's complete Provider and Model ID from the Environment. Never use a Project default as an implicit evaluation runtime.

Workflow

  • A Pilot is a one-Run-per-Case evaluation used to improve the Benchmark. Unselected Pilot results never enter the Scoreboard; the selected result becomes the Formal Baseline after Freeze.
  • Freeze means the Benchmark revision and evaluation settings stop changing.
  • A Formal Baseline is the accepted result of the selected complete valid Pilot revision, recorded after that exact revision is frozen on one unchanged Agent State version.

Follow this order:

  1. Validate the Test Agent, target capability, resolved evaluation Runtime, and evaluation access.
  2. Write a Capability Contract that defines the observable process to measure, common weaker behavior, and the general Agent State improvement the Benchmark should train.
  3. Plan the complete initial Case set and point allocation. For each Case, privately state the intended behavior, a plausible shortcut for a strong Test Agent, and how the Case distinguishes them. Write and leak-check the complete initial Benchmark.
  4. Complete one valid evaluation for every planned Case. Together these results form Pilot iteration 1; finish this complete set before refining any Case.
  5. For later Pilot iterations, use scores and Traces to reconstruct how the Test Agent solved each Case. A single iteration may refine multiple Cases or difficulty dimensions; rerun every affected Case.
  6. Freeze the first valid Pilot revision that meets the desired baseline score. If none does within the requested valid-iteration limit, restore and freeze the lowest-scoring valid Pilot revision. The desired score steers refinement; the publish gate is fixed at 85 on the 0..100 scale, so a frozen revision that scores below 85 is published even when it misses the desired score.
  7. Freeze the selected revision and record its complete one-Run-per-Case Pilot result as the Formal Baseline when every cell is valid, the Agent State version remains unchanged, and no known design defect remains. Do not rerun or backfill it. The Formal score does not determine validity.

Setup and access

Require a Test Agent id, target capability, desired baseline score on the fixed 0..100 scale, and a positive Pilot iteration limit. Derive a short semantic Benchmark id if needed. Resolve (provider, model_id) once before the first Pilot: use a user-supplied complete pair when present, otherwise inherit the current Builder Session's Provider and Model ID from the Environment. Reject a half pair or an unavailable inherited value. Read thinking_level from the Test Agent's model.thinking_level in agent_state/system_config.yaml, using the normal Agent-config default medium only when that field is absent. Do not read thinking_level from a Trace and do not inspect Project configuration.

The current Session must provide run_subagent, and the current Agent must have agent-evaluation installed. If either is missing, stop and explain what is needed.

Use the Environment's App Data Dir and the explicit Test Agent id:

text
TEST_AGENT_DIR = <app_data_dir>/agents/<test_agent_id>
BENCHMARK_DIR = <app_data_dir>/benchmarks/<benchmark_id>
SCOREBOARD = <benchmark_dir>/scoreboard.yaml

A Benchmark lives beside agents/, not inside one: it belongs to the Project and may evaluate several Agents. test_agent_id names the one this request evaluates, and every Evaluation records it.

Access only the specified Test Agent and Benchmark: the Agent State, complete Benchmark, and Test Traces or artifacts from valid evaluations. Do not access other Agents, Project secrets, or Evaluator State, Workspace, or Trace.

Read the Agent State version from the top-level version in agent_state/system_config.yaml; use 1 only when it is absent.

Build the Benchmark

text
<benchmark_id>/
├── benchmark_config.toml
├── scoreboard.yaml
└── CASE-<nnn>-<semantic-name>/
    ├── statement/
    │   ├── README.md
    │   └── <optional-public-materials>
    └── rubric/
        └── README.md

Each Case contains:

  • statement/, which is public to the Test Agent and defines the objective, available materials, and required artifact.
  • rubric/, which is private and defines observable scoring items, points, and Gold answers.

Both directories require a README.md and may contain supporting files. Do not put Gold answers for evaluated instances, hidden mappings, or private scoring conditions in statement/.

Create benchmark_config.toml with title, description, runs = 1, and status = "draft". Benchmark design always uses one Run per Case; do not ask for or accept another Run count. status = "draft" tells the Web App that the Benchmark is still being built — it shows the Benchmark masked, and nobody can use or open it until the status is published. A draft ends in one of two states: published once a Formal Baseline scoring below 85 is recorded, or failed when calibration produces no valid Pilot result to freeze or when the lowest-scoring valid revision still scores 85 or above at the iteration limit. Initialize scoreboard.yaml with evaluations: [].

Pass the resolved (provider, model_id) explicitly in every Pilot Evaluator request, starting with the first cell. Freeze that pair and the Test Agent's configured thinking_level for the complete Benchmark workflow. Every scored Evaluator result must report the requested pair and the same configured thinking level. A mismatch invalidates the matrix.

Before planning Cases, state the Capability Contract:

  • the public evidence available to the Test Agent;
  • the observable decisions, intermediate artifacts, and checks the capability requires;
  • the weaker behaviors or shortcuts the Benchmark should distinguish; and
  • the reusable Agent State behavior that could improve the measured capability.

Before writing each Case, privately record the required behavior, a plausible shortcut for a strong Test Agent, the chosen difficulty, the different scored decision or artifact each behavior should produce, and why the distinction measures the target capability. Design the Case so the measured capability affects the score. Do not optimize the Statement to help the Test Agent succeed or copy this design rationale into it.

The Statement presents the task, not the Benchmark's teaching or design intent. It describes the objective, available materials, option meanings, output format, and necessary constraints. It must not prescribe the reasoning sequence, identify decisive evidence, name the shortcut, or reveal private scoring preferences. When an auditable artifact is needed, request concise supporting evidence without prescribing how to obtain it.

Keep the evaluation contract well-defined, but do not require the public Statement to uniquely determine the Gold. Public information may be incomplete or conflicting, and the Rubric may encode a private decision standard or preference. Fix that private standard before evaluating the revision and never change its Gold after seeing the evaluated answer. The standard must remain tied to the target capability: it should express a stable reusable policy, priority, inference boundary, or other behavior that a better Agent State could apply across instances. Do not use a capability-irrelevant random hidden mapping merely to lower the score, and do not disclose every decisive premise or priority merely to make the public task complete.

The first complete revision is an exploratory probe. Use its Pilot to learn how the Test Agent interprets the tasks, forms candidate rules, and uses shortcuts; refine the Benchmark before treating it as calibrated. A later revision may intentionally add information gaps, conflicts, private preferences, or other capability-relevant distinctions in response to an earlier Trace, provided the next revision's Rubric is fixed before dispatch.

Every Case Rubric has a fixed maximum of 100 points, with observable scoring items and meaningful partial credit. Allocate most points within each Case to decisions or concise artifacts on which the intended behavior and plausible shortcut differ. Keep generic format compliance, evidence enumeration, and analysis completeness from creating a high score floor unless those are themselves the target capability. Allocate points from capability coverage before the first Pilot. Do not change scoring items solely to satisfy the desired score; when a redesign changes coverage, re-plan that Case's 100-point allocation before evaluating the revised Case set. When final choices do not distinguish the intended behavior from a shortcut, score a concise auditable artifact, but define only its required content or format—not the method used to produce it.

Before the first dispatch of every new or changed Case revision, run a consistency review:

  • Confirm that the current Statement is internally coherent. Intentional conflicts must be presented as conflicts between sources, rules, or positions rather than as contradictory claims by the Benchmark itself.
  • Confirm that the current Rubric is consistent with the current Statement and fixed private standard. It must be self-contained and must not refer to an earlier revision or missing context.
  • Confirm that every scoring item applies to the Case's actual requested output and relies only on premises that are defined, provided, or explicitly private under the fixed standard.

This review does not require the public Statement to contain enough information to reproduce the private standard or uniquely derive every Gold answer. Unchanged Cases do not need another review during that iteration. Keep this review in Builder analysis and Trace; fix defects in the Case rather than creating a separate audit artifact.

Also compare all public files with the private Rubric. Confirm that no public file reveals Gold answers, private scoring conditions, or hints that identify the intended solution. This is the leak check.

Delegate evaluation

For each Case × Run cell, call run_subagent with the request below. Dispatch independent cells in parallel up to available concurrency.

text
Use the `agent-evaluation` Skill. Run the specified Test Agent on the specified Case exactly once, then score that single execution.
protocol_version: 1
case_id: <case_id>
run: <1_based_run_index>
expected_version: <test_agent_state_version>
test_agent_id: <test_agent_id>
benchmark_id: <benchmark_id>
provider: <provider>
model_id: <model_id>

Inspect the complete streamed and final worker response. Before reading status, score, or any other protocol field, verify that the worker-authored text is exactly one plain protocol YAML document. Narration, headings, code fences, summaries, or scoring details are not valid protocol. Ask the same Evaluator to resend only the clean YAML from its existing result; do not rerun the Test Agent for a formatting repair and do not extract YAML from the invalid response yourself. Transport metadata added by run_subagent is not worker-authored text. A wrong or missing Test Agent artifact is a valid scored result and must not be retried.

For every scored result, require non-empty agent_id, provider, model_id, and thinking_level. Require agent_id to equal the requested Test Agent. Require the model pair to equal the explicitly resolved pair and the thinking level to equal the Test Agent configuration read before dispatch. Reject a Pilot result whose cells report mixed or mismatched runtimes. The Evaluator verifies provider/model from the root Trace and reports thinking from the unchanged Target Agent configuration; it does not require Trace metadata for thinking.

Correct and resend an invalid_request. For benchmark_invalid, repair and rerun the affected Case during Pilot. For version_changed, discard the current Pilot result and restart after the Agent version is stable.

For evaluation_failed, keep the same Benchmark revision and cell. Diagnose the failure and retry only when evidence proves the Test Agent did not start and the retry applies a new, specific repair. Do not set a numeric retry limit or repeat an unchanged launch. Stop when no new safe repair remains, external configuration is required, or it is unclear whether the Test Agent started. Never treat an evaluation failure as score zero.

Show full SKILL.md (1,263 more words)Show less

Refine the Benchmark

Treat the first draft as a hypothesis. The first valid result from every planned Case together forms Pilot iteration 1. A later iteration starts after a difficulty refinement and completes when every affected Case has a valid new result. Request corrections, validity repairs, and evaluation reruns stay in the current iteration and do not consume the requested iteration budget. Use the recorded Agent State version and fixed evaluation runtime.

Keep unselected Pilot results out of the Scoreboard. During calibration, retain only one temporary restorable copy: the lowest-scoring complete valid revision seen so far, including its one-Run-per-Case result. Store it outside the Project's benchmarks/, replace it only when a lower valid revision completes, and never retain invalid revisions.

Use the Pilot to find the current Test Agent's capability boundary.

Before editing, distinguish a validity repair from a difficulty refinement. A validity repair fixes an unusable task or scoring contract and stays in the current Pilot iteration. A difficulty refinement changes what the valid Benchmark measures and completes the next iteration after every affected Case has a valid result.

Before editing, estimate how much of the score the planned refinements can affect. If the range is too small to materially approach the desired score, revise more affected Cases, use more than one difficulty dimension, or replace low-signal Cases.

Prefer refinements that create one or more scored separating decisions. A refinement may change the public task or evidence, introduce or preserve a reasonable information gap or conflict, or apply a fixed private standard. Adding another explicit rule, exception, source, or checklist is not a difficulty increase when the observed strategy can still follow it to the Gold. A Rubric-only refinement is allowed but not preferred when the public task already contains the relevant information, the current Rubric fails to distinguish merely mentioning it from handling it correctly, and the Builder can explain which reusable capability the new scoring distinction measures. Do not add points merely because the previous Test Agent omitted a phrase. Fix the revised Rubric before dispatch and treat it as a changed Case revision.

For each refinement iteration:

  1. Observed strategy. Reconstruct the Test Agent's actual solution method from its score, artifact, and Trace.
  2. Missing behavior. Identify the general behavior that the observed strategy skipped or simplified. Repair missing evidence, arbitrary mappings, ambiguity, or scoring defects before increasing difficulty.
  3. Separating prediction. Before dispatch, predict the decision or artifact the observed strategy will produce, the different result the desired behavior will produce, and the score range affected. If both behaviors are expected to reach the same scored result, choose another refinement.
  4. Update any number of diagnosed Cases or difficulty dimensions, run the consistency review and leak check for each changed revision, and rerun every affected Case.

Reuse a Pilot result only when the Case revision, scoring, Agent State version, and evaluation runtime are unchanged.

An information gap or supported alternative is not automatically a design defect. Treat it as a defect only when the task or fixed private standard is incoherent, changes after evaluation, leaks the answer, or no reusable Agent behavior could plausibly improve the score.

More rows, fields, distractors, files, near-duplicate examples, or explicit rule layers do not increase difficulty when the observed strategy still solves the Case. Base refinements on observed behavior and fix the Gold before each evaluation.

Freeze immediately when a complete valid Pilot iteration meets the desired baseline score and no known design defect remains. Do not run another difficulty refinement merely to create more score margin. Otherwise continue through the requested valid-iteration limit. If the desired score is still unmet, restore the temporary lowest-scoring valid revision and proceed to Freeze. The desired baseline score is a calibration target, not the publish gate: the gate is fixed at 85 on the 0..100 scale, and any frozen valid revision scoring below 85 is published, however far it stays from the desired score. Report calibration_failed only when no valid Pilot revision can be produced, evaluation failures prevent a valid selection, or the lowest-scoring valid revision still scores 85 or above at the iteration limit — a Test Agent that already scores that high leaves the Benchmark nothing to measure. Missing the desired score alone is never a failure.

Freeze and record the Formal Baseline

After selecting the Pilot revision, restore that exact revision and its complete result if needed. Run a complete consistency review and final leak check across every Case. If the review finds a defect, repair it and produce a complete valid one-Run-per-Case Pilot result for the repaired revision before selecting and freezing it. Freeze the Benchmark and record the current Agent State version. Do not launch a fresh Formal matrix, rerun the selected Pilot, or backfill it to another Run count.

Accept the selected Pilot result as the Formal Baseline when every Case has exactly one valid Run, every cell reports the frozen evaluation runtime, the Agent State version remains unchanged, the private scoring standard remained fixed, and every score loss reflects the Capability Contract. Record the Formal Baseline even when its score does not meet the desired baseline score; only a score of 85 or above blocks it.

Report calibration_failed only when no valid revision remains, evaluation failures prevent a complete selected Pilot result, or the selected revision scores 85 or above. Never record a partial, abandoned, invalid, or non-selected Pilot result as the Formal Baseline.

Record and finish

After validation, obtain the current UTC timestamp from the environment, for example with date -u +"%Y-%m-%dT%H:%M:%SZ", rather than inferring UTC from a displayed local time. Append only the accepted Formal Baseline to scoreboard.yaml using exactly this structure:

yaml
evaluations:
  - time: <ISO-8601 timestamp>
    agent_id: <test_agent_id>
    version: <Agent State version>
    provider: <provider>
    model_id: <model_id>
    thinking_level: <thinking_level>
    summary_title: >-
      <public title>
    summary: >-
      <public summary>
    score: <average of the Case scores>
    cost: <average of known Case costs, or null when every Case cost is null>
    duration_ms: <average of the Case durations>
    cases:
      - case: <case_id>
        score: <average of the Run scores>
        cost: <average of known Run costs, or null when every Run cost is null>
        duration_ms: <average of the Run durations>
        runs:
          - score: <Run score>
            cost: <Run cost or null>
            duration_ms: <Run duration>
            session_id: <Test Session id>

After writing, parse the complete scoreboard.yaml and verify the appended Evaluation, including its agent_id, before reporting success or continuing.

Once the Formal Baseline is verified, set status = "published" in benchmark_config.toml. Change only that line, keep title, description and runs as they are, and parse the file again to confirm it is valid TOML. When the run ends in calibration_failed, set status = "failed" the same way — change only that line, keep the other fields, and parse the file again — so the Web App tells the user that this Benchmark failed to calibrate and has to be deleted and created again. Never leave a failed Benchmark on draft, and never write failed because the desired baseline score was missed: a Formal Baseline below 85 is published.

Every Run and Case score is on the fixed 0..100 scale. Do not write max_score. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore null values when averaging cost and write null only when all contributing costs are unknown; round score averages to two decimal places, cost averages to six decimal places, and duration_ms averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an aggregate object or use case_id, mean_score, mean_cost, or mean_duration_ms.

Report the Benchmark path, configuration, Agent State version, Evaluation average and Case Run scores, Test Session ids, and known limitations. Include one compact row per Pilot iteration with its score, diagnosed capability gap, difficulty adjustment, and freeze or stop decision. Identify which one-Run-per-Case Pilot result was recorded as the Formal Baseline.

After the accepted Formal Baseline is recorded, delete the temporary lowest-revision copy and other Builder calibration scaffolding. Keep the frozen Benchmark, Scoreboard, evaluation Workspaces, and score-linked Traces.

Do not reveal Rubrics, Gold answers, latent rules, per-item scores, or private scoring information. Stop after reporting the Baseline; do not modify the Test Agent or begin optimization.

© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugins/agent-tuning/skills/benchmark-design of Prism-Shadow/penguin-harness.

Open the folder on GitHubat commit d56d9ce

Compare with similar skills

Benchmark Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Benchmark Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Benchmark Design this skillPrism-Shadow/penguin-harness2.5k—~5.7kAutomated safety check: PassApache-2.0
Benchmarkaffaan-m/ECC274k3 repos~654Automated safety check: PassMIT
Benchmarkaffaan-m/ECC274k—~412Automated safety check: PassMIT
Benchmarkaffaan-m/ECC274k—~330Automated safety check: PassMIT
Gstack Performance Benchmarkgarrytan/gstack136k—~7.2kAutomated safety check: NotesMIT
Product Capabilityaffaan-m/ECC274k2 repos~1.1kAutomated safety check: PassMIT

Similar skills

  • Benchmark

    affaan-m/ECC

    Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…

    274k GitHub starsUsed in 3 repos~654 tokens
    Frontend & DesignAuto-check passed
  • Benchmark

    affaan-m/ECC

    このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.

    274k GitHub stars~412 tokensUpdated 2 days ago
    Auto-check passed
  • Benchmark

    affaan-m/ECC

    使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。

    274k GitHub stars~330 tokensUpdated 2 days ago
    Auto-check passed
  • Establishes page load, Core Web Vitals and resource-size baselines, then compares before and after on every pull request to track performance trends over time.

    136k GitHub stars~7.2k tokensUpdated yesterday
    Frontend & DesignAuto-check: notes
  • Product Capability

    affaan-m/ECC

    Translate PRD intent, roadmap asks, or product discussions into an implementation-ready capability plan that exposes constraints, invariants, interfaces, and unresolved decisions before…

    274k GitHub starsUsed in 2 repos~1.1k tokens
    Product & Project ManagementAuto-check passed
  • Benchmark

    androidx/androidx

    Benchmarking and improving the performance of Jetpack Compose.

    6.1k GitHub stars~1.1k tokensUpdated today
    MobileAuto-check passed

More from Prism-Shadow/penguin-harness

All 31 skills in this repo
  • A2ui

    Prism-Shadow/penguin-harness

    Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…

    2.5k GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Penguin Harness Dev

    Prism-Shadow/penguin-harness

    A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…

    2.5k GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Bento Slides

    Prism-Shadow/penguin-harness

    Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.

    2.5k GitHub stars~1.6k tokensUpdated yesterday
    Auto-check passed
  • Penguin Harness Manual Test

    Prism-Shadow/penguin-harness

    A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…

    2.5k GitHub stars~1.3k tokensUpdated yesterday
    Auto-check passed
  • Penguin Harness Frontend

    Prism-Shadow/penguin-harness

    A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…

    2.5k GitHub stars~6.4k tokensUpdated yesterday
    Auto-check passed
  • Browser Automation

    Prism-Shadow/penguin-harness

    Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…

    2.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check: warnings

Questions about Benchmark Design

What does Benchmark Design do?

Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline. Benchmark Design is an agent skill from Prism-Shadow/penguin-harness. Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.

How do I install Benchmark Design in Claude Code?

Run `npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a claude-code`. Or copy the skill folder (plugins/agent-tuning/skills/benchmark-design in Prism-Shadow/penguin-harness) into .claude/skills/benchmark-design in your project. Claude Code loads it when a task matches its description.

How do I install Benchmark Design in Codex?

Run `npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a codex`. Or copy the skill folder (plugins/agent-tuning/skills/benchmark-design in Prism-Shadow/penguin-harness) into .agents/skills/benchmark-design in your project. Codex loads it when a task matches its description.

Can I use Benchmark Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-design, .gemini/skills/benchmark-design, .github/skills/benchmark-design and .opencode/skills/benchmark-design in your project.

What does Benchmark Design need to run?

SKILL.md names no scripts, command-line tools or credentials: Benchmark Design is instructions for the agent only.

Does Benchmark Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Benchmark Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Benchmark Design use?

Benchmark Design is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Benchmark Design use?

About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Benchmark Design?

Skills that share tags, products or a category with Benchmark Design: Benchmark (affaan-m/ECC, 274k stars), Benchmark (affaan-m/ECC, 274k stars), Benchmark (affaan-m/ECC, 274k stars) and Gstack Performance Benchmark (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Benchmark Design?

Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,450 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.

Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.