Benchmark
affaan-m/ECC
Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…
Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.
$ npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Prism-Shadow/penguin-harness benchmark-design --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agent-tuning/skills/benchmark-design .claude/skills/benchmark-design && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "benchmark-design" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/benchmark-design into .claude/skills/benchmark-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-design", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/benchmark-designType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Prism-Shadow/penguin-harness benchmark-design --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/agent-tuning/skills/benchmark-design .agents/skills/benchmark-design && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "benchmark-design" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/benchmark-design into .agents/skills/benchmark-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-design", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Prism-Shadow/penguin-harness benchmark-design --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/agent-tuning/skills/benchmark-design .cursor/skills/benchmark-design && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "benchmark-design" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/benchmark-design into .cursor/skills/benchmark-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-design", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Prism-Shadow/penguin-harness.git --path plugins/agent-tuning/skills/benchmark-design--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Prism-Shadow/penguin-harness benchmark-design --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/agent-tuning/skills/benchmark-design .gemini/skills/benchmark-design && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "benchmark-design" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/benchmark-design into .gemini/skills/benchmark-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-design", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Prism-Shadow/penguin-harness benchmark-designInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/agent-tuning/skills/benchmark-design .github/skills/benchmark-design && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-design" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/benchmark-design into .github/skills/benchmark-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-design", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Prism-Shadow/penguin-harness benchmark-design --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/agent-tuning/skills/benchmark-design .opencode/skills/benchmark-design && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "benchmark-design" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/benchmark-design into .opencode/skills/benchmark-design/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "benchmark-design", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
benchmark-designDesign and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.
Benchmark Design is an agent skill from Prism-Shadow/penguin-harness. Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.
Its SKILL.md is about 5.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.
7 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Benchmark Design loads about 5.7k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 3,118 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 3,118 words, ~5,658 tokens.
.claude/skills/benchmark-design/SKILL.md (or your agent's skills folder).Build a multi-Case Benchmark for one Test Agent, calibrate its difficulty with one Run per Case, and record the selected frozen Pilot as the Formal Baseline.
This Skill changes the Benchmark, never the Test Agent. It does not run or score the Test Agent. Delegate every evaluation with run_subagent, and tell each worker to use agent-evaluation. Stop after the Baseline; do not begin optimization.
If the request does not identify a Test Agent, target capability, desired baseline score, and Pilot iteration limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them. Treat the current Agent as the Builder. A user-specified evaluation (provider, model_id) takes priority; otherwise inherit the current Builder Session's complete Provider and Model ID from the Environment. Never use a Project default as an implicit evaluation runtime.
Follow this order:
0..100 scale, so a frozen revision that scores below 85 is published even when it misses the desired score.Require a Test Agent id, target capability, desired baseline score on the fixed 0..100 scale, and a positive Pilot iteration limit. Derive a short semantic Benchmark id if needed. Resolve (provider, model_id) once before the first Pilot: use a user-supplied complete pair when present, otherwise inherit the current Builder Session's Provider and Model ID from the Environment. Reject a half pair or an unavailable inherited value. Read thinking_level from the Test Agent's model.thinking_level in agent_state/system_config.yaml, using the normal Agent-config default medium only when that field is absent. Do not read thinking_level from a Trace and do not inspect Project configuration.
The current Session must provide run_subagent, and the current Agent must have agent-evaluation installed. If either is missing, stop and explain what is needed.
Use the Environment's App Data Dir and the explicit Test Agent id:
TEST_AGENT_DIR = <app_data_dir>/agents/<test_agent_id>
BENCHMARK_DIR = <app_data_dir>/benchmarks/<benchmark_id>
SCOREBOARD = <benchmark_dir>/scoreboard.yamlA Benchmark lives beside agents/, not inside one: it belongs to the Project and may evaluate several Agents. test_agent_id names the one this request evaluates, and every Evaluation records it.
Access only the specified Test Agent and Benchmark: the Agent State, complete Benchmark, and Test Traces or artifacts from valid evaluations. Do not access other Agents, Project secrets, or Evaluator State, Workspace, or Trace.
Read the Agent State version from the top-level version in agent_state/system_config.yaml; use 1 only when it is absent.
<benchmark_id>/
├── benchmark_config.toml
├── scoreboard.yaml
└── CASE-<nnn>-<semantic-name>/
├── statement/
│ ├── README.md
│ └── <optional-public-materials>
└── rubric/
└── README.mdEach Case contains:
statement/, which is public to the Test Agent and defines the objective, available materials, and required artifact.rubric/, which is private and defines observable scoring items, points, and Gold answers.Both directories require a README.md and may contain supporting files. Do not put Gold answers for evaluated instances, hidden mappings, or private scoring conditions in statement/.
Create benchmark_config.toml with title, description, runs = 1, and status = "draft". Benchmark design always uses one Run per Case; do not ask for or accept another Run count. status = "draft" tells the Web App that the Benchmark is still being built — it shows the Benchmark masked, and nobody can use or open it until the status is published. A draft ends in one of two states: published once a Formal Baseline scoring below 85 is recorded, or failed when calibration produces no valid Pilot result to freeze or when the lowest-scoring valid revision still scores 85 or above at the iteration limit. Initialize scoreboard.yaml with evaluations: [].
Pass the resolved (provider, model_id) explicitly in every Pilot Evaluator request, starting with the first cell. Freeze that pair and the Test Agent's configured thinking_level for the complete Benchmark workflow. Every scored Evaluator result must report the requested pair and the same configured thinking level. A mismatch invalidates the matrix.
Before planning Cases, state the Capability Contract:
Before writing each Case, privately record the required behavior, a plausible shortcut for a strong Test Agent, the chosen difficulty, the different scored decision or artifact each behavior should produce, and why the distinction measures the target capability. Design the Case so the measured capability affects the score. Do not optimize the Statement to help the Test Agent succeed or copy this design rationale into it.
The Statement presents the task, not the Benchmark's teaching or design intent. It describes the objective, available materials, option meanings, output format, and necessary constraints. It must not prescribe the reasoning sequence, identify decisive evidence, name the shortcut, or reveal private scoring preferences. When an auditable artifact is needed, request concise supporting evidence without prescribing how to obtain it.
Keep the evaluation contract well-defined, but do not require the public Statement to uniquely determine the Gold. Public information may be incomplete or conflicting, and the Rubric may encode a private decision standard or preference. Fix that private standard before evaluating the revision and never change its Gold after seeing the evaluated answer. The standard must remain tied to the target capability: it should express a stable reusable policy, priority, inference boundary, or other behavior that a better Agent State could apply across instances. Do not use a capability-irrelevant random hidden mapping merely to lower the score, and do not disclose every decisive premise or priority merely to make the public task complete.
The first complete revision is an exploratory probe. Use its Pilot to learn how the Test Agent interprets the tasks, forms candidate rules, and uses shortcuts; refine the Benchmark before treating it as calibrated. A later revision may intentionally add information gaps, conflicts, private preferences, or other capability-relevant distinctions in response to an earlier Trace, provided the next revision's Rubric is fixed before dispatch.
Every Case Rubric has a fixed maximum of 100 points, with observable scoring items and meaningful partial credit. Allocate most points within each Case to decisions or concise artifacts on which the intended behavior and plausible shortcut differ. Keep generic format compliance, evidence enumeration, and analysis completeness from creating a high score floor unless those are themselves the target capability. Allocate points from capability coverage before the first Pilot. Do not change scoring items solely to satisfy the desired score; when a redesign changes coverage, re-plan that Case's 100-point allocation before evaluating the revised Case set. When final choices do not distinguish the intended behavior from a shortcut, score a concise auditable artifact, but define only its required content or format—not the method used to produce it.
Before the first dispatch of every new or changed Case revision, run a consistency review:
This review does not require the public Statement to contain enough information to reproduce the private standard or uniquely derive every Gold answer. Unchanged Cases do not need another review during that iteration. Keep this review in Builder analysis and Trace; fix defects in the Case rather than creating a separate audit artifact.
Also compare all public files with the private Rubric. Confirm that no public file reveals Gold answers, private scoring conditions, or hints that identify the intended solution. This is the leak check.
For each Case × Run cell, call run_subagent with the request below. Dispatch independent cells in parallel up to available concurrency.
Use the `agent-evaluation` Skill. Run the specified Test Agent on the specified Case exactly once, then score that single execution.
protocol_version: 1
case_id: <case_id>
run: <1_based_run_index>
expected_version: <test_agent_state_version>
test_agent_id: <test_agent_id>
benchmark_id: <benchmark_id>
provider: <provider>
model_id: <model_id>Inspect the complete streamed and final worker response. Before reading status, score, or any other protocol field, verify that the worker-authored text is exactly one plain protocol YAML document. Narration, headings, code fences, summaries, or scoring details are not valid protocol. Ask the same Evaluator to resend only the clean YAML from its existing result; do not rerun the Test Agent for a formatting repair and do not extract YAML from the invalid response yourself. Transport metadata added by run_subagent is not worker-authored text. A wrong or missing Test Agent artifact is a valid scored result and must not be retried.
For every scored result, require non-empty agent_id, provider, model_id, and thinking_level. Require agent_id to equal the requested Test Agent. Require the model pair to equal the explicitly resolved pair and the thinking level to equal the Test Agent configuration read before dispatch. Reject a Pilot result whose cells report mixed or mismatched runtimes. The Evaluator verifies provider/model from the root Trace and reports thinking from the unchanged Target Agent configuration; it does not require Trace metadata for thinking.
Correct and resend an invalid_request. For benchmark_invalid, repair and rerun the affected Case during Pilot. For version_changed, discard the current Pilot result and restart after the Agent version is stable.
For evaluation_failed, keep the same Benchmark revision and cell. Diagnose the failure and retry only when evidence proves the Test Agent did not start and the retry applies a new, specific repair. Do not set a numeric retry limit or repeat an unchanged launch. Stop when no new safe repair remains, external configuration is required, or it is unclear whether the Test Agent started. Never treat an evaluation failure as score zero.
Treat the first draft as a hypothesis. The first valid result from every planned Case together forms Pilot iteration 1. A later iteration starts after a difficulty refinement and completes when every affected Case has a valid new result. Request corrections, validity repairs, and evaluation reruns stay in the current iteration and do not consume the requested iteration budget. Use the recorded Agent State version and fixed evaluation runtime.
Keep unselected Pilot results out of the Scoreboard. During calibration, retain only one temporary restorable copy: the lowest-scoring complete valid revision seen so far, including its one-Run-per-Case result. Store it outside the Project's benchmarks/, replace it only when a lower valid revision completes, and never retain invalid revisions.
Use the Pilot to find the current Test Agent's capability boundary.
Before editing, distinguish a validity repair from a difficulty refinement. A validity repair fixes an unusable task or scoring contract and stays in the current Pilot iteration. A difficulty refinement changes what the valid Benchmark measures and completes the next iteration after every affected Case has a valid result.
Before editing, estimate how much of the score the planned refinements can affect. If the range is too small to materially approach the desired score, revise more affected Cases, use more than one difficulty dimension, or replace low-signal Cases.
Prefer refinements that create one or more scored separating decisions. A refinement may change the public task or evidence, introduce or preserve a reasonable information gap or conflict, or apply a fixed private standard. Adding another explicit rule, exception, source, or checklist is not a difficulty increase when the observed strategy can still follow it to the Gold. A Rubric-only refinement is allowed but not preferred when the public task already contains the relevant information, the current Rubric fails to distinguish merely mentioning it from handling it correctly, and the Builder can explain which reusable capability the new scoring distinction measures. Do not add points merely because the previous Test Agent omitted a phrase. Fix the revised Rubric before dispatch and treat it as a changed Case revision.
For each refinement iteration:
Reuse a Pilot result only when the Case revision, scoring, Agent State version, and evaluation runtime are unchanged.
An information gap or supported alternative is not automatically a design defect. Treat it as a defect only when the task or fixed private standard is incoherent, changes after evaluation, leaks the answer, or no reusable Agent behavior could plausibly improve the score.
More rows, fields, distractors, files, near-duplicate examples, or explicit rule layers do not increase difficulty when the observed strategy still solves the Case. Base refinements on observed behavior and fix the Gold before each evaluation.
Freeze immediately when a complete valid Pilot iteration meets the desired baseline score and no known design defect remains. Do not run another difficulty refinement merely to create more score margin. Otherwise continue through the requested valid-iteration limit. If the desired score is still unmet, restore the temporary lowest-scoring valid revision and proceed to Freeze. The desired baseline score is a calibration target, not the publish gate: the gate is fixed at 85 on the 0..100 scale, and any frozen valid revision scoring below 85 is published, however far it stays from the desired score. Report calibration_failed only when no valid Pilot revision can be produced, evaluation failures prevent a valid selection, or the lowest-scoring valid revision still scores 85 or above at the iteration limit — a Test Agent that already scores that high leaves the Benchmark nothing to measure. Missing the desired score alone is never a failure.
After selecting the Pilot revision, restore that exact revision and its complete result if needed. Run a complete consistency review and final leak check across every Case. If the review finds a defect, repair it and produce a complete valid one-Run-per-Case Pilot result for the repaired revision before selecting and freezing it. Freeze the Benchmark and record the current Agent State version. Do not launch a fresh Formal matrix, rerun the selected Pilot, or backfill it to another Run count.
Accept the selected Pilot result as the Formal Baseline when every Case has exactly one valid Run, every cell reports the frozen evaluation runtime, the Agent State version remains unchanged, the private scoring standard remained fixed, and every score loss reflects the Capability Contract. Record the Formal Baseline even when its score does not meet the desired baseline score; only a score of 85 or above blocks it.
Report calibration_failed only when no valid revision remains, evaluation failures prevent a complete selected Pilot result, or the selected revision scores 85 or above. Never record a partial, abandoned, invalid, or non-selected Pilot result as the Formal Baseline.
After validation, obtain the current UTC timestamp from the environment, for example with date -u +"%Y-%m-%dT%H:%M:%SZ", rather than inferring UTC from a displayed local time. Append only the accepted Formal Baseline to scoreboard.yaml using exactly this structure:
evaluations:
- time: <ISO-8601 timestamp>
agent_id: <test_agent_id>
version: <Agent State version>
provider: <provider>
model_id: <model_id>
thinking_level: <thinking_level>
summary_title: >-
<public title>
summary: >-
<public summary>
score: <average of the Case scores>
cost: <average of known Case costs, or null when every Case cost is null>
duration_ms: <average of the Case durations>
cases:
- case: <case_id>
score: <average of the Run scores>
cost: <average of known Run costs, or null when every Run cost is null>
duration_ms: <average of the Run durations>
runs:
- score: <Run score>
cost: <Run cost or null>
duration_ms: <Run duration>
session_id: <Test Session id>After writing, parse the complete scoreboard.yaml and verify the appended Evaluation, including its agent_id, before reporting success or continuing.
Once the Formal Baseline is verified, set status = "published" in benchmark_config.toml. Change only that line, keep title, description and runs as they are, and parse the file again to confirm it is valid TOML. When the run ends in calibration_failed, set status = "failed" the same way — change only that line, keep the other fields, and parse the file again — so the Web App tells the user that this Benchmark failed to calibrate and has to be deleted and created again. Never leave a failed Benchmark on draft, and never write failed because the desired baseline score was missed: a Formal Baseline below 85 is published.
Every Run and Case score is on the fixed 0..100 scale. Do not write max_score. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore null values when averaging cost and write null only when all contributing costs are unknown; round score averages to two decimal places, cost averages to six decimal places, and duration_ms averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an aggregate object or use case_id, mean_score, mean_cost, or mean_duration_ms.
Report the Benchmark path, configuration, Agent State version, Evaluation average and Case Run scores, Test Session ids, and known limitations. Include one compact row per Pilot iteration with its score, diagnosed capability gap, difficulty adjustment, and freeze or stop decision. Identify which one-Run-per-Case Pilot result was recorded as the Formal Baseline.
After the accepted Formal Baseline is recorded, delete the temporary lowest-revision copy and other Builder calibration scaffolding. Keep the frozen Benchmark, Scoreboard, evaluation Workspaces, and score-linked Traces.
Do not reveal Rubrics, Gold answers, latent rules, per-item scores, or private scoring information. Stop after reporting the Baseline; do not modify the Test Agent or begin optimization.
© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/agent-tuning/skills/benchmark-design of Prism-Shadow/penguin-harness.
Open the folder on GitHubat commit d56d9ce
Benchmark Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Benchmark Design this skillPrism-Shadow/penguin-harness | 2.5k | — | ~5.7k | Automated safety check: Pass | Apache-2.0 | |
| Benchmarkaffaan-m/ECC | 274k | 3 repos | ~654 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 274k | — | ~412 | Automated safety check: Pass | MIT | |
| Benchmarkaffaan-m/ECC | 274k | — | ~330 | Automated safety check: Pass | MIT | |
| Gstack Performance Benchmarkgarrytan/gstack | 136k | — | ~7.2k | Automated safety check: Notes | MIT | |
| Product Capabilityaffaan-m/ECC | 274k | 2 repos | ~1.1k | Automated safety check: Pass | MIT |
affaan-m/ECC
Measure performance baselines and detect regressions across browser Core Web Vitals (LCP, INP, CLS, page weight), API endpoint latency percentiles, and build/test feedback times, with before/after…
affaan-m/ECC
このスキルを使用して、パフォーマンスベースラインを測定し、PR前後の回帰を検出し、スタック代替案を比較します. An agent skill from affaan-m/ECC.
affaan-m/ECC
使用此技能测量性能基线,检测PR前后的回归,并比较堆栈替代方案。
garrytan/gstack
Establishes page load, Core Web Vitals and resource-size baselines, then compares before and after on every pull request to track performance trends over time.
affaan-m/ECC
Translate PRD intent, roadmap asks, or product discussions into an implementation-ready capability plan that exposes constraints, invariants, interfaces, and unresolved decisions before…
androidx/androidx
Benchmarking and improving the performance of Jetpack Compose.
Prism-Shadow/penguin-harness
Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…
Prism-Shadow/penguin-harness
A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…
Prism-Shadow/penguin-harness
Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.
Prism-Shadow/penguin-harness
A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…
Prism-Shadow/penguin-harness
A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…
Prism-Shadow/penguin-harness
Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…
Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline. Benchmark Design is an agent skill from Prism-Shadow/penguin-harness. Design and calibrate a multi-Case capability Benchmark and establish a traceable Formal Baseline.
Run `npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a claude-code`. Or copy the skill folder (plugins/agent-tuning/skills/benchmark-design in Prism-Shadow/penguin-harness) into .claude/skills/benchmark-design in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a codex`. Or copy the skill folder (plugins/agent-tuning/skills/benchmark-design in Prism-Shadow/penguin-harness) into .agents/skills/benchmark-design in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill benchmark-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/benchmark-design, .gemini/skills/benchmark-design, .github/skills/benchmark-design and .opencode/skills/benchmark-design in your project.
SKILL.md names no scripts, command-line tools or credentials: Benchmark Design is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Benchmark Design is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.7k tokens (SKILL.md is roughly 23k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Benchmark Design: Benchmark (affaan-m/ECC, 274k stars), Benchmark (affaan-m/ECC, 274k stars), Benchmark (affaan-m/ECC, 274k stars) and Gstack Performance Benchmark (garrytan/gstack, 136k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,450 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.
Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.