Claw Score
openclaw/openclaw
Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.
Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-optimization --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-optimization .claude/skills/agent-optimization && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "agent-optimization" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-optimization into .claude/skills/agent-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-optimization", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-optimizationType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-optimization --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .agents/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-optimization .agents/skills/agent-optimization && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "agent-optimization" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-optimization into .agents/skills/agent-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-optimization", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-optimization --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-optimization .cursor/skills/agent-optimization && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "agent-optimization" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-optimization into .cursor/skills/agent-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-optimization", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/Prism-Shadow/penguin-harness.git --path plugins/agent-tuning/skills/agent-optimization--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-optimization --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-optimization .gemini/skills/agent-optimization && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "agent-optimization" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-optimization into .gemini/skills/agent-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-optimization", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install Prism-Shadow/penguin-harness agent-optimizationInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .github/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-optimization .github/skills/agent-optimization && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "agent-optimization" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-optimization into .github/skills/agent-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-optimization", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install Prism-Shadow/penguin-harness agent-optimization --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/Prism-Shadow/penguin-harness.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/plugins/agent-tuning/skills/agent-optimization .opencode/skills/agent-optimization && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "agent-optimization" agent skill from https://github.com/Prism-Shadow/penguin-harness/tree/main/plugins/agent-tuning/skills/agent-optimization into .opencode/skills/agent-optimization/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "agent-optimization", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
agent-optimizationImprove an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.
Agent Optimization is an agent skill from Prism-Shadow/penguin-harness. Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.
Its SKILL.md is about 3.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: 🐧 Unified and Stable RSI Platform. The licence is Apache-2.0.
8 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit d56d9ce. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md (its code samples are yaml).
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Agent Optimization loads about 3.3k tokens when it runs. Until then it costs about 29 tokens; SKILL.md has 1,699 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from Prism-Shadow/penguin-harness at commit d56d9ce, republished under its Apache-2.0 licence (© Prism-Shadow). 1,699 words, ~3,325 tokens.
.claude/skills/agent-optimization/SKILL.md (or your agent's skills folder).Improve one Test Agent through an evidence → hypothesis → Candidate → evaluation → accept or rollback loop. Use public Statements, scores, and Test Traces as black-box feedback. Delegate every evaluation to an agent-evaluation subagent; never run or score the Test Agent directly.
If the request does not identify the Test Agent, frozen Benchmark, desired target score, positive Run count, and round limit, ask for the missing inputs. When they are already supplied, proceed without asking the user to restate them.
Require an explicit Test Agent, a frozen Benchmark with a complete valid Formal Baseline, a desired target score, a positive runs value, and a positive round limit. The Benchmark's benchmark_config.toml must say status = "published"; a draft Benchmark is still being built and is not frozen, and a failed one never finished calibrating, so stop and explain in either case. runs is the number of Runs per Case for every Candidate in this optimization Session. Freeze it for the Session; do not infer it from benchmark_config.toml or the Formal Baseline. Read the evaluation (provider, model_id, thinking_level) from the complete Evaluation that matches the current Agent State; do not require the user to repeat it. An Evaluation without any part of this runtime is incomplete and cannot be used as a Reference. The top-level Session must provide run_subagent, and the current Agent must have the agent-evaluation Skill. If a prerequisite is missing, stop and explain what is needed. Do not create the missing Agent, Benchmark, or Baseline, and do not evaluate the Test Agent directly.
A Reference is the Agent State currently kept as best, together with its complete Evaluation on the frozen Benchmark.
Each round starts from the Reference and tests a bounded, general Candidate. Evaluate every Candidate on the frozen Case set with the requested runs count and the Reference evaluation runtime. The initial Formal Baseline has one Run per Case; do not rerun or backfill it to the requested count. Compare each Candidate's stored top-level average directly with the current Reference score even when their Run counts differ. Accept the Candidate only when the change is admissible, its Evaluation is complete and valid, and its top-level score is strictly higher than the Reference Evaluation's score. An accepted Candidate and its Evaluation become the next Reference; otherwise restore the previous Reference. Stop early when the Reference reaches the desired target; otherwise run no more than the requested number of complete valid Candidate rounds.
Resolve paths from the Environment's App Data Dir without recursively discovering the Project:
PROJECT_DIR = <app_data_dir>
PROJECT_ID = <basename_of_project_dir>
PENGUIN_HOME = <parent_of_project_dir>
TARGET = <app_data_dir>/agents/<test_agent_id>
STATE = <target>/agent_state
TRACES = <target>/traces
BENCHMARK = <app_data_dir>/benchmarks/<benchmark_id>
SCOREBOARD = <benchmark>/scoreboard.yaml
SNAPSHOTS = <target>/snapshotsThe Benchmark is Project-level rather than owned by the Test Agent: it sits beside agents/ and may evaluate several Agents. test_agent_id names the one this Session optimizes, and every Evaluation records it. Use only the Evaluations whose agent_id is that Agent as a Reference or for diagnosis.
Inspect only the requested Test Agent and Benchmark: the Agent State, public Statements, Scoreboard, and score-linked Test Traces or artifacts from the Baseline and this optimization, including rejected Candidates.
Do not inspect Rubrics, Gold answers, private scoring conditions, Evaluator State, Workspace, or Trace, other Agents, or Project secrets. If private evaluation information enters the Optimizer context, restore the active Candidate and stop as contaminated.
Modify only the Test Agent State and the versioned snapshot required to protect it. Do not change the frozen Benchmark, Test Traces, or Project configuration. The only Benchmark write is appending a complete accepted Candidate Evaluation to scoreboard.yaml.
For each round:
runs count.runs[].score on the fixed 0..100 scale; use the Evaluation's top-level average score only for whole-version comparison. Use public Statements, score-linked Test Traces, and prior accepted or rejected attempts to identify observable behaviors that general Agent State changes could improve. Use repeated Runs to distinguish stable behavior from variation.runs matrix in parallel under the evaluation rules below and assemble all returned cells. Do not modify the Candidate while any cell is in flight.score is strictly higher than the Reference Evaluation's score. Otherwise restore the Reference. Record separately whether the predicted Case behavior changed; a higher Evaluation score accepts the Candidate even when the stated hypothesis was not supported.A round counts only after one Candidate has a complete valid Evaluation. Corrected requests, validity repairs, and evaluation retries do not consume the round limit. A complete valid Evaluation of a rejected Candidate does count.
Create one Candidate per round from the current Reference. Put behavioral guidance in AGENTS.md, reusable target-owned capabilities in a focused Skill, and runtime limits in safe system_config.yaml fields. Do not edit system_prompt unless requested, modify library-provided Skills for target-specific behavior, or change model.thinking_level; the Reference Scoreboard fixes the evaluation thinking level.
Candidate version numbers only increase. Start with Reference version + 1 and never reuse a rejected version. Before changing the Agent State, save the original contents and record any files the Candidate creates.
Before changing each Reference State, ensure <target>/snapshots/v<Reference version>.tar.gz exists. Reuse it when present. Otherwise create it yourself before editing by atomically archiving agent_state/ while excluding .vault.toml; validate the archived version and never overwrite an existing same-version snapshot. If snapshot creation fails, stop before changing Agent State and report the failure.
Keep the exact original-file record for fast in-round rollback.
If the Candidate is rejected or cannot be evaluated, restore the Reference files and version, remove files created by the Candidate, and verify the restoration. If another process changes the Agent State, stop without overwriting it.
For each frozen Case, dispatch exactly the requested number of Run cells, using one-based Run indices 1..runs. Call run_subagent for each cell with:
Use the `agent-evaluation` Skill. Run the specified Test Agent on the specified Case exactly once, then score that single execution.
protocol_version: 1
case_id: <case_id>
run: <1_based_run_index>
expected_version: <test_agent_state_version>
test_agent_id: <test_agent_id>
benchmark_id: <benchmark_id>
provider: <provider>
model_id: <model_id>Inspect the complete streamed and final worker response. Before reading status, score, or any other protocol field, verify that the worker-authored text is exactly one plain protocol YAML document. Narration, headings, code fences, summaries, or scoring details are not valid protocol. Ask the same Evaluator to resend only the clean YAML from its existing result; do not rerun the Test Agent for a formatting repair and do not extract YAML from the invalid response yourself. Transport metadata added by run_subagent is not worker-authored text. If private evaluation information appears, follow the contamination rule above.
For every scored result, require its agent_id to equal the requested Test Agent and its actual provider, model_id, and thinking_level to equal the Reference runtime. A mismatch invalidates the Candidate matrix and stops optimization; never compare or record scores produced under a different runtime.
Correct and resend an invalid_request. Stop on version_changed or benchmark_invalid.
For evaluation_failed, keep the same Candidate and incomplete matrix. Ask the same Evaluator to diagnose and repair the failed cell, then rerun only that cell when evidence proves the Test Agent did not start. Every retry must apply a new, specific repair; never repeat an unchanged request or launch, and do not impose a numeric retry limit while distinct safe repairs remain. Do not inspect private Evaluator State or abandon the Candidate to design the next version. Stop when no new safe repair remains, external configuration is required, or it is unclear whether the Test Agent started.
Append each complete accepted Candidate Evaluation to scoreboard.yaml immediately after acceptance and verify the stored Agent id, version, score, matrix, and Session ids before continuing. Obtain the current UTC timestamp from the environment, for example with date -u +"%Y-%m-%dT%H:%M:%SZ", rather than inferring UTC from a displayed local time. Use the same field names as the Baseline:
- time: <ISO-8601 timestamp>
agent_id: <test_agent_id>
version: <Candidate version>
provider: <provider>
model_id: <model_id>
thinking_level: <thinking_level>
summary_title: >-
<public title>
summary: >-
<public summary>
score: <average of the Case scores>
cost: <average of known Case costs, or null when every Case cost is null>
duration_ms: <average of the Case durations>
cases:
- case: <case_id>
score: <average of the Run scores>
cost: <average of known Run costs, or null when every Run cost is null>
duration_ms: <average of the Run durations>
runs:
- score: <Run score>
cost: <Run cost or null>
duration_ms: <Run duration>
session_id: <Test Session id>After writing, parse the complete scoreboard.yaml and verify the appended Evaluation, including its agent_id, before reporting success or continuing.
Every Run and Case score is on the fixed 0..100 scale. Do not write max_score. Calculate and write every Case and Evaluation average directly in the Scoreboard: ignore null values when averaging cost and write null only when all contributing costs are unknown; round score averages to two decimal places, cost averages to six decimal places, and duration_ms averages to the nearest integer. These stored values are authoritative—do not add a server, frontend, script, or consistency check that recomputes or validates them. Do not add an aggregate object or use case_id, mean_score, mean_cost, or mean_duration_ms. Do not record rejected Candidates in the Scoreboard.
Report the Baseline and every fully evaluated Candidate with its score, Run count, version, change, decision, and Test Session ids. Make the one-Run Formal Baseline and requested Candidate runs count explicit. For each Candidate, distinguish the acceptance decision from whether its stated hypothesis was supported by the predicted Case behavior. Include the final retained version, stop reason, and known limitations. Never report a score for an Agent State that was not evaluated.
© Prism-Shadow, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in plugins/agent-tuning/skills/agent-optimization of Prism-Shadow/penguin-harness.
Open the folder on GitHubat commit d56d9ce
Agent Optimization next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Agent Optimization this skillPrism-Shadow/penguin-harness | 2.5k | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | |
| Claw Scoreopenclaw/openclaw | 392k | — | ~2.5k | Automated safety check: Pass | MIT | |
| Internal Linksthedaviddias/Front-End-Checklist | 74k | — | ~770 | Automated safety check: Pass | MIT | |
| Observe Traceruvnet/ruflo | 74k | — | ~522 | Automated safety check: Notes | MIT | |
| External Linksthedaviddias/Front-End-Checklist | 74k | — | ~773 | Automated safety check: Pass | MIT | |
| Arize Linkgithub/awesome-copilot | 40k | 1 repos | ~1.1k | Automated safety check: Pass | MIT |
openclaw/openclaw
Audit or refresh OpenClaw maturity scorecard docs from root taxonomy, maturity scores, and QA evidence artifacts without using maintainer discrawl data or committed inventory reports.
thedaviddias/Front-End-Checklist
A skill your agent uses when auditing a site's internal link structure, identifying pages that need more incoming links, generating contextual linking opportunities between related content, or…
ruvnet/ruflo
Trace agent execution by collecting spans and building a trace tree for a task
thedaviddias/Front-End-Checklist
A skill your agent uses when auditing content pages for citation quality, suggesting authoritative sources to link for factual claims, or reviewing whether a page's external link attributes…
github/awesome-copilot
Generates deep links to the Arize UI for traces, spans, sessions, datasets, labeling queues, evaluators, and annotation configs.
thedaviddias/Front-End-Checklist
A skill your agent uses when auditing a page's link elements for crawlability, reviewing JavaScript-heavy SPAs where navigation may not use <a href tags, or checking that dynamically generated links…
Prism-Shadow/penguin-harness
Make a reply easier to read and act on with rich blocks inside ordinary Markdown — a choice the user picks from, a form that collects several answers, a procedure as steps with warnings in place, a…
Prism-Shadow/penguin-harness
A skill your agent uses when developing PenguinHarness itself — changing packages/{core,server,web,cli,desktop,landing,docs,skills}, the built-in model catalog, the installers or the release…
Prism-Shadow/penguin-harness
Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON.
Prism-Shadow/penguin-harness
A skill your agent uses when standing PenguinHarness up to try a change by hand — launching the Web App, the desktop shell, the landing page, the docs site or the component gallery to click through…
Prism-Shadow/penguin-harness
A skill your agent uses when changing the PenguinHarness Web App (packages/web) or the shared UI package — adding or restyling any UI, picking a status colour, adding an icon, laying out a row or a…
Prism-Shadow/penguin-harness
Drive the PenguinHarness agent browser — the desktop app's built-in browser or the user's own Chrome — from the shell with penguin browser: open pages, read them as simplified HTML or text, act with…
Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark. Agent Optimization is an agent skill from Prism-Shadow/penguin-harness. Improve an Agent State through versioned scores and score-linked Traces from a frozen Benchmark.
Run `npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a claude-code`. Or copy the skill folder (plugins/agent-tuning/skills/agent-optimization in Prism-Shadow/penguin-harness) into .claude/skills/agent-optimization in your project. Claude Code loads it when a task matches its description.
Run `npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a codex`. Or copy the skill folder (plugins/agent-tuning/skills/agent-optimization in Prism-Shadow/penguin-harness) into .agents/skills/agent-optimization in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add Prism-Shadow/penguin-harness --skill agent-optimization -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/agent-optimization, .gemini/skills/agent-optimization, .github/skills/agent-optimization and .opencode/skills/agent-optimization in your project.
SKILL.md names no scripts, command-line tools or credentials: Agent Optimization is instructions for the agent only.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Agent Optimization is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.3k tokens (SKILL.md is roughly 13k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Agent Optimization: Claw Score (openclaw/openclaw, 392k stars), Internal Links (thedaviddias/Front-End-Checklist, 74k stars), Observe Trace (ruvnet/ruflo, 74k stars) and External Links (thedaviddias/Front-End-Checklist, 74k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Prism-Shadow (a GitHub organization) maintains it in Prism-Shadow/penguin-harness, which has 2,455 GitHub stars. The repository holds 31 skills in this directory. The repository was last updated on October 7, 2026.
Source: Prism-Shadow/penguin-harness on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.